State of Web Access2026

No sector has been louder about AI access policy — and the robots.txt data shows it. 64% restrict AI crawlers by name. But technically, newspapers are among the least defended sites in the survey: antibot appears on just 8% of sites.

State of Web Access 2026
Newspapers · Barrier Profile · 2026

CAPTCHA and JavaScript define Newspapers' active barrier layer

CAPTCHA and JavaScript define Newspapers' active barrier layer
Value
WAF96%
Antibot8%
CAPTCHA21%
JavaScript36%
Rate Limiting15%
TLS10%

WAF coverage at 96% is near-universal. JavaScript rendering (36%) and CAPTCHA (21%) are the sector's defining active barriers, reflecting the value publishers place on controlling how their content is accessed. Antibot tools feature in only 8% of sites — publishers prioritise access policy via robots.txt over technical bot fingerprinting.

100 sites
State of Web Access 2026
Newspapers · Barrier & Policy Adoption · 2026

JavaScript tops Newspapers' active secondary defences

JavaScript tops Newspapers' active secondary defences
Value
WAF96%
Antibot8%
CAPTCHA21%
JavaScript36%
Rate Limiting15%
TLS10%
Has robots.txt90%
Blocks AI crawlers64%

WAF is deployed by 96% of Newspaper sites. JavaScript rendering leads active secondary barriers at 36%, followed by CAPTCHA at 21% and Rate Limiting at 15%. Nine in ten sites publish a robots.txt file and 64% use it to block AI crawlers — a policy commitment that far outpaces technical barrier deployment in this sector.

100 sites
State of Web Access 2026
Newspapers · Barrier Count per Site · 2026

Single-barrier sites are Newspapers' most common configuration

Single-barrier sites are Newspapers' most common configuration
% of sites
0 barriers2%
1 barrier45%
2 barriers27%
3 barriers20%
4 barriers3%
5 barriers3%
6 barriers0%

45% of Newspaper sites carry exactly one barrier — almost always WAF operating alone. Two-barrier configurations (27%) typically pair WAF with JavaScript rendering. Just 6% of sites reach four or more barriers, concentrated among the largest national titles with commercially sensitive content.

100 sites
State of Web Access 2026
Newspapers · Recommended Zyte API Tier · 2026

Nearly half of Newspaper sites are Simple to access

Nearly half of Newspaper sites are Simple to access
Mean recommended tier
All sites1.77

A mean recommended tier of 1.77 and a median of Easy place Newspapers around the middle of the sectors surveyed. 46% of sites are Simple to access; 17% require Moderate or above.

100 sites
State of Web Access 2026
Newspapers · Zyte API Tier Distribution · 2026

Most Newspaper sites are Simple or Easy to access

Most Newspaper sites are Simple or Easy to access
% of sites
1 - Simple46%
2 - Easy37%
3 - Moderate12%
4 - Complex4%
5 - Advanced1%

Simple (46%) is the most common tier. Simple and Easy together account for 83% of sites, while 12% land at Moderate and 5% at Complex or Advanced.

100 sites
State of Web Access 2026
Newspapers · robots.txt Coverage · 2026 · 100 Sites

Nine in ten Newspaper sites publish a robots.txt

Nine in ten Newspaper sites publish a robots.txt
% of sites
Publish a valid robots.txt90%

robots.txt adoption in Newspapers is among the highest of any sector surveyed, with 90% of sites publishing a valid file. Just 10% leave crawl policy undefined — a significantly lower rate than the global average of around 24%.

100 sites
State of Web Access 2026
Newspapers · Named Bot Count per Site · 2026

Half of all Newspaper sites name 16 or more bots

Half of all Newspaper sites name 16 or more bots
% of sites
011%
113%
2–49%
5–99%
10–158%
16+50%

The median Newspaper site names 15 distinct bots in its robots.txt and the mean of 33.8 reflects the long tail of sites maintaining exhaustive bot lists. Half of all Newspaper sites name 16 or more bots by name. This level of specificity is rare across sectors and reflects deliberate, actively maintained crawl policy in newsrooms.

100 sites
State of Web Access 2026
Newspapers · Bot Category Mentions in robots.txt · 2026 · 100 Sites

AI crawlers top Newspaper bot categories in robots.txt

AI crawlers top Newspaper bot categories in robots.txt
% of sites
AI crawlers66%
SEO / search indexing29%
SEO audit tools40%
Commercial scrapers39%
Social / content scrapers10%
Archive / research crawlers21%
Monitoring / uptime bots2%

66% of Newspaper sites name at least one AI crawler in their robots.txt — the most mentioned bot category in the sector. SEO audit tools (40%) and commercial scrapers (39%) follow closely. Figures are percentages of all 100 Newspaper sites, including the 10% with no robots.txt at all.

100 sites
State of Web Access 2026
Newspapers · robots.txt Rule Posture Types · 2026 · 100 Sites

Uniform and path-mixed rules split Newspaper robots.txt files evenly

Uniform and path-mixed rules split Newspaper robots.txt files evenly
% of sites
Uniform per-agent rules41%
Path-level mixed rules42%
Allow-only (no Disallow rules)0%
Mixed AI posture13%
Blanket Disallow: /3%

Of the 90% of Newspaper sites with a valid robots.txt, posture splits roughly evenly between uniform per-agent rules (41%) and path-level mixed rules (42%). A further 13% maintain a mixed AI posture — blocking some AI crawlers while explicitly allowing others. Only 3% issue a blanket Disallow: /. Figures are percentages of all 100 sites.

100 sites
State of Web Access 2026
Newspapers · Bot Agent Disallow vs Allow · 2026 · 100 Sites Surveyed

GPTBot and CCBot face the most Disallow rules in Newspaper robots.txt

GPTBot and CCBot face the most Disallow rules in Newspaper robots.txt
DisallowedAllowed (whitelisted)
*86%36%
GPTBot51%3%
Google-Extended46%6%
CCBot51%1%
PerplexityBot38%6%
anthropic-ai42%2%
ClaudeBot41%2%
ChatGPT-User34%8%
Claude-Web36%3%
Bytespider36%2%
OAI-SearchBot25%12%
cohere-ai34%2%
Applebot-Extended33%2%
Amazonbot29%3%
Diffbot30%1%
FacebookBot27%2%
omgilibot27%1%
Meta-ExternalAgent25%2%
omgili26%1%
YouBot22%2%
PetalBot20%1%
Timpibot20%1%
TurnitinBot20%
Scrapy18%1%
ImagesiftBot18%1%

Values are percentages of all 100 Newspaper sites surveyed. The wildcard (*) at 86% Disallow reflects a default-restrict posture applied to uncategorised crawlers. Among named bots, GPTBot and CCBot lead blocks at 51% each. OAI-SearchBot (12% Allow) and ChatGPT-User (8% Allow) attract the most explicit whitelisting — a pattern of restricting training crawlers while selectively permitting search-facing agents.

100 sites surveyed · 90 with robots.txt
State of Web Access 2026
Media & Publishing · Access Barriers · 2026

Newspapers lead their cohort on AI blocking, not on technical barriers

Newspapers lead their cohort on AI blocking, not on technical barriers
WAFAntibotCAPTCHAJavaScriptRate LimitTLSrobots.txtBlocks AITier
Advertising & Marketing97%23%32%31%21%14%87%6%Tier 1
Mass Media98%12%21%49%16%8%92%43%Tier 2
Newspapers96%8%21%36%15%10%90%64%Tier 2
Public Relations95%10%14%25%14%14%76%6%Tier 1
Publishing95%16%24%45%21%11%92%49%Tier 2
Writing & Editing98%14%33%31%11%6%79%11%Tier 1

Newspapers' 64% AI-blocking rate is the highest in the Media & Publishing cohort — around 15 points above Publishing, the next-highest sub-category. JavaScript adoption (36%) also leads the group. Antibot (8%) and TLS (10%) sit at or near the cohort floor, consistent with a sub-category that favours robots.txt policy over middleware deployment.

11,100 landing pages