State of Web Access2026

75.8% of top landing pages publish a robots.txt, but just 30.7% address specific bots by name — most rely on wildcard rules alone.

State of Web Access 2026
robots.txt Coverage

Three in four websites publish a robots.txt file

Three in four websites publish a robots.txt file
have a valid robots.txt
robots.txt Coverage75.8%

75.8% of landing pages return a valid robots.txt — a higher baseline than any individual technical defence. Everything that follows covers the 8,419 sites that have one.

11,100 landing pages
State of Web Access 2026
robots.txt Sophistication with Valid robots.txt

Most sites with robots.txt have made a deliberate access decision

Most sites with robots.txt have made a deliberate access decision
BasicIntermediateAdvancedExpert
11.2%23.5%56.8%8.5%

56.9% of sites with valid robots.txt are Advanced — they've made a one-sided access decision (either a whitelist or a blacklist, but not both). Only 11.2% have published nothing beyond wildcard rules. The 8.5% with Expert status actively manage access from both ends simultaneously.

8,419 sites with valid robots.txt
State of Web Access 2026
robots.txt Named Agents

Seven in ten sites name no crawler specifically — but the most attentive track hundreds by name

Seven in ten sites name no crawler specifically — but the most attentive track hundreds by name
Value
069.3%
16.3%
2–49.4%
5–95.3%
10–193.9%
20–493.8%
50+2%

30.7% of all sites (3,410) name at least one specific User-agent in their robots.txt rather than relying on the wildcard catch-all. Among those, granularity varies widely: most name fewer than five agents, but 9.7% address ten or more distinct crawlers. The most elaborate file in the dataset names 1,774 separate agents.

11,100 landing pages
State of Web Access 2026
robots.txt Sophistication · By Industry

Content industries lead on granular bot management — two thirds of newspapers name 5 or more crawlers

Content industries lead on granular bot management — two thirds of newspapers name 5 or more crawlers
Value
Newspapers66%
Mass Media48%
Publishing48%
Online Services44%
Travel & Tourism36%
Retail35%
Sports33%
Entertainment31%
Automotive30%
Real Estate30%
Music27%
Computer & Video Games26%
Photography24%
Airlines23%
Beauty & Cosmetics23%

66% of newspaper sites explicitly name five or more distinct crawlers by User-agent — more than any other industry, and more than double the rate of travel and retail. Mass media and publishing cluster just behind. High specificity reflects both the commercial stakes and legal risk of content being used for AI training.

11,100 landing pages
State of Web Access 2026
robots.txt Rule Architecture

Most robots.txt files use one-directional rules per bot — but 3 in 10 of all sites mix allows and blocks for the same agent

Most robots.txt files use one-directional rules per bot — but 3 in 10 of all sites mix allows and blocks for the same agent
Value
Uniform per-agent rules (no path mixing)45.9%
Path-level mixed rules for same agent29.9%
Allow-only (no disallow rules)3.1%
Mixed AI posture (block some, allow others)1.9%
Blanket Disallow: /1.4%

46% of all sites use uniform per-agent rules — all-block or all-allow, no path exceptions. 30% mix allows and blocks for the same bot at path level: blocked on /admin/ but welcome on /blog/. A meaningful 3.1% publish no Disallow rules at all — a deliberate open posture.

11,100 landing pages
State of Web Access 2026
robots.txt Bot Categories · % of All Sites Mentioning Any Agent

SEO bots lead by volume — AI crawlers now rank second

SEO bots lead by volume — AI crawlers now rank second
Value
SEO / search indexing19.2%
AI crawlers13.7%
SEO audit tools11.3%
Commercial scrapers7.1%
Social / content scrapers6.6%
Archive / research crawlers5.3%
Monitoring / uptime bots0.7%

The AI debate is loud but SEO tools command more robots.txt attention. AI crawlers rank second — named by 13.7% of all sites — ahead of SEO audit tools. See RT11 for individual commercial scraper detail.

11,100 landing pages
State of Web Access 2026
Commercial Scrapers · robots.txt Disallow Frequency

Diffbot is the most blocked commercial scraper — site-copier tools appear as a template cluster

Diffbot is the most blocked commercial scraper — site-copier tools appear as a template cluster
ValueCategory
Diffbot 321 sitesDeliberate, targeted block
WebCopier 207 sitesTemplate-propagated
WebZip 201 sitesTemplate-propagated
WebStripper 198 sitesTemplate-propagated
Offline Explorer 197 sitesTemplate-propagated
Teleport 190 sitesTemplate-propagated
SiteSnagger 189 sitesTemplate-propagated
HTTrack 175 sitesTemplate-propagated
Larbin 170 sitesTemplate-propagated
aiHitBot 105 sitesDeliberate, targeted block
EmailCollector 104 sitesDeliberate, targeted block
EmailSiphon 103 sitesDeliberate, targeted block
TrendictionBot 91 sitesDeliberate, targeted block
Sidetrade indexer 86 sitesDeliberate, targeted block
CazoodleBot 85 sitesDeliberate, targeted block

Diffbot is blocked by 321 sites — nearly twice any other agent. The site-copier cluster (WebCopier, WebZip, HTTrack, etc.) appears together so consistently that shared blacklist templates are the likely explanation. The data aggregators and email harvesters at 85–105 sites each appear independently — deliberate, targeted blocks.

11,100 landing pages · 8,419 with valid robots.txt
State of Web Access 2026
Crawl Delay · Value Distribution

10 seconds is the default — used by nearly half of all delay-setting sites

10 seconds is the default — used by nearly half of all delay-setting sites
Value
1s14.8%
2s4.3%
3s5.3%
5s12.5%
10s44.6%
20s2.8%
30s4.8%
60s+3.7%

The 10s spike almost certainly reflects a copied CMS default, not a calculated threshold. At 30s and above, comprehensive crawling within any commercial timeframe becomes impractical.

1,373 sites with crawl-delay directive
State of Web Access 2026
Crawl Delay · Average Per Agent

SEO audit tools are throttled hardest — social, search, and AI crawlers follow at a fraction of the delay

SEO audit tools are throttled hardest — social, search, and AI crawlers follow at a fraction of the delay
ValueCategory
MJ12bot 352s average crawl delaySEO audit tools
AhrefsBot 291s average crawl delaySEO audit tools
BLEXBot 83s average crawl delaySEO audit tools
Twitterbot 66s average crawl delaySocial / content
Facebot 48s average crawl delaySocial / content
Yandex 39s average crawl delaySEO / search indexing
SemrushBot 39s average crawl delaySEO audit tools
Pinterestbot 38s average crawl delaySocial / content
AmazonBot 29s average crawl delayCommercial scrapers
FacebookExternalHit 28s average crawl delaySocial / content
OAI-SearchBot 28s average crawl delayAI crawlers
BaiduSpider 27s average crawl delaySEO / search indexing
PerplexityBot 23s average crawl delayAI crawlers
PetalBot 23s average crawl delaySEO / search indexing
DotBot 21s average crawl delaySEO audit tools
MSNBot 21s average crawl delaySEO / search indexing
Slurp 18s average crawl delaySEO / search indexing
AhrefsSiteAudit 18s average crawl delaySEO audit tools
ClaudeBot 16s average crawl delayAI crawlers
Bytespider 16s average crawl delayAI crawlers
CCBot 15s average crawl delayAI crawlers
BingBot 14s average crawl delaySEO / search indexing
GPTBot 14s average crawl delayAI crawlers
ChatGPT-User 14s average crawl delayAI crawlers
Googlebot 14s average crawl delaySEO / search indexing
ia_archiver 12s average crawl delayArchive / research

SEO audit tools are throttled to near-impracticality — passive exclusion dressed as courtesy. Social bots are slowed; search engines receive crawl-budget management. AI agents sit at the bottom — inconvenienced, not neutralised.

Sites with per-agent crawl-delay directives