
Three in four websites publish a robots.txt file
| have a valid robots.txt | |
|---|---|
| robots.txt Coverage | 75.8% |
75.8% of landing pages return a valid robots.txt — a higher baseline than any individual technical defence. Everything that follows covers the 8,419 sites that have one.
The sophistication spectrum

Most sites with robots.txt have made a deliberate access decision
| Basic | Intermediate | Advanced | Expert | |
|---|---|---|---|---|
| 11.2% | 23.5% | 56.8% | 8.5% |
56.9% of sites with valid robots.txt are Advanced — they've made a one-sided access decision (either a whitelist or a blacklist, but not both). Only 11.2% have published nothing beyond wildcard rules. The 8.5% with Expert status actively manage access from both ends simultaneously.

Seven in ten sites name no crawler specifically — but the most attentive track hundreds by name
| Value | |
|---|---|
| 0 | 69.3% |
| 1 | 6.3% |
| 2–4 | 9.4% |
| 5–9 | 5.3% |
| 10–19 | 3.9% |
| 20–49 | 3.8% |
| 50+ | 2% |
30.7% of all sites (3,410) name at least one specific User-agent in their robots.txt rather than relying on the wildcard catch-all. Among those, granularity varies widely: most name fewer than five agents, but 9.7% address ten or more distinct crawlers. The most elaborate file in the dataset names 1,774 separate agents.

Content industries lead on granular bot management — two thirds of newspapers name 5 or more crawlers
| Value | |
|---|---|
| Newspapers | 66% |
| Mass Media | 48% |
| Publishing | 48% |
| Online Services | 44% |
| Travel & Tourism | 36% |
| Retail | 35% |
| Sports | 33% |
| Entertainment | 31% |
| Automotive | 30% |
| Real Estate | 30% |
| Music | 27% |
| Computer & Video Games | 26% |
| Photography | 24% |
| Airlines | 23% |
| Beauty & Cosmetics | 23% |
66% of newspaper sites explicitly name five or more distinct crawlers by User-agent — more than any other industry, and more than double the rate of travel and retail. Mass media and publishing cluster just behind. High specificity reflects both the commercial stakes and legal risk of content being used for AI training.
Rule architecture

Most robots.txt files use one-directional rules per bot — but 3 in 10 of all sites mix allows and blocks for the same agent
| Value | |
|---|---|
| Uniform per-agent rules (no path mixing) | 45.9% |
| Path-level mixed rules for same agent | 29.9% |
| Allow-only (no disallow rules) | 3.1% |
| Mixed AI posture (block some, allow others) | 1.9% |
| Blanket Disallow: / | 1.4% |
46% of all sites use uniform per-agent rules — all-block or all-allow, no path exceptions. 30% mix allows and blocks for the same bot at path level: blocked on /admin/ but welcome on /blog/. A meaningful 3.1% publish no Disallow rules at all — a deliberate open posture.
The wider bot ecosystem

SEO bots lead by volume — AI crawlers now rank second
| Value | |
|---|---|
| SEO / search indexing | 19.2% |
| AI crawlers | 13.7% |
| SEO audit tools | 11.3% |
| Commercial scrapers | 7.1% |
| Social / content scrapers | 6.6% |
| Archive / research crawlers | 5.3% |
| Monitoring / uptime bots | 0.7% |
The AI debate is loud but SEO tools command more robots.txt attention. AI crawlers rank second — named by 13.7% of all sites — ahead of SEO audit tools. See RT11 for individual commercial scraper detail.

Diffbot is the most blocked commercial scraper — site-copier tools appear as a template cluster
| Value | Category | |
|---|---|---|
| Diffbot | 321 sites | Deliberate, targeted block |
| WebCopier | 207 sites | Template-propagated |
| WebZip | 201 sites | Template-propagated |
| WebStripper | 198 sites | Template-propagated |
| Offline Explorer | 197 sites | Template-propagated |
| Teleport | 190 sites | Template-propagated |
| SiteSnagger | 189 sites | Template-propagated |
| HTTrack | 175 sites | Template-propagated |
| Larbin | 170 sites | Template-propagated |
| aiHitBot | 105 sites | Deliberate, targeted block |
| EmailCollector | 104 sites | Deliberate, targeted block |
| EmailSiphon | 103 sites | Deliberate, targeted block |
| TrendictionBot | 91 sites | Deliberate, targeted block |
| Sidetrade indexer | 86 sites | Deliberate, targeted block |
| CazoodleBot | 85 sites | Deliberate, targeted block |
Diffbot is blocked by 321 sites — nearly twice any other agent. The site-copier cluster (WebCopier, WebZip, HTTrack, etc.) appears together so consistently that shared blacklist templates are the likely explanation. The data aggregators and email harvesters at 85–105 sites each appear independently — deliberate, targeted blocks.
Crawl delay

10 seconds is the default — used by nearly half of all delay-setting sites
| Value | |
|---|---|
| 1s | 14.8% |
| 2s | 4.3% |
| 3s | 5.3% |
| 5s | 12.5% |
| 10s | 44.6% |
| 20s | 2.8% |
| 30s | 4.8% |
| 60s+ | 3.7% |
The 10s spike almost certainly reflects a copied CMS default, not a calculated threshold. At 30s and above, comprehensive crawling within any commercial timeframe becomes impractical.

SEO audit tools are throttled hardest — social, search, and AI crawlers follow at a fraction of the delay
| Value | Category | |
|---|---|---|
| MJ12bot | 352s average crawl delay | SEO audit tools |
| AhrefsBot | 291s average crawl delay | SEO audit tools |
| BLEXBot | 83s average crawl delay | SEO audit tools |
| Twitterbot | 66s average crawl delay | Social / content |
| Facebot | 48s average crawl delay | Social / content |
| Yandex | 39s average crawl delay | SEO / search indexing |
| SemrushBot | 39s average crawl delay | SEO audit tools |
| Pinterestbot | 38s average crawl delay | Social / content |
| AmazonBot | 29s average crawl delay | Commercial scrapers |
| FacebookExternalHit | 28s average crawl delay | Social / content |
| OAI-SearchBot | 28s average crawl delay | AI crawlers |
| BaiduSpider | 27s average crawl delay | SEO / search indexing |
| PerplexityBot | 23s average crawl delay | AI crawlers |
| PetalBot | 23s average crawl delay | SEO / search indexing |
| DotBot | 21s average crawl delay | SEO audit tools |
| MSNBot | 21s average crawl delay | SEO / search indexing |
| Slurp | 18s average crawl delay | SEO / search indexing |
| AhrefsSiteAudit | 18s average crawl delay | SEO audit tools |
| ClaudeBot | 16s average crawl delay | AI crawlers |
| Bytespider | 16s average crawl delay | AI crawlers |
| CCBot | 15s average crawl delay | AI crawlers |
| BingBot | 14s average crawl delay | SEO / search indexing |
| GPTBot | 14s average crawl delay | AI crawlers |
| ChatGPT-User | 14s average crawl delay | AI crawlers |
| Googlebot | 14s average crawl delay | SEO / search indexing |
| ia_archiver | 12s average crawl delay | Archive / research |
SEO audit tools are throttled to near-impracticality — passive exclusion dressed as courtesy. Social bots are slowed; search engines receive crawl-budget management. AI agents sit at the bottom — inconvenienced, not neutralised.