
Nearly 4 in 10 sites effectively block AI crawlers — most do so without naming them
| Explicitly blocks AI by name | Blocks all bots — AI caught by default | Explicitly allows AI crawlers | robots.txt — no AI position | No robots.txt | |
|---|---|---|---|---|---|
| 12.7% | 27.2% | 1.8% | 34.1% | 24.2% |
12.7% of sites name AI crawlers in a Disallow rule. A further 27.2% block all bots via wildcard — AI is caught by default, not by intent. Only 1.8% actively welcome AI with explicit Allow rules. The remaining 34.1% hold robots.txt but take no position on AI.
By bot

GPTBot and CCBot face the most blocks — but search-linked AI agents attract more allows
| Disallowed | Allowed (whitelisted) | |
|---|---|---|
| GPTBot | 8.4% | 1.7% |
| CCBot | 7.3% | 0.6% |
| ClaudeBot | 6.3% | 1% |
| Google-Extended | 5.6% | 1.1% |
| Bytespider | 5.5% | 0.3% |
| meta-externalagent | 4.4% | 0.5% |
| ChatGPT-User | 4.3% | 1.7% |
| PerplexityBot | 4.2% | 1.5% |
| Applebot-Extended | 3.9% | 0.4% |
| OAI-SearchBot | 2.9% | 1.7% |
| TurnitinBot | 2.1% | — |
The Disallow/Allow ratio reveals commercial intent: ChatGPT-User and OAI-SearchBot attract blocks and explicit whitelists in roughly equal measure, reflecting their value as AI-powered search referrers. CCBot and Bytespider are blocked almost without exception — sites that name them have only one thing to say.

robots.txt rules favour Disallow over Allow
| Disallow:Allow | |
|---|---|
| OAI-SearchBot | 2:1 |
| ChatGPT-User | 3:1 |
| PerplexityBot | 3:1 |
| GPTBot | 5:1 |
| Google-Extended | 5:1 |
| ClaudeBot | 6:1 |
| meta-externalagent | 8:1 |
| Applebot-Extended | 10:1 |
| CCBot | 12:1 |
| Bytespider | 17:1 |
| TurnitinBot | Disallow:Allow = ∞ (233 blocks, 0 allows) |
Blocks outnumber allows for every named AI agent. OAI-SearchBot is the most contested at 2:1; Bytespider and CCBot are blocked almost without exception. TurnitinBot stands alone: 233 blocks, zero allows.
By industry

Newspapers are the most prolific explicit blockers of specific AI crawlers
| Value | |
|---|---|
| Newspapers | 64% |
| Publishing | 49% |
| Mass Media | 43% |
| Online Services | 37% |
| Sports | 33% |
| Entertainment | 30% |
| Retail | 30% |
| Distance Learning | 26% |
| Music | 25% |
| Travel & Tourism | 25% |
| Visual Art | 24% |
| Automotive | 23% |
| Computer & Video Games | 23% |
| Food & Beverages | 23% |
| Wellness | 22% |
64% of newspaper sites name at least one AI crawler — GPTBot, ClaudeBot, CCBot and others — in a Disallow rule, well ahead of publishing and mass media. The pattern is economic: content producers are protecting the text AI companies need for training.

Services and logistics show near-zero explicit AI crawler restrictions
| Value | |
|---|---|
| Restaurants | 3% |
| Transport & Logistics | 3% |
| Public Utility | 3% |
| Printing | 3% |
| Mail & Package Delivery | 3% |
| Insurance | 3% |
| Furniture | 3% |
| Events Services | 3% |
| Fishery | 2% |
| Human Resources | 2% |
| Venture Capital | 2% |
| Accounting & Auditing | 1% |
| Facilities Services | 1% |
| Outsourcing | 0% |
| Market Research | 0% |
Categories with no text asset to protect barely register. Market research, outsourcing and facilities services cluster at or near zero — their value is in services, not publishable content.

Real estate and travel sites lead on explicitly welcoming AI crawlers
| Value | |
|---|---|
| Real Estate | 13% |
| Travel & Tourism | 13% |
| Airlines | 10% |
| Newspapers | 10% |
| Hospitality | 8% |
| Telecom | 8% |
| Software & Development | 7% |
| Retail | 7% |
| Automotive | 6% |
| Publishing | 6% |
| Wellness | 6% |
| Design | 5% |
| Finance | 5% |
| Furniture | 5% |
| Mass Media | 5% |
13% of real estate and travel & tourism sites explicitly allow at least one AI crawler. Newspapers — which also lead on named AI blocking — sit joint third at 10%, making them the category most engaged with AI access policy in both directions.