State of Web Access2026

13% of top landing pages now restrict AI crawlers in robots.txt. GPTBot faces the most named blocks (8.4% of all sites); TurnitinBot draws 233 Disallow rules and zero Allows.

State of Web Access 2026
AI Crawler Access Posture

Nearly 4 in 10 sites effectively block AI crawlers — most do so without naming them

Nearly 4 in 10 sites effectively block AI crawlers — most do so without naming them
Explicitly blocks AI by nameBlocks all bots — AI caught by defaultExplicitly allows AI crawlersrobots.txt — no AI positionNo robots.txt
12.7%27.2%1.8%34.1%24.2%

12.7% of sites name AI crawlers in a Disallow rule. A further 27.2% block all bots via wildcard — AI is caught by default, not by intent. Only 1.8% actively welcome AI with explicit Allow rules. The remaining 34.1% hold robots.txt but take no position on AI.

11,100 landing pages
State of Web Access 2026
AI Crawler Access · robots.txt Directives

GPTBot and CCBot face the most blocks — but search-linked AI agents attract more allows

GPTBot and CCBot face the most blocks — but search-linked AI agents attract more allows
DisallowedAllowed (whitelisted)
GPTBot8.4%1.7%
CCBot7.3%0.6%
ClaudeBot6.3%1%
Google-Extended5.6%1.1%
Bytespider5.5%0.3%
meta-externalagent4.4%0.5%
ChatGPT-User4.3%1.7%
PerplexityBot4.2%1.5%
Applebot-Extended3.9%0.4%
OAI-SearchBot2.9%1.7%
TurnitinBot2.1%

The Disallow/Allow ratio reveals commercial intent: ChatGPT-User and OAI-SearchBot attract blocks and explicit whitelists in roughly equal measure, reflecting their value as AI-powered search referrers. CCBot and Bytespider are blocked almost without exception — sites that name them have only one thing to say.

11,100 landing pages · 8,419 with valid robots.txt
State of Web Access 2026
AI Crawler Access · Disallow:Allow Ratio

robots.txt rules favour Disallow over Allow

robots.txt rules favour Disallow over Allow
Disallow:Allow
OAI-SearchBot2:1
ChatGPT-User3:1
PerplexityBot3:1
GPTBot5:1
Google-Extended5:1
ClaudeBot6:1
meta-externalagent8:1
Applebot-Extended10:1
CCBot12:1
Bytespider17:1
TurnitinBotDisallow:Allow = ∞ (233 blocks, 0 allows)

Blocks outnumber allows for every named AI agent. OAI-SearchBot is the most contested at 2:1; Bytespider and CCBot are blocked almost without exception. TurnitinBot stands alone: 233 blocks, zero allows.

11,100 landing pages · 8,419 with valid robots.txt
State of Web Access 2026
Explicit AI Crawler Restrictions · Top 15 Categories

Newspapers are the most prolific explicit blockers of specific AI crawlers

Newspapers are the most prolific explicit blockers of specific AI crawlers
Value
Newspapers64%
Publishing49%
Mass Media43%
Online Services37%
Sports33%
Entertainment30%
Retail30%
Distance Learning26%
Music25%
Travel & Tourism25%
Visual Art24%
Automotive23%
Computer & Video Games23%
Food & Beverages23%
Wellness22%

64% of newspaper sites name at least one AI crawler — GPTBot, ClaudeBot, CCBot and others — in a Disallow rule, well ahead of publishing and mass media. The pattern is economic: content producers are protecting the text AI companies need for training.

11,100 landing pages
State of Web Access 2026
Explicit AI Crawler Restrictions · Bottom 15 Categories

Services and logistics show near-zero explicit AI crawler restrictions

Services and logistics show near-zero explicit AI crawler restrictions
Value
Restaurants3%
Transport & Logistics3%
Public Utility3%
Printing3%
Mail & Package Delivery3%
Insurance3%
Furniture3%
Events Services3%
Fishery2%
Human Resources2%
Venture Capital2%
Accounting & Auditing1%
Facilities Services1%
Outsourcing0%
Market Research0%

Categories with no text asset to protect barely register. Market research, outsourcing and facilities services cluster at or near zero — their value is in services, not publishable content.

11,100 landing pages
State of Web Access 2026
Explicit AI Crawler Allows · Top 15 Categories

Real estate and travel sites lead on explicitly welcoming AI crawlers

Real estate and travel sites lead on explicitly welcoming AI crawlers
Value
Real Estate13%
Travel & Tourism13%
Airlines10%
Newspapers10%
Hospitality8%
Telecom8%
Software & Development7%
Retail7%
Automotive6%
Publishing6%
Wellness6%
Design5%
Finance5%
Furniture5%
Mass Media5%

13% of real estate and travel & tourism sites explicitly allow at least one AI crawler. Newspapers — which also lead on named AI blocking — sit joint third at 10%, making them the category most engaged with AI access policy in both directions.

11,100 landing pages