The debate about AI companies and web content has generated an enormous amount of heat.
Publishers say AI companies are taking their content without compensation. AI companies say they are indexing the public web as search engines always have. Lawyers are arguing about what training data is, and what robots.txt actually means in a legal context.
Zyte's State of Web Access 2026 does not address those questions. But it does offer something the debate has largely lacked: a ground-level measurement of what site operators are actually doing about it, across 11,100 of the world's most popular landing pages.
How sites block AI with a text file

The headline numbers?
- 12.7% of all sites now name at least one named AI crawler in a
Disallowrule - an explicit, named policy decision targeting AI access specifically. - A further 27.2% block all automated traffic via wildcard, catching AI crawlers by default without naming them.
- Only 1.8% actively welcome AI with explicit
Allowrules.
The data reveals not a uniform response but a fractured one - with clear distinctions between training crawlers and search-linked agents, between content industries and service industries, and between blocking and welcoming that sometimes coexist on the same site.
GPTBot leads the block list, but the ratios tell a different story
Among the eleven named AI agents in our dataset, GPTBot - OpenAI's primary training crawler - faces the most named Disallow rules: 8.4% of all 11,100 sites.

CCBot, operated by Common Crawl and used historically as a training data source, sits at 7.3%.
ClaudeBot (Anthropic) is at 6.3%, Google-Extended at 5.6%, and Bytespider (ByteDance) at 5.5%.
But the Disallow count alone obscures a more interesting pattern.
AI search is more favoured than AI crawling

The Disallow:Allow ratio separates the agents into two groups with fundamentally different commercial positions.
OAI-SearchBot - the crawler behind ChatGPT's web search feature - is blocked at a 2:1 ratio: the most balanced in the dataset. ChatGPT-User and PerplexityBot are blocked at roughly 3:1.
These agents power AI search products that can send referral traffic to publishers - a dynamic sites recognise and some have decided is worth accommodating.
Training crawlers carry no such proposition. GPTBot is blocked at 5:1. CCBot at 12:1. Bytespider at 17:1. Sites that name these agents in robots.txt have almost uniformly one thing to say.
The training/search distinction
The divergence between training crawlers and search-linked agents reflects a commercial logic that publishers have articulated explicitly in the legal cases and licensing discussions of the past two years: search engines send traffic; AI seeks training content.
A site that blocks GPTBot but allows or tolerates OAI-SearchBot is drawing a precise distinction - between a crawler that builds a product competing with the publisher and one that distributes the publisher's content to users. That distinction shows up cleanly in the ratio data.
Google-Extended - Google's opt-out string for Gemini training and Vertex AI - sits at 5:1, in line with other pure training crawlers. Applebot-Extended at 10:1. Neither carries a search referral proposition compelling enough to shift the ratio. The pattern holds: agents tied to AI search products face balanced responses; agents associated with model training face near-uniform exclusion.
Newspapers are most likely to block AI agents
The industry distribution of explicit AI blocking is dominated by content producers.

Newspapers lead: 64% of newspaper sites name at least one AI crawler in a Disallow rule. Publishing sits at 49%, mass media at 43%.
These are the sectors most directly threatened by AI systems trained on their output and capable of generating competing content at scale without the editorial infrastructure.
Sports at 33% and entertainment at 30% reflect similar dynamics - content that has commercial value beyond the original site, that can be repurposed or summarised, and that publishers have decided they prefer not to feed into training pipelines without compensation.
At the other end of the spectrum, market research and outsourcing register zero explicit AI blocking. Transport, logistics, printing, and utilities cluster at 2–3%. These sectors have no text asset that requires protection from AI training - their value is in services, not publishable content, and the economics of AI access policy simply don't apply.

Many sectors actively welcome some AI bots
The most striking finding in the allow data is that newspapers - which lead on explicit AI blocking at 64% - also sit joint third on explicit AI allowing, at 10%.

Real estate and travel and tourism lead the allow side at 13% each, airlines at 10%.
This is no contradiction but the search-linked distinction in practice: a newspaper that blocks GPTBot and CCBot while explicitly allowing OAI-SearchBot and PerplexityBot is executing a precise commercial decision about which AI agents return value. Newspapers are the industry most actively engaged with AI access policy in both directions simultaneously.
Real estate and travel leading on explicit allows reflects a different calculation: these are inventory-heavy sectors where AI-powered search could drive significant referral and booking traffic. A travel site that allows PerplexityBot is betting that AI search will surface its inventory to users who then click through. The economics of that wager look different from the economics facing a news publisher.
Catch-all wildcard rules block the most AI crawlers
The 12.7% who name AI crawlers explicitly represent only one part of the picture. The larger group - 27.2% of all sites - blocks AI crawlers without naming them, via wildcard User-agent: * with a blanket Disallow.
These sites have not made a specific policy decision about AI. They have made a general decision to close their robots.txt to all automated access, and AI crawlers are caught as part of that.
Whether this reflects intentional AI exclusion or simply inherited configuration - a Disallow: / that predates the AI crawler debate by a decade - cannot be determined from the file itself.
The combined effective block rate is 39.9%. Nearly four in ten of the world's most popular sites are inaccessible to a well-behaved AI crawler that respects robots.txt.











