Retail’s biggest marketplaces are built to be found. Well-known everything-stores are among the most visited commercial sites on the planet, and their whole model rests on being browsed at enormous scale.
So our research is telling: retail sites are the second-hardest industry for data gathering bots to access, according to Zyte’s State of Web Access 2026, second only to apparel and fashion.
The average retail site scores 2.60 out of five on Zyte’s five-tier scraping complexity scale, ahead of every sector except fashion and well above the research-wide average of 1.58.
What we're looking at
In this research, “retail” is the general-merchandise catch-all: the giant marketplaces and department-store platforms that sell across categories rather than a single vertical.
Within Zyte’s overall Retail & E-commerce group, they rank second of 18 sub-sectors for access complexity, ahead of consumer electronics, jewelry and luxury and the big vertical stores.

For scrapers, reading a product page on a big marketplace could take a full browser, higher-tier infrastructure, and tooling that can cope with a rate limiter alongside heavy JavaScript.
The marketplaces have reason to build that way: they are a rich source of price and catalog comparison data in e-commerce, and their access controls are built to match.
Why are the big marketplaces so hard to scrape?
Retail’s defense leans on rendering and scale rather than user-facing challenges:
A web application firewall covers 96% of retail sites, close to universal.
JavaScript rendering is required on 51%, the highest active barrier in the sector (those beside firewalls), because many marketplaces are built as heavy client-side applications.
Rate limiting runs on 33% of sites.
TLS fingerprinting is in use on 20% of sites.
Antibot systems are used by 19%.
The barrier retail avoids using is the revealing one. CAPTCHA appears on just 15% of sites, below the cross-industry average. A CAPTCHA in a checkout flow costs conversions, so the marketplaces keep their defenses out of the shopper’s way and lean on quiet detection instead.

Most retail sites stack two or more barriers
Retail rarely relies on a single control. The most common posture is a two-barrier stack, typically a Web Application Firewall paired with JavaScript rendering, and the flagship platforms layer well beyond that.
Because the barriers sit at different levels, from the connection to the rendered page, reaching a marketplace reliably means handling several at once rather than clearing one and moving on.

What does it cost to scrape a marketplace?
Zyte API’s complexity tier breakdown puts a price on that. Zyte API’s pricing charges for the difficulty a site actually presents rather than a flat rate, so customers pay the lowest rate each site requires.
Just over half of retail sites sit at Simple or Easy. But the tail is heavy: nearly half rate Moderate or above, and one in five reach Complex or Advanced.

Set against a research-wide average of 1.58, retail’s mean of 2.60 is second only to fashion.
Why retail also blocks AI crawlers
Here, retail parts company with fashion. While a fashion brand is largely indifferent to whether a chatbot reads its pages, the big marketplaces are not.
Twenty-nine percent of retail sites name an AI crawler to block in their robots.txt, seven times the rate in fashion, and 85% publish a robots.txt at all, among the highest coverage of any sector.
Where they name agents, AI crawlers such as GPTBot draw as much attention as the search bots.

The reason is what a marketplace holds: vast structured catalogs, pricing, reviews and product questions, exactly the material AI shopping assistants and models want to ingest.
A marketplace has content that AI is after, so it fences AI out deliberately, on top of the technical barriers aimed at price collection.
Retail, in other words, fights on two fronts at once: price and catalog extraction on the infrastructure side, and AI access on the policy side.
What this means if you work with retail data
For anyone scraping retail or wider e-commerce data at scale, the study points to a clear playbook:
Expect JavaScript rendering as the baseline, since a full browser is needed on more than half of sites.
The heaviest defenses sit on the flagship marketplaces, so budget for the higher tiers where the data matters most.
Rate limiting and TLS checks are common enough that a plain HTTP client will not carry you far on the big platforms.
Read the robots.txt closely: retail’s high AI-block rate means named-agent policy is more active here than in most sectors.
Retail marketplaces are the second hardest sector on the web to reach at scale, and the ranking is no accident. They hold the web’s richest commercial data, and they guard it against price scrapers and AI crawlers alike.
Key takeaways: Retail access, in five numbers
No general marketplace is easy to collect from at scale:
2.60 of 5 - retail’s mean access-complexity tier, second only to fashion.
51% of retail sites require JavaScript rendering, the highest active barrier in the sector.
96% run a web application firewall, close to universal.
29% block AI crawlers in robots.txt, seven times the rate in fashion.
One in five retail sites rate Complex or Advanced to access.






_HFpro5d6k3.png&w=256&q=75)
_E4PyVpfAxa.png&w=256&q=75)


-(1).png&w=1920&q=75)
-(1)_VZGHqxCgXV.png&w=1920&q=75)