Web data

Articles from the Zyte blog about Web data.

Zyte MCP chat box connected to fetch, search and extract tools.
Data gathering for AI

Introducing Zyte MCP: Powerful web data gathering for your agent

Zyte MCP brings live web access, structured extraction and Scrapy Cloud operations into coding agents such as Claude Code, Cursor and Codex.

Valter Sciarrillo7 min read
A classic toy robot waves a protest banner showing "I'm not a robot"
Access handling

CAPTCHA is more than just a puzzle piece

The famously awkward security mechanism has quietly become an invisible behavioral scoring system. Two vendors enable 89% of CAPTCHAs, and the smallest sites use it most.

Robert Andrews
One Scrapy Spider, Three Browser Setups: Playwright, Patchright, and Zyte
How To

One Scrapy Spider, Three Browser Setups: Playwright, Patchright, and Zyte

One Scrapy spider, three browser setups: stock Playwright, Patchright via the new PLAYWRIGHT_BROWSER_PROVIDER hook, and a remote browser on Zyte over CDP. Same spider and selectors throughout, only the browser changes.

John Rooney
A spyglass reveals a robot secretly using a web browser with its hand on the mouse.
Access handling

Inside antibot, the arms race you can't see

Only 18.5% of top sites run dedicated antibot, but every one of them chose to. Why it's the most intentional barrier in the stack, and the strongest signal of a hardened site.

Robert Andrews
A firewall surrounds a server computer
Access handling

Firewalls run the web, just not on purpose

Web Application Firewalls run on 92.4% of the top websites, but most arrived bundled with a CDN and were never tuned to resist automated access.

Robert Andrews
Text file struggling under the weight of European privacy burden.
Web data collection legality

Why Europe’s new AI scraping guidelines miss the mark

New guidelines on generative AI scraping aim to protect user privacy. But turning a rudimentary, 30-year-old web server standard into a legal barrier will disenfranchise users and create accidental monopolies.

Sally-Anne Hinfey
spidermon-part-2
Web data collection

Spider monitoring made easy

How do you know you're collecting all the data you need? And how can you be sure it's actually what you were expecting? Use Spidermon.

John Rooney
Data is the new water
Web data collection

Data is the new water

For 20 years, we were told “data is the new oil”. That’s no longer true. In the era of the fluid web, information is more fundamental, cleaner and abundant than that.

Robert Andrews10 min read
Brand visibility in the digital era: How web data help brands see the full picture
Web data collection

Brand visibility in the digital era: How web data help brands see the full picture

Discover how web data helps brands improve visibility, track competitors, monitor availability, and analyze reviews to win on the digital shelf.

Theresia Tanzil5 min read
How online retailers use web data to compete on price, promotion, and availability
Web data collection

How online retailers use web data to compete on price, promotion, and availability

Discover how retailers leverage web data to optimize pricing, track competitor stock, detect trends, and improve sales performance.

Theresia Tanzil5 min read
K1
Anti-ban

The recipe for a request: Scaling data extraction through investigation

Learn how an investigative mindset helps scale data extraction from single requests to millions daily by building resilient, efficient scraping systems.

Kieron Spearing5 min read
How web data turns e-commerce listings into retail intelligence
Web data collection

How web data turns e-commerce listings into retail intelligence

Discover how web data enables digital shelf analytics vendors to track prices, availability, and product trends at scale—fueling real-time retail intelligence and competitive advantage.

Theresia Tanzil5 min read