Field notes from the world of data extraction.
Articles, interviews and analysis on how data is gathered, used and fought over — written by the people closest to it.

Running Playwright at scale: connecting to the Zyte CDP browser
We've released our CDP Browser, ideal for those with existing Playwright scripts looking for better access, as well as developers who need granular control over a browser, but don't want the hassle of running locally.

Four in 10 sites block AI bots with robots.txt
Nearly four in 10 of the world's top sites are now closed to well-behaved AI crawlers. Here's what operators actually do with robots.txt, and which AI agents they block or welcome.

Fable 5.1 shipped, and GLM-5.3-Flash turned out to be someone I already met
A personal take on Claude Fable 5.1 and GLM-5.3-Flash: real benchmarks, a live extraction test, and the model I guessed before Zhipu confirmed it.

Web data in a reactive notebook: an introduction to marimo
marimo is a reactive Python notebook that reruns only affected cells. Learn to scrape web data with Zyte API, chart prices, and cache costly API calls.

75% of the web uses robots.txt - here's how
Three in four of the world's top sites publish a robots.txt, yet few name individual crawlers. A look inside the web's advisory layer: coverage, sophistication and limits.

Introducing Zyte CDP support: Your browser automation, our infrastructure
Want to use traditional browser automation frameworks without infrastructure headaches? Tap our new CDP support to run scripts on Zyte’s powerful infrastructure.

Why Europe’s new AI scraping guidelines miss the mark
New guidelines on generative AI scraping aim to protect user privacy. But turning a rudimentary, 30-year-old web server standard into a legal barrier will disenfranchise users and create accidental monopolies.

WebFetch is ‘lossy’ by design: give your coding agent a better fetch in one command
Your coding agent built-in webfetch tool is not the best and it's hampering your research and coding workflows, fix it with one CLI tool and never face blocks again.

Revealed: How retailers use tech to block the bots
The web’s big general marketplaces exist to be browsed by everyone. But they are also the second most defended sector online, shutting the door on AI crawlers.

The web is being priced, not blocked
Zyte's State of Web Access research finds new barriers making web scraping more difficult. We interview the researcher who says that doesn't mean the opportunity is over.

The State of Web Access
View the free State of Web Access report to learn just how difficult modern web scraping has become.

Fashion websites are the hardest to size up
While publishers fight over AI crawlers, fashion has built some of the most heavily defended shop windows on the web. So, which controls are in fashion, in fashion?