Field notes from the world of data extraction.

Articles, interviews and analysis on how data is gathered, used and fought over — written by the people closest to it.

playwright-zyte-cdp
How To

Running Playwright at scale: connecting to the Zyte CDP browser

We've released our CDP Browser, ideal for those with existing Playwright scripts looking for better access, as well as developers who need granular control over a browser, but don't want the hassle of running locally.

John Rooney
A toy robot walking toward a blocking robots.txt file
Data gathering for AI

Four in 10 sites block AI bots with robots.txt

Nearly four in 10 of the world's top sites are now closed to well-behaved AI crawlers. Here's what operators actually do with robots.txt, and which AI agents they block or welcome.

Robert Andrews
fable cover

Fable 5.1 shipped, and GLM-5.3-Flash turned out to be someone I already met

A personal take on Claude Fable 5.1 and GLM-5.3-Flash: real benchmarks, a live extraction test, and the model I guessed before Zhipu confirmed it.

Ayan Pahwa
zytexmarimo-cover

Web data in a reactive notebook: an introduction to marimo

marimo is a reactive Python notebook that reruns only affected cells. Learn to scrape web data with Zyte API, chart prices, and cache costly API calls.

Ayan Pahwa
Toy robot walking toward a wall
Future of the web

75% of the web uses robots.txt - here's how

Three in four of the world's top sites publish a robots.txt, yet few name individual crawlers. A look inside the web's advisory layer: coverage, sophistication and limits.

Robert Andrews
Browser automation scripts - Puppeteer and Playwright
Product Update

Introducing Zyte CDP support: Your browser automation, our infrastructure

Want to use traditional browser automation frameworks without infrastructure headaches? Tap our new CDP support to run scripts on Zyte’s powerful infrastructure.

Valter Sciarrillo
Text file struggling under the weight of European privacy burden.
Web data collection legality

Why Europe’s new AI scraping guidelines miss the mark

New guidelines on generative AI scraping aim to protect user privacy. But turning a rudimentary, 30-year-old web server standard into a legal barrier will disenfranchise users and create accidental monopolies.

Sally-Anne Hinfey
webfetch workflow

WebFetch is ‘lossy’ by design: give your coding agent a better fetch in one command

Your coding agent built-in webfetch tool is not the best and it's hampering your research and coding workflows, fix it with one CLI tool and never face blocks again.

Ayan Pahwa
SoWA retail industry analysis
Access handling

Revealed: How retailers use tech to block the bots

The web’s big general marketplaces exist to be browsed by everyone. But they are also the second most defended sector online, shutting the door on AI crawlers.

Robert Andrews
Theresia Tanzil, State of Web Access portrait
Scraping strategy

The web is being priced, not blocked

Zyte's State of Web Access research finds new barriers making web scraping more difficult. We interview the researcher who says that doesn't mean the opportunity is over.

Robert Andrews
The State of Web Access report cover

The State of Web Access

View the free State of Web Access report to learn just how difficult modern web scraping has become.

SoWA fsahion icon
Access handling

Fashion websites are the hardest to size up

While publishers fight over AI crawlers, fashion has built some of the most heavily defended shop windows on the web. So, which controls are in fashion, in fashion?

Robert Andrews