Zyte Blog
Field notes from the world of data extraction.
Articles, interviews and analysis on how data is gathered, used and fought over — written by the people closest to it.
Browse
Hot topics
Latest
What we've been publishing

The page your agent scrapes is now an attack surface. Is it ready for the hostile web?
Real web is becoming hostile for AI Agents. The page your agent scrapes now could be a potential attack surface. Read more and join Zyte's virtual meet-up to see it in action.

Running Playwright at scale: connecting to the Zyte CDP browser
We've released our CDP Browser, ideal for those with existing Playwright scripts looking for better access, as well as developers who need granular control over a browser, but don't want the hassle of running locally.

Four in 10 sites block AI bots with robots.txt
Nearly four in 10 of the world's top sites are now closed to well-behaved AI crawlers. Here's what operators actually do with robots.txt, and which AI agents they block or welcome.

Fable 5.1 shipped, and GLM-5.3-Flash turned out to be someone I already met
A personal take on Claude Fable 5.1 and GLM-5.3-Flash: real benchmarks, a live extraction test, and the model I guessed before Zhipu confirmed it.

Web data in a reactive notebook: an introduction to marimo
marimo is a reactive Python notebook that reruns only affected cells. Learn to scrape web data with Zyte API, chart prices, and cache costly API calls.

75% of the web uses robots.txt - here's how
Three in four of the world's top sites publish a robots.txt, yet few name individual crawlers. A look inside the web's advisory layer: coverage, sophistication and limits.

Introducing Zyte CDP support: Your browser automation, our infrastructure
Want to use traditional browser automation frameworks without infrastructure headaches? Tap our new CDP support to run scripts on Zyte’s powerful infrastructure.

Why Europe’s new AI scraping guidelines miss the mark
New guidelines on generative AI scraping aim to protect user privacy. But turning a rudimentary, 30-year-old web server standard into a legal barrier will disenfranchise users and create accidental monopolies.

WebFetch is ‘lossy’ by design: give your coding agent a better fetch in one command
Your coding agent built-in webfetch tool is not the best and it's hampering your research and coding workflows, fix it with one CLI tool and never face blocks again.

Revealed: How retailers use tech to block the bots
The web’s big general marketplaces exist to be browsed by everyone. But they are also the second most defended sector online, shutting the door on AI crawlers.

The web is being priced, not blocked
Zyte's State of Web Access research finds new barriers making web scraping more difficult. We interview the researcher who says that doesn't mean the opportunity is over.

The State of Web Access
View the free State of Web Access report to learn just how difficult modern web scraping has become.
Zyte YouTube
Watch & learn

Video · Zyte YouTube
AI generated these Scrapy projects - why I won't ship them
July 7, 2026

Video · Zyte YouTube
Screenshot webpages with this Claude Skill and Zyte API
March 6, 2026

Video · Zyte YouTube
0% Hallucination? RAG + Web Scraping (Step-by-Step)
March 5, 2026

Video · Zyte YouTube
Generate HTML Parsing code the right way with Scrapy & Web Scraping Copilot
February 23, 2026

Video · Zyte YouTube
Zyte API Sessions - flexible cookie management maintaining control
February 6, 2026
Events & community
Where Zyte shows up
24 Sep 2026Ship Agents That Survive the Real Web
Webinar · Zoom, with Humanbound
7–8 Oct 2026Extract Summit 2026 — Austin, TX
In-person conference · Austin, TX
10–11 Nov 2026Extract Summit 2026 — Dublin
In-person conference · Dublin, Ireland
On demand2026 Web Scraping Industry Report by Zyte
Webinar · On demand










