Field notes from the world of data extraction.
Articles, interviews and analysis on how data is gathered, used and fought over — written by the people closest to it.

The web is synthetic. Now look at the human trail
Domagoj Marić explores how AI, web scraping, and OSINT turn scattered personal data into profiles, scams, and security risks at Extract Summit.

The State of Web Access: How different industries behave to bots
Zyte's large-scale audit of web access controls shows how different kinds of businesses exhibit different policies.

A $70-a-month production app replaced a $5,000-a-month platform: What production-grade AI coding looks like
See how Fran Muñoz used production-grade AI coding, specification-driven development, and 3,600 tests to replace a $5,000-a-month platform with a $70 app.
Podcast Episode 10 is out: Skills, packages & updates
Episode 10 of the Zyte podcast is out, and it is one of the widest-ranging conversations we have recorded this year. John Rooney sits down with Neha Setia Nagpal and Ayan Pahwa for an unfiltered, all-in-style chat covering everything the team has been building, reading, and arguing about lately.

Chrome has a new potential fingerprint vector
navigator.cpuPerformance is coming to Chrome in August, v152. What does it mean for fingerprinting and will it mean changes to scraping stacks?
Harness Engineering, part 4: giving your agent a custom fetch tool that survives the real web
Your AI agent is as powerful as the tools it has access to. Here's a tutorial on how you can create your own custom agent tool using Claude Agent SDK for zyte which makes getting structured data from web a breeze.

Meet scrapy-spidey-sense: A preflight check for Scrapy spiders
`scrapy-spidey-sense` is an open-source command-line interface (CLI) that checks a Scrapy project before the crawl begins. It performs local, static analysis, reports the production-readiness basics that are present or missing, assigns a score, and connects each finding to a practical fix or relevant documentation

Rendering Javascript pages without giving up Scrapy
How to render dynamic content and work with a browser through playwright and scrapy.

Spider monitoring made easy
How do you know you're collecting all the data you need? And how can you be sure it's actually what you were expecting? Use Spidermon.

I joined the Real Python podcast to talk harnesses and Scrapy
I was recently a guest on the Real Python podcast where host Christopher Bailey opened with the question that has been following me around all year: which matters more, the model or the harness around it? Let's look into it.

Building maintainable spiders with scrapy-poet
Separating your extract and parsing logic out help increase the maintainability and extensibility of your projects, and scrapy-poet makes it easy.

Modern Scrapy for experienced developers: A new series
How to create production ready Scrapy projects to scrape the modern web. In this article we start the process of creating our spider, look at settings, and build a pipeline to help keep our data quality high.