Field notes from the world of data extraction.

Articles, interviews and analysis on how data is gathered, used and fought over — written by the people closest to it.

Dom and Neha talking about synthetic web and extract summit

The web is synthetic. Now look at the human trail

Domagoj Marić explores how AI, web scraping, and OSINT turn scattered personal data into profiles, scams, and security risks at Extract Summit.

Neha Setia Nagpal
State of Web Access - industries
Access handling

The State of Web Access: How different industries behave to bots

Zyte's large-scale audit of web access controls shows how different kinds of businesses exhibit different policies.

Robert Andrews
Fran Muñoz and Neha Setia Nagpal beside the words “From $5,000 to $70: Production-grade AI coding.”

A $70-a-month production app replaced a $5,000-a-month platform: What production-grade AI coding looks like

See how Fran Muñoz used production-grade AI coding, specification-driven development, and 3,600 tests to replace a $5,000-a-month platform with a $70 app.

Neha Setia Nagpal
podcast-10

Podcast Episode 10 is out: Skills, packages & updates

Episode 10 of the Zyte podcast is out, and it is one of the widest-ranging conversations we have recorded this year. John Rooney sits down with Neha Setia Nagpal and Ayan Pahwa for an unfiltered, all-in-style chat covering everything the team has been building, reading, and arguing about lately.

John Rooney
fingerprint-changes

Chrome has a new potential fingerprint vector

navigator.cpuPerformance is coming to Chrome in August, v152. What does it mean for fingerprinting and will it mean changes to scraping stacks?

John Rooney
zyte-agent-tool

Harness Engineering, part 4: giving your agent a custom fetch tool that survives the real web

Your AI agent is as powerful as the tools it has access to. Here's a tutorial on how you can create your own custom agent tool using Claude Agent SDK for zyte which makes getting structured data from web a breeze.

Ayan Pahwa
spidey-sense-1

Meet scrapy-spidey-sense: A preflight check for Scrapy spiders

`scrapy-spidey-sense` is an open-source command-line interface (CLI) that checks a Scrapy project before the crawl begins. It performs local, static analysis, reports the production-readiness basics that are present or missing, assigns a score, and connects each finding to a practical fix or relevant documentation

Neha Setia Nagpal
scrapy-series-4
Scraping strategy

Rendering Javascript pages without giving up Scrapy

How to render dynamic content and work with a browser through playwright and scrapy.

John Rooney
spidermon-part-2
Web data collection

Spider monitoring made easy

How do you know you're collecting all the data you need? And how can you be sure it's actually what you were expecting? Use Spidermon.

John Rooney
ayan-pahwa-real-python-podcast

I joined the Real Python podcast to talk harnesses and Scrapy

I was recently a guest on the Real Python podcast where host Christopher Bailey opened with the question that has been following me around all year: which matters more, the model or the harness around it? Let's look into it.

Ayan Pahwa
scrapy-series-2

Building maintainable spiders with scrapy-poet

Separating your extract and parsing logic out help increase the maintainability and extensibility of your projects, and scrapy-poet makes it easy.

John Rooney
scrapy-series-1

Modern Scrapy for experienced developers: A new series

How to create production ready Scrapy projects to scrape the modern web. In this article we start the process of creating our spider, look at settings, and build a pipeline to help keep our data quality high.

John Rooney