Field notes from the world of data extraction.
Articles, interviews and analysis on how data is gathered, used and fought over — written by the people closest to it.

Enhancing AI model performance with fresh web data
No-one likes an out-of-touch AI assistant. Fortunately, rapid refreshing can keep AI models aware of the very latest public information.

What's becoming of web scraping developers in the age of AI agents?
AI agents can generate code, suggest selectors, and draft crawl logic. What they can't do is design the system that decides when to stop, what to trust, and how to recover when the web pushes back. That job still belongs to a human.

Building superior AI models with quality web data
Training data quality can make or break AI model effectiveness. So how are engineers sourcing the web’s best input?

Web scraping on an iPhone? Yes, really!
When you can scrape the web by API, a world of possibility opens up. Yes, you can extract live web data using iOS Shortcuts.

Introducing Zyte Web Data for Claude Code: Production-ready scraping from a prompt
Developers are embracing agentic coding tools - but data engineers need tools with specialist scraping skills.

Copilot like a pro: Eight tips that supercharged my workflow
AI-assisted coding is a revelation. But are you getting the most out of your IDE’s sidebar sidekick?

Automate deployment of your web scraper on VPS with Ubuntu 24.04 cloud-init
Your VPS is ready, but now you need to work through the same sequence you have run a dozen times before: apt update, apt install python3-pip, pip install scrapy, playwright install chromium, the Chromium dependency list that never installs cleanly on the first try, Redis, possibly Postgres, whatever else this particular project needs.

What multi-agent orchestration looks like in a large-scale web scraping project
Multi-agent orchestration is having its moment. The diagrams are everywhere now. Boxes for planners, boxes for hands, boxes for daemons, arrows to a shared brain, a human floating at the top. They keep getting prettier. The part where the web pushes back is still the part nobody draws.

Announcing powerful new spending controls and usage insights for Zyte API
Consign bill-shock to the trashcan. New custom spending limits and usage insights put data-gatherers in control.

Meet the new-look Zyte Domain Health Hub: Your command center for data extraction performance
Monitor your data-gathering pipelines like a boss - and act on domain issues in real-time.

NotAnInterview: “I Have Superpowers Now"
The problem was a project with 12,000 websites to crawl, and there’s no world where you write custom spiders for 12,000 websites, not with a human team and certainly not sustainably. So Javier built a workflow: a set of AI prompts that could analyze a website, figure out its structure, and generate a crawl configuration that a generic spider could then use.

Building a self-hosted browser scraping service (is it more hassle than its worth?)
If you want to understand exactly how a browser scraping service works at the infrastructure level, or you have a steady workload that you want running on hardware you already own, building one yourself teaches you things that matter. Here's how I did it




_HFpro5d6k3.png&w=256&q=75)
_E4PyVpfAxa.png&w=256&q=75)


-(1).png&w=1920&q=75)
-(1)_VZGHqxCgXV.png&w=1920&q=75)