Field notes from the world of data extraction.

Articles, interviews and analysis on how data is gathered, used and fought over — written by the people closest to it.

3 Rules of Modern Web Scraping

3 Rules of Modern Web Scraping

John Rooney10 min read
Zyte Blog — field notes from the world of data extraction

The Modern Scrapy Developer's Guide (Part 3): Auto-Generating Page Objects with the Web Scraping Copilot

In this guide, we'll show you how to use Web Scraping Copilot (our VS Code extension) to automatically write 100% of your Items, Page Objects, and even your unit tests.

5 min read
Zyte Blog — field notes from the world of data extraction
Scraping practice

The Modern Scrapy Developer's Guide (Part 2): Page Objects with scrapy-poet

In this guide, we'll fix this by refactoring our spider to a professional, modern standard using Scrapy Items and Page Objects (via crapy-poet). We will completely separate our crawling logic from our parsing logic.

John Rooney5 min read
Zyte Blog — field notes from the world of data extraction
Scraping practice

The Modern Scrapy Developer's Guide (Part 1): Building Your First Spider

In this definitive guide, we will walk you through, step-by-step, how to build a real, multi-page crawling spider. You will go from an empty folder to a clean JSON file of structured data in about 15 minutes

John Rooney4 min read
How to build a daily industry news digest

Zyte’s 2025 review: Year of locks and unlocks

2025 reshaped web scraping. From AI-assisted extraction and escalating bot defenses to clearer legal frameworks and cheaper APIs, Zyte reviews the forces redefining access to web data—and what comes next.

Robert Andrews5 min read
AI’s legal frontier: What Europe’s privacy regulators say about scraping personal data
Web data collection legality

AI’s legal frontier: What Europe’s privacy regulators say about scraping personal data

Explore how EU privacy regulators view AI web scraping, lawful bases like legitimate interest, risks of collecting personal data, and compliance best practices.

Victoria Vlahoyiannis5 min read
Zyte leads the pack in Proxyway’s 2025 Web Scraping API Report
Web scraping APIs

Zyte leads the pack in Proxyway’s 2025 Web Scraping API Report

Proxyway’s 2025 Web Scraping API Report ranks Zyte #1 for unblocking success, speed, cost efficiency, and AI-powered data extraction. See the full breakdown.

Robert Andrews10 min read
Zyte Blog — field notes from the world of data extraction
How To

The Modern Web Scraping Method You NEED to Know

Learn how to scrape data in json format from a websites API

John Rooney10 min read
AI and the web: What 2025 changed and what comes next
Anti-ban

Beyond the block: The front line of data access

A deep dive into the evolving battle for web data access—featuring insights from Castle, Scrapoxy, and Zyte at Extract Summit 2025. Learn how AI, anti-bots, economics, and authentication standards like Web Bot Auth are transforming scraping, security, and the future of the open internet.

Robert Andrews10 min read
How to build a daily industry news digest
Web data collection

How to build a daily industry news digest

Learn how data analyst Anshika Khandelwal automated a daily AI funding news digest using n8n and Zyte API. Discover how to pull articles, classify funding stories, and deliver a curated newsletter that saves 10+ hours per week.

Robert Andrews5 min read
AI and the web: What 2025 changed and what comes next
Web data collection

Scraping a synthetic web: Dead Internet Theory meets web data extraction

AI-generated content now dominates the web. Explore the rise of synthetic internet traffic, how bots shape online discourse, and how data experts can fight back.

Domagoj Marić10 min read
Gemini 3.0 Pro is the new best model for writing scrapers
AI-assisted data extraction

Gemini 3.0 Pro is the new best model for writing scrapers

Gemini 3.0 Pro outperforms GPT-5, Claude, and other leading LLMs in Zyte’s Web Scraping Copilot benchmarks, delivering the highest code accuracy and lowest complexity. See full results, pros, cons, and recommendations for production workflows.

Konstantin Lopukhin10 min read