PINGDOM_CHECK

#ExtractSummit2026 The world's largest web scraping conference returns. Austin Oct 7–8 · Dublin Nov 10–11.

Register now
Data Services
Pricing
Login
Try Zyte APIContact Sales
  • Unblocking and Extraction

    Zyte API

    The ultimate API for web scraping. Avoid website bans and access a headless browser or AI Parsing

    Ban Handling

    Headless Browser

    AI Extraction

    SERP

    Enterprise

    DocumentationSupport

    Hosting and Deployment

    Scrapy Cloud

    Run, monitor, and control your Scrapy spiders however you want to.

    Coding Agent Add-Ons

    Agentic Web Data

    Plugins that give coding agents the context to build production Scrapy projects. Starts with Claude Code.

  • Data Services
  • Pricing
  • Browse

    • BlogArticles, podcasts, videos
    • Case studiesCustomer outcomes
    • White papersIn-depth reports
    • DocumentationGuides & API reference
    • EventsConferences, webinars, recordings

    Subscribe

    • NewsletterSwiftly delivered
    • Discord communityExtract Data community
  • Product and E-commerce

    From e-commerce and online marketplaces

    Data for AI

    Collect and structure web data to feed AI

    Job Posting

    From job boards and recruitment websites

    Real Estate

    From Listings portals and specialist websites

    News and Article

    From online publishers and news websites

    Search

    Search engine results page data (SERP)

    Social Media

    From social media platforms online

  • Meet Zyte

    Our story, people and values

    Contact us

    Get in touch

    Support

    Knowledge base and raise support tickets

    Terms and Policies

    Accept our terms and policies

    Open Source

    Our open source projects and contributions

    Web Data Compliance

    Guidelines and resources for compliant web data collection

    Join the team building the future of web data
    We're Hiring
    Trust Center
    Security, compliance & certifications
Login
Try Zyte APIContact Sales
All articles
AI65, 65 articles
Data quality13, 13 articles
Developer interest57, 57 articles
Integration2, 2 articles
Open-source41, 41 articles
Proxies29, 29 articles
Scraping practice19, 19 articles
Scraping strategy28, 28 articles
Web data60, 60 articles
Web scraping APIs36, 36 articles
Scrapy47, 47 articles
Scrapy Cloud14, 14 articles
Web Scraping Copilot11, 11 articles
Zyte API57, 57 articles
AI & Machine Learning3, 3 articles
Automotive2, 2 articles
E-commerce & retail27, 27 articles
Entertainment & Streaming2, 2 articles
Financial Services8, 8 articles
Government2, 2 articles
Market Research & Intelligence3, 3 articles
Media & publishing8, 8 articles
Real Estate2, 2 articles
Recruitment & HR3, 3 articles
Transportation & Logistics2, 2 articles
Travel & hospitality2, 2 articles
Extract Summit25, 25 articles
PyCon1, 1 articles
iPaaS2, 2 articles
Large language model24, 24 articles
MCP3, 3 articles
Python88, 88 articles
Web Scraping Industry Report14, 14 articles

Appearance

Discord Community
BlogScraping strategyPodcast Ep08 - Scrapy, Python and mushroom soup
ArticleInsight and analysisScraping strategyScraping practiceZyte API

Podcast Ep08 - Scrapy, Python and mushroom soup

Scrapy's core handles crawling well and deliberately leaves almost everything else out: no bundled browser, no opinion about how you shape your data, no built-in answer for every anti-bot wrinkle. What it gives you instead is a clean way to add those things at the edges, exactly when you need them and never before.

John Rooney · Developer Engagement Manager

July 20, 2026

Podcast Ep08 - Scrapy, Python and mushroom soup

Scrapy does less than people expect, and that is the whole idea. The core handles crawling well and deliberately leaves almost everything else out: no bundled browser, no opinion about how you shape your data, no built-in answer for every anti-bot wrinkle. What it gives you instead is a clean way to add those things at the edges, exactly when you need them and never before. That decision, lean at the center and easy to extend everywhere else, is the thread John, Neha, and Ayan keep pulling on in episode 8, and it is why a framework first released in 2008 still fits the way people scrape today.

They arrive at it from three different directions, a runaway crawl at IBM, a wall of JavaScript that broke a testing script, an ecommerce automation job that quietly turned into a career, but they land in the same place. In an era of stealth browsers, TLS fingerprinting, and AI agents that write their own spiders, a small core you can extend beats a big framework that tries to decide everything for you.

The shift hiding inside the vocabulary

Somewhere in the last few years, web scraping quietly became data extraction, and the team argues the change is more than a rebrand. It reframes the whole job around the data and the pipeline it feeds rather than the act of grabbing pages. That reframing is also why John drifted away from Scrapy for a while, back when a site's back-end API was sitting right there in the network tab and a full framework felt like overkill. What pulled him back has everything to do with how cheap and how common anti-bot systems have become, and he walks through exactly where that line gets crossed. We went deeper on one piece of that puzzle in our write-up on the pluggable browser providers now in scrapy-playwright, which is the lean-core idea taken right down to the choice of browser binary, but the episode is where the reasoning behind it lands.

Moments worth pressing play for

A few of the threads the three of them get into:

  • The Django comparison that finally makes Scrapy's lean, extensible design click.
  • Why John reaches for exactly four plugins on every production-grade build, and which one he thinks almost nobody uses.
  • The anti-bot tradeoff that looks completely different once you see it from the vendor's side, and what that tells you about where the gaps are.
  • How keeping the core small is quietly what makes self-healing, agent-driven spiders possible in the first place.
  • Neha on Pydantic versus Spidermon, framed as a QA engineer at the end of the line against an auditor watching the whole crawl.
  • The single-file Scrapy trick that makes no sense until an agent needs it, and the decades-old Linux tool that lets a headless browser run on a server.
  • A detour into food science, the "law of extraction," and why it has no business circling back to the theme as neatly as it does.

Watch or listen

Prefer audio?

Apple Podcasts

Direct player

Either way, the parts where the three of them disagree are usually where the useful stuff is. And if you would rather skip the fingerprinting arms race entirely and get straight to clean data, you can try Zyte API on a free trial and let the unblocking run in the background.

Try Zyte API

Build your first scraper in minutes

Free trial, no credit card. From a single request to production in an afternoon.

Get started
Scraping strategyScraping practiceZyte API

John Rooney

Developer Engagement Manager

John is the Developer Engagement Manager at Zyte, working closely with the community, creating content and helping developers learn web scraping, Zyte products an much more. He has spoken at Extract Summit's and also creates the workshop's for the events.

  • X (Twitter)
  • LinkedIn
More from this author

In this article

  • The shift hiding inside the vocabulary
  • Moments worth pressing play for
  • Watch or listen

Follow

Get the latest

Zyte and the data web in your inbox — or wherever you already are.

Subscribe

Or follow elsewhere

Continue reading

How to build your first Scrapy extension
Scraping strategy

How to build your first Scrapy extension

Why my Scrapy project plays a triumphant fanfare when a crawl finishes clean and a sad trombone when it doesn't, and how I finally learned how to build Scrapy extensions (it's easy)

Ayan Pahwa·June 18, 2026

The Community · Newsletter

The best of Zyte and the data web, in your inbox.

One curated edition — new articles, product updates, and the stories shaping the data web. No noise.

Services

Zyte Data

Coding tools & hacks straight to your inbox. Bi-weekly dosage of all things code.

Talk to us

Web Scraping API

Zyte API

Coding tools & hacks straight to your inbox. Bi-weekly dosage of all things code.

Sign Up

Developers

Zyte Developers

Coding tools & hacks straight to your inbox. Bi-weekly dosage of all things code.

Join Us
    • Zyte API
    • Ban Handling
    • AI Extraction
    • SERP
    • Enterprise
    • Scrapy Cloud
    • Agentic Web Data
    • Pricing
    • Product & E-commerce
    • Data for AI
    • Job Posting
    • Real Estate
    • News & Articles
    • Search
    • Social Media
    • Blog
    • Learn
    • Case Studies
    • Webinars
    • White Papers
    • Join our community
    • Documentation
    • Meet Zyte
    • Contact us
    • Jobs
    • Support
    • Terms and Policies
    • Trust Center
    • Do not sell
    • Cookie settings
    • Web Data Compliance
    • Open Source
    • What is Web Scraping
    • Web Scraping in Python: Ultimate Guide
    • Stop getting blocked, start scraping
  • EWDCI logoMost loved workplace certificateZyte rewardISO 27001 iconG2 rewardG2 rewardG2 reward
    XFacebookInstagramYouTubeLinkedInDiscord

    © Zyte Group Limited 2026