Introducing Zyte MCP: Powerful web data gathering for your agent

Zyte MCP brings live web access, structured extraction and Scrapy Cloud operations into coding agents such as Claude Code, Cursor and Codex.

Valter Sciarrillo · Product Marketing

7 min read ·

Introducing Zyte MCP: Powerful web data gathering for your agent

AI agents are increasingly where engineers plan, write and debug. But their work can often be thwarted at the first protected web page. The agent times out, mistakes a consent screen for the content or falls back on stale training data.

Even when it gets the page, a web data team still has to jump between code, documentation, dashboards and logs to run and diagnose the crawl.

Today, we’re introducing Zyte MCP Server: a hosted Model Context Protocol server that makes Zyte API and Scrapy Cloud available to MCP-compatible AI and coding agents.

We are exposing all of Zyte, inside your coding agent.

What Zyte MCP makes possible

Model Context Protocol, or MCP, gives AI applications a standard way to discover and invoke external tools. Zyte MCP uses that interface to bring reliable web access, structured extraction and crawl operations into environments such as Claude Code, Cursor, Codex, GitHub Copilot and VS Code.

Zyte MCP is not a new extraction engine. Under the hood, Zyte API still retrieves and extracts the data, while Scrapy Cloud runs spiders. MCP changes where you use those capabilities: inside the agent where the engineering work is already happening. Your agent will be able to:

There is no separate MCP fee. Usage is metered through Zyte API and Scrapy Cloud as it is today.

Why it matters

Agents get the page they actually need

Built-in browsing works for ordinary pages; it is less dependable when a site uses anti-bot protection or requires browser rendering. Zyte MCP gives the agent a specialist web-access tool, so an answer can be grounded in the current page rather than a refusal, consent notice or model memory.

The connection’s real power comes from the underlying Zyte API, independently rated the industry’s #1 solution for unblocking and speed. With 320,000 access strategies and adaptive unblocking under the hood, Zyte API has the highest success rate in the game, now available through fluid, conversational interaction via MCP.

The work after the fetch stays in the agent

A fetch-only tool helps with prototypes. Production web data also needs jobs to be run, monitored and diagnosed.

Zyte MCP lets an agent inspect Scrapy Cloud state, logs, statistics and item samples in the same conversation where an engineer reviews the code. That shortens the path from “this crawl looks wrong” to an evidence-based diagnosis.

Clean content uses fewer tokens

Raw HTML contains navigation, scripts, styling and repeated page furniture that add cost without helping the model reason.

Zyte’s MCP can return clean Markdown instead of messy HTML.

In an internal benchmark of 35 HTML cleanup and HTML-to-Markdown tools plus three hosted services across 299 real pages, Zyte API’s shipped Markdown implementation retained 95% of the main content while removing 85% of unwanted page chrome and reducing HTML token volume by about 97%.

Five ways to use it

With Zyte MCP Server enabled in your agent as a connector, you can use Zyte fluidly and conversationally, with prompts like these:

1. Ground research in the live web

“Search for the latest official sources about [topic]. Fetch the five best pages, including any ordinary browsing cannot read, and cite a URL for every claim.”

The agent can search for sources, retrieve clean content and synthesise an answer from the pages themselves.

2. Extract current product data

“Extract the name, price, currency, availability and SKU from this product URL. Return JSON and flag any field that is missing from the page.”

The agent receives structured fields rather than parsing noisy HTML in its context window.

3. Diagnose a failed crawl

“Find the latest incomplete run for this spider. Compare it with the previous successful job, inspect the relevant errors and show three affected item samples.”

The agent gathers the evidence from Scrapy Cloud and helps distinguish an access problem from a spider or coverage problem.

4. Turn a working spider into a recurring job

“Start this deployed spider with these arguments. If the item sample is correct, schedule it for 06:00 every weekday.”

The agent carries out the routine operational steps through structured tools, subject to the user’s approval controls — developers stay in control throughout.

5. Understand where your spend is going

“Where did most of my spend go in the last billing period? Break it down by product and by domain, and highlight any anomalies.”

The agent can pull usage and billing signals into a quick summary, so you can spot cost drivers and unexpected spikes without manually digging through dashboards.

Compose it with the rest of your stack

MCP-enabled tools can work together around one outcome. For example:

  • Pair GitHub and Zyte MCP to review a spider change, start the deployed version and inspect the result.
  • Zyte already has plugins for Claude Code, Codex, Visual Studio Code and GitHub, plus a range of agent-neutral agent skills. Use your own tools to build spiders, fetch the most defended pages and run your web scraping operations without leaving your coding agent.
  • Send a daily crawl-health summary to Slack or another messaging tool, including only jobs that need attention.
  • Fetch difficult-to-access sources and save a cited research brief to a knowledge base.
  • Search, fetch and extract web data before passing an approved sample to an analytics or storage destination.

How to connect

Find our MCP documentation in the Zyte API docs.

Zyte MCP is hosted, so there is no server package to maintain.

Add this remote Streamable HTTP endpoint to a compatible MCP client:

1https://mcp.zyte.com/v1/mcp
Copy

A typical configuration is:

1{
2  "mcpServers": {
3    "zyte": {
4      "type": "http",
5      "url": "https://mcp.zyte.com/v1/mcp"
6    }
7  }
8}
Copy

For Claude Code, the expected command is:

1claude mcp add --transport http zyte https://mcp.zyte.com/v1/mcp
Copy

On first connection, complete the Zyte sign-in flow. OAuth gives the client a short-lived, scoped token instead of placing a permanent Zyte API key in the agent’s configuration.

Then ask for an outcome rather than a tool name:

“Fetch this page and return the main content as Markdown.”

“Extract the current product price and availability from this URL.”

“List my recent failed Scrapy Cloud jobs and explain the most common error.”

What’s next

Check out the docs for full details of capabilities, setup instructions, billing and usage.

More tools and capabilities for our MCP server are in the pipeline.

And we have ambitions for agentic web data beyond even MCP.

Fluid data gathering is already here. Connect your client to mcp.zyte.com/v1/mcp and ask your agent to fetch a page it could not read before.

Try Zyte API

Build your first scraper in minutes

Free trial, no credit card. From a single request to production in an afternoon.

Get started

Valter Sciarrillo

Product Marketing

Valter Sciarrillo - Product Marketing @ Zyte

More from this author

The Community · Newsletter

The best of Zyte and the data web, in your inbox.

One curated edition — new articles, product updates, and the stories shaping the data web. No noise.