AI agents are increasingly where engineers plan, write and debug. But their work can often be thwarted at the first protected web page. The agent times out, mistakes a consent screen for the content or falls back on stale training data.
Even when it gets the page, a web data team still has to jump between code, documentation, dashboards and logs to run and diagnose the crawl.
Today, we’re introducing Zyte MCP Server: a hosted Model Context Protocol server that makes Zyte API and Scrapy Cloud available to MCP-compatible AI and coding agents.
We are exposing all of Zyte, inside your coding agent.
What Zyte MCP makes possible
Model Context Protocol, or MCP, gives AI applications a standard way to discover and invoke external tools. Zyte MCP uses that interface to bring reliable web access, structured extraction and crawl operations into environments such as Claude Code, Cursor, Codex, GitHub Copilot and VS Code.
Zyte MCP is not a new extraction engine. Under the hood, Zyte API still retrieves and extracts the data, while Scrapy Cloud runs spiders. MCP changes where you use those capabilities: inside the agent where the engineering work is already happening. Your agent will be able to:
- Fetch protected and JavaScript-heavy pages through Zyte’s access infrastructure.
- Return clean Markdown, with browser rendering, screenshots and geolocation available where needed.
- Search the web and consume structured results.
- Extract structured fields instead of processing a whole page.
- Start, stop and monitor Scrapy Cloud jobs from existing deployed spider versions.
- Inspect job states, statistics, logs and item samples to diagnose problems.
- Create and manage recurring schedules.
- Ask about usage, current spend and success rate by domain in natural language.
There is no separate MCP fee. Usage is metered through Zyte API and Scrapy Cloud as it is today.
Why it matters
Agents get the page they actually need
Built-in browsing works for ordinary pages; it is less dependable when a site uses anti-bot protection or requires browser rendering. Zyte MCP gives the agent a specialist web-access tool, so an answer can be grounded in the current page rather than a refusal, consent notice or model memory.
The connection’s real power comes from the underlying Zyte API, independently rated the industry’s #1 solution for unblocking and speed. With 320,000 access strategies and adaptive unblocking under the hood, Zyte API has the highest success rate in the game, now available through fluid, conversational interaction via MCP.
The work after the fetch stays in the agent
A fetch-only tool helps with prototypes. Production web data also needs jobs to be run, monitored and diagnosed.
Zyte MCP lets an agent inspect Scrapy Cloud state, logs, statistics and item samples in the same conversation where an engineer reviews the code. That shortens the path from “this crawl looks wrong” to an evidence-based diagnosis.
Clean content uses fewer tokens
Raw HTML contains navigation, scripts, styling and repeated page furniture that add cost without helping the model reason.
Zyte’s MCP can return clean Markdown instead of messy HTML.
In an internal benchmark of 35 HTML cleanup and HTML-to-Markdown tools plus three hosted services across 299 real pages, Zyte API’s shipped Markdown implementation retained 95% of the main content while removing 85% of unwanted page chrome and reducing HTML token volume by about 97%.
Five ways to use it
With Zyte MCP Server enabled in your agent as a connector, you can use Zyte fluidly and conversationally, with prompts like these:
1. Ground research in the live web
“Search for the latest official sources about [topic]. Fetch the five best pages, including any ordinary browsing cannot read, and cite a URL for every claim.”
The agent can search for sources, retrieve clean content and synthesise an answer from the pages themselves.
2. Extract current product data
“Extract the name, price, currency, availability and SKU from this product URL. Return JSON and flag any field that is missing from the page.”
The agent receives structured fields rather than parsing noisy HTML in its context window.
3. Diagnose a failed crawl
“Find the latest incomplete run for this spider. Compare it with the previous successful job, inspect the relevant errors and show three affected item samples.”
The agent gathers the evidence from Scrapy Cloud and helps distinguish an access problem from a spider or coverage problem.
4. Turn a working spider into a recurring job
“Start this deployed spider with these arguments. If the item sample is correct, schedule it for 06:00 every weekday.”
The agent carries out the routine operational steps through structured tools, subject to the user’s approval controls — developers stay in control throughout.
5. Understand where your spend is going
“Where did most of my spend go in the last billing period? Break it down by product and by domain, and highlight any anomalies.”
The agent can pull usage and billing signals into a quick summary, so you can spot cost drivers and unexpected spikes without manually digging through dashboards.
Compose it with the rest of your stack
MCP-enabled tools can work together around one outcome. For example:
- Pair GitHub and Zyte MCP to review a spider change, start the deployed version and inspect the result.
- Zyte already has plugins for Claude Code, Codex, Visual Studio Code and GitHub, plus a range of agent-neutral agent skills. Use your own tools to build spiders, fetch the most defended pages and run your web scraping operations without leaving your coding agent.
- Send a daily crawl-health summary to Slack or another messaging tool, including only jobs that need attention.
- Fetch difficult-to-access sources and save a cited research brief to a knowledge base.
- Search, fetch and extract web data before passing an approved sample to an analytics or storage destination.
How to connect
Find our MCP documentation in the Zyte API docs.
Zyte MCP is hosted, so there is no server package to maintain.
Add this remote Streamable HTTP endpoint to a compatible MCP client:
1https://mcp.zyte.com/v1/mcpA typical configuration is:
1{
2 "mcpServers": {
3 "zyte": {
4 "type": "http",
5 "url": "https://mcp.zyte.com/v1/mcp"
6 }
7 }
8}For Claude Code, the expected command is:
1claude mcp add --transport http zyte https://mcp.zyte.com/v1/mcpOn first connection, complete the Zyte sign-in flow. OAuth gives the client a short-lived, scoped token instead of placing a permanent Zyte API key in the agent’s configuration.
Then ask for an outcome rather than a tool name:
“Fetch this page and return the main content as Markdown.”
“Extract the current product price and availability from this URL.”
“List my recent failed Scrapy Cloud jobs and explain the most common error.”
What’s next
Check out the docs for full details of capabilities, setup instructions, billing and usage.
More tools and capabilities for our MCP server are in the pipeline.
And we have ambitions for agentic web data beyond even MCP.
Fluid data gathering is already here. Connect your client to mcp.zyte.com/v1/mcp and ask your agent to fetch a page it could not read before.

