Skip to content

A hosted web scraping MCP server for AI agents

Connect your agent to https://mcp.spicrawl.com/mcp and it can read any URL as Markdown or JSON, run batch jobs, reuse logged-in sessions and check its own usage, without you writing scraping code.

What the Spicrawl MCP server does

The server gives an AI agent Spicrawl as native tools. It is hosted at https://mcp.spicrawl.com/mcp, speaks MCP Streamable HTTP, and authenticates with your own API key as a bearer token. Every tool call runs under that key exactly as if you had called the API yourself: the same limits and the same credits.

Set it up in one command

The Spicrawl CLI adds the MCP server and the Spicrawl agent skill to your client, and shows the changes before applying them:

bash
npm install -g @spicrawl/cli
spicrawl init --client claude   # or: --client cursor, --client codex

Or add it by hand. Keep the key in the SPICRAWL_API_KEY environment variable and reference it from the config, so the key never lands in a committed file:

claude mcp add --transport http spicrawl https://mcp.spicrawl.com/mcp \
  --header "Authorization: Bearer $SPICRAWL_API_KEY"

# Or, for a team, commit .mcp.json (each person's own key is used):
{
  "mcpServers": {
    "spicrawl": {
      "type": "http",
      "url": "https://mcp.spicrawl.com/mcp",
      "headers": { "Authorization": "Bearer ${SPICRAWL_API_KEY}" }
    }
  }
}

Check the connection with claude mcp list, codex mcp list or your client’s MCP settings. You should see 25 spicrawl_* tools.

The tools, grouped by job

JobTools
Read one pagespicrawl_scrape: Markdown (default), text, HTML or JSON, with optional rendering, tag filters, extraction and browser actions
Many URLsspicrawl_batch_submit, _status, _list, _results, _task_content, _add_items, _close, _retry, _cancel
Logged-in pagesspicrawl_session_create, _list, _get, _context, _release, _delete
Debug a failurespicrawl_requests_list, spicrawl_request_get
Credits and usagespicrawl_usage, _summary, _reconciliation (need the read scope)
Look things upspicrawl_docs_search, spicrawl_docs_read, spicrawl_docs_index

The tool for connecting to a remote browser over CDP is listed but coming soon.

A first task to try

prompt
Use Spicrawl to read https://example.com/pricing as markdown.
List each plan with its monthly price. If the page comes back empty,
retry with rendering.

The agent calls spicrawl_scrape with the URL and format: "markdown", and adds render: true only if the first result is empty. A full walkthrough is in how to scrape a website with Claude Code.

Keep agents cheap and safe

  • Cheapest first. A plain request is 1 credit, browser rendering 3, full Chromium 8. Agents should escalate one step at a time and stop at the first that works.
  • Cap every call. Set max_cost and a request that would cost more is refused before it runs.
  • Failures are free. Errors and blocked pages cost 0 credits.
  • Test without spending. spicrawl_test_ keys run against a sandbox and never spend live credits.
  • Limit what a key can do. Keys get the scrape, batch and sessions scopes by default; grant read only if the agent should see usage.

What it does not do

  • No crawling. Spicrawl does not follow links or discover URLs. Give the agent the URLs, or a sitemap’s list, and use a batch job for many pages.
  • Coming soon: managed proxies, stealth mode, a remote browser connection, and extraction from a plain-language prompt. Today you can bring your own proxy at no extra cost, and extract data with CSS selectors or a JSON Schema.

How it compares with other scraping MCP servers

Most scraping APIs now ship an MCP server. What each offers, from our comparison pages, which cite each vendor’s own docs:

ProviderTheir MCP server
FirecrawlHosted MCP server that works without a key, with OAuth sign-in or with an API key.
ZenRowsHosted MCP server, plus an agent toolkit of skills and plugins for Claude Code, Cursor and Codex.
Context.devHosted MCP server that signs in with OAuth instead of an API key, plus an agent skill.
ScrapingBeeHosted MCP server with tools for fetching pages, text, screenshots, search results and usage.
Bright DataMIT-licensed MCP server, hosted or run locally. A free default mode scrapes pages as Markdown and searches the web; Pro mode adds 60+ tools.
Jina ReaderHosted MCP server with tools to read URLs, take screenshots and search. Reading works without a key.
ApifyOfficial MCP server that lets agents discover and run ready-made scrapers (Actors).
BrowserbaseHosted MCP server built on Stagehand that controls a cloud browser with plain-language commands.

Choose Spicrawl’s server when your agent mostly reads pages it already has URLs for, and you want clean Markdown, sessions for logged-in pages and failures that cost nothing. Choose one that crawls or searches when the agent has to find pages first.

Frequently asked questions

What is a web scraping MCP server?

It is a Model Context Protocol server that gives an AI agent tools to fetch web pages. Instead of writing HTTP code, the agent calls a tool such as spicrawl_scrape with a URL and gets the page back as Markdown, text, HTML or JSON.

Which AI agents and editors does it work with?

Any MCP client that can connect to a remote server over Streamable HTTP and send an Authorization header, including Claude Code, Cursor, VS Code and the Codex CLI. Claude Desktop can connect through the mcp-remote bridge.

Is the Spicrawl MCP server free?

Spicrawl is free during its beta, with no credit card. Tool calls use the same credits as the API: 1 for a plain request, 3 for browser rendering and 8 for full Chromium. Failed and blocked requests cost 0.

Can the MCP server crawl a whole website?

No. Spicrawl does not discover URLs or follow links. Give the agent the URLs you want, or a list from a sitemap, and use a batch job for many pages.

How do I keep my API key safe?

Keep the key in the SPICRAWL_API_KEY environment variable and reference it from your client configuration, for example ${SPICRAWL_API_KEY} in Claude Code’s .mcp.json or ${env:SPICRAWL_API_KEY} in Cursor. Never commit the key itself, and revoke it in the dashboard if it leaks.

Does it work with logged-in pages?

Yes, for accounts you own or may use. Create a session with spicrawl_session_create, log in with browser actions, then pass the session ID on later scrapes so cookies and browser storage carry over.