A hosted web scraping MCP server for AI agents
Connect your agent to https://mcp.spicrawl.com/mcp and it can read any URL as Markdown or JSON, run batch jobs, reuse logged-in sessions and check its own usage, without you writing scraping code.
What the Spicrawl MCP server does
The server gives an AI agent Spicrawl as native tools. It is hosted at https://mcp.spicrawl.com/mcp, speaks MCP Streamable HTTP, and authenticates with your own API key as a bearer token. Every tool call runs under that key exactly as if you had called the API yourself: the same limits and the same credits.
Set it up in one command
The Spicrawl CLI adds the MCP server and the Spicrawl agent skill to your client, and shows the changes before applying them:
npm install -g @spicrawl/cli
spicrawl init --client claude # or: --client cursor, --client codexOr add it by hand. Keep the key in the SPICRAWL_API_KEY environment variable and reference it from the config, so the key never lands in a committed file:
claude mcp add --transport http spicrawl https://mcp.spicrawl.com/mcp \
--header "Authorization: Bearer $SPICRAWL_API_KEY"
# Or, for a team, commit .mcp.json (each person's own key is used):
{
"mcpServers": {
"spicrawl": {
"type": "http",
"url": "https://mcp.spicrawl.com/mcp",
"headers": { "Authorization": "Bearer ${SPICRAWL_API_KEY}" }
}
}
}// .cursor/mcp.json (or ~/.cursor/mcp.json for every project)
{
"mcpServers": {
"spicrawl": {
"url": "https://mcp.spicrawl.com/mcp",
"headers": { "Authorization": "Bearer ${env:SPICRAWL_API_KEY}" }
}
}
}// .vscode/mcp.json: VS Code asks for the key once and stores it securely
{
"inputs": [
{ "type": "promptString", "id": "spicrawl-api-key", "description": "Spicrawl API key", "password": true }
],
"servers": {
"spicrawl": {
"type": "http",
"url": "https://mcp.spicrawl.com/mcp",
"headers": { "Authorization": "Bearer ${input:spicrawl-api-key}" }
}
}
}codex mcp add spicrawl --url https://mcp.spicrawl.com/mcp --bearer-token-env-var SPICRAWL_API_KEY
# Equivalent ~/.codex/config.toml:
[mcp_servers.spicrawl]
url = "https://mcp.spicrawl.com/mcp"
bearer_token_env_var = "SPICRAWL_API_KEY"// claude_desktop_config.json: bridges to the hosted server with mcp-remote (needs Node.js)
{
"mcpServers": {
"spicrawl": {
"command": "npx",
"args": ["-y", "mcp-remote", "https://mcp.spicrawl.com/mcp", "--header", "Authorization:${SPICRAWL_AUTH}"],
"env": { "SPICRAWL_AUTH": "Bearer spicrawl_live_..." }
}
}
}Check the connection with claude mcp list, codex mcp list or your client’s MCP settings. You should see 25 spicrawl_* tools.
The tools, grouped by job
| Job | Tools |
|---|---|
| Read one page | spicrawl_scrape: Markdown (default), text, HTML or JSON, with optional rendering, tag filters, extraction and browser actions |
| Many URLs | spicrawl_batch_submit, _status, _list, _results, _task_content, _add_items, _close, _retry, _cancel |
| Logged-in pages | spicrawl_session_create, _list, _get, _context, _release, _delete |
| Debug a failure | spicrawl_requests_list, spicrawl_request_get |
| Credits and usage | spicrawl_usage, _summary, _reconciliation (need the read scope) |
| Look things up | spicrawl_docs_search, spicrawl_docs_read, spicrawl_docs_index |
The tool for connecting to a remote browser over CDP is listed but coming soon.
A first task to try
Use Spicrawl to read https://example.com/pricing as markdown.
List each plan with its monthly price. If the page comes back empty,
retry with rendering.The agent calls spicrawl_scrape with the URL and format: "markdown", and adds render: true only if the first result is empty. A full walkthrough is in how to scrape a website with Claude Code.
Keep agents cheap and safe
- Cheapest first. A plain request is 1 credit, browser rendering 3, full Chromium 8. Agents should escalate one step at a time and stop at the first that works.
- Cap every call. Set
max_costand a request that would cost more is refused before it runs. - Failures are free. Errors and blocked pages cost 0 credits.
- Test without spending.
spicrawl_test_keys run against a sandbox and never spend live credits. - Limit what a key can do. Keys get the
scrape,batchandsessionsscopes by default; grantreadonly if the agent should see usage.
What it does not do
- No crawling. Spicrawl does not follow links or discover URLs. Give the agent the URLs, or a sitemap’s list, and use a batch job for many pages.
- Coming soon: managed proxies, stealth mode, a remote browser connection, and extraction from a plain-language prompt. Today you can bring your own proxy at no extra cost, and extract data with CSS selectors or a JSON Schema.
How it compares with other scraping MCP servers
Most scraping APIs now ship an MCP server. What each offers, from our comparison pages, which cite each vendor’s own docs:
| Provider | Their MCP server |
|---|---|
| Firecrawl | Hosted MCP server that works without a key, with OAuth sign-in or with an API key. |
| ZenRows | Hosted MCP server, plus an agent toolkit of skills and plugins for Claude Code, Cursor and Codex. |
| Context.dev | Hosted MCP server that signs in with OAuth instead of an API key, plus an agent skill. |
| ScrapingBee | Hosted MCP server with tools for fetching pages, text, screenshots, search results and usage. |
| Bright Data | MIT-licensed MCP server, hosted or run locally. A free default mode scrapes pages as Markdown and searches the web; Pro mode adds 60+ tools. |
| Jina Reader | Hosted MCP server with tools to read URLs, take screenshots and search. Reading works without a key. |
| Apify | Official MCP server that lets agents discover and run ready-made scrapers (Actors). |
| Browserbase | Hosted MCP server built on Stagehand that controls a cloud browser with plain-language commands. |
Choose Spicrawl’s server when your agent mostly reads pages it already has URLs for, and you want clean Markdown, sessions for logged-in pages and failures that cost nothing. Choose one that crawls or searches when the agent has to find pages first.
Frequently asked questions
What is a web scraping MCP server?
It is a Model Context Protocol server that gives an AI agent tools to fetch web pages. Instead of writing HTTP code, the agent calls a tool such as spicrawl_scrape with a URL and gets the page back as Markdown, text, HTML or JSON.
Which AI agents and editors does it work with?
Any MCP client that can connect to a remote server over Streamable HTTP and send an Authorization header, including Claude Code, Cursor, VS Code and the Codex CLI. Claude Desktop can connect through the mcp-remote bridge.
Is the Spicrawl MCP server free?
Spicrawl is free during its beta, with no credit card. Tool calls use the same credits as the API: 1 for a plain request, 3 for browser rendering and 8 for full Chromium. Failed and blocked requests cost 0.
Can the MCP server crawl a whole website?
No. Spicrawl does not discover URLs or follow links. Give the agent the URLs you want, or a list from a sitemap, and use a batch job for many pages.
How do I keep my API key safe?
Keep the key in the SPICRAWL_API_KEY environment variable and reference it from your client configuration, for example ${SPICRAWL_API_KEY} in Claude Code’s .mcp.json or ${env:SPICRAWL_API_KEY} in Cursor. Never commit the key itself, and revoke it in the dashboard if it leaks.
Does it work with logged-in pages?
Yes, for accounts you own or may use. Create a session with spicrawl_session_create, log in with browser actions, then pass the session ID on later scrapes so cookies and browser storage carry over.