Skip to content

How to scrape a website with Claude Code

Set up Spicrawl's MCP tools and skill, then ask Claude Code to read known URLs while it works.

By Spicrawl teamPublished 6 min read

Pixel-art illustration of a robot labelled Claude Code typing at a laptop next to a terminal reading scrape website, extract data, save as json, with arrows from a web page at example.com to a window of scraped JSON data.
On this page

Claude Code can use live web pages while it works: a documentation page for an implementation, a published price for a comparison, or a set of URLs for research. The page still needs fetching and turning into something the agent can read.

This guide connects Claude Code to Spicrawl's hosted MCP server and agent skill. You provide the URLs; Spicrawl returns each page as Markdown, HTML, text or JSON. It does not crawl a site or discover URLs for you.

What you need

  • Claude Code and a free Spicrawl beta account.
  • An API key available as SPICRAWL_API_KEY in the shell that starts Claude Code. Set it using your shell's secret storage or, for a quick local session, run read -rs SPICRAWL_API_KEY && export SPICRAWL_API_KEY and paste your own key (it stays out of your screen and shell history). Do not put the real key in a project file or commit it.
  • Node.js 16 or later if you use the npm CLI installer below.

The fastest way: one command

From your project's root directory, install the published CLI and run its Claude Code setup:

bash
npm install -g @spicrawl/cli
spicrawl init --client claude

init shows the changes and asks before applying them. It adds the hosted MCP server to the project's .mcp.json and installs the Spicrawl skill at .claude/skills/spicrawl/SKILL.md. Its project configuration refers to ${SPICRAWL_API_KEY}, so each developer supplies their own key. Add the short CLAUDE.md rules below as a separate step: the CLI's documented default installs the MCP configuration and skill, not those rules. In a non-interactive shell, use the documented --yes flag.

Or install the plugin

The Spicrawl plugin marketplace is available if you prefer a Claude Code plugin. Its documented commands are:

bash
claude plugin marketplace add OfficialSpicrawl/agent-plugins
claude plugin install spicrawl@spicrawl-plugins

The plugin bundles the MCP server and skill. Claude Code asks for your key when you enable it. Add the project rules below if you want the same guidance in CLAUDE.md. The Spicrawl Claude Code guide describes this route.

Or set it up by hand

Add the MCP server

For your own use across projects, run this in a shell where SPICRAWL_API_KEY is set:

bash
claude mcp add --transport http --scope user spicrawl https://mcp.spicrawl.com/mcp \
  --header "Authorization: Bearer $SPICRAWL_API_KEY"

The shell expands the key and Claude Code stores the header in your user configuration outside the repository. For a team, create a .mcp.json at the project root instead:

json
{
  "mcpServers": {
    "spicrawl": {
      "type": "http",
      "url": "https://mcp.spicrawl.com/mcp",
      "headers": { "Authorization": "Bearer ${SPICRAWL_API_KEY}" }
    }
  }
}

Claude Code expands ${SPICRAWL_API_KEY} from each teammate's environment. If you use claude mcp add --scope project, the shell expands the header before Claude writes .mcp.json; replace the expanded value with the variable reference before committing that file. The Claude Code MCP reference explains scopes and environment variable expansion.

Check it works

Run claude mcp list, or open /mcp in Claude Code, and check the Spicrawl connection. The hosted server documents 25 spicrawl_* tools, including spicrawl_scrape, batch and session tools. Restart Claude Code if you added the configuration during an existing session.

Install the skill

bash
mkdir -p .claude/skills/spicrawl
curl -fsSL https://app.spicrawl.com/skill.md -o .claude/skills/spicrawl/SKILL.md

This gives Claude Code the Spicrawl skill for choosing tools and writing API calls. For all your projects, the docs also support ~/.claude/skills/spicrawl/.

Add rules to CLAUDE.md

Add a short project rule so Claude asks for the right output and checks failures:

markdown
## Web data (Spicrawl)

Use the spicrawl_* MCP tools to read URLs I provide. Ask spicrawl_scrape for
format: "markdown" when you need page text. Use format: "json" when you must
check the target site's status or credits. Retry an error only if retryable
is true. Start with a plain request, set max_cost, and render only if needed.
Never print or commit SPICRAWL_API_KEY.

This is an excerpt. The full rules also cover target statuses, error codes, caching and batches. The MCP tool uses format and render; the REST API uses response_format and js_render.

Try it

Start with one known, public URL:

Use spicrawl_scrape to read https://example.com as Markdown. Summarise the page in two sentences and cite the URL. If you need to confirm the site's status, request JSON and check status before answering.

The equivalent REST request is:

bash
curl -sS -X POST https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url":"https://example.com","response_format":"markdown","max_cost":1}'

Here are three more tasks to try with URLs you are allowed to read:

Read the public pricing URL I give you. Return JSON with each plan name, price, billing period and source URL. If a field is missing, use null rather than guessing. Turn off the cache if I need the current price.

Read these 30 public documentation URLs as one batch job, as Markdown. Give me the job status and the URLs that failed.

Read this JavaScript-heavy public page. Start with a plain fetch. If it is an empty loading shell, try rendering, then summarise the content and tell me which step worked.

Pages that need more

JavaScript pages. Start with the plain request. If it returns a loading shell, try render: true on spicrawl_scrape (or js_render: true in the REST API). mode: "auto" can choose a browser after a plain fetch, but check the allowed max_cost before using it. See best practices.

Pages behind a login. If you are authorised to use the site, create a session, log in with browser actions, then pass the same session_id on later scrapes. Sessions retain cookies and browser storage. Do not put credentials in prompts or logs.

Clicks and scrolling. Browser actions can click, type, select, scroll and wait within a request. They are useful when content appears only after an interaction. Spicrawl's remote browser connection is coming soon; these actions run through scrape requests.

Keeping it cheap and reliable

  • Ask for Markdown when you need page text. Its main-content filter is on by default; use include_tags or exclude_tags when you need a narrower section.
  • Start with a plain fetch (1 credit), then render if needed (3 credits). Full Chromium costs 8 credits. Set max_cost to the step you intend before each request. Blocked requests and API errors cost 0; a real 404 or 410 page is billable.
  • Check X-Target-Status on REST responses. API HTTP 200 says the request completed, not that the target page returned 200. With the MCP tool, request format: "json" when you need the target status or cost; Markdown output has no response headers.
  • Keep the cache on for stable pages. Turn it off for time-sensitive prices or stock. Reuse content already in your agent's context instead of fetching it again.

Troubleshooting

SymptomWhat to check
Spicrawl tools do not appearUse claude mcp list or /mcp, confirm the server is https://mcp.spicrawl.com/mcp, and restart Claude Code after changing its configuration.
The MCP server reports 401Check that SPICRAWL_API_KEY is set in the shell that starts Claude Code and that the key has not been revoked. Do not print the key while debugging.
The target site returns 403Check the target status rather than the API status. Retry once where the error says to, then escalate one step; use your own proxy if you have one.
A tool rejects js_renderUse render for spicrawl_scrape; js_render is the REST field.
An API request failsRead its code and diagnostics.hint. Retry only when retryable is true, waiting for retry_after_seconds when present.

Spicrawl accepts known URLs, not a whole site to discover. If your task starts with a sitemap, fetch and select its URLs first, then send those URLs to individual scrape calls or a batch job.

Frequently asked questions

Can Claude Code scrape websites?

Yes. Connect Spicrawl's hosted MCP server, then ask Claude Code to use spicrawl_scrape on a URL you provide. It can read the returned Markdown or JSON while working on your task.

Is Spicrawl free to use with Claude Code?

Spicrawl is free during its beta and does not require a credit card. Successful plain requests use 1 credit, browser rendering 3 and full Chromium 8; blocked requests and errors cost 0.

Does the same setup work with Cursor or Codex?

Yes. The CLI accepts --client cursor or --client codex. See the Spicrawl Cursor and Codex setup guides for the configuration each client uses.

Can Claude Code crawl a whole site with Spicrawl?

Spicrawl does not discover URLs or follow links. Give it the URLs you want, perhaps from a sitemap you fetched, and use a batch job when you have many of them.

Can it read pages behind a login?

Yes, when you are authorised to access the pages. Use a Spicrawl session to keep cookies and browser storage, perform the login with browser actions, and reuse the session ID for later requests.

Sources

Product details were checked against each company’s own website and docs. Products change: if something here is out of date, email support@spicrawl.com.