Skip to content

Web scraping in VS Code with Copilot

Add the Spicrawl MCP server to VS Code. Copilot in agent mode can then read web pages as clean Markdown.

Set it up

Three steps give Copilot the tools, the know-how to write Spicrawl code, and rules for using both. spicrawl init --client vscode does the first two for you.

  1. Create .vscode/mcp.json in the project. VS Code uses the servers key, not mcpServers. The password input makes VS Code ask for your key once and store it, so the file is safe to commit:

    .vscode/mcp.json
    {
      "inputs": [
        {
          "type": "promptString",
          "id": "spicrawl-api-key",
          "description": "Spicrawl API key",
          "password": true
        }
      ],
      "servers": {
        "spicrawl": {
          "type": "http",
          "url": "https://mcp.spicrawl.com/mcp",
          "headers": { "Authorization": "Bearer ${input:spicrawl-api-key}" }
        }
      }
    }
  2. Run MCP: List Servers from the Command Palette and start spicrawl. Enter your key when asked.

  3. Install the skill, and commit it:

    bash
    spicrawl skill install --client vscode
    
    # Without the CLI:
    mkdir -p .github/skills/spicrawl
    curl -fsSL https://app.spicrawl.com/skill.md -o .github/skills/spicrawl/SKILL.md
  4. Save this as .github/copilot-instructions.md, or add it to the file you already have:

    .github/copilot-instructions.md
    # Web data (Spicrawl)
    
    Use Spicrawl to read web pages: the spicrawl_* MCP tools in chat, or
    POST https://api.spicrawl.com/v1/scrape with `Authorization: Bearer $SPICRAWL_API_KEY` in code.
    Docs: https://docs.spicrawl.com/llms.txt (append .md to any page URL for Markdown).
    
    - Ask for markdown: `response_format: "markdown"` (API default is html); on spicrawl_scrape, `format: "markdown"`.
    - Check the site's status, not only the HTTP status: `X-Target-Status` header, or `status`
      in the JSON envelope (spicrawl_scrape with `format: "json"`). 200 is the page; 404/410
      mean it does not exist; 403/429/503 mean the site refused, so escalate.
    - On an error, switch on `code`. Retry only when `retryable` is true, after
      `retry_after_seconds`. Read `diagnostics.hint` and change what it names first.
      Never retry ERR::REQUEST::*, ERR::AUTH::* or ERR::LIMIT::QUOTA_EXCEEDED.
    - Escalate one step at a time and stop at the first that works:
      plain fetch (1 credit) -> `js_render: true` (3) -> add the user's own `proxy`
      if they have one. Empty or skeleton content means render; ERR::UPSTREAM::CHALLENGE
      or a 403 target status means retry once, then the user's own proxy.
    - Set `max_cost` on every request to the price of the step you intend.
    - Keep the cache on (default). Set `cache: false` only for prices, stock or other live data.
    - Trim tokens with `main_content_only` (on by default for markdown), `include_tags`, `exclude_tags`.
    - For more than 20 URLs, use a batch job (spicrawl_batch_submit / POST /v1/batch).
      Batch items return html, markdown or text (spicrawl_batch_submit defaults to markdown,
      POST /v1/batch to html). Extraction, screenshots, actions and sessions are refused
      with a 400 before anything is charged: scrape those URLs one by one.
    - Never print or commit SPICRAWL_API_KEY. Log `X-Request-Id` for failures.
  5. Open the Chat view, switch to Agent mode and open the tools picker. spicrawl lists 25 spicrawl_* tools.

The Copilot CLI and the Copilot cloud agent each read their own MCP config. The docs page covers both.

A first task to try

Open Copilot Chat in Agent mode and ask:

prompt
Use Spicrawl to read https://example.com/pricing as markdown. List each plan with its monthly price.
If the page comes back empty, retry with rendering.

The agent should call spicrawl_scrape with format: "markdown". It should add render: true only if the first result is empty.

What it can do

The server gives the agent 25 spicrawl_* tools. Each call runs under your API key, with the same limits and credits as a direct API call.

  • Read one page: spicrawl_scrape returns Markdown (the default), text, HTML, JSON or PDF.
  • Read many pages: spicrawl_batch_submit queues up to 10,000 URLs. Follow it with spicrawl_batch_status and spicrawl_batch_results.
  • Pages behind a login: spicrawl_session_create keeps cookies and storage across calls.
  • Debug a failure: spicrawl_requests_list and spicrawl_request_get.
  • Check usage: spicrawl_usage_summary and the other usage tools. These need a key with the read scope.
  • Look things up: spicrawl_docs_search and spicrawl_docs_read search and read the Spicrawl docs.

One tool, spicrawl_browser_connect_url, is listed but coming soon. On spicrawl_scrape, the rendering argument is render, not the API’s js_render.

Good to know

  • No crawling. Spicrawl does not follow links or find pages. Give the agent the URLs.
  • Credits per request. A plain fetch is 1 credit. JavaScript rendering is 3. Full-browser rendering is 8.
  • Failures are free. A failed request costs 0 credits.
  • The cache saves time, not credits. A cache hit costs the same as the fetch that stored it.
  • Cap each call. Set max_cost, and a request that could cost more is refused before it runs.
  • Coming soon: managed proxies, stealth mode, a remote browser and AI extraction from a plain-language prompt. Today you can bring your own proxy.

Frequently asked questions

Why does VS Code use "servers" and not "mcpServers"?

That is VS Code’s own format. Use the servers key in .vscode/mcp.json, as in the example on this page.

Where is my key stored?

The inputs entry with password: true makes VS Code ask for the key once and store it. The file only holds ${input:spicrawl-api-key}, so you can commit it.

Does it work with the Copilot CLI and the cloud agent?

Yes, but each reads its own MCP config. The Copilot CLI uses copilot mcp add. The cloud agent is set up by a repository admin under Settings → Copilot → MCP servers. The docs page covers both.

What does spicrawl init --client vscode write?

It writes .vscode/mcp.json and the skill at .github/skills/spicrawl/SKILL.md. It does not write the instructions file or the Copilot CLI or cloud agent config.

How much does each page cost?

A plain fetch costs 1 credit, JavaScript rendering 3 and full-browser rendering 8. Failed requests cost 0. A cache hit costs the same as the fetch that stored it. Ask the agent to set max_cost on each request so it never spends more than you expect.