Web scraping in Windsurf
Windsurf is now called Devin Desktop. Add the Spicrawl MCP server and its agent can read web pages as clean Markdown.
Set it up
Devin Desktop has two agents. Devin Local is the default for new tabs. Cascade is the legacy agent. They read MCP servers from different files, so set up the one you use.
Set your API key in the environment Devin Desktop starts from:
bashexport SPICRAWL_API_KEY=spicrawl_live_... # add to your shell profile, then restart Devin DesktopDevin Local: add the server in local scope. It is saved to
.devin/mcp_config.local.json, which Devin gitignores. Check it withdevin mcp list.bashdevin mcp add spicrawl https://mcp.spicrawl.com/mcp -H "Authorization: Bearer $SPICRAWL_API_KEY"Do not use
-s project. It writes the shared, committed file with your key in it.Cascade: open the Cascade panel, click the
...menu and open the MCP config file. Add the server. Cascade fills in the key, so the file holds no key:~/.config/devin/mcp_config.json{ "mcpServers": { "spicrawl": { "serverUrl": "https://mcp.spicrawl.com/mcp", "headers": { "Authorization": "Bearer ${env:SPICRAWL_API_KEY}" } } } }Cascade allows 100 tools across all servers. Spicrawl uses 25 of them.
Install the skill:
bashspicrawl skill install --client codex # Without the CLI: mkdir -p .agents/skills/spicrawl curl -fsSL https://app.spicrawl.com/skill.md -o .agents/skills/spicrawl/SKILL.mdSave this rule as
.devin/rules/spicrawl.md(.windsurf/rules/still works). The agent reads it when a task involves web data..devin/rules/spicrawl.md--- trigger: model_decision description: Fetching, scraping or extracting data from web pages with Spicrawl --- # Web data (Spicrawl) Use Spicrawl to read web pages: the spicrawl_* MCP tools in chat, or POST https://api.spicrawl.com/v1/scrape with `Authorization: Bearer $SPICRAWL_API_KEY` in code. Docs: https://docs.spicrawl.com/llms.txt (append .md to any page URL for Markdown). - Ask for markdown: `response_format: "markdown"` (API default is html); on spicrawl_scrape, `format: "markdown"`. - Check the site's status, not only the HTTP status: `X-Target-Status` header, or `status` in the JSON envelope (spicrawl_scrape with `format: "json"`). 200 is the page; 404/410 mean it does not exist; 403/429/503 mean the site refused, so escalate. - On an error, switch on `code`. Retry only when `retryable` is true, after `retry_after_seconds`. Read `diagnostics.hint` and change what it names first. Never retry ERR::REQUEST::*, ERR::AUTH::* or ERR::LIMIT::QUOTA_EXCEEDED. - Escalate one step at a time and stop at the first that works: plain fetch (1 credit) -> `js_render: true` (3) -> add the user's own `proxy` if they have one. Empty or skeleton content means render; ERR::UPSTREAM::CHALLENGE or a 403 target status means retry once, then the user's own proxy. - Set `max_cost` on every request to the price of the step you intend. - Keep the cache on (default). Set `cache: false` only for prices, stock or other live data. - Trim tokens with `main_content_only` (on by default for markdown), `include_tags`, `exclude_tags`. - For more than 20 URLs, use a batch job (spicrawl_batch_submit / POST /v1/batch). Batch items return html, markdown or text (spicrawl_batch_submit defaults to markdown, POST /v1/batch to html). Extraction, screenshots, actions and sessions are refused with a 400 before anything is charged: scrape those URLs one by one. - Never print or commit SPICRAWL_API_KEY. Log `X-Request-Id` for failures.
A first task to try
Open a new agent tab and ask:
Use Spicrawl to read https://example.com/pricing as markdown. List each plan with its monthly price.
If the page comes back empty, retry with rendering.The agent should call spicrawl_scrape with format: "markdown". It should add render: true only if the first result is empty.
What it can do
The server gives the agent 25 spicrawl_* tools. Each call runs under your API key, with the same limits and credits as a direct API call.
- Read one page:
spicrawl_scrapereturns Markdown (the default), text, HTML, JSON or PDF. - Read many pages:
spicrawl_batch_submitqueues up to 10,000 URLs. Follow it withspicrawl_batch_statusandspicrawl_batch_results. - Pages behind a login:
spicrawl_session_createkeeps cookies and storage across calls. - Debug a failure:
spicrawl_requests_listandspicrawl_request_get. - Check usage:
spicrawl_usage_summaryand the other usage tools. These need a key with thereadscope. - Look things up:
spicrawl_docs_searchandspicrawl_docs_readsearch and read the Spicrawl docs.
One tool, spicrawl_browser_connect_url, is listed but coming soon. On spicrawl_scrape, the rendering argument is render, not the API’s js_render.
Good to know
- No crawling. Spicrawl does not follow links or find pages. Give the agent the URLs.
- Credits per request. A plain fetch is 1 credit. JavaScript rendering is 3. Full-browser rendering is 8.
- Failures are free. A failed request costs 0 credits.
- The cache saves time, not credits. A cache hit costs the same as the fetch that stored it.
- Cap each call. Set
max_cost, and a request that could cost more is refused before it runs. - Coming soon: managed proxies, stealth mode, a remote browser and AI extraction from a plain-language prompt. Today you can bring your own proxy.
Frequently asked questions
Is Windsurf the same as Devin Desktop?
Yes. Windsurf was renamed Devin Desktop on 2026-06-02. It is the same editor, and these steps work for both names.
Which agent should I set up, Devin Local or Cascade?
Devin Local is the default agent for new tabs. Cascade is the legacy agent. They read MCP servers from different files, so set up the one you use.
Does Spicrawl count toward the Cascade tool limit?
Yes. Cascade allows 100 tools across all servers, and Spicrawl uses 25. Disable other servers if you are near the limit.
My team blocks the server. What should I do?
A team allowlist is active. Ask the admin to allowlist the Server ID spicrawl (case-sensitive).
How much does each page cost?
A plain fetch costs 1 credit, JavaScript rendering 3 and full-browser rendering 8. Failed requests cost 0. A cache hit costs the same as the fetch that stored it. Ask the agent to set max_cost on each request so it never spends more than you expect.