LangChain web scraping with Spicrawl
Connect a LangChain agent to the Spicrawl MCP server. The agent can then read web pages as Markdown.
Set it up
Install
langchain-mcp-adaptersandlangchain[openai].Set
SPICRAWL_API_KEYand your model provider’s key (OPENAI_API_KEYhere).Create a
MultiServerMCPClientwith onespicrawlentry."transport": "streamable_http"andheadersset up the connection.Keep only the tools you need, and build a LangGraph agent with
create_agent.
Example: read a pricing page
import asyncio
import os
from langchain.agents import create_agent
from langchain_mcp_adapters.client import MultiServerMCPClient
async def main() -> None:
client = MultiServerMCPClient(
{
"spicrawl": {
"transport": "streamable_http",
"url": "https://mcp.spicrawl.com/mcp",
"headers": {"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
}
}
)
keep = {"spicrawl_scrape", "spicrawl_batch_submit"}
tools = [t for t in await client.get_tools() if t.name in keep]
agent = create_agent("openai:gpt-4.1", tools)
result = await agent.ainvoke(
{"messages": "Read https://example.com/pricing as markdown and list the plans."}
)
print(result["messages"][-1].content)
asyncio.run(main())By default, each tool call opens its own MCP session. That is fine for spicrawl_scrape, which is stateless. The server has no OAuth, so always send the key in headers.
Keep the tool list short
All 25 tool schemas are sent on every model call. That raises prompt cost and can confuse the model. Keep spicrawl_scrape, and for many URLs add spicrawl_batch_submit, spicrawl_batch_status and spicrawl_batch_results.
If your code already knows which URL to fetch, you may not need MCP at all. A plain function that calls POST /v1/scrape costs no tool-schema tokens and is easier to test.
What it can do
The server gives the agent 25 spicrawl_* tools. Each call runs under your API key, with the same limits and credits as a direct API call.
- Read one page:
spicrawl_scrapereturns Markdown (the default), text, HTML, JSON or PDF. - Read many pages:
spicrawl_batch_submitqueues up to 10,000 URLs. Follow it withspicrawl_batch_statusandspicrawl_batch_results. - Pages behind a login:
spicrawl_session_createkeeps cookies and storage across calls. - Debug a failure:
spicrawl_requests_listandspicrawl_request_get. - Check usage:
spicrawl_usage_summaryand the other usage tools. These need a key with thereadscope. - Look things up:
spicrawl_docs_searchandspicrawl_docs_readsearch and read the Spicrawl docs.
One tool, spicrawl_browser_connect_url, is listed but coming soon. On spicrawl_scrape, the rendering argument is render, not the API’s js_render.
Good to know
- No crawling. Spicrawl does not follow links or find pages. Give the agent the URLs.
- Credits per request. A plain fetch is 1 credit. JavaScript rendering is 3. Full-browser rendering is 8.
- Failures are free. A failed request costs 0 credits.
- The cache saves time, not credits. A cache hit costs the same as the fetch that stored it.
- Cap each call. Set
max_cost, and a request that could cost more is refused before it runs. - Coming soon: managed proxies, stealth mode, a remote browser and AI extraction from a plain-language prompt. Today you can bring your own proxy.
Frequently asked questions
Which packages do I need?
langchain-mcp-adapters and langchain[openai]. Set SPICRAWL_API_KEY and OPENAI_API_KEY before you run the example.
Should I give the agent all 25 tools?
Usually not. All tool schemas are sent on every model call. Keep spicrawl_scrape, and for many URLs spicrawl_batch_submit, spicrawl_batch_status and spicrawl_batch_results.
Does each tool call open a new session?
Yes, by default. That is fine for spicrawl_scrape, which is stateless. See the adapter’s README for client.session("spicrawl") if you want one session per run.
Should I use MCP or call the API directly?
Use MCP when the model decides what to fetch. When your code already knows the URL, a plain function that calls POST /v1/scrape is simpler and costs no tool-schema tokens.
How much does each page cost?
A plain fetch costs 1 credit, JavaScript rendering 3 and full-browser rendering 8. Failed requests cost 0. A cache hit costs the same as the fetch that stored it. Ask the agent to set max_cost on each request so it never spends more than you expect.