Skip to content

OpenAI Agents SDK web scraping

Connect an OpenAI Agents SDK agent to the Spicrawl MCP server. The agent can then read web pages as Markdown.

Set it up

  1. Install openai-agents.

  2. Set SPICRAWL_API_KEY and OPENAI_API_KEY.

  3. Open MCPServerStreamableHttp with the URL and the Authorization header in params.

  4. Use create_static_tool_filter to limit which tools the model sees, and pass the server to your Agent.

Example: read a pricing page

openai_agents_example.py
import asyncio
import os

from agents import Agent, Runner
from agents.mcp import MCPServerStreamableHttp, create_static_tool_filter


async def main() -> None:
    async with MCPServerStreamableHttp(
        name="spicrawl",
        params={
            "url": "https://mcp.spicrawl.com/mcp",
            "headers": {"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
        },
        cache_tools_list=True,
        tool_filter=create_static_tool_filter(
            allowed_tool_names=["spicrawl_scrape", "spicrawl_batch_submit"]
        ),
    ) as server:
        agent = Agent(
            name="Page reader",
            instructions="Use the Spicrawl tools to read web pages.",
            mcp_servers=[server],
        )
        result = await Runner.run(
            agent,
            "Read https://example.com/pricing as markdown and list the plans.",
        )
        print(result.final_output)


asyncio.run(main())

The server has no OAuth, so always send the key in headers. If you get a 401, check that SPICRAWL_API_KEY is set in the process that runs the agent.

Keep the tool list short

All 25 tool schemas are sent on every model call. Keep spicrawl_scrape, and for many URLs add spicrawl_batch_submit, spicrawl_batch_status and spicrawl_batch_results.

If your code already knows which URL to fetch, a plain function that calls POST /v1/scrape is simpler. It costs no tool-schema tokens and is easier to test.

What it can do

The server gives the agent 25 spicrawl_* tools. Each call runs under your API key, with the same limits and credits as a direct API call.

  • Read one page: spicrawl_scrape returns Markdown (the default), text, HTML, JSON or PDF.
  • Read many pages: spicrawl_batch_submit queues up to 10,000 URLs. Follow it with spicrawl_batch_status and spicrawl_batch_results.
  • Pages behind a login: spicrawl_session_create keeps cookies and storage across calls.
  • Debug a failure: spicrawl_requests_list and spicrawl_request_get.
  • Check usage: spicrawl_usage_summary and the other usage tools. These need a key with the read scope.
  • Look things up: spicrawl_docs_search and spicrawl_docs_read search and read the Spicrawl docs.

One tool, spicrawl_browser_connect_url, is listed but coming soon. On spicrawl_scrape, the rendering argument is render, not the API’s js_render.

Good to know

  • No crawling. Spicrawl does not follow links or find pages. Give the agent the URLs.
  • Credits per request. A plain fetch is 1 credit. JavaScript rendering is 3. Full-browser rendering is 8.
  • Failures are free. A failed request costs 0 credits.
  • The cache saves time, not credits. A cache hit costs the same as the fetch that stored it.
  • Cap each call. Set max_cost, and a request that could cost more is refused before it runs.
  • Coming soon: managed proxies, stealth mode, a remote browser and AI extraction from a plain-language prompt. Today you can bring your own proxy.

Frequently asked questions

Which package do I need?

openai-agents. Set SPICRAWL_API_KEY and OPENAI_API_KEY before you run the example.

How do I limit which tools the model sees?

Pass tool_filter=create_static_tool_filter(allowed_tool_names=[...]). The example keeps spicrawl_scrape and spicrawl_batch_submit.

Why do I get a 401 when the agent connects?

The header is missing or the key is wrong. Build the header in code from SPICRAWL_API_KEY and check the variable is set in the process that runs the agent.

What happens when a tool call fails?

The error reaches the model as an MCP tool error with the API’s code. Failed calls cost 0 credits.

How much does each page cost?

A plain fetch costs 1 credit, JavaScript rendering 3 and full-browser rendering 8. Failed requests cost 0. A cache hit costs the same as the fetch that stored it. Ask the agent to set max_cost on each request so it never spends more than you expect.