Skip to content

Batch web scraping API

Send up to 10,000 URLs in one call. Get every page back as clean data.

Submit the URLs
batch.py
import requestsjob = requests.post(    "https://api.spicrawl.com/v1/batch",    headers={"Authorization": "Bearer spicrawl_live_…"},    json={        "urls": ["https://example.com/one", "https://example.com/two"],        "js_render": True,    },).json()# Poll GET /v1/batch/{id} until status is "completed"print(job["id"], job["total_items"])

How it works

  1. Send your URLs to POST /v1/batch. You get a job ID back at once.
  2. Check the job until its status is completed, failed or cancelled.
  3. Read the results as JSON Lines, one page at a time, with X-Next-Cursor.
curl https://api.spicrawl.com/v1/batch \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "catalog-refresh",
    "urls": ["https://example.com/products/1", "https://example.com/products/2"],
    "response_format": "markdown"
  }'

Keep your key in SPICRAWL_API_KEY. Your key needs the batch scope.

What each item can do

  • Formats: html (the default), markdown, text or json.
  • Clean-up: main_content_only, include_tags and exclude_tags work on Markdown and text.
  • JavaScript: set js_render: true for pages that need a browser.
  • Per-URL settings: send items instead of urls to give each URL its own settings and an external_id.

Good to know

  • Not on batch: extraction, screenshots, browser actions and sessions. A job that asks for them is refused with a 400 before anything is charged. Use single scrapes for those pages.
  • Results expire 72 hours after the job finishes.
  • No crawling. Spicrawl does not find URLs or follow links. You send the list.
  • Pricing: free during the beta. Each successful item uses the same credits as a single scrape.

Frequently asked questions

How many URLs can one batch job have?

Up to 10,000 URLs per call. Send them as a list in urls, or as items when each URL needs its own settings or an external_id you can match on.

What formats does a batch job return?

Each item comes back as html (the default), markdown, text or json. Markdown and text use the same conversion as a single scrape.

How long are batch results kept?

For 72 hours after the job finishes. Read them with GET /v1/batch/{id}/results, which returns JSON Lines one page at a time.

Can a batch job extract fields or click buttons?

No. Extraction, screenshots, browser actions and sessions do not run on batch. A job that asks for them is refused with a 400 before anything is charged. Use single scrapes for those pages.

Does batch crawl a website?

No. Spicrawl does not find URLs or follow links. You send the list of URLs you want.