Batch web scraping API
Send up to 10,000 URLs in one call. Get every page back as clean data.
import requestsjob = requests.post( "https://api.spicrawl.com/v1/batch", headers={"Authorization": "Bearer spicrawl_live_…"}, json={ "urls": ["https://example.com/one", "https://example.com/two"], "js_render": True, },).json()# Poll GET /v1/batch/{id} until status is "completed"print(job["id"], job["total_items"])How it works
- Send your URLs to
POST /v1/batch. You get a job ID back at once. - Check the job until its status is
completed,failedorcancelled. - Read the results as JSON Lines, one page at a time, with
X-Next-Cursor.
curl https://api.spicrawl.com/v1/batch \
-H "Authorization: Bearer $SPICRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "catalog-refresh",
"urls": ["https://example.com/products/1", "https://example.com/products/2"],
"response_format": "markdown"
}'import os, time, requests
API = "https://api.spicrawl.com"
H = {"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"}
job = requests.post(f"{API}/v1/batch", headers=H, json={
"urls": ["https://example.com/products/1", "https://example.com/products/2"],
"response_format": "markdown",
}, timeout=60).json()
while job["status"] not in ("completed", "failed", "cancelled"):
time.sleep(5)
job = requests.get(f"{API}/v1/batch/{job['id']}", headers=H, timeout=30).json()
results = requests.get(f"{API}/v1/batch/{job['id']}/results", headers=H, timeout=120)
print(results.text) # JSON Lines: one result per linespicrawl batch submit urls.txt --format markdown --wait
spicrawl batch results <job-id> --all -o results.jsonlKeep your key in SPICRAWL_API_KEY. Your key needs the batch scope.
What each item can do
- Formats:
html(the default),markdown,textorjson. - Clean-up:
main_content_only,include_tagsandexclude_tagswork on Markdown and text. - JavaScript: set
js_render: truefor pages that need a browser. - Per-URL settings: send
itemsinstead ofurlsto give each URL its own settings and anexternal_id.
Good to know
- Not on batch: extraction, screenshots, browser actions and sessions. A job that asks for them is refused with a 400 before anything is charged. Use single scrapes for those pages.
- Results expire 72 hours after the job finishes.
- No crawling. Spicrawl does not find URLs or follow links. You send the list.
- Pricing: free during the beta. Each successful item uses the same credits as a single scrape.
Frequently asked questions
How many URLs can one batch job have?
Up to 10,000 URLs per call. Send them as a list in urls, or as items when each URL needs its own settings or an external_id you can match on.
What formats does a batch job return?
Each item comes back as html (the default), markdown, text or json. Markdown and text use the same conversion as a single scrape.
How long are batch results kept?
For 72 hours after the job finishes. Read them with GET /v1/batch/{id}/results, which returns JSON Lines one page at a time.
Can a batch job extract fields or click buttons?
No. Extraction, screenshots, browser actions and sessions do not run on batch. A job that asks for them is refused with a 400 before anything is charged. Use single scrapes for those pages.
Does batch crawl a website?
No. Spicrawl does not find URLs or follow links. You send the list of URLs you want.