Extract structured data from any website
Pick the fields you need. Get them back as clean JSON.
# Wireless Mouse — Model X**$49.00** · In stock · 4.6 (1,284 reviews)Low-latency 2.4 GHz receiver, 70-day battery lifeand silent switches.| Spec | Value || --------- | ------------ || Connection| 2.4 GHz, BT5 || Battery | 70 days |Menus, footers, ads and images removed. About a tenth of the tokens of raw HTML. Token counts are from an example page.
The full reply: status, final URL, engine, warnings, and your data. Token counts are from an example page.
Three ways to get fields: free page info, CSS rules, or a JSON Schema with selectors. Token counts are from an example page.
A list of steps to run on the page. The same steps work on every engine.
How it works
- Name your fields and give each one a CSS selector.
- Send them in
extractwith the page URL. - Read the JSON. Your fields are under
data.
curl https://api.spicrawl.com/v1/scrape \
-H "Authorization: Bearer $SPICRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com/products/42",
"extract": { "title": "h1", "price": ".price", "next_page": "a.next@href" }
}'import os, requests
r = requests.post(
"https://api.spicrawl.com/v1/scrape",
headers={"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
json={
"url": "https://example.com/products/42",
"extract": {"title": "h1", "price": ".price", "next_page": "a.next@href"},
},
timeout=120,
)
env = r.json()
print(env["data"], env.get("empty_fields"))spicrawl scrape https://example.com/products/42 \
--extract '{"title":"h1","price":".price","next_page":"a.next@href"}'{
"status": 200,
"data": {
"title": "Walnut Desk",
"price": "$349.00"
},
"empty_fields": ["next_page"]
}empty_fields lists fields that matched nothing. It is often the first sign that a site changed its layout.
Three ways to get data
| Way | What you send | What you get |
|---|---|---|
autoparse | "autoparse": true | The page’s own JSON-LD, OpenGraph, microdata, metadata and app state |
| Selector map | Field names and CSS selectors | Exactly those fields, as text |
| JSON Schema | A schema whose fields have selectors | Typed, checked values, so "$1,234.56" becomes 1234.56 |
Good to know
- No extra credits for extraction. You pay only for the request: 1 credit plain, 3 with JavaScript.
- Coming soon: extraction from a plain-language prompt. Today, use selectors, a schema or autoparse.
- Not on batch. Extraction runs on single scrapes only.
- Your responsibility: only extract data you have the right to use. See our terms.
Frequently asked questions
How do I extract data from a website as JSON?
Send extract with a map of field names to CSS selectors, such as "title": "h1" and "price": ".price". The values come back under data in the JSON response.
What does autoparse do?
It returns what the page already says about itself, with no selectors: JSON-LD, OpenGraph, Twitter Card, microdata, RDFa, head metadata and embedded app state.
Can it return numbers instead of text?
Yes. Use a JSON Schema whose properties have a selector. Values are converted to the declared type, so "$1,234.56" becomes 1234.56, and the result is checked against the schema.
How do I know when a site changes its layout?
Check empty_fields in the response. It lists the fields whose selector matched nothing, which is often the first sign that the markup changed.
Can I describe the fields in plain words instead of selectors?
Not yet. Extraction with a model prompt is coming soon. Today, use selectors, a JSON Schema or autoparse.