Skip to content

Extract structured data from any website

Pick the fields you need. Get them back as clean JSON.

example.com/product/1
Wireless Mouse — Model X
$49.00In stock★ 4.6 (1,284)
Low-latency 2.4 GHz receiver, 70-day battery life and silent switches.
SpecValue
Connection2.4 GHz, BT5Battery70 days
<script type="application/ld+json">
result.md
# Wireless Mouse — Model X**$49.00** · In stock · 4.6 (1,284 reviews)Low-latency 2.4 GHz receiver, 70-day battery lifeand silent switches.| Spec      | Value        || --------- | ------------ || Connection| 2.4 GHz, BT5 || Battery   | 70 days      |
Reading the page
Raw HTML12,480 tokens
Markdown0 tokens

Menus, footers, ads and images removed. About a tenth of the tokens of raw HTML. Token counts are from an example page.

How it works

  1. Name your fields and give each one a CSS selector.
  2. Send them in extract with the page URL.
  3. Read the JSON. Your fields are under data.
curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/products/42",
    "extract": { "title": "h1", "price": ".price", "next_page": "a.next@href" }
  }'
json
{
  "status": 200,
  "data": {
    "title": "Walnut Desk",
    "price": "$349.00"
  },
  "empty_fields": ["next_page"]
}

empty_fields lists fields that matched nothing. It is often the first sign that a site changed its layout.

Three ways to get data

WayWhat you sendWhat you get
autoparse"autoparse": trueThe page’s own JSON-LD, OpenGraph, microdata, metadata and app state
Selector mapField names and CSS selectorsExactly those fields, as text
JSON SchemaA schema whose fields have selectorsTyped, checked values, so "$1,234.56" becomes 1234.56

Good to know

  • No extra credits for extraction. You pay only for the request: 1 credit plain, 3 with JavaScript.
  • Coming soon: extraction from a plain-language prompt. Today, use selectors, a schema or autoparse.
  • Not on batch. Extraction runs on single scrapes only.
  • Your responsibility: only extract data you have the right to use. See our terms.

Frequently asked questions

How do I extract data from a website as JSON?

Send extract with a map of field names to CSS selectors, such as "title": "h1" and "price": ".price". The values come back under data in the JSON response.

What does autoparse do?

It returns what the page already says about itself, with no selectors: JSON-LD, OpenGraph, Twitter Card, microdata, RDFa, head metadata and embedded app state.

Can it return numbers instead of text?

Yes. Use a JSON Schema whose properties have a selector. Values are converted to the declared type, so "$1,234.56" becomes 1234.56, and the result is checked against the schema.

How do I know when a site changes its layout?

Check empty_fields in the response. It lists the fields whose selector matched nothing, which is often the first sign that the markup changed.

Can I describe the fields in plain words instead of selectors?

Not yet. Extraction with a model prompt is coming soon. Today, use selectors, a JSON Schema or autoparse.