Skip to content

Scrape JavaScript websites

Send a URL that needs JavaScript. Get the page back as clean Markdown or JSON.

Spicrawl control panel
Auto escalate
Browser
Fingerprint
Proxy
Managed proxies: coming soon
Session
Actions
Block assets
Network capture
Cache
POST /v1/scrape
{  "url": "https://shop.example/product/1",  "js_render": true,  "cache": true,  "response_format": "markdown"}
Example response headers
X-Target-Status
200
X-Engine
obscura
X-Proxy-Source
direct
Cache-State
hit
X-Credits-Charged
0
Browser

Render JavaScript in a real browser, or run it headful with a real display for sites that check for headless. A pinned engine never escalates, so headful turns auto mode off.

Interactive example, not a live request. Credits are shown for reference; the beta is free.

A rendered page, in the format you need

See how page content becomes Markdown, JSON or HTML. These examples show the output, not a live request.

example.com/product/1
Wireless Mouse — Model X
$49.00In stock★ 4.6 (1,284)
Low-latency 2.4 GHz receiver, 70-day battery life and silent switches.
SpecValue
Connection2.4 GHz, BT5Battery70 days
<script type="application/ld+json">
result.md
# Wireless Mouse — Model X**$49.00** · In stock · 4.6 (1,284 reviews)Low-latency 2.4 GHz receiver, 70-day battery lifeand silent switches.| Spec      | Value        || --------- | ------------ || Connection| 2.4 GHz, BT5 || Battery   | 70 days      |
Reading the page
Raw HTML12,480 tokens
Markdown0 tokens

Menus, footers, ads and images removed. About a tenth of the tokens of raw HTML. Token counts are from an example page.

How it works

  1. Send the page URL to POST /v1/scrape with js_render: true.
  2. Wait for the content. Set wait_for to a CSS selector, such as .price.
  3. Get the rendered page in the format you choose. The examples below return Markdown.
curl -sS -D - https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/products/42", "js_render": true, "wait_for": ".price", "response_format": "markdown"}'

Keep your key in SPICRAWL_API_KEY. Never put it in browser code or a file you commit.

When to render JavaScript

A plain request downloads HTML without running scripts. If the result is an empty shell or a “Loading…” message, the page may need JavaScript. Set js_render: true to run it in a browser before reading it.

Use a plain request for content already in the HTML. Rendering is also needed for browser actions, resource blocking and network capture.

Wait for the right content

OptionWhat it does
wait_forWait for a CSS selector to appear before capture.
wait_for_timeoutSet the wait limit from 0 to 120,000 ms. Requires wait_for; 0 uses the engine default.
waitAdd a fixed delay from 0 to 30,000 ms after load. Prefer a selector when you can.
block_resourcesChoose which assets to skip. Images and fonts are blocked by default, with exceptions for actions and captures.
response_formatChoose Markdown, JSON or HTML for the page content.

An explicit resource list replaces the default. Use ["none"] to block nothing. Do not block scripts when the page needs them to build its content.

Check the result

http
HTTP/2 200
content-type: text/markdown; charset=utf-8
x-target-status: 200
x-engine: obscura
x-credits-charged: 3

X-Engine tells you which engine ran. X-Target-Status is the site's status, while X-Credits-Charged shows the charge.

If wait_for never matches, the page may come back as it stood with X-Warning: RENDER_DEGRADED. Check the content and warning before using it.

Choose an engine and check the cost

  • Free during beta. Credits show the cost of each request.
  • Browser rendering: js_render: true normally uses Obscura, at 3 credits per successful request.
  • Chromium: costs 8 credits. If Obscura is not deployed, an unpinned render may use Chromium and return X-Warning: ENGINE_SUBSTITUTED.
  • Pin an engine with engine: "obscura" to avoid substitution. Your plan must allow it; a pinned engine never escalates.
  • Failures cost 0. Waiting and resource blocking add no extra charge. Set max_cost to cap a request.

Good to know

  • Choose rendering or auto mode. mode: "auto" cannot be combined with js_render or engine. It starts with fetch and may move to Obscura; it never reaches Chromium.
  • Browser flags need a browser. Sending wait_for, actions or resource blocking with the fetch engine returns an error.
  • No crawling. Send the URLs you want; Spicrawl does not find pages or follow links.
  • Coming soon: managed proxies, stealth mode, a remote browser and extraction from a plain-language prompt.
  • Your responsibility: only fetch pages you have the right to use. Read our terms.

Frequently asked questions

How do I scrape a JavaScript website?

Send the page URL to POST https://api.spicrawl.com/v1/scrape with js_render set to true. The browser runs the scripts before Spicrawl reads the page. Choose response_format for Markdown, JSON or HTML.

How do I wait for content to load?

Set wait_for to a CSS selector, such as .price. Use wait_for_timeout to set a timeout up to 120000 milliseconds. If the selector never matches, the page may come back as it stood with X-Warning: RENDER_DEGRADED. Check the warning before using the result.

How much does JavaScript rendering cost?

Spicrawl is free during beta. Successful Obscura renders use 3 credits; Chromium uses 8. If js_render is true without a pinned engine, Chromium may serve the request when Obscura is not deployed. Failed requests use 0 credits. Read X-Engine and X-Credits-Charged to check what ran.

Can I use auto mode with JavaScript rendering?

No. mode: auto cannot be combined with js_render or engine. Auto mode starts with a plain request and moves to Obscura when the result is unusable. For pages you know need JavaScript, use js_render: true instead.

Can it crawl a JavaScript website?

No. Spicrawl does not find URLs or follow links. Send the page URLs you want to read. Only fetch pages you have the right to use.