Scrape websites behind a login
Log in once and stay logged in. Read your account pages with one API.
{ "url": "https://shop.example/product/1", "js_render": true, "session_id": "01J9ZQ4M7R3T8VX2K5N6P0B1CD", "cache": true, "response_format": "markdown"}- X-Target-Status
- 200
- X-Engine
- obscura
- X-Proxy-Source
- direct
- Cache-State
- bypass
- X-Credits-Charged
- 3
Start with a plain request and move up to a real browser only if the site says no. You are charged for the engine that worked, and a blocked or refused request costs nothing.
Render JavaScript in a real browser, or run it headful with a real display for sites that check for headless. A pinned engine never escalates, so headful turns auto mode off.
Plain requests carry a real browser’s TLS and HTTP/2 fingerprint, which gets past checks that never run JavaScript.
Use a direct connection or your own proxy with no proxy surcharge. Managed proxies and built-in country targeting are coming soon, outside the current beta.
Keep cookies and storage between requests, so a login carries over. Session requests are never served from the cache.
Click, type, select, scroll and run JavaScript before the page is read. The same steps run on every browser engine.
Skip images, fonts and media while rendering. Pages load faster and you move fewer bytes.
Record every request the page itself made while it loaded: the API calls behind the page, not just its HTML.
On by default for up to 48 hours. A repeat of the same request is served from the cache at no charge.
Keep the login between requests
A session shares one cookie jar across the pages you read. This example shows the flow, not a live request.
const s = await fetch("https://api.spicrawl.com/v1/sessions", { method: "POST", headers: { Authorization: "Bearer spicrawl_live_…" },}).then((r) => r.json());// Log in with browser actions on this session first.await scrape({ url: dashboardUrl, session_id: s.id, js_render: true });How it works
- Create a session. Keep the ID from the response.
- Log in with browser actions. Send the session ID with the login request.
- Read account pages. Reuse the same ID, one request at a time.
- Release the session when you are done.
Your API key needs the sessions scope to create and manage sessions, and scrape to use one. Keep it in SPICRAWL_API_KEY.
1. Create a session
curl https://api.spicrawl.com/v1/sessions \
-H "Authorization: Bearer $SPICRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{"engine": "chromium", "ttl_seconds": 7200}'import os, requests
API = "https://api.spicrawl.com"
headers = {"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"}
s = requests.post(f"{API}/v1/sessions", headers=headers,
json={"engine": "chromium", "ttl_seconds": 7200}, timeout=30)
s.raise_for_status()
session_id = s.json()["id"]const API = "https://api.spicrawl.com";
const headers = {
Authorization: `Bearer ${process.env.SPICRAWL_API_KEY}`,
"Content-Type": "application/json",
};
const s = await fetch(`${API}/v1/sessions`, {
method: "POST", headers,
body: JSON.stringify({ engine: "chromium", ttl_seconds: 7200 }),
});
if (!s.ok) throw new Error(await s.text());
const { id: sessionId } = await s.json();spicrawl sessions create --engine chromium --ttl 7200The response includes id, engine, expires_at and hard_expires_at. It never includes cookies. The example below uses a sample ID; replace it with yours.
2. Log in with actions
Send this body to POST /v1/scrape with your bearer token. Change the URL and selectors to match the site. Use credentials for an account you own or may use.
{
"url": "https://example.com/login",
"session_id": "01J9ZQ4M7R3T8VX2K5N6P0B1CD",
"js_render": true,
"actions": [
{
"fill": {
"selector": "input[name=email]",
"value": "YOUR_EMAIL"
}
},
{
"fill": {
"selector": "input[name=password]",
"value": "YOUR_PASSWORD",
"secret": true
}
},
{
"click": "button[type=submit]"
},
{
"wait_for": ".account-menu"
}
]
}Replace the email and password placeholders securely at runtime. The password action uses secret: true. Keep js_render: true: a session ID alone does not allow browser-only fields.
3. Read pages with the same login
After the login succeeds, pass session_id on each scrape. Cookies and storage carry over. Python and TypeScript reuse the variables from the create step.
curl https://api.spicrawl.com/v1/scrape \
-H "Authorization: Bearer $SPICRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com/account/orders", "session_id": "01J9ZQ4M7R3T8VX2K5N6P0B1CD", "js_render": true, "response_format": "markdown"}'r = requests.post(f"{API}/v1/scrape", headers=headers, json={
"url": "https://example.com/account/orders", "session_id": session_id,
"js_render": True, "response_format": "markdown",
}, timeout=120)
r.raise_for_status()
print(r.text)const r = await fetch(`${API}/v1/scrape`, {
method: "POST", headers,
body: JSON.stringify({
url: "https://example.com/account/orders", session_id: sessionId,
js_render: true, response_format: "markdown",
}),
});
if (!r.ok) throw new Error(await r.text());
console.log(await r.text());spicrawl scrape https://example.com/account/orders --session 01J9ZQ4M7R3T8VX2K5N6P0B1CD --render --format markdownSession requests always bypass the result cache. Check Cache-State: bypass; the page is fetched fresh.
4. Release the session
curl -X POST https://api.spicrawl.com/v1/sessions/01J9ZQ4M7R3T8VX2K5N6P0B1CD/release \
-H "Authorization: Bearer $SPICRAWL_API_KEY"Release deletes cookies and storage. It does not keep the login. Later requests on that ID return 410 ERR::SESSION::RELEASED.
Session options
| Option | What to know |
|---|---|
engine | Fixed for the session's life. A scrape cannot choose a different engine. Check the create response if you leave this out. |
ttl_seconds | Default 1,800 seconds; minimum 30. A successful scrape extends the lifetime, never past hard_expires_at. |
session_context | Restore cookies and storage saved from an active session. The saved context is limited to 1 MiB. |
To save a login before release, read GET /v1/sessions/{id}/context. Store that state safely, then pass it as session_context when creating a new session.
What it costs
Free during beta. Creating, reading and releasing sessions costs nothing. Each scrape uses the session engine's price: fetch 1 credit, Obscura 3, Chromium 8. Failures, including a busy or expired session, cost 0.
Good to know
- One request per session at a time. A second render returns
409 ERR::SESSION::BUSYwithRetry-After. Use one session per worker for parallel work. - Expired sessions need a new login. An unknown, deleted or malformed ID returns an error; the request never runs without its session.
- Your own proxy: use a fixed exit for sites that challenge a login when its IP changes. Managed session exits are coming soon.
- No crawling. You supply the URLs. Managed proxies, stealth, a remote browser and extraction from a prompt are coming soon.
- Your responsibility: only use accounts and pages you have the right to access. Read our terms.
Frequently asked questions
How do I scrape a website behind a login?
Create a session, log in with browser actions, then send its ID as session_id on each scrape. Use js_render: true for actions and other browser-only fields. Only use accounts and pages you have the right to access.
What does a session keep?
A session keeps cookies and browser storage across requests. The engine is fixed for its life. The normal session response never contains cookies; the context endpoint returns the saved state while the session is active.
Can I send many requests on one session at once?
No. Send one request at a time per session. A second request while a render runs returns 409 ERR::SESSION::BUSY with Retry-After. For work in parallel, create one session per worker.
What happens when I release a session?
Release deletes its cookies and storage. It does not keep the login for later. A scrape using the released ID returns 410 ERR::SESSION::RELEASED. Save the context before release if you need to restore it in a new session.
How much do sessions cost?
Creating, reading and releasing sessions costs nothing. Each scrape uses the session engine price: 1 credit for fetch, 3 for Obscura or 8 for Chromium. Failed requests cost 0. Spicrawl is free during beta.