Skip to content

Scrape websites behind a login

Log in once and stay logged in. Read your account pages with one API.

Spicrawl control panel
Auto escalate
Browser
Fingerprint
Proxy
Managed proxies: coming soon
Session
Actions
Block assets
Network capture
Cache
POST /v1/scrape
{  "url": "https://shop.example/product/1",  "js_render": true,  "session_id": "01J9ZQ4M7R3T8VX2K5N6P0B1CD",  "cache": true,  "response_format": "markdown"}
Example response headers
X-Target-Status
200
X-Engine
obscura
X-Proxy-Source
direct
Cache-State
bypass
X-Credits-Charged
3
Sessions

Keep cookies and storage between requests, so a login carries over. Session requests are never served from the cache.

Interactive example, not a live request. Credits are shown for reference; the beta is free.

Keep the login between requests

A session shares one cookie jar across the pages you read. This example shows the flow, not a live request.

Create a session
session.js
const s = await fetch("https://api.spicrawl.com/v1/sessions", {  method: "POST",  headers: { Authorization: "Bearer spicrawl_live_…" },}).then((r) => r.json());// Log in with browser actions on this session first.await scrape({ url: dashboardUrl, session_id: s.id, js_render: true });

How it works

  1. Create a session. Keep the ID from the response.
  2. Log in with browser actions. Send the session ID with the login request.
  3. Read account pages. Reuse the same ID, one request at a time.
  4. Release the session when you are done.

Your API key needs the sessions scope to create and manage sessions, and scrape to use one. Keep it in SPICRAWL_API_KEY.

1. Create a session

curl https://api.spicrawl.com/v1/sessions \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"engine": "chromium", "ttl_seconds": 7200}'

The response includes id, engine, expires_at and hard_expires_at. It never includes cookies. The example below uses a sample ID; replace it with yours.

2. Log in with actions

Send this body to POST /v1/scrape with your bearer token. Change the URL and selectors to match the site. Use credentials for an account you own or may use.

json
{
  "url": "https://example.com/login",
  "session_id": "01J9ZQ4M7R3T8VX2K5N6P0B1CD",
  "js_render": true,
  "actions": [
    {
      "fill": {
        "selector": "input[name=email]",
        "value": "YOUR_EMAIL"
      }
    },
    {
      "fill": {
        "selector": "input[name=password]",
        "value": "YOUR_PASSWORD",
        "secret": true
      }
    },
    {
      "click": "button[type=submit]"
    },
    {
      "wait_for": ".account-menu"
    }
  ]
}

Replace the email and password placeholders securely at runtime. The password action uses secret: true. Keep js_render: true: a session ID alone does not allow browser-only fields.

3. Read pages with the same login

After the login succeeds, pass session_id on each scrape. Cookies and storage carry over. Python and TypeScript reuse the variables from the create step.

curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/account/orders", "session_id": "01J9ZQ4M7R3T8VX2K5N6P0B1CD", "js_render": true, "response_format": "markdown"}'

Session requests always bypass the result cache. Check Cache-State: bypass; the page is fetched fresh.

4. Release the session

bash
curl -X POST https://api.spicrawl.com/v1/sessions/01J9ZQ4M7R3T8VX2K5N6P0B1CD/release \
  -H "Authorization: Bearer $SPICRAWL_API_KEY"

Release deletes cookies and storage. It does not keep the login. Later requests on that ID return 410 ERR::SESSION::RELEASED.

Session options

OptionWhat to know
engineFixed for the session's life. A scrape cannot choose a different engine. Check the create response if you leave this out.
ttl_secondsDefault 1,800 seconds; minimum 30. A successful scrape extends the lifetime, never past hard_expires_at.
session_contextRestore cookies and storage saved from an active session. The saved context is limited to 1 MiB.

To save a login before release, read GET /v1/sessions/{id}/context. Store that state safely, then pass it as session_context when creating a new session.

What it costs

Free during beta. Creating, reading and releasing sessions costs nothing. Each scrape uses the session engine's price: fetch 1 credit, Obscura 3, Chromium 8. Failures, including a busy or expired session, cost 0.

Good to know

  • One request per session at a time. A second render returns 409 ERR::SESSION::BUSY with Retry-After. Use one session per worker for parallel work.
  • Expired sessions need a new login. An unknown, deleted or malformed ID returns an error; the request never runs without its session.
  • Your own proxy: use a fixed exit for sites that challenge a login when its IP changes. Managed session exits are coming soon.
  • No crawling. You supply the URLs. Managed proxies, stealth, a remote browser and extraction from a prompt are coming soon.
  • Your responsibility: only use accounts and pages you have the right to access. Read our terms.

Frequently asked questions

How do I scrape a website behind a login?

Create a session, log in with browser actions, then send its ID as session_id on each scrape. Use js_render: true for actions and other browser-only fields. Only use accounts and pages you have the right to access.

What does a session keep?

A session keeps cookies and browser storage across requests. The engine is fixed for its life. The normal session response never contains cookies; the context endpoint returns the saved state while the session is active.

Can I send many requests on one session at once?

No. Send one request at a time per session. A second request while a render runs returns 409 ERR::SESSION::BUSY with Retry-After. For work in parallel, create one session per worker.

What happens when I release a session?

Release deletes its cookies and storage. It does not keep the login for later. A scrape using the released ID returns 410 ERR::SESSION::RELEASED. Save the context before release if you need to restore it in a new session.

How much do sessions cost?

Creating, reading and releasing sessions costs nothing. Each scrape uses the session engine price: 1 credit for fetch, 3 for Obscura or 8 for Chromium. Failed requests cost 0. Spicrawl is free during beta.