Convert any website to clean Markdown for LLMs
Send a URL, get the page’s main content back as Markdown: headings, lists, links, tables and code, without the navigation, scripts and footer that waste an LLM’s tokens.
How it works
- Send a URL to the scrape endpoint with
response_format: "markdown". - Spicrawl fetches the page, rendering JavaScript if you ask, and keeps the main content.
- You get Markdown back as the response body, ready for a prompt, a RAG index or a file.
curl https://api.spicrawl.com/v1/scrape \
-H "Authorization: Bearer $SPICRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com/blog/pricing-update", "response_format": "markdown"}'import os, requests
r = requests.post(
"https://api.spicrawl.com/v1/scrape",
headers={"Authorization": f"Bearer {os.environ['SPICRAWL_API_KEY']}"},
json={"url": "https://example.com/blog/pricing-update", "response_format": "markdown"},
timeout=120,
)
r.raise_for_status()
markdown = r.text
print(r.headers["X-Target-Status"]) # the site's own statusconst r = await fetch("https://api.spicrawl.com/v1/scrape", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.SPICRAWL_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({ url: "https://example.com/blog/pricing-update", response_format: "markdown" }),
});
if (!r.ok) throw new Error(JSON.stringify(await r.json()));
const markdown = await r.text();spicrawl scrape https://example.com/blog/pricing-update --format markdownSet the key once with SPICRAWL_API_KEY and never write it into code you commit. The API’s default format is HTML, so always set response_format.
From HTML to Markdown
A typical page wraps a few paragraphs in navigation, scripts and footer markup. Spicrawl keeps the content and its structure:
<header><nav>… 40 links …</nav></header>
<main>
<article>
<h1>Pricing update for 2026</h1>
<p>From 1 March, the Starter plan includes 50,000 requests per month.</p>
<table>
<tr><th>Plan</th><th>Price</th></tr>
<tr><td>Starter</td><td>$29</td></tr>
<tr><td>Scale</td><td>$199</td></tr>
</table>
</article>
</main>
<footer>… cookie banner, scripts, tracking …</footer># Pricing update for 2026
From 1 March, the Starter plan includes 50,000 requests per month.
| Plan | Price |
|---|---|
| Starter | $29 |
| Scale | $199 |Headings, lists, links, tables and code blocks survive the conversion. Layout, styling and interactive parts do not. In our measurements of six real pages, Spicrawl’s Markdown used 53% to 97% fewer tokens than the same pages’ HTML.
Options that cut tokens
| Field | What it does |
|---|---|
response_format | markdown keeps headings, links and tables; text drops all Markdown syntax. |
main_content_only | On by default for Markdown: navigation, footers and asides are removed. Set false to keep the whole page. |
include_tags | CSS selectors to keep. Only matching parts of the page are converted. |
exclude_tags | CSS selectors to remove, such as cookie banners or related-post lists. |
js_render | Load the page in a browser first, for content built by JavaScript. |
links | Also return the page’s links as a separate list, in a JSON response. |
The tag filters run before conversion, so anything you remove never reaches your model:
{
"url": "https://example.com/docs/install",
"response_format": "markdown",
"include_tags": ["article", ".docs-content"],
"exclude_tags": [".cookie-banner", "aside.related"]
}JavaScript pages and PDFs
Start with a plain request (1 credit). If the result is an empty shell or a loading message, send it again with js_render: true (3 credits) so the page runs in a browser before conversion. Set max_cost to the step you intend, so a request never costs more than you expect.
When the URL is a PDF, Spicrawl parses it to plain text at the normal price. Headings and tables are not rebuilt, and scanned or encrypted PDFs return no text.
Convert many pages
A batch job takes up to 10,000 URLs per call, runs them in parallel and can return each one as Markdown. Results are kept for 72 hours.
curl https://api.spicrawl.com/v1/batch \
-H "Authorization: Bearer $SPICRAWL_API_KEY" \
-H "Content-Type: application/json" \
-d '{"urls": ["https://example.com/1", "https://example.com/2"],
"response_format": "markdown"}'Spicrawl converts the URLs you send. It does not crawl a site or discover pages for you, so collect the URLs first, for example from the site’s sitemap.
From an AI agent
Agents can call the same conversion through the Spicrawl MCP server: the spicrawl_scrape tool returns Markdown by default. The spicrawl CLI does the same from a terminal or a script.
A converter API or an HTML-to-Markdown library?
If you already have the HTML, an open-source library such as Turndown converts it in your own code for free. You still have to fetch the page, run its JavaScript, get past pages that refuse simple requests, and strip the navigation yourself. Spicrawl does those steps in one request, which matters most when you convert pages from many different sites.
Good to know
- Check the site’s status. The HTTP status is Spicrawl’s. The page’s own status is in the
X-Target-Statusheader: a site that answers 403 still comes back as HTTP 200, and costs 0 credits. - Check a sample. Main-content isolation works on most pages, but some navigation can survive on unusual layouts. Look at one page of each type before indexing thousands.
- Pricing. Free during the beta, with no credit card. A plain request is 1 credit, browser rendering 3, full Chromium 8. Markdown conversion adds nothing, and failed or blocked requests cost 0.
Frequently asked questions
How do I convert a website to Markdown?
Send a POST request to https://api.spicrawl.com/v1/scrape with the page URL and response_format set to markdown. The response body is the Markdown itself. You can also use the spicrawl CLI (spicrawl scrape URL --format markdown) or the spicrawl_scrape MCP tool from an AI agent.
Does it remove navigation, ads and footers?
Yes. For Markdown, main-content isolation is on by default, so navigation, footers and asides are stripped. You can narrow it further with include_tags and exclude_tags, or set main_content_only to false to keep the whole page. Some navigation can still survive on unusual layouts, so check a sample of each page type.
Can it convert JavaScript-heavy pages?
Yes. Set js_render to true and the page is loaded in a browser before conversion. Start with a plain request, which costs 1 credit, and render only when the plain result is empty or incomplete; rendering costs 3 credits.
Can it convert PDFs to Markdown?
A PDF is parsed to plain text when you ask for markdown or text, at the normal price. Headings and tables are not rebuilt, and scanned or encrypted PDFs return no text.
How much does it cost?
Spicrawl is free during its beta, with no credit card. Requests use credits: 1 for a plain request, 3 for browser rendering and 8 for full Chromium. Converting to Markdown costs nothing extra, and failed or blocked requests cost 0.
Can I convert many pages at once?
Yes. A batch job takes up to 10,000 URLs per call and can return each one as Markdown. Spicrawl does not discover pages for you: send the URLs you want, for example from a sitemap.