Skip to content

Convert any website to clean Markdown for LLMs

Send a URL, get the page’s main content back as Markdown: headings, lists, links, tables and code, without the navigation, scripts and footer that waste an LLM’s tokens.

https://
Format: Markdown Main content only

How it works

  1. Send a URL to the scrape endpoint with response_format: "markdown".
  2. Spicrawl fetches the page, rendering JavaScript if you ask, and keeps the main content.
  3. You get Markdown back as the response body, ready for a prompt, a RAG index or a file.
curl https://api.spicrawl.com/v1/scrape \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/blog/pricing-update", "response_format": "markdown"}'

Set the key once with SPICRAWL_API_KEY and never write it into code you commit. The API’s default format is HTML, so always set response_format.

From HTML to Markdown

A typical page wraps a few paragraphs in navigation, scripts and footer markup. Spicrawl keeps the content and its structure:

html
<header><nav>… 40 links …</nav></header>
<main>
  <article>
    <h1>Pricing update for 2026</h1>
    <p>From 1 March, the Starter plan includes 50,000 requests per month.</p>
    <table>
      <tr><th>Plan</th><th>Price</th></tr>
      <tr><td>Starter</td><td>$29</td></tr>
      <tr><td>Scale</td><td>$199</td></tr>
    </table>
  </article>
</main>
<footer>… cookie banner, scripts, tracking …</footer>
markdown
# Pricing update for 2026

From 1 March, the Starter plan includes 50,000 requests per month.

| Plan | Price |
|---|---|
| Starter | $29 |
| Scale | $199 |

Headings, lists, links, tables and code blocks survive the conversion. Layout, styling and interactive parts do not. In our measurements of six real pages, Spicrawl’s Markdown used 53% to 97% fewer tokens than the same pages’ HTML.

Options that cut tokens

FieldWhat it does
response_formatmarkdown keeps headings, links and tables; text drops all Markdown syntax.
main_content_onlyOn by default for Markdown: navigation, footers and asides are removed. Set false to keep the whole page.
include_tagsCSS selectors to keep. Only matching parts of the page are converted.
exclude_tagsCSS selectors to remove, such as cookie banners or related-post lists.
js_renderLoad the page in a browser first, for content built by JavaScript.
linksAlso return the page’s links as a separate list, in a JSON response.

The tag filters run before conversion, so anything you remove never reaches your model:

json
{
  "url": "https://example.com/docs/install",
  "response_format": "markdown",
  "include_tags": ["article", ".docs-content"],
  "exclude_tags": [".cookie-banner", "aside.related"]
}

JavaScript pages and PDFs

Start with a plain request (1 credit). If the result is an empty shell or a loading message, send it again with js_render: true (3 credits) so the page runs in a browser before conversion. Set max_cost to the step you intend, so a request never costs more than you expect.

When the URL is a PDF, Spicrawl parses it to plain text at the normal price. Headings and tables are not rebuilt, and scanned or encrypted PDFs return no text.

Convert many pages

A batch job takes up to 10,000 URLs per call, runs them in parallel and can return each one as Markdown. Results are kept for 72 hours.

bash
curl https://api.spicrawl.com/v1/batch \
  -H "Authorization: Bearer $SPICRAWL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"urls": ["https://example.com/1", "https://example.com/2"],
       "response_format": "markdown"}'

Spicrawl converts the URLs you send. It does not crawl a site or discover pages for you, so collect the URLs first, for example from the site’s sitemap.

From an AI agent

Agents can call the same conversion through the Spicrawl MCP server: the spicrawl_scrape tool returns Markdown by default. The spicrawl CLI does the same from a terminal or a script.

A converter API or an HTML-to-Markdown library?

If you already have the HTML, an open-source library such as Turndown converts it in your own code for free. You still have to fetch the page, run its JavaScript, get past pages that refuse simple requests, and strip the navigation yourself. Spicrawl does those steps in one request, which matters most when you convert pages from many different sites.

Good to know

  • Check the site’s status. The HTTP status is Spicrawl’s. The page’s own status is in the X-Target-Status header: a site that answers 403 still comes back as HTTP 200, and costs 0 credits.
  • Check a sample. Main-content isolation works on most pages, but some navigation can survive on unusual layouts. Look at one page of each type before indexing thousands.
  • Pricing. Free during the beta, with no credit card. A plain request is 1 credit, browser rendering 3, full Chromium 8. Markdown conversion adds nothing, and failed or blocked requests cost 0.

Frequently asked questions

How do I convert a website to Markdown?

Send a POST request to https://api.spicrawl.com/v1/scrape with the page URL and response_format set to markdown. The response body is the Markdown itself. You can also use the spicrawl CLI (spicrawl scrape URL --format markdown) or the spicrawl_scrape MCP tool from an AI agent.

Does it remove navigation, ads and footers?

Yes. For Markdown, main-content isolation is on by default, so navigation, footers and asides are stripped. You can narrow it further with include_tags and exclude_tags, or set main_content_only to false to keep the whole page. Some navigation can still survive on unusual layouts, so check a sample of each page type.

Can it convert JavaScript-heavy pages?

Yes. Set js_render to true and the page is loaded in a browser before conversion. Start with a plain request, which costs 1 credit, and render only when the plain result is empty or incomplete; rendering costs 3 credits.

Can it convert PDFs to Markdown?

A PDF is parsed to plain text when you ask for markdown or text, at the normal price. Headings and tables are not rebuilt, and scanned or encrypted PDFs return no text.

How much does it cost?

Spicrawl is free during its beta, with no credit card. Requests use credits: 1 for a plain request, 3 for browser rendering and 8 for full Chromium. Converting to Markdown costs nothing extra, and failed or blocked requests cost 0.

Can I convert many pages at once?

Yes. A batch job takes up to 10,000 URLs per call and can return each one as Markdown. Spicrawl does not discover pages for you: send the URLs you want, for example from a sitemap.