# Convert a web page to clean Markdown

Paste a public URL to get its main content as Markdown for an LLM or a RAG index, with the route used, the exit country and the token count. The same call works from the API and over MCP.

Updated 1 Oct 2026 · Tested on 1 Oct 2026 · https://skryp.dev/tools/url-to-markdown

```python
# Convert one page to Markdown with the Skryp API and print the result the tool shows.
# pip install httpx    export SKRYP_API_KEY=...
import json
import os

import httpx

API = os.environ.get("SKRYP_API_URL", "https://api.skryp.dev")
page = httpx.post(
    f"{API}/v1/scrape",
    headers={"Authorization": f"Bearer {os.environ['SKRYP_API_KEY']}"},
    json={"url": "https://en.wikipedia.org/wiki/Markdown", "formats": ["markdown"]},
    timeout=120,
).json()

r = page["receipt"]
print(json.dumps({
    "url": page["url"], "title": page.get("title"), "tokens": page["tokens"],
    "route": r["route"], "exit": r["egress"]["country"], "verified": r["egress"]["verified"], "credits": r["credits"],
    "markdown": page["markdown"][:2400],
}, ensure_ascii=False, indent=1))
```

Output (ran 1 Oct 2026; Python 3.12.14, httpx 0.28.1; 1.3 s):

```text
{
 "url": "https://en.wikipedia.org/wiki/Markdown",
 "title": "Markdown - Wikipedia",
 "tokens": 4291,
 "route": "http:direct",
 "exit": "za",
 "verified": true,
 "credits": 1,
 "markdown": "From Wikipedia, the free encyclopedia\n\nFor the marketing term, see [Price markdown](https://en.wikipedia.org/wiki/Price_markdown \"Price markdown\").\n\n- [icon](https://en.wikipedia.org/wiki/File:Question_book-new.svg) This article **relies excessively on [references](https://en.wikipedia.org/wiki/Wikipedia:Verifiability \"Wikipedia:Verifiability\") to [primary sources](https://en.wikipedia.org/wiki/Wikipedia:No_original_research \"Wikipedia:No original research\")**. Please improve this article by adding [secondary or tertiary sources](https://en.wikipedia.org/wiki/Wikipedia:No_original_research \"Wikipedia:No original research\"). *Find sources:* [\"Markdown\"](https://www.google.com/search?as_eq=wikipedia&q=%22Markdown%22) – [news](https://www.google.com/search?tbm=nws&q=%22Markdown%22+-wikipedia&tbs=ar:1) **·** [newspapers](https://www.google.com/search?q=%22Markdown%22&tbs=bkt:s&tbm=bks) **·** [books](https://www.google.com/search?tbs=bks:1&q=%22Markdown%22+-wikipedia) **·** [scholar](https://scholar.google.com/scholar?q=%22Markdown%22) **·** [JSTOR](https://www.jstor.org/action/doBasicSearch?Query=%22Markdown%22&acc=on&wc=on) *(September 2025)* *([Learn how and when to remove this message](https://en.wikipedia.org/wiki/Help:Maintenance_template_removal \"Help:Maintenance template removal\"))*\n\n| Markdown                                                                                                                                                                                                 | |\n| -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |\n|  | |\n| [Filename extensions](https://en.wikipedia.org/wiki/Filename_extension \"Filename extension\")                                                                                                             | `.md`, `.markdown`[1][2]                                                                                                                                                                                 |\n| [Internet media type](https://en.wikipedia.org/wiki/Media_ty"
}
```

Signed in, the tool runs your page live on your workspace (1 credit on Skryp's own network, more if the page needs a residential IP; failed pages cost nothing [checked 1 Oct 2026]). Without an account it shows the recorded example and takes you to sign-up with your address kept.

## What is kept and what is removed

Skryp keeps the page's main content and drops what a reader skips:

- **Kept:** headings, paragraphs, lists, tables, code blocks, and links with their addresses.
- **Removed:** navigation, footers, cookie banners and scripts; image embeds (a linked image keeps its alt text as the link label); tracking parameters such as `utm_source` in links.
- **Measured:** on 170 public pages the median page came back as 1,800 tokens of Markdown, against 65,378 tokens of HTML as served [checked 1 Oct 2026]. See [How many tokens is a web page?](https://skryp.dev/benchmarks/tokens-per-web-page) for every page type.

Very long pages stay on Skryp's side behind a reference that your agent reads in sections, so its context holds only what it asked for.

## The same call from your code or your agent

This is the request the tool makes, as a script. It produced the recorded example above:

```python
# Convert one page to Markdown with the Skryp API and print the result the tool shows.
# pip install httpx    export SKRYP_API_KEY=...
import json
import os

import httpx

API = os.environ.get("SKRYP_API_URL", "https://api.skryp.dev")
page = httpx.post(
    f"{API}/v1/scrape",
    headers={"Authorization": f"Bearer {os.environ['SKRYP_API_KEY']}"},
    json={"url": "https://en.wikipedia.org/wiki/Markdown", "formats": ["markdown"]},
    timeout=120,
).json()

r = page["receipt"]
print(json.dumps({
    "url": page["url"], "title": page.get("title"), "tokens": page["tokens"],
    "route": r["route"], "exit": r["egress"]["country"], "verified": r["egress"]["verified"], "credits": r["credits"],
    "markdown": page["markdown"][:2400],
}, ensure_ascii=False, indent=1))
```

From an AI agent connected over MCP, the same thing is one tool call: `scrape` with the page's `url`. To connect Claude Code, Codex or Cursor, see the [Quickstart](https://skryp.dev/docs/quickstart). The [Scrape API](https://skryp.dev/product/scrape) page lists the other formats: clean HTML, raw HTML, links, the page's own data and screenshots.

## Limits

- **Pages behind a sign-in** need a saved login: your own account, connected once in the dashboard. See [Live browser and saved logins](https://skryp.dev/product/browser).
- **PDFs** are converted to text; scanned pages inside them are images and are not read.
- **Documentation sites:** in our benchmark Skryp returned 20 of 23 documentation pages usable, where Firecrawl returned all 23 [checked 1 Oct 2026]. Check a few pages of your docs before indexing a whole site.
- **Some sites block automated readers** on every route at times. The receipt says so and the page costs nothing.

## Questions

**How do you convert a web page to Markdown?**

Paste its address into a converter such as the one above, or call a scraping API that returns Markdown. Converter libraries (markdownify in Python, Turndown in JavaScript) also work on HTML you already have, but they keep navigation and other clutter unless you clean the page first.

**How do you convert a website to Markdown for an LLM?**

List its pages with a sitemap or a crawl, convert each page's main content to Markdown, and keep the source address with each document so answers can cite it. Main-content Markdown is a fraction of the tokens of raw HTML, which lowers the cost of every prompt.


## Evidence

- On the same 170 public pages, the median output was 1,800 tokens per page from Skryp and 3,520 from Firecrawl. (observed test, checked 1 Oct 2026; Same benchmark run; Markdown output. Limits: Fewer tokens is not automatically better: completeness was slightly lower for Skryp (median 0.98 vs 1.00). PDFs were the exception: on the five PDFs both services read, they returned about the same length (Skryp shorter on three, longer on two; median ratio 0.95).)
- Blocked and failed pages cost 0 credits. (observed test, checked 1 Oct 2026; All page reads through the API, MCP and dashboard. Limits: Applies to page reads; a model extraction that runs is charged when it runs.)
- On the same 170 public pages, Firecrawl returned 159 usable pages and Skryp 157; median completeness was 1.00 for Firecrawl and 0.98 for Skryp; on documentation pages Firecrawl returned 23 of 23 usable and Skryp 20 of 23. (observed test, checked 1 Oct 2026; 170 public pages across articles, docs, shops, listings, JavaScript apps, PDFs, protected and regional pages. Limits: One run. Firecrawl at default settings. Skryp's own benchmark; publish the URL list and scoring code with it.)
