# How many tokens is a web page?

A typical web page is far bigger as HTML than as text. Across 170 public pages, the median page was 65,378 tokens of HTML as served, 3,520 tokens as Markdown from Firecrawl and 1,800 from Skryp; the main text alone was about 400. Shop and listing pages shrink most; JavaScript apps and PDFs are the exceptions.

Updated 1 Oct 2026 · Tested on 1 Oct 2026 · https://skryp.dev/benchmarks/tokens-per-web-page

## The answer in numbers

How many tokens a page costs depends on what your agent reads. On the median page of the 170 we measured:

- **HTML as served:** 65,378 tokens [checked 1 Oct 2026], what a plain fetch hands an agent before any cleaning.
- **Markdown from Firecrawl:** 3,520 tokens [checked 1 Oct 2026].
- **Markdown from Skryp:** 1,800 tokens.
- **Main text only:** about 400 tokens, from a neutral cleaned copy that exists for 105 of the pages.

At an example price of $3 per million input tokens, reading 1,000 such pages costs about $196 as HTML, $10.56 as Firecrawl's Markdown and $5.40 as Skryp's. The price is an illustration; the ratios are what carries over to your model.

## By page type

| Page type | Pages | HTML as served | Skryp Markdown | Firecrawl Markdown | Main text only |
| --- | ---: | ---: | ---: | ---: | ---: |
| News articles | 24 | 85,092 (20) | 2,471 | 3,682 | 734 |
| Documentation | 23 | 135,013 | 1,065 | 2,261 | 394 |
| Shop pages (global) | 13 | 347,746 (7) | 1,072 | 2,724 | 233 |
| Shop pages (South Africa) | 22 | 78,493 | 1,760 | 3,713 | 178 |
| Country-priced pages | 32 | 50,006 (24) | 1,292 | 2,790 | — |
| Listings and search results | 14 | 49,920 (6) | 1,606 | 5,397 | 467 |
| JavaScript apps | 5 | 15,491 (3) | 6,549 | 15,213 | 360 |
| Bot-protected pages | 15 | 196,442 (6) | 2,568 | 13,489 | 226 |
| PDFs | 7 | — | 17,771 | 15,660 | — |
| Our own websites | 15 | 12,175 | 1,372 | 1,464 | 577 |
| All pages | 170 | 65,378 (126) | 1,800 | 3,520 | 408 (105) |

Median tokens per page (cl100k). Markdown medians cover each service's usable pages. A number in brackets is how many pages the median covers when fewer than all: blocked pages, PDFs and fetch failures have no HTML count, and the main-text reference exists only where a full browser render could be cleaned.

Global shop pages carry the heaviest HTML, a median of 347,746 tokens on the 7 that returned it [checked 1 Oct 2026]. Documentation is heavy as HTML (135,013) yet among the shortest as Markdown (1,065 from Skryp). JavaScript apps are the reverse case: the HTML that arrives is a small shell, and the content appears only after scripts run, so a reader that does not render the page gets almost nothing of what a visitor sees.

## How much smaller, page by page

| Page type | HTML ÷ Skryp Markdown | Pages | Skryp ÷ Firecrawl Markdown | Pages |
| --- | ---: | ---: | ---: | ---: |
| News articles | 30.3× | 20 | 0.58× | 24 |
| Documentation | 68.3× | 20 | 0.62× | 20 |
| Shop pages (global) | 123.3× | 7 | 0.43× | 13 |
| Shop pages (South Africa) | 73.2× | 21 | 0.46× | 21 |
| Country-priced pages | 20.1× | 24 | 0.69× | 32 |
| Listings and search results | 100.2× | 6 | 0.25× | 9 |
| JavaScript apps | 0.5× | 3 | 0.43× | 5 |
| Bot-protected pages | 62.9× | 6 | 0.61× | 9 |
| PDFs | — | — | 0.95× | 5 |
| Our own websites | 19.3× | 15 | 0.93× | 15 |
| All pages | 34× | 122 | 0.6× | 153 |

Median of the per-page ratio, over pages where both measures exist (the Pages columns). Below 1× means the first is smaller.

Across the 122 pages where both counts exist, the HTML was a median 34 times larger than Skryp's Markdown. Listings and global shop pages shrank more than 100 times. On the 153 pages both services read, Skryp's Markdown was a median 0.6 times the length of Firecrawl's [checked 1 Oct 2026], and smaller still on shop, listing and JavaScript-heavy pages.

PDFs are the exception. There is no HTML to compare, and the two services returned about the same length: on the five PDFs both read, Skryp's text was shorter on three and longer on two, a median ratio of 0.95. Both convert the whole document, so a PDF's token count is mostly the document itself.

## Fewer tokens is not automatically better

Leaner output drops navigation, repeated menus and cookie banners, and occasionally something a reader needed. Over the same run, the median share of each page's main content that came back was 0.98 for Skryp and 1.00 for Firecrawl, and on documentation pages Firecrawl returned 23 of 23 usable pages to Skryp's 20 [checked 1 Oct 2026]. The main-text-only copy is smaller still, but it is a cleaned extract that can lose tables, code and lists.

Choose by the task. To answer questions from a page, lean Markdown is usually enough. To pull specific fields such as a price or a stock level, ask for those fields directly: [Extract](https://skryp.dev/product/extract) returns them as JSON, which is smaller than any whole-page format.

## Count the tokens of your own page

This script counts one page both ways with the same tokenizer as the study. It ran as shown:

```python
# Count a page's tokens as raw HTML and as Markdown from Skryp, with the same tokenizer the study used.
# pip install httpx tiktoken    export SKRYP_API_KEY=...
import os

import httpx
import tiktoken

URL = "https://docs.python.org/3/tutorial/controlflow.html"
API = os.environ.get("SKRYP_API_URL", "https://api.skryp.dev")
enc = tiktoken.get_encoding("cl100k_base")

html = httpx.get(URL, follow_redirects=True, timeout=30).text
page = httpx.post(
    f"{API}/v1/scrape",
    headers={"Authorization": f"Bearer {os.environ['SKRYP_API_KEY']}"},
    json={"url": URL, "formats": ["markdown"]},
    timeout=120,
).json()

html_tokens = len(enc.encode(html, disallowed_special=()))
md_tokens = len(enc.encode(page["markdown"], disallowed_special=()))
receipt = page["receipt"]
print(f"HTML as served:  {html_tokens:>7,} tokens")
print(f"Skryp Markdown:  {md_tokens:>7,} tokens  ({html_tokens / md_tokens:.0f}x smaller)")
credits = receipt["credits"]
print(f"Route {receipt['route']}, exit {receipt['egress']['country'].upper()}, cache {receipt['cache']}, "
      f"{credits} credit{'' if credits == 1 else 's'}")
```

Output (ran 1 Oct 2026; Python 3.12.14, httpx 0.28.1, tiktoken 0.14.0; 1.4 s):

```text
HTML as served:   41,309 tokens
Skryp Markdown:    6,242 tokens  (7x smaller)
Route http:direct, exit ZA, cache miss, 1 credit
```

The [Scrape API](https://skryp.dev/product/scrape) and the MCP `scrape` tool also return a `tokens` field with each page, so an agent can see the cost of what it read. To try a page without writing code, use [URL to Markdown](https://skryp.dev/tools/url-to-markdown).

## Method

- **Pages:** 170 public pages in ten groups: news articles, documentation, shop pages from South Africa and elsewhere, country-priced pages requested from several countries, listings and search results, JavaScript apps, bot-protected pages, PDFs, and 15 pages from websites we run. Those 15 are counted; their addresses are withheld.
- **Markdown:** one run through each service on 1 October 2026 from a South African connection, with each service's default output (Firecrawl's main-content Markdown, Skryp's lean Markdown). Medians cover each service's usable pages.
- **Usable:** the page's content was present and complete enough to answer from, judged against a reference copy of the page.
- **HTML:** one plain GET per address later the same day, with desktop Chrome headers and no JavaScript. Pages that were blocked, failed or were PDFs have no HTML count and are left out of that median.
- **Main text:** a full browser render cleaned with trafilatura, an open-source extractor that is neither service's converter.
- **Tokens:** cl100k_base, OpenAI's tokenizer for GPT-4-class models. Other models split text differently, so their counts differ; the ratios between formats hold up better than the absolute numbers.
- **Data:** [download the CSV](https://skryp.dev/benchmarks/tokens-per-web-page/data.csv), one row per page with every count.

## What this does not show

- It is one run, from one place, on one day. Pages change, and a different set of 170 would give different medians.
- It counts what each service returned at its default settings, not the best either can do with tuning.
- Token counts measure size, not quality. Completeness and the usable count above are the quality measures, and they favour Firecrawl slightly.
- Skryp ran this study on its own product. The data and the counting method are published so you can check them.

## Questions

**How do you convert HTML to Markdown?**

Use a converter library such as markdownify (Python) or Turndown (JavaScript), or an API that fetches the page and returns Markdown. Converters keep everything on the page, including navigation; scraping APIs usually keep only the main content, which is why their output is much shorter.

**How many tokens is a web page?**

On our 170 public pages, the median page was 65,378 tokens of HTML as served, about 3,520 as Firecrawl's Markdown and 1,800 as Skryp's. The page's main text alone was about 400 tokens.

**Why does Markdown save so many tokens?**

HTML carries markup, scripts, styles and navigation that an agent does not need. A scraper's Markdown keeps the content; on pages where Skryp returned usable Markdown, the HTML was a median 34 times larger.


## Evidence

- Fetched with a plain HTTP GET on 1 October 2026, the median benchmark page with an HTML count was 65,378 tokens of HTML; on the 122 pages where Skryp also returned usable Markdown, the HTML was a median 34 times larger. (observed test, checked 1 Oct 2026; 126 of the benchmark's 170 pages had an HTML count; the others were blocked or errors (30), failed to load (9) or were PDFs (5). Limits: HTML as served, before JavaScript runs: for JavaScript apps it is a small shell without the content. Fetched hours after the Markdown benchmark, so some pages changed in between. Country variants of one URL share a single fetch.)
- On the same 170 public pages, the median output was 1,800 tokens per page from Skryp and 3,520 from Firecrawl. (observed test, checked 1 Oct 2026; Same benchmark run; Markdown output. Limits: Fewer tokens is not automatically better: completeness was slightly lower for Skryp (median 0.98 vs 1.00). PDFs were the exception: on the five PDFs both services read, they returned about the same length (Skryp shorter on three, longer on two; median ratio 0.95).)
- On the same 170 public pages, Firecrawl returned 159 usable pages and Skryp 157; median completeness was 1.00 for Firecrawl and 0.98 for Skryp; on documentation pages Firecrawl returned 23 of 23 usable and Skryp 20 of 23. (observed test, checked 1 Oct 2026; 170 public pages across articles, docs, shops, listings, JavaScript apps, PDFs, protected and regional pages. Limits: One run. Firecrawl at default settings. Skryp's own benchmark; publish the URL list and scoring code with it.)
- On a 170-URL mixed benchmark run on 1 October 2026, 6 URLs no longer existed (both services got a 404); of the 164 live pages, Skryp returned 157 usable and Firecrawl 159. (observed test, checked 1 Oct 2026; 170 public URLs in ten page types (articles, documentation, shops in South Africa and elsewhere, country-priced pages, listings, JavaScript apps, bot-protected pages, PDFs and our own websites), each read once by each service. Limits: One run from a South African connection before the move to a server; results from a data-centre connection may differ. Skryp's seven misses on live pages: three documentation pages answered Skryp with HTTP 503 during the run, two pages behind a Cloudflare challenge, a shop home page where Skryp kept only a banner and the cookie notice, and one page neither service read. Firecrawl's five: four listing and social pages and the same page neither read.)
- On the 153 public pages both services returned as usable on 1 October 2026, Skryp's Markdown was a median 40% shorter than Firecrawl's, and shorter on 135 of them. (observed test, checked 1 Oct 2026; Same benchmark run; Markdown output; pages where both services returned usable content. Limits: Fewer tokens is not automatically better: Skryp kept slightly less of each page's reference text (median completeness 0.98 vs 1.00). PDFs came out about the same length. One run on 170 pages we chose.)
