Skip to content
Skryp

How many tokens is a web page?

Updated 1 Oct 2026 · Tested on 1 Oct 2026

A typical web page is far bigger as HTML than as text. Across 170 public pages, the median page was 65,378 tokens of HTML as served, 3,520 tokens as Markdown from Firecrawl and 1,800 from Skryp; the main text alone was about 400. Shop and listing pages shrink most; JavaScript apps and PDFs are the exceptions.

The answer in numbers

How many tokens a page costs depends on what your agent reads. On the median page of the 170 we measured:

  • HTML as served: 65,378 tokens[1 Oct 2026], what a plain fetch hands an agent before any cleaning.
  • Markdown from Firecrawl: 3,520 tokens[1 Oct 2026].
  • Markdown from Skryp: 1,800 tokens.
  • Main text only: about 400 tokens, from a neutral cleaned copy that exists for 105 of the pages.

At an example price of $3 per million input tokens, reading 1,000 such pages costs about $196 as HTML, $10.56 as Firecrawl's Markdown and $5.40 as Skryp's. The price is an illustration; the ratios are what carries over to your model.

By page type

Page typePagesHTML as servedSkryp MarkdownFirecrawl MarkdownMain text only
News articles2485,092 (20)2,4713,682734
Documentation23135,0131,0652,261394
Shop pages (global)13347,746 (7)1,0722,724233
Shop pages (South Africa)2278,4931,7603,713178
Country-priced pages3250,006 (24)1,2922,790—
Listings and search results1449,920 (6)1,6065,397467
JavaScript apps515,491 (3)6,54915,213360
Bot-protected pages15196,442 (6)2,56813,489226
PDFs7—17,77115,660—
Our own websites1512,1751,3721,464577
All pages17065,378 (126)1,8003,520408 (105)
Median tokens per page (cl100k). Markdown medians cover each service's usable pages. A number in brackets is how many pages the median covers when fewer than all: blocked pages, PDFs and fetch failures have no HTML count, and the main-text reference exists only where a full browser render could be cleaned.

Global shop pages carry the heaviest HTML, a median of 347,746 tokens on the 7 that returned it[1 Oct 2026]. Documentation is heavy as HTML (135,013) yet among the shortest as Markdown (1,065 from Skryp). JavaScript apps are the reverse case: the HTML that arrives is a small shell, and the content appears only after scripts run, so a reader that does not render the page gets almost nothing of what a visitor sees.

How much smaller, page by page

Page typeHTML ÷ Skryp MarkdownPagesSkryp ÷ Firecrawl MarkdownPages
News articles30.3×200.58×24
Documentation68.3×200.62×20
Shop pages (global)123.3×70.43×13
Shop pages (South Africa)73.2×210.46×21
Country-priced pages20.1×240.69×32
Listings and search results100.2×60.25×9
JavaScript apps0.5×30.43×5
Bot-protected pages62.9×60.61×9
PDFs——0.95×5
Our own websites19.3×150.93×15
All pages34×1220.6×153
Median of the per-page ratio, over pages where both measures exist (the Pages columns). Below 1× means the first is smaller.

Across the 122 pages where both counts exist, the HTML was a median 34 times larger than Skryp's Markdown. Listings and global shop pages shrank more than 100 times. On the 153 pages both services read, Skryp's Markdown was a median 0.6 times the length of Firecrawl's[1 Oct 2026], and smaller still on shop, listing and JavaScript-heavy pages.

PDFs are the exception. There is no HTML to compare, and the two services returned about the same length: on the five PDFs both read, Skryp's text was shorter on three and longer on two, a median ratio of 0.95. Both convert the whole document, so a PDF's token count is mostly the document itself.

Fewer tokens is not automatically better

Leaner output drops navigation, repeated menus and cookie banners, and occasionally something a reader needed. Over the same run, the median share of each page's main content that came back was 0.98 for Skryp and 1.00 for Firecrawl, and on documentation pages Firecrawl returned 23 of 23 usable pages to Skryp's 20[1 Oct 2026]. The main-text-only copy is smaller still, but it is a cleaned extract that can lose tables, code and lists.

Choose by the task. To answer questions from a page, lean Markdown is usually enough. To pull specific fields such as a price or a stock level, ask for those fields directly: Extract returns them as JSON, which is smaller than any whole-page format.

Count the tokens of your own page

This script counts one page both ways with the same tokenizer as the study. It ran as shown:

tokens-per-web-page/01-count-tokens.py
# Count a page's tokens as raw HTML and as Markdown from Skryp, with the same tokenizer the study used.
# pip install httpx tiktoken    export SKRYP_API_KEY=...
import os

import httpx
import tiktoken

URL = "https://docs.python.org/3/tutorial/controlflow.html"
API = os.environ.get("SKRYP_API_URL", "https://api.skryp.dev")
enc = tiktoken.get_encoding("cl100k_base")

html = httpx.get(URL, follow_redirects=True, timeout=30).text
page = httpx.post(
    f"{API}/v1/scrape",
    headers={"Authorization": f"Bearer {os.environ['SKRYP_API_KEY']}"},
    json={"url": URL, "formats": ["markdown"]},
    timeout=120,
).json()

html_tokens = len(enc.encode(html, disallowed_special=()))
md_tokens = len(enc.encode(page["markdown"], disallowed_special=()))
receipt = page["receipt"]
print(f"HTML as served:  {html_tokens:>7,} tokens")
print(f"Skryp Markdown:  {md_tokens:>7,} tokens  ({html_tokens / md_tokens:.0f}x smaller)")
credits = receipt["credits"]
print(f"Route {receipt['route']}, exit {receipt['egress']['country'].upper()}, cache {receipt['cache']}, "
      f"{credits} credit{'' if credits == 1 else 's'}")
OutputRan 1 Oct 2026 · Python 3.12.14, httpx 0.28.1, tiktoken 0.14.0 · 1.4 s
HTML as served:   41,309 tokens
Skryp Markdown:    6,242 tokens  (7x smaller)
Route http:direct, exit ZA, cache miss, 1 credit

The Scrape API and the MCP scrape tool also return a tokens field with each page, so an agent can see the cost of what it read. To try a page without writing code, use URL to Markdown.

Method

  • Pages: 170 public pages in ten groups: news articles, documentation, shop pages from South Africa and elsewhere, country-priced pages requested from several countries, listings and search results, JavaScript apps, bot-protected pages, PDFs, and 15 pages from websites we run. Those 15 are counted; their addresses are withheld.
  • Markdown: one run through each service on 1 October 2026 from a South African connection, with each service's default output (Firecrawl's main-content Markdown, Skryp's lean Markdown). Medians cover each service's usable pages.
  • Usable: the page's content was present and complete enough to answer from, judged against a reference copy of the page.
  • HTML: one plain GET per address later the same day, with desktop Chrome headers and no JavaScript. Pages that were blocked, failed or were PDFs have no HTML count and are left out of that median.
  • Main text: a full browser render cleaned with trafilatura, an open-source extractor that is neither service's converter.
  • Tokens: cl100k_base, OpenAI's tokenizer for GPT-4-class models. Other models split text differently, so their counts differ; the ratios between formats hold up better than the absolute numbers.
  • Data: download the CSV, one row per page with every count.

What this does not show

  • It is one run, from one place, on one day. Pages change, and a different set of 170 would give different medians.
  • It counts what each service returned at its default settings, not the best either can do with tuning.
  • Token counts measure size, not quality. Completeness and the usable count above are the quality measures, and they favour Firecrawl slightly.
  • Skryp ran this study on its own product. The data and the counting method are published so you can check them.

Questions

How do you convert HTML to Markdown?
Use a converter library such as markdownify (Python) or Turndown (JavaScript), or an API that fetches the page and returns Markdown. Converters keep everything on the page, including navigation; scraping APIs usually keep only the main content, which is why their output is much shorter.
How many tokens is a web page?
On our 170 public pages, the median page was 65,378 tokens of HTML as served, about 3,520 as Firecrawl's Markdown and 1,800 as Skryp's. The page's main text alone was about 400 tokens.
Why does Markdown save so many tokens?
HTML carries markup, scripts, styles and navigation that an agent does not need. A scraper's Markdown keeps the content; on pages where Skryp returned usable Markdown, the HTML was a median 34 times larger.

Evidence

  1. Fetched with a plain HTTP GET on 1 October 2026, the median benchmark page with an HTML count was 65,378 tokens of HTML; on the 122 pages where Skryp also returned usable Markdown, the HTML was a median 34 times larger.

    observed test · checked 1 Oct 2026 · 126 of the benchmark's 170 pages had an HTML count; the others were blocked or errors (30), failed to load (9) or were PDFs (5).

    Limits: HTML as served, before JavaScript runs: for JavaScript apps it is a small shell without the content. Fetched hours after the Markdown benchmark, so some pages changed in between. Country variants of one URL share a single fetch.

  2. On the same 170 public pages, the median output was 1,800 tokens per page from Skryp and 3,520 from Firecrawl.

    observed test · checked 1 Oct 2026 · Same benchmark run; Markdown output.

    Limits: Fewer tokens is not automatically better: completeness was slightly lower for Skryp (median 0.98 vs 1.00). PDFs were the exception: on the five PDFs both services read, they returned about the same length (Skryp shorter on three, longer on two; median ratio 0.95).

  3. On the same 170 public pages, Firecrawl returned 159 usable pages and Skryp 157; median completeness was 1.00 for Firecrawl and 0.98 for Skryp; on documentation pages Firecrawl returned 23 of 23 usable and Skryp 20 of 23.

    observed test · checked 1 Oct 2026 · 170 public pages across articles, docs, shops, listings, JavaScript apps, PDFs, protected and regional pages.

    Limits: One run. Firecrawl at default settings. Skryp's own benchmark; publish the URL list and scoring code with it.

  4. On a 170-URL mixed benchmark run on 1 October 2026, 6 URLs no longer existed (both services got a 404); of the 164 live pages, Skryp returned 157 usable and Firecrawl 159.

    observed test · checked 1 Oct 2026 · 170 public URLs in ten page types (articles, documentation, shops in South Africa and elsewhere, country-priced pages, listings, JavaScript apps, bot-protected pages, PDFs and our own websites), each read once by each service.

    Limits: One run from a South African connection before the move to a server; results from a data-centre connection may differ. Skryp's seven misses on live pages: three documentation pages answered Skryp with HTTP 503 during the run, two pages behind a Cloudflare challenge, a shop home page where Skryp kept only a banner and the cookie notice, and one page neither service read. Firecrawl's five: four listing and social pages and the same page neither read.

  5. On the 153 public pages both services returned as usable on 1 October 2026, Skryp's Markdown was a median 40% shorter than Firecrawl's, and shorter on 135 of them.

    observed test · checked 1 Oct 2026 · Same benchmark run; Markdown output; pages where both services returned usable content.

    Limits: Fewer tokens is not automatically better: Skryp kept slightly less of each page's reference text (median completeness 0.98 vs 1.00). PDFs came out about the same length. One run on 170 pages we chose.