The best web scraping tools, tested
Updated 1 Oct 2026 · Tested on 1 Oct 2026
The right scraper depends on the pages and who runs it. We ran eight tools on the same two pages on 1 October 2026: libraries that read HTML missed everything JavaScript built, browser libraries got it all with hand-written selectors, and page-to-Markdown tools got it all without selectors. Below: what each is, its price and who it suits. Skryp is ours.
Eight tools, the same two pages
We ran eight tools on two pages: an ordinary HTML listing of 11 books, and a page that builds its 10 quotes with JavaScript after it loads. The libraries were given the CSS selectors for the items; the other four return the whole page as text, and we checked whether each item was in it. The answer key is the page's real items, read with a browser. The script and what it printed on 1 October 2026:
# Eight scraping tools on the same two pages: a plain HTML listing (11 books) and a page that builds its
# 10 quotes with JavaScript. Each is scored on how many of the page's items its output contains.
# pip install requests beautifulsoup4 scrapy playwright selenium crawl4ai httpx
# export SKRYP_API_KEY=... FIRECRAWL_API_KEY=... (Playwright's and Selenium's browsers installed)
# timeout: 900
import asyncio
import json
import os
import subprocess
import sys
import time
from pathlib import Path
import httpx
import requests
from bs4 import BeautifulSoup
STATIC = "https://books.toscrape.com/catalogue/category/books/travel_2/index.html"
JS = "https://quotes.toscrape.com/js/"
SKRYP = os.environ.get("SKRYP_API_URL", "https://api.skryp.dev")
def timed(fn, *a):
t0 = time.monotonic()
try:
return fn(*a), time.monotonic() - t0
except Exception as e: # a tool that fails is a result too
return f"error: {type(e).__name__}", time.monotonic() - t0
# --- libraries: you write the selectors ---
def bs4_items(url):
soup = BeautifulSoup(requests.get(url, timeout=30).content, "html.parser")
return [a["title"] for a in soup.select("article.product_pod h3 a")] + [q.get_text() for q in soup.select(".quote .text")]
def scrapy_items(url):
spider = Path(__file__).with_name("_books_spider.py")
out = subprocess.run([sys.executable, "-m", "scrapy", "runspider", str(spider), "-a", f"url={url}", "-O", "-:json",
"--nolog"], capture_output=True, text=True, timeout=120).stdout
return [next(iter(d.values())) for d in json.loads(out or "[]")]
def playwright_items(url):
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
b = p.chromium.launch()
page = b.new_page()
page.goto(url)
page.wait_for_selector("article.product_pod, .quote")
items = page.eval_on_selector_all("article.product_pod h3 a", "els => els.map(e => e.title)")
items += page.eval_on_selector_all(".quote .text", "els => els.map(e => e.textContent)")
b.close()
return items
def selenium_items(url):
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as ec
from selenium.webdriver.support.ui import WebDriverWait
o = webdriver.ChromeOptions()
o.add_argument("--headless=new")
d = webdriver.Chrome(options=o)
try:
d.get(url)
WebDriverWait(d, 20).until(ec.presence_of_element_located((By.CSS_SELECTOR, "article.product_pod, .quote")))
return ([e.get_attribute("title") for e in d.find_elements(By.CSS_SELECTOR, "article.product_pod h3 a")]
+ [e.text for e in d.find_elements(By.CSS_SELECTOR, ".quote .text")])
finally:
d.quit()
# --- tools that return the whole page as Markdown or text: no selectors ---
def crawl4ai_text(url):
from crawl4ai import AsyncWebCrawler
async def go():
async with AsyncWebCrawler(verbose=False) as c:
return str((await c.arun(url)).markdown or "")
return asyncio.run(go())
def jina_text(url):
return httpx.get(f"https://r.jina.ai/{url}", timeout=60).text
def firecrawl_text(url):
r = httpx.post("https://api.firecrawl.dev/v2/scrape", headers={"Authorization": f"Bearer {os.environ['FIRECRAWL_API_KEY']}"},
json={"url": url, "formats": ["markdown"]}, timeout=120).json()
return r["data"]["markdown"]
def skryp_text(url):
r = httpx.post(f"{SKRYP}/v1/scrape", headers={"Authorization": f"Bearer {os.environ['SKRYP_API_KEY']}"},
json={"url": url}, timeout=120).json()
return r["markdown"]
def norm(s):
return " ".join(s.replace("“", '"').replace("”", '"').replace("′", "'").replace("’", "'").split())
# the pages' real items, read with a browser, are the answer key
truth = {url: [norm(x)[:40] for x in playwright_items(url)] for url in (STATIC, JS)}
print(f"answer key: {len(truth[STATIC])} books on the HTML page, {len(truth[JS])} quotes on the JavaScript page\n")
TOOLS = [("requests + BeautifulSoup", bs4_items), ("Scrapy", scrapy_items), ("Playwright", playwright_items),
("Selenium", selenium_items), ("Crawl4AI", crawl4ai_text), ("Jina Reader", jina_text),
("Firecrawl", firecrawl_text), ("Skryp (ours)", skryp_text)]
print(f"{'tool':<26}{'HTML page':>11}{'JS page':>9}{'seconds':>9}")
for name, fn in TOOLS:
cells, secs = [], 0.0
for url in (STATIC, JS):
out, s = timed(fn, url)
secs += s
if isinstance(out, str) and out.startswith("error"):
cells.append(out[7:][:9])
continue
text = norm(" ".join(out) if isinstance(out, list) else out)
cells.append(f"{sum(1 for t in truth[url] if t in text)}/{len(truth[url])}")
print(f"{name:<26}{cells[0]:>11}{cells[1]:>9}{secs:>9.1f}")answer key: 11 books on the HTML page, 10 quotes on the JavaScript page
tool HTML page JS page seconds
requests + BeautifulSoup 11/11 0/10 5.0
Scrapy 11/11 0/10 4.4
Playwright 11/11 10/10 5.3
Selenium 11/11 10/10 11.1
Crawl4AI 10/11 10/10 5.4
Jina Reader 11/11 10/10 1.5
Firecrawl 11/11 10/10 2.3
Skryp (ours) 11/11 10/10 8.8What it shows:
- Plain HTTP libraries cannot see what JavaScript builds. requests with BeautifulSoup, and Scrapy, read every book on the HTML page and none of the quotes, because the quotes are not in the HTML the server sends.
- Browser libraries get both, if you write the selectors. Playwright and Selenium read all 21 items. Selenium took twice as long here, from Python; in our browser library race for Node.js, the three drivers' median task times were within half a second of each other[1 Oct 2026].
- Page-to-Markdown tools need no selectors. Crawl4AI, Jina Reader, Firecrawl and Skryp returned both pages as text an agent or a parser can use. Crawl4AI's output shortened one long book title, so that title did not match.
Two easy pages show the basics, not how a tool copes with blocked sites, other countries or scale. For hard pages, see our 170-page benchmark, where Skryp read 157 and Firecrawl 159 of 164 live pages[1 Oct 2026].
The tools
Open-source libraries
requests and BeautifulSoup (Python). Fetch the HTML and pick out what you need. Free, fast, and enough for pages that arrive complete. They cannot run JavaScript. See web scraping with Python.
Scrapy (Python). A framework for crawling many pages, with queues, retries and export built in. Free. Like requests, it reads the HTML the server sends, so JavaScript-built content is missing.
Playwright (JavaScript, Python, Java and .NET) and Selenium (Java, Python, C#, Ruby and JavaScript). They drive real browsers, so they see what a person sees, and they can click, type and scroll[1 Oct 2026]. Free; you run and maintain the browsers. See Playwright vs Puppeteer vs Selenium.
Crawl4AI (Python). Opens pages in a browser and returns Markdown for language models. Open source under Apache-2.0, free to run yourself, or hosted[1 Oct 2026].
Hosted services
Jina Reader. Put r.jina.ai/ in front of a URL and get the page as text. Works without a key within per-IP limits; a key comes with 10 million free tokens, then it is priced per token[1 Oct 2026].
Firecrawl. An API and MCP server that returns pages as Markdown or JSON, with crawl, map and search. Plans from $16 a month billed annually for 5,000 credits, and 1,000 free credits a month[1 Oct 2026]; the core is open source under AGPL-3.0[1 Oct 2026]. See Skryp vs Firecrawl.
Skryp (ours). An API and MCP server that returns pages as Markdown or checked JSON, from a country you choose, with a receipt for each result: the route, the exit country and the credits. 1,000 free credits a month; plans from $9 for 20,000[1 Oct 2026]. It was the slowest of the four page-to-Markdown tools on these two pages.
Services we did not run for this page
ScrapingBee and Bright Data's Web Unlocker sell per-request access with JavaScript rendering and proxies. We compared their published prices with Skryp's in Skryp vs ScrapingBee and Skryp vs Bright Data, without running them. No-code scraping tools are not covered here.
Which to choose
| If you | Choose |
|---|---|
| Read a few sites whose pages arrive complete | requests and BeautifulSoup |
| Crawl one site at scale and can run servers | Scrapy |
| Need to click, log in or scroll, and can maintain browsers | Playwright or Selenium |
| Want pages as Markdown on your own machine, for free | Crawl4AI |
| Want one page as text now, without an account | Jina Reader |
| Want an API or MCP server for your agent, without running browsers | Firecrawl or Skryp |
| Need pages as seen from a country, with proof of where each request came from | Skryp |
Prices and limits are from each provider's pages on 1 October 2026.
Questions
- What is the best free web scraping tool?
- For pages that arrive complete, requests with BeautifulSoup or Scrapy. For pages built with JavaScript, Playwright or Selenium, or Crawl4AI if you want Markdown. All are open source. Jina Reader works without a key within rate limits.
- Why did requests and Scrapy miss the quotes?
- The quotes on that page are added by JavaScript after it loads. requests and Scrapy read the HTML the server sends, which does not contain them; a browser, or a service that runs one, does.
- Which tool is best for AI agents?
- One that returns clean text or JSON over an API or MCP server, so the agent does not run browsers: Firecrawl, Jina Reader or Skryp. Skryp also verifies the country each request came from and shows the cost of each result.
- Is Skryp independent?
- No. Skryp is ours. The script and its output are on this page so you can run the same test yourself.
Evidence
On 1 October 2026, on an HTML listing of 11 books and a page that builds 10 quotes with JavaScript, requests with BeautifulSoup and Scrapy read all 11 books and none of the quotes; Playwright, Selenium, Jina Reader, Firecrawl and Skryp read all 21 items; Crawl4AI read 20, its Markdown having shortened one long title.
observed test · checked 1 Oct 2026 · Two practice pages (books.toscrape.com and quotes.toscrape.com/js), one run per tool from South Africa.
Limits: Two easy pages: they show whether a tool runs JavaScript, not how it copes with blocked sites, other countries or scale. Times include starting each tool.
Selenium's main project ships language bindings for Java, Python, C#, Ruby and JavaScript and drives Chrome, Edge, Firefox and Safari; Playwright is available for JavaScript/TypeScript, Python, Java and .NET and drives Chromium, Firefox and WebKit plus branded Chrome and Edge; Puppeteer is a Node.js library that drives Chrome and Firefox.
provider docs · checked 1 Oct 2026 · selenium.dev/downloads and /documentation/webdriver/browsers, playwright.dev/docs/languages and /docs/browsers, pptr.dev/supported-browsers
Limits: Community bindings exist for other languages; only each project's official ones are listed.
On four JavaScript-heavy tasks (a rendered page, a page that renders after 10 seconds, infinite scroll and a table loaded by a click), each run three times on 1 October 2026, Playwright 1.63, Puppeteer 24.43 and Selenium 4.50 for Node.js all got the data 12 of 12 times; their median task times were within half a second of each other, and median browser start was 0.11 s (Playwright), 0.37 s (Selenium) and 0.40 s (Puppeteer).
observed test · checked 1 Oct 2026 · quotes.toscrape.com and scrapethissite.com practice pages; headless Chromium-based browsers on a MacBook in South Africa.
Limits: Practice sites, one machine, one connection; real sites with bot protection or heavy pages will differ. Playwright's start time uses its lightweight headless shell.
Crawl4AI is open source (Apache-2.0): free as a Python library or as your own server, or hosted as Crawl4AI Cloud, pay as you go with the first $10 free.
provider docs · checked 1 Oct 2026 · github.com/unclecode/crawl4ai
Limits: Not run in this comparison.
Jina Reader converts a page to LLM-friendly text when you put r.jina.ai/ in front of its URL; it works without an API key under per-IP rate limits, and a new API key comes with 10 million free tokens, after which it is priced per token.
provider docs · checked 1 Oct 2026 · jina.ai/reader
Limits: Not run in this comparison; the rate-limit numbers per tier were not recorded.
Firecrawl's pricing page (prices effective 4 September 2026) lists Free (1,000 credits a month), Hobby ($19 a month, or $16 billed annually, 5,000 credits), Standard ($83 a month billed annually, 100,000 credits), Growth ($333, 500,000) and Scale ($599, 1,000,000). One credit is one page on a basic scrape, crawl or map; a page that responds with an error status such as 403 or 404 costs 1 credit, and a scrape that returns no result is not charged.
provider docs · checked 1 Oct 2026 · firecrawl.dev/pricing
Limits: Prices change; re-read on the day a page citing them is published.
Firecrawl is open source under the AGPL-3.0 licence (its SDKs under MIT) and can be self-hosted; the hosted cloud at firecrawl.dev has additional features.
provider docs · checked 1 Oct 2026 · github.com/firecrawl/firecrawl
Limits: Which features the self-hosted version lacks is listed by Firecrawl, not tested here.
Skryp's plans from 1 October 2026: Free (1,000 credits a month), Hobby ($9 a month for 20,000 credits), Standard ($49 for 150,000), Growth ($149 for 600,000) and Scale ($399 for 2,000,000); unused plan credits expire at the end of each month, and top-ups, which never expire (also after a plan is cancelled), cost $10 for 20,000 credits on Free and the plan's rate on a plan. A page costs 1 credit through Skryp's own network, 5 through a residential or mobile proxy (10 with JavaScript rendering) and 25 for advanced unblocking, only when it succeeds; model extraction adds 5; failed, blocked and cached results cost 0, and a missing page (404 or 410) costs at most 1.
observed test · checked 1 Oct 2026 · Skryp's own prices (the service's rate card).
Limits: Prices in US dollars, charged in rand through Paystack at R17 to the US dollar (Chad, 2 Oct 2026).
On a 170-URL mixed benchmark run on 1 October 2026, 6 URLs no longer existed (both services got a 404); of the 164 live pages, Skryp returned 157 usable and Firecrawl 159.
observed test · checked 1 Oct 2026 · 170 public URLs in ten page types (articles, documentation, shops in South Africa and elsewhere, country-priced pages, listings, JavaScript apps, bot-protected pages, PDFs and our own websites), each read once by each service.
Limits: One run from a South African connection before the move to a server; results from a data-centre connection may differ. Skryp's seven misses on live pages: three documentation pages answered Skryp with HTTP 503 during the run, two pages behind a Cloudflare challenge, a shop home page where Skryp kept only a banner and the cookie notice, and one page neither service read. Firecrawl's five: four listing and social pages and the same page neither read.
Related
- A Firecrawl alternative you can switch to without rewriting your code
- A ScrapingBee alternative with a receipt for every request
- A Bright Data alternative for teams that want results, not proxy plans
- Web scraping with Python: requests, BeautifulSoup, Playwright or an API
- Playwright vs Puppeteer vs Selenium for web scraping