Web scraping with Python: requests, BeautifulSoup, Playwright or an API
Updated 1 Oct 2026 · Tested on 1 Oct 2026
For static pages, fetch with requests and parse with BeautifulSoup; for pages that build their content with JavaScript, render them with Playwright; when you need many sites, countries or blocked pages, an API is usually cheaper than maintaining that yourself. Each step below has code that runs.
Pick the tool by the page
| The page | What works | Example |
|---|---|---|
| Plain HTML: the data is in the page source | requests + BeautifulSoup | 1, 2 |
| Built by JavaScript after it loads | Playwright (a real browser) | 3 |
| Many different sites, other countries, bot protection, schedules | A scraping API | 4 |
To check which kind of page you have, compare the browser's View source with what you see. If the data is not in the source, the page builds it with JavaScript.
All the code below ran as shown, against books.toscrape.com and quotes.toscrape.com, two sites made for practising scraping. The files are ready to run on their own.
1. A static page: requests and BeautifulSoup
requests downloads the page; BeautifulSoup turns the HTML into a tree you can query with CSS selectors.
# Fetch a static page with requests and pull fields out of it with BeautifulSoup.
# pip install requests beautifulsoup4
import requests
from bs4 import BeautifulSoup
URL = "https://books.toscrape.com/catalogue/category/books/travel_2/index.html"
response = requests.get(URL, headers={"User-Agent": "my-scraper/1.0 (you@example.com)"}, timeout=30)
response.raise_for_status()
# pass the raw bytes: BeautifulSoup reads the page's own charset (response.text would guess, and "£" becomes "£")
soup = BeautifulSoup(response.content, "html.parser")
books = []
for card in soup.select("article.product_pod"):
books.append({
"title": card.h3.a["title"],
"price": card.select_one(".price_color").get_text(strip=True),
"in_stock": "In stock" in card.select_one(".availability").get_text(),
})
print(f"{len(books)} books on the page")
for book in books[:5]:
print(book)11 books on the page
{'title': "It's Only the Himalayas", 'price': '£45.17', 'in_stock': True}
{'title': 'Full Moon over Noah’s Ark: An Odyssey to Mount Ararat and Beyond', 'price': '£49.43', 'in_stock': True}
{'title': 'See America: A Celebration of Our National Parks & Treasured Sites', 'price': '£48.87', 'in_stock': True}
{'title': 'Vagabonding: An Uncommon Guide to the Art of Long-Term World Travel', 'price': '£36.94', 'in_stock': True}
{'title': 'Under the Tuscan Sun', 'price': '£37.33', 'in_stock': True}Three habits in that script save trouble later:
- Identify yourself. A
User-Agentwith a contact address tells the site who is calling, and some sites refuse the default one. - Fail loudly.
raise_for_status()stops on a 403 or 404 instead of parsing an error page as if it were data. - Let BeautifulSoup read the bytes. This site does not declare its character set in its headers, so
response.textguesses wrong and "£" comes out as "£". Passingresponse.contentlets BeautifulSoup read the encoding from the page itself.
2. Many pages: follow "next" and save a CSV
Listings spread records over several pages. Follow the "next" link until there is none, with a hard cap and a pause between requests.
# Follow "next" links across a listing, politely, and save every record to a CSV file.
# pip install requests beautifulsoup4
import csv
import time
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup
url = "https://books.toscrape.com/catalogue/category/books/mystery_3/index.html"
session = requests.Session()
session.headers["User-Agent"] = "my-scraper/1.0 (you@example.com)"
rows, pages = [], 0
while url and pages < 5: # a hard cap, so a bug cannot crawl forever
soup = BeautifulSoup(session.get(url, timeout=30).content, "html.parser")
pages += 1
for card in soup.select("article.product_pod"):
rows.append({"title": card.h3.a["title"], "price": card.select_one(".price_color").get_text(strip=True),
"url": urljoin(url, card.h3.a["href"])})
next_link = soup.select_one("li.next a")
url = urljoin(url, next_link["href"]) if next_link else None
time.sleep(1) # one request a second is plenty for a small site
with open("mystery-books.csv", "w", newline="", encoding="utf-8") as f:
writer = csv.DictWriter(f, fieldnames=["title", "price", "url"])
writer.writeheader()
writer.writerows(rows)
print(f"{len(rows)} books from {pages} pages saved to mystery-books.csv")
print(open("mystery-books.csv", encoding="utf-8").read().splitlines()[1])32 books from 2 pages saved to mystery-books.csv
Sharp Objects,£47.82,https://books.toscrape.com/catalogue/sharp-objects_997/index.htmlA Session reuses the connection between requests, urljoin turns relative links into full addresses, and the cap means a bug in the "next" selector cannot run forever. One request a second is plenty for a small site.
3. Pages built with JavaScript: Playwright
quotes.toscrape.com/js sends an almost empty page and builds the quotes in the browser. requests receives none of them; Playwright runs the page in Chromium and sees all ten.
# A page that builds its content with JavaScript: requests sees none of it, a real browser does.
# pip install playwright python -m playwright install chromium
import requests
from playwright.sync_api import sync_playwright
URL = "https://quotes.toscrape.com/js/"
html = requests.get(URL, timeout=30).text
print(f"requests: {html.count('class=\"quote\"')} quotes in the HTML it received")
with sync_playwright() as p:
browser = p.chromium.launch()
page = browser.new_page()
page.goto(URL)
page.wait_for_selector("div.quote") # wait for the content, not for a fixed number of seconds
quotes = [
{"text": q.locator(".text").inner_text(), "author": q.locator(".author").inner_text()}
for q in page.locator("div.quote").all()
]
browser.close()
print(f"playwright: {len(quotes)} quotes")
print(quotes[0])requests: 0 quotes in the HTML it received
playwright: 10 quotes
{'text': '“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”', 'author': 'Albert Einstein'}Wait for the content you need with wait_for_selector, not for a fixed number of seconds: it is faster when the page is quick and does not break when it is slow. Playwright needs its browser installed once with python -m playwright install chromium.
4. When an API is simpler
Your own code is the cheapest option for a few sites you know. It gets expensive when you need many different sites, pages that block automated readers, prices as seen from another country, or a job that runs on a schedule: then you are maintaining browsers, proxies and parsers. A scraping API does that work per request.
Here is the same JavaScript page through Skryp's extraction API, asking for the fields with a JSON schema:
# The same JavaScript page through a scraping API: one request, and the API decides whether a browser is needed.
# pip install httpx export SKRYP_API_KEY=...
import os
import httpx
API = os.environ.get("SKRYP_API_URL", "https://api.skryp.dev")
response = httpx.post(
f"{API}/v1/extract",
headers={"Authorization": f"Bearer {os.environ['SKRYP_API_KEY']}"},
json={
"url": "https://quotes.toscrape.com/js/",
"schema": {
"type": "object",
"properties": {"quotes": {"type": "array", "items": {"type": "object", "properties": {
"text": {"type": "string"}, "author": {"type": "string"}}}}},
},
},
timeout=120,
).json()
quotes = response["data"]["quotes"]
print(f"{len(quotes)} quotes, extracted by {response['method']}, values checked against the page: {response['checks']}")
print(quotes[0])
receipt = response["receipt"]
credits = receipt["credits"] # the page's route price (0 when it came from cache) + 5 when a model did the extracting
print(f"route {receipt['route']}, cache {receipt['cache']}, {credits} credit{'' if credits == 1 else 's'}")10 quotes, extracted by model, values checked against the page: {'quotes': 'on_page'}
{'text': 'The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.', 'author': 'Albert Einstein'}
route browser:direct, cache miss, 6 creditsThe API decided the page needed a browser (the route on the receipt), returned the quotes as JSON, and checked every value against the page. It cost 6 credits: 1 for the page on Skryp's own network and 5 because a model did the extracting. There is a free monthly allowance, and failed pages cost nothing. See Pricing.
The same page, four ways
This script fetches the JavaScript page with each approach and compares what came back:
# One JavaScript-built page fetched four ways: what each returns, how long it takes, and what fails.
# pip install requests beautifulsoup4 playwright httpx export SKRYP_API_KEY=...
import os
import time
import httpx
import requests
from bs4 import BeautifulSoup
from playwright.sync_api import sync_playwright
URL = "https://quotes.toscrape.com/js/"
API = os.environ.get("SKRYP_API_URL", "https://api.skryp.dev")
results = []
def run(name, fn):
t0 = time.monotonic()
found, size = fn()
results.append((name, found, size, time.monotonic() - t0))
def plain_requests():
html = requests.get(URL, timeout=30).text
return html.count('class="quote"'), len(html)
def beautifulsoup():
soup = BeautifulSoup(requests.get(URL, timeout=30).text, "html.parser")
return len(soup.select("div.quote")), len(soup.get_text())
def playwright_browser():
with sync_playwright() as p:
browser = p.chromium.launch()
page = browser.new_page()
page.goto(URL)
page.wait_for_selector("div.quote")
found, size = page.locator("div.quote").count(), len(page.content())
browser.close()
return found, size
def scraping_api():
page = httpx.post(f"{API}/v1/scrape", headers={"Authorization": f"Bearer {os.environ['SKRYP_API_KEY']}"},
json={"url": URL, "formats": ["markdown"], "max_age_s": 0}, timeout=120).json()
return page["markdown"].count("by "), len(page["markdown"])
for name, fn in (("requests", plain_requests), ("requests + BeautifulSoup", beautifulsoup),
("Playwright", playwright_browser), ("Skryp API (Markdown)", scraping_api)):
run(name, fn)
print(f"{'method':<26}{'quotes found':>13}{'characters':>12}{'seconds':>9}")
for name, found, size, secs in results:
print(f"{name:<26}{found:>13}{size:>12,}{secs:>9.1f}")method quotes found characters seconds
requests 0 5,806 0.7
requests + BeautifulSoup 0 161 0.7
Playwright 10 8,940 2.4
Skryp API (Markdown) 10 1,525 6.4| Approach | Good for | Watch out for |
|---|---|---|
requests | Fast reads of plain HTML and JSON endpoints | Gets nothing from pages built by JavaScript |
| BeautifulSoup | Picking fields out of HTML you have | Only as good as the HTML you give it |
| Playwright | Anything a browser can show, clicks and forms | Slower; you run and update the browsers |
| A scraping API | Many sites, other countries, blocked pages, schedules | A cost per page; you depend on a provider |
Scraping responsibly
- Read the site's terms and its
robots.txtbefore you start, and respect what they ask. Some sites offer an official API: use it when it has the data. - Go slowly. A pause between requests and a cap on pages keep you from loading a small site.
- Store only what you need. Avoid collecting personal data unless you have a lawful reason and a plan for it.
- Cache what you fetched, so a re-run of your parser does not hit the site again.
Where to go next
- Drive a real browser step by step: Web scraping with Selenium and Python.
- Choose between browser libraries: Playwright vs Puppeteer vs Selenium.
- The same techniques in Node.js: Web scraping with JavaScript.
Questions
- How do you scrape a website with Python?
- Download the page with requests, parse the HTML with BeautifulSoup and pick out the elements you need with CSS selectors. If the content is built by JavaScript, load the page in Playwright instead. Follow pagination with a pause between requests, save the results to CSV or JSON, and check the site's terms first.
- Scrapy or BeautifulSoup?
- BeautifulSoup only parses HTML you have already downloaded; Scrapy is a crawling framework that downloads pages, follows links, retries failures and exports results. Use requests and BeautifulSoup for a handful of pages, and Scrapy when you crawl whole sites.
- Is Python or JavaScript better for web scraping?
- Both work. Python has requests, BeautifulSoup, Scrapy and Playwright; JavaScript has Cheerio, Puppeteer and Playwright and suits teams already on Node.js. Use the language your project already uses: the browser tools behave much the same in both.
- What is web scraping in Python?
- Writing Python code that downloads web pages and extracts data from them, usually into a spreadsheet, a database or an AI model's context. The usual tools are requests to download, BeautifulSoup to parse, and Playwright or Selenium when a page needs a real browser.