Skip to content
Skryp

Web scraping with Python: requests, BeautifulSoup, Playwright or an API

Updated 1 Oct 2026 · Tested on 1 Oct 2026

For static pages, fetch with requests and parse with BeautifulSoup; for pages that build their content with JavaScript, render them with Playwright; when you need many sites, countries or blocked pages, an API is usually cheaper than maintaining that yourself. Each step below has code that runs.

Pick the tool by the page

The pageWhat worksExample
Plain HTML: the data is in the page sourcerequests + BeautifulSoup1, 2
Built by JavaScript after it loadsPlaywright (a real browser)3
Many different sites, other countries, bot protection, schedulesA scraping API4

To check which kind of page you have, compare the browser's View source with what you see. If the data is not in the source, the page builds it with JavaScript.

All the code below ran as shown, against books.toscrape.com and quotes.toscrape.com, two sites made for practising scraping. The files are ready to run on their own.

1. A static page: requests and BeautifulSoup

requests downloads the page; BeautifulSoup turns the HTML into a tree you can query with CSS selectors.

web-scraping-python/01-requests-beautifulsoup.py
# Fetch a static page with requests and pull fields out of it with BeautifulSoup.
# pip install requests beautifulsoup4
import requests
from bs4 import BeautifulSoup

URL = "https://books.toscrape.com/catalogue/category/books/travel_2/index.html"

response = requests.get(URL, headers={"User-Agent": "my-scraper/1.0 (you@example.com)"}, timeout=30)
response.raise_for_status()
# pass the raw bytes: BeautifulSoup reads the page's own charset (response.text would guess, and "£" becomes "£")
soup = BeautifulSoup(response.content, "html.parser")

books = []
for card in soup.select("article.product_pod"):
    books.append({
        "title": card.h3.a["title"],
        "price": card.select_one(".price_color").get_text(strip=True),
        "in_stock": "In stock" in card.select_one(".availability").get_text(),
    })

print(f"{len(books)} books on the page")
for book in books[:5]:
    print(book)
OutputRan 1 Oct 2026 · Python 3.12.14, beautifulsoup4 4.15.0, requests 2.34.2 · 1.1 s
11 books on the page
{'title': "It's Only the Himalayas", 'price': '£45.17', 'in_stock': True}
{'title': 'Full Moon over Noah’s Ark: An Odyssey to Mount Ararat and Beyond', 'price': '£49.43', 'in_stock': True}
{'title': 'See America: A Celebration of Our National Parks & Treasured Sites', 'price': '£48.87', 'in_stock': True}
{'title': 'Vagabonding: An Uncommon Guide to the Art of Long-Term World Travel', 'price': '£36.94', 'in_stock': True}
{'title': 'Under the Tuscan Sun', 'price': '£37.33', 'in_stock': True}

Three habits in that script save trouble later:

  • Identify yourself. A User-Agent with a contact address tells the site who is calling, and some sites refuse the default one.
  • Fail loudly. raise_for_status() stops on a 403 or 404 instead of parsing an error page as if it were data.
  • Let BeautifulSoup read the bytes. This site does not declare its character set in its headers, so response.text guesses wrong and "£" comes out as "£". Passing response.content lets BeautifulSoup read the encoding from the page itself.

2. Many pages: follow "next" and save a CSV

Listings spread records over several pages. Follow the "next" link until there is none, with a hard cap and a pause between requests.

web-scraping-python/02-pagination-to-csv.py
# Follow "next" links across a listing, politely, and save every record to a CSV file.
# pip install requests beautifulsoup4
import csv
import time
from urllib.parse import urljoin

import requests
from bs4 import BeautifulSoup

url = "https://books.toscrape.com/catalogue/category/books/mystery_3/index.html"
session = requests.Session()
session.headers["User-Agent"] = "my-scraper/1.0 (you@example.com)"
rows, pages = [], 0

while url and pages < 5:  # a hard cap, so a bug cannot crawl forever
    soup = BeautifulSoup(session.get(url, timeout=30).content, "html.parser")
    pages += 1
    for card in soup.select("article.product_pod"):
        rows.append({"title": card.h3.a["title"], "price": card.select_one(".price_color").get_text(strip=True),
                     "url": urljoin(url, card.h3.a["href"])})
    next_link = soup.select_one("li.next a")
    url = urljoin(url, next_link["href"]) if next_link else None
    time.sleep(1)  # one request a second is plenty for a small site

with open("mystery-books.csv", "w", newline="", encoding="utf-8") as f:
    writer = csv.DictWriter(f, fieldnames=["title", "price", "url"])
    writer.writeheader()
    writer.writerows(rows)

print(f"{len(rows)} books from {pages} pages saved to mystery-books.csv")
print(open("mystery-books.csv", encoding="utf-8").read().splitlines()[1])
OutputRan 1 Oct 2026 · Python 3.12.14, beautifulsoup4 4.15.0, requests 2.34.2 · 3.8 s
32 books from 2 pages saved to mystery-books.csv
Sharp Objects,£47.82,https://books.toscrape.com/catalogue/sharp-objects_997/index.html

A Session reuses the connection between requests, urljoin turns relative links into full addresses, and the cap means a bug in the "next" selector cannot run forever. One request a second is plenty for a small site.

3. Pages built with JavaScript: Playwright

quotes.toscrape.com/js sends an almost empty page and builds the quotes in the browser. requests receives none of them; Playwright runs the page in Chromium and sees all ten.

web-scraping-python/03-playwright-javascript-page.py
# A page that builds its content with JavaScript: requests sees none of it, a real browser does.
# pip install playwright    python -m playwright install chromium
import requests
from playwright.sync_api import sync_playwright

URL = "https://quotes.toscrape.com/js/"

html = requests.get(URL, timeout=30).text
print(f"requests: {html.count('class=\"quote\"')} quotes in the HTML it received")

with sync_playwright() as p:
    browser = p.chromium.launch()
    page = browser.new_page()
    page.goto(URL)
    page.wait_for_selector("div.quote")  # wait for the content, not for a fixed number of seconds
    quotes = [
        {"text": q.locator(".text").inner_text(), "author": q.locator(".author").inner_text()}
        for q in page.locator("div.quote").all()
    ]
    browser.close()

print(f"playwright: {len(quotes)} quotes")
print(quotes[0])
OutputRan 1 Oct 2026 · Python 3.12.14, playwright 1.63.0, requests 2.34.2 · 6.0 s
requests: 0 quotes in the HTML it received
playwright: 10 quotes
{'text': '“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”', 'author': 'Albert Einstein'}

Wait for the content you need with wait_for_selector, not for a fixed number of seconds: it is faster when the page is quick and does not break when it is slow. Playwright needs its browser installed once with python -m playwright install chromium.

4. When an API is simpler

Your own code is the cheapest option for a few sites you know. It gets expensive when you need many different sites, pages that block automated readers, prices as seen from another country, or a job that runs on a schedule: then you are maintaining browsers, proxies and parsers. A scraping API does that work per request.

Here is the same JavaScript page through Skryp's extraction API, asking for the fields with a JSON schema:

web-scraping-python/04-skryp-api.py
# The same JavaScript page through a scraping API: one request, and the API decides whether a browser is needed.
# pip install httpx    export SKRYP_API_KEY=...
import os

import httpx

API = os.environ.get("SKRYP_API_URL", "https://api.skryp.dev")
response = httpx.post(
    f"{API}/v1/extract",
    headers={"Authorization": f"Bearer {os.environ['SKRYP_API_KEY']}"},
    json={
        "url": "https://quotes.toscrape.com/js/",
        "schema": {
            "type": "object",
            "properties": {"quotes": {"type": "array", "items": {"type": "object", "properties": {
                "text": {"type": "string"}, "author": {"type": "string"}}}}},
        },
    },
    timeout=120,
).json()

quotes = response["data"]["quotes"]
print(f"{len(quotes)} quotes, extracted by {response['method']}, values checked against the page: {response['checks']}")
print(quotes[0])
receipt = response["receipt"]
credits = receipt["credits"]  # the page's route price (0 when it came from cache) + 5 when a model did the extracting
print(f"route {receipt['route']}, cache {receipt['cache']}, {credits} credit{'' if credits == 1 else 's'}")
OutputRan 1 Oct 2026 · Python 3.12.14, httpx 0.28.1 · 11 s
10 quotes, extracted by model, values checked against the page: {'quotes': 'on_page'}
{'text': 'The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.', 'author': 'Albert Einstein'}
route browser:direct, cache miss, 6 credits

The API decided the page needed a browser (the route on the receipt), returned the quotes as JSON, and checked every value against the page. It cost 6 credits: 1 for the page on Skryp's own network and 5 because a model did the extracting. There is a free monthly allowance, and failed pages cost nothing. See Pricing.

The same page, four ways

This script fetches the JavaScript page with each approach and compares what came back:

web-scraping-python/05-same-page-four-ways.py
# One JavaScript-built page fetched four ways: what each returns, how long it takes, and what fails.
# pip install requests beautifulsoup4 playwright httpx    export SKRYP_API_KEY=...
import os
import time

import httpx
import requests
from bs4 import BeautifulSoup
from playwright.sync_api import sync_playwright

URL = "https://quotes.toscrape.com/js/"
API = os.environ.get("SKRYP_API_URL", "https://api.skryp.dev")
results = []


def run(name, fn):
    t0 = time.monotonic()
    found, size = fn()
    results.append((name, found, size, time.monotonic() - t0))


def plain_requests():
    html = requests.get(URL, timeout=30).text
    return html.count('class="quote"'), len(html)


def beautifulsoup():
    soup = BeautifulSoup(requests.get(URL, timeout=30).text, "html.parser")
    return len(soup.select("div.quote")), len(soup.get_text())


def playwright_browser():
    with sync_playwright() as p:
        browser = p.chromium.launch()
        page = browser.new_page()
        page.goto(URL)
        page.wait_for_selector("div.quote")
        found, size = page.locator("div.quote").count(), len(page.content())
        browser.close()
    return found, size


def scraping_api():
    page = httpx.post(f"{API}/v1/scrape", headers={"Authorization": f"Bearer {os.environ['SKRYP_API_KEY']}"},
                      json={"url": URL, "formats": ["markdown"], "max_age_s": 0}, timeout=120).json()
    return page["markdown"].count("by "), len(page["markdown"])


for name, fn in (("requests", plain_requests), ("requests + BeautifulSoup", beautifulsoup),
                 ("Playwright", playwright_browser), ("Skryp API (Markdown)", scraping_api)):
    run(name, fn)

print(f"{'method':<26}{'quotes found':>13}{'characters':>12}{'seconds':>9}")
for name, found, size, secs in results:
    print(f"{name:<26}{found:>13}{size:>12,}{secs:>9.1f}")
OutputRan 1 Oct 2026 · Python 3.12.14, beautifulsoup4 4.15.0, httpx 0.28.1, playwright 1.63.0, requests 2.34.2 · 10 s
method                     quotes found  characters  seconds
requests                              0       5,806      0.7
requests + BeautifulSoup              0         161      0.7
Playwright                           10       8,940      2.4
Skryp API (Markdown)                 10       1,525      6.4
ApproachGood forWatch out for
requestsFast reads of plain HTML and JSON endpointsGets nothing from pages built by JavaScript
BeautifulSoupPicking fields out of HTML you haveOnly as good as the HTML you give it
PlaywrightAnything a browser can show, clicks and formsSlower; you run and update the browsers
A scraping APIMany sites, other countries, blocked pages, schedulesA cost per page; you depend on a provider

Scraping responsibly

  • Read the site's terms and its robots.txt before you start, and respect what they ask. Some sites offer an official API: use it when it has the data.
  • Go slowly. A pause between requests and a cap on pages keep you from loading a small site.
  • Store only what you need. Avoid collecting personal data unless you have a lawful reason and a plan for it.
  • Cache what you fetched, so a re-run of your parser does not hit the site again.

Where to go next

Questions

How do you scrape a website with Python?
Download the page with requests, parse the HTML with BeautifulSoup and pick out the elements you need with CSS selectors. If the content is built by JavaScript, load the page in Playwright instead. Follow pagination with a pause between requests, save the results to CSV or JSON, and check the site's terms first.
Scrapy or BeautifulSoup?
BeautifulSoup only parses HTML you have already downloaded; Scrapy is a crawling framework that downloads pages, follows links, retries failures and exports results. Use requests and BeautifulSoup for a handful of pages, and Scrapy when you crawl whole sites.
Is Python or JavaScript better for web scraping?
Both work. Python has requests, BeautifulSoup, Scrapy and Playwright; JavaScript has Cheerio, Puppeteer and Playwright and suits teams already on Node.js. Use the language your project already uses: the browser tools behave much the same in both.
What is web scraping in Python?
Writing Python code that downloads web pages and extracts data from them, usually into a spreadsheet, a database or an AI model's context. The usual tools are requests to download, BeautifulSoup to parse, and Playwright or Selenium when a page needs a real browser.