# Web scraping with Selenium and Python

Selenium drives a real browser, so it can scrape pages that only render with JavaScript. This guide sets it up, waits for content correctly, follows pagination and extracts data, and shows when a lighter tool or an API does the same job for less.

Updated 1 Oct 2026 · Tested on 1 Oct 2026 · https://skryp.dev/guides/selenium-web-scraping

## Set up

```bash
pip install selenium
```

That is all: since Selenium 4.6, Selenium Manager downloads the ChromeDriver that matches your installed Chrome the first time you start a browser. The examples run Chrome headless (no window); remove the `--headless=new` option to watch them work.

Every example below ran as shown, against [quotes.toscrape.com](https://quotes.toscrape.com), a site made for practising scraping.

## 1. Wait for the content, then extract it

The [JavaScript version](https://quotes.toscrape.com/js/) of the quotes site builds its content in the browser, so a plain HTTP request gets none of it. Selenium loads the page in Chrome; an explicit wait holds until the quotes exist.

```python
# Open a JavaScript-built page in Chrome with Selenium, wait for the content, and extract it.
# pip install selenium   (Selenium Manager downloads the matching ChromeDriver on first run)
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait

options = webdriver.ChromeOptions()
options.add_argument("--headless=new")
driver = webdriver.Chrome(options=options)
try:
    driver.get("https://quotes.toscrape.com/js/")
    # wait until the quotes exist, up to 10 seconds; never sleep for a fixed time
    WebDriverWait(driver, 10).until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "div.quote")))
    quotes = [
        {"text": q.find_element(By.CSS_SELECTOR, ".text").text, "author": q.find_element(By.CSS_SELECTOR, ".author").text}
        for q in driver.find_elements(By.CSS_SELECTOR, "div.quote")
    ]
finally:
    driver.quit()  # always close the browser, even when something fails

print(f"{len(quotes)} quotes")
print(quotes[0])
```

Output (ran 1 Oct 2026; Python 3.12.14, selenium 4.50.0; 5.4 s):

```text
10 quotes
{'text': '“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”', 'author': 'Albert Einstein'}
```

`WebDriverWait` with `presence_of_all_elements_located` waits as long as needed and no longer, up to the limit you set. A fixed `time.sleep()` either wastes time or fails on a slow day. The `try`/`finally` makes sure the browser closes even when something goes wrong; a scraper that leaks browsers runs out of memory.

## 2. Click through pages

To follow pagination, click "Next" and wait for the old page to disappear before reading the new one. Reading too early returns the previous page's quotes twice.

```python
# Click "Next" through a JavaScript listing, waiting for each new page to replace the old one.
# pip install selenium
from selenium import webdriver
from selenium.common.exceptions import NoSuchElementException
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait

options = webdriver.ChromeOptions()
options.add_argument("--headless=new")
driver = webdriver.Chrome(options=options)
wait = WebDriverWait(driver, 10)
authors = []
try:
    driver.get("https://quotes.toscrape.com/js/")
    for page in range(1, 4):  # first three pages
        wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "div.quote")))
        authors += [a.text for a in driver.find_elements(By.CSS_SELECTOR, "small.author")]
        print(f"page {page}: {len(authors)} quotes so far")
        try:
            first = driver.find_element(By.CSS_SELECTOR, "div.quote")
            driver.find_element(By.CSS_SELECTOR, "li.next a").click()
            wait.until(EC.staleness_of(first))  # the old page is gone: the next one is loading
        except NoSuchElementException:
            break  # no "Next" link: last page
finally:
    driver.quit()

print(f"{len(set(authors))} different authors, e.g. {sorted(set(authors))[:3]}")
```

Output (ran 1 Oct 2026; Python 3.12.14, selenium 4.50.0; 4.3 s):

```text
page 1: 10 quotes so far
page 2: 20 quotes so far
page 3: 30 quotes so far
20 different authors, e.g. ['Albert Einstein', 'Allen Saunders', 'André Gide']
```

`staleness_of(first)` waits until an element from the old page is gone from the document, which is a reliable sign that the next page has replaced it.

## 3. Fill in and submit a form

Selenium can do what a visitor does: choose options, type, submit. The site's [search page](https://quotes.toscrape.com/search.aspx) has two linked lists: picking an author reloads the page with that author's tags.

```python
# Choose options in a search form and submit it, as a person would.
# pip install selenium
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import Select, WebDriverWait

options = webdriver.ChromeOptions()
options.add_argument("--headless=new")
options.add_argument("--window-size=1280,900")  # headless windows are small; a covered button cannot be clicked
driver = webdriver.Chrome(options=options)
wait = WebDriverWait(driver, 10)
try:
    driver.get("https://quotes.toscrape.com/search.aspx")
    Select(driver.find_element(By.ID, "author")).select_by_visible_text("Albert Einstein")
    # choosing an author reloads the page with that author's tags in the second list
    wait.until(lambda d: len(Select(d.find_element(By.ID, "tag")).options) > 1)
    Select(driver.find_element(By.ID, "tag")).select_by_visible_text("inspirational")
    wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, "input[type=submit]"))).click()
    results = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "div.quote span.content")))
    texts = [r.text for r in results]
finally:
    driver.quit()

print(f"{len(texts)} {'quote' if len(texts) == 1 else 'quotes'} by Albert Einstein tagged inspirational")
print(texts[0][:100])
```

Output (ran 1 Oct 2026; Python 3.12.14, selenium 4.50.0; 3.6 s):

```text
1 quote by Albert Einstein tagged inspirational
“There are only two ways to live your life. One is as though nothing is a miracle. The other is as t
```

The first version of this script failed with `ElementClickInterceptedException`: in a small headless window, another element covered the submit button. Setting a window size and waiting until the button is clickable fixed it.

## Errors you will meet

| Error | What it usually means | Fix |
| --- | --- | --- |
| `NoSuchElementException` | The element is not there yet, or the selector is wrong | Wait for it with `WebDriverWait`; check the selector in the browser's developer tools |
| `TimeoutException` | The wait ran out: the content never appeared | Check the page in a visible browser; it may need a click, a scroll or a sign-in |
| `ElementClickInterceptedException` | Something covers the element: a banner, a header, a small window | Set a window size, close the banner, wait for `element_to_be_clickable` |
| `StaleElementReferenceException` | The page changed after you found the element | Find it again after each navigation or reload |

## When Selenium is overkill

A browser is the heaviest tool you can use. Before starting one, look at the page source: many pages that render with JavaScript ship their data in it. The quotes page carries every quote as a JSON array in a script tag, so `requests` and a regular expression read the same data in under a second:

```python
# Before reaching for Selenium, look in the page source: this "JavaScript" page carries its data as JSON.
# pip install requests
import json
import re

import requests

html = requests.get("https://quotes.toscrape.com/js/", timeout=30).text
raw = re.search(r"var data = (\[.*?\]);", html, re.S).group(1)
quotes = json.loads(raw)

print(f"{len(quotes)} quotes straight from the page source, no browser")
print({"text": quotes[0]["text"], "author": quotes[0]["author"]["name"]})
```

Output (ran 1 Oct 2026; Python 3.12.14, requests 2.34.2; 0.8 s):

```text
10 quotes straight from the page source, no browser
{'text': '“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”', 'author': 'Albert Einstein'}
```

Also check the browser's network tab for the JSON request the page makes; calling that directly is often simpler still. And when you need many sites, other countries or pages that block automated browsers, a scraping API runs the browser for you: see [step 4 of the Python guide](https://skryp.dev/guides/web-scraping-python#4-when-an-api-is-simpler) and [Skryp's live browser](https://skryp.dev/product/browser).

## Selenium, Playwright or Puppeteer?

Selenium has the widest language support and the longest history. Playwright waits for elements automatically and drives Chromium, Firefox and WebKit from one API; Puppeteer is the lightest option for Chrome in Node.js. We ran the same pages through all three in [Playwright vs Puppeteer vs Selenium](https://skryp.dev/guides/playwright-vs-puppeteer-vs-selenium).

## Questions

**How do you use Selenium for web scraping?**

Install it with pip install selenium, start Chrome with webdriver.Chrome(), open the page with driver.get(), wait for the elements you need with WebDriverWait, then read them with find_elements and CSS selectors. Close the browser with driver.quit() when you are done.

**Is Selenium good for web scraping?**

It is good for pages that need a real browser: content built by JavaScript, clicks, forms and pagination. It is slower and heavier than a plain HTTP request, so check the page source and network requests first; many pages do not need a browser at all.

**Scrapy or Selenium?**

Scrapy is a fast crawling framework for pages whose data is in the HTML; Selenium drives a browser for pages that need JavaScript or interaction. Use Scrapy for crawling many plain pages and Selenium (or Playwright) for the pages that need a browser. They can be combined.

**Selenium or BeautifulSoup?**

They do different jobs. BeautifulSoup parses HTML you already have; Selenium loads pages in a browser and can click and type. For plain HTML use requests and BeautifulSoup; for pages built by JavaScript use Selenium, and you can still hand driver.page_source to BeautifulSoup to parse.
