Web scraping with Selenium and Python
Updated 1 Oct 2026 · Tested on 1 Oct 2026
Selenium drives a real browser, so it can scrape pages that only render with JavaScript. This guide sets it up, waits for content correctly, follows pagination and extracts data, and shows when a lighter tool or an API does the same job for less.
Set up
pip install seleniumThat is all: since Selenium 4.6, Selenium Manager downloads the ChromeDriver that matches your installed Chrome the first time you start a browser. The examples run Chrome headless (no window); remove the --headless=new option to watch them work.
Every example below ran as shown, against quotes.toscrape.com, a site made for practising scraping.
1. Wait for the content, then extract it
The JavaScript version of the quotes site builds its content in the browser, so a plain HTTP request gets none of it. Selenium loads the page in Chrome; an explicit wait holds until the quotes exist.
# Open a JavaScript-built page in Chrome with Selenium, wait for the content, and extract it.
# pip install selenium (Selenium Manager downloads the matching ChromeDriver on first run)
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
options = webdriver.ChromeOptions()
options.add_argument("--headless=new")
driver = webdriver.Chrome(options=options)
try:
driver.get("https://quotes.toscrape.com/js/")
# wait until the quotes exist, up to 10 seconds; never sleep for a fixed time
WebDriverWait(driver, 10).until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "div.quote")))
quotes = [
{"text": q.find_element(By.CSS_SELECTOR, ".text").text, "author": q.find_element(By.CSS_SELECTOR, ".author").text}
for q in driver.find_elements(By.CSS_SELECTOR, "div.quote")
]
finally:
driver.quit() # always close the browser, even when something fails
print(f"{len(quotes)} quotes")
print(quotes[0])10 quotes
{'text': '“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”', 'author': 'Albert Einstein'}WebDriverWait with presence_of_all_elements_located waits as long as needed and no longer, up to the limit you set. A fixed time.sleep() either wastes time or fails on a slow day. The try/finally makes sure the browser closes even when something goes wrong; a scraper that leaks browsers runs out of memory.
2. Click through pages
To follow pagination, click "Next" and wait for the old page to disappear before reading the new one. Reading too early returns the previous page's quotes twice.
# Click "Next" through a JavaScript listing, waiting for each new page to replace the old one.
# pip install selenium
from selenium import webdriver
from selenium.common.exceptions import NoSuchElementException
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
options = webdriver.ChromeOptions()
options.add_argument("--headless=new")
driver = webdriver.Chrome(options=options)
wait = WebDriverWait(driver, 10)
authors = []
try:
driver.get("https://quotes.toscrape.com/js/")
for page in range(1, 4): # first three pages
wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "div.quote")))
authors += [a.text for a in driver.find_elements(By.CSS_SELECTOR, "small.author")]
print(f"page {page}: {len(authors)} quotes so far")
try:
first = driver.find_element(By.CSS_SELECTOR, "div.quote")
driver.find_element(By.CSS_SELECTOR, "li.next a").click()
wait.until(EC.staleness_of(first)) # the old page is gone: the next one is loading
except NoSuchElementException:
break # no "Next" link: last page
finally:
driver.quit()
print(f"{len(set(authors))} different authors, e.g. {sorted(set(authors))[:3]}")page 1: 10 quotes so far
page 2: 20 quotes so far
page 3: 30 quotes so far
20 different authors, e.g. ['Albert Einstein', 'Allen Saunders', 'André Gide']staleness_of(first) waits until an element from the old page is gone from the document, which is a reliable sign that the next page has replaced it.
3. Fill in and submit a form
Selenium can do what a visitor does: choose options, type, submit. The site's search page has two linked lists: picking an author reloads the page with that author's tags.
# Choose options in a search form and submit it, as a person would.
# pip install selenium
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import Select, WebDriverWait
options = webdriver.ChromeOptions()
options.add_argument("--headless=new")
options.add_argument("--window-size=1280,900") # headless windows are small; a covered button cannot be clicked
driver = webdriver.Chrome(options=options)
wait = WebDriverWait(driver, 10)
try:
driver.get("https://quotes.toscrape.com/search.aspx")
Select(driver.find_element(By.ID, "author")).select_by_visible_text("Albert Einstein")
# choosing an author reloads the page with that author's tags in the second list
wait.until(lambda d: len(Select(d.find_element(By.ID, "tag")).options) > 1)
Select(driver.find_element(By.ID, "tag")).select_by_visible_text("inspirational")
wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, "input[type=submit]"))).click()
results = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, "div.quote span.content")))
texts = [r.text for r in results]
finally:
driver.quit()
print(f"{len(texts)} {'quote' if len(texts) == 1 else 'quotes'} by Albert Einstein tagged inspirational")
print(texts[0][:100])1 quote by Albert Einstein tagged inspirational
“There are only two ways to live your life. One is as though nothing is a miracle. The other is as tThe first version of this script failed with ElementClickInterceptedException: in a small headless window, another element covered the submit button. Setting a window size and waiting until the button is clickable fixed it.
Errors you will meet
| Error | What it usually means | Fix |
|---|---|---|
NoSuchElementException | The element is not there yet, or the selector is wrong | Wait for it with WebDriverWait; check the selector in the browser's developer tools |
TimeoutException | The wait ran out: the content never appeared | Check the page in a visible browser; it may need a click, a scroll or a sign-in |
ElementClickInterceptedException | Something covers the element: a banner, a header, a small window | Set a window size, close the banner, wait for element_to_be_clickable |
StaleElementReferenceException | The page changed after you found the element | Find it again after each navigation or reload |
When Selenium is overkill
A browser is the heaviest tool you can use. Before starting one, look at the page source: many pages that render with JavaScript ship their data in it. The quotes page carries every quote as a JSON array in a script tag, so requests and a regular expression read the same data in under a second:
# Before reaching for Selenium, look in the page source: this "JavaScript" page carries its data as JSON.
# pip install requests
import json
import re
import requests
html = requests.get("https://quotes.toscrape.com/js/", timeout=30).text
raw = re.search(r"var data = (\[.*?\]);", html, re.S).group(1)
quotes = json.loads(raw)
print(f"{len(quotes)} quotes straight from the page source, no browser")
print({"text": quotes[0]["text"], "author": quotes[0]["author"]["name"]})10 quotes straight from the page source, no browser
{'text': '“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”', 'author': 'Albert Einstein'}Also check the browser's network tab for the JSON request the page makes; calling that directly is often simpler still. And when you need many sites, other countries or pages that block automated browsers, a scraping API runs the browser for you: see step 4 of the Python guide and Skryp's live browser.
Selenium, Playwright or Puppeteer?
Selenium has the widest language support and the longest history. Playwright waits for elements automatically and drives Chromium, Firefox and WebKit from one API; Puppeteer is the lightest option for Chrome in Node.js. We ran the same pages through all three in Playwright vs Puppeteer vs Selenium.
Questions
- How do you use Selenium for web scraping?
- Install it with pip install selenium, start Chrome with webdriver.Chrome(), open the page with driver.get(), wait for the elements you need with WebDriverWait, then read them with find_elements and CSS selectors. Close the browser with driver.quit() when you are done.
- Is Selenium good for web scraping?
- It is good for pages that need a real browser: content built by JavaScript, clicks, forms and pagination. It is slower and heavier than a plain HTTP request, so check the page source and network requests first; many pages do not need a browser at all.
- Scrapy or Selenium?
- Scrapy is a fast crawling framework for pages whose data is in the HTML; Selenium drives a browser for pages that need JavaScript or interaction. Use Scrapy for crawling many plain pages and Selenium (or Playwright) for the pages that need a browser. They can be combined.
- Selenium or BeautifulSoup?
- They do different jobs. BeautifulSoup parses HTML you already have; Selenium loads pages in a browser and can click and type. For plain HTML use requests and BeautifulSoup; for pages built by JavaScript use Selenium, and you can still hand driver.page_source to BeautifulSoup to parse.