Web crawling vs web scraping: what's the difference?
Updated 1 Oct 2026 · Tested on 1 Oct 2026
Web crawling finds pages by following links or sitemaps; web scraping extracts data from the pages you have. Most data projects do both: crawl to discover the URLs, then scrape each one. The example below maps a small site and then extracts the same fields from every page.
The difference
| Web crawling | Web scraping | |
|---|---|---|
| Goal | Find pages | Get data out of pages |
| Starts from | One address, a sitemap or a domain | Addresses you already have |
| Produces | A list of URLs | Records: names, prices, dates, text |
| Typical tools | Scrapy, Crawlee, a site map or crawl API, search engines' bots | BeautifulSoup, Cheerio, Playwright, an extraction API |
| What goes wrong | Loops, duplicate URLs, crawling too fast, wandering off the site | Layout changes, pages built by JavaScript, missing fields |
Search engines are crawlers: Googlebot and Bingbot follow links across the web to find pages to index. A price tracker is a scraper: it reads the same product pages again and pulls out the price. Most data projects need a bit of both.
All the code below ran as shown, against quotes.toscrape.com and books.toscrape.com, sites made for practising scraping.
Crawling: discover the pages
A crawler starts somewhere, reads the links on each page, and queues the ones it has not seen. The result is addresses, not data:
# Crawling: start at one page and follow links to discover others. The output is a list of URLs, not data.
# pip install requests beautifulsoup4
import time
from collections import deque
from urllib.parse import urljoin, urlparse
import requests
from bs4 import BeautifulSoup
START = "https://quotes.toscrape.com/"
seen, queue, found = {START}, deque([START]), []
while queue and len(found) < 25: # stop after 25 pages
url = queue.popleft()
soup = BeautifulSoup(requests.get(url, timeout=30).content, "html.parser")
found.append(url)
for a in soup.select("a[href]"):
link = urljoin(url, a["href"]).split("#")[0]
if urlparse(link).netloc == urlparse(START).netloc and link not in seen and "/login" not in link:
seen.add(link)
queue.append(link)
time.sleep(0.5)
authors = [u for u in found if "/author/" in u]
print(f"crawled {len(found)} pages and saw {len(seen)} links on the site")
print(f"{len(authors)} of the crawled pages are author pages, e.g. {authors[:2]}")crawled 25 pages and saw 116 links on the site
5 of the crawled pages are author pages, e.g. ['https://quotes.toscrape.com/author/Albert-Einstein', 'https://quotes.toscrape.com/author/J-K-Rowling']Three rules keep a crawler out of trouble: stay on the site you meant to crawl, never queue an address twice, and stop at a limit you chose in advance. Pause between requests, and read the site's robots.txt before you start: it says which paths the owner does not want crawled. Many sites also publish a sitemap that lists their pages directly, which is faster and lighter than following links.
Scraping: get the data out
A scraper takes known pages and extracts the same fields from each. Here, three of the author pages the crawler found:
# Scraping: take pages you already know and extract fields from each. The output is data.
# pip install requests beautifulsoup4
import time
import requests
from bs4 import BeautifulSoup
AUTHOR_PAGES = [
"https://quotes.toscrape.com/author/Albert-Einstein/",
"https://quotes.toscrape.com/author/J-K-Rowling/",
"https://quotes.toscrape.com/author/Jane-Austen/",
]
for url in AUTHOR_PAGES:
soup = BeautifulSoup(requests.get(url, timeout=30).content, "html.parser")
print({
"name": soup.select_one("h3.author-title").get_text(strip=True),
"born": soup.select_one(".author-born-date").get_text(strip=True),
"born_in": soup.select_one(".author-born-location").get_text(strip=True).removeprefix("in "),
})
time.sleep(0.5){'name': 'Albert Einstein', 'born': 'March 14, 1879', 'born_in': 'Ulm, Germany'}
{'name': 'J.K. Rowling', 'born': 'July 31, 1965', 'born_in': 'Yate, South Gloucestershire, England, The United Kingdom'}
{'name': 'Jane Austen', 'born': 'December 16, 1775', 'born_in': 'Steventon Rectory, Hampshire, The United Kingdom'}The selectors (h3.author-title, .author-born-date) are the fragile part: if the site changes its layout, they find nothing. That is why scrapers need monitoring that notices when a field comes back empty.
Both together
Through an API, crawling and scraping are two calls. Map lists a site's URLs from its sitemaps and pages in one request; Extract then pulls the same fields from each page:
# Both together through an API: map a site's URLs (crawling), then extract the same fields from each (scraping).
# pip install httpx export SKRYP_API_KEY=...
import os
import httpx
API = os.environ.get("SKRYP_API_URL", "https://api.skryp.dev")
auth = {"Authorization": f"Bearer {os.environ['SKRYP_API_KEY']}"}
site = httpx.post(f"{API}/v1/map", headers=auth, json={"url": "https://books.toscrape.com/", "limit": 200}, timeout=120).json()
urls = [link["url"] for link in site["links"]] # each link also says where it was found: the sitemap or a page
books = [u for u in urls if "/catalogue/" in u and u.endswith("/index.html") and "/category/" not in u]
print(f"map: {len(urls)} URLs from the home page and sitemaps, {len(books)} of them book pages")
schema = {"type": "object", "properties": {"title": {"type": "string"}, "price": {"type": "string"}, "availability": {"type": "string"}}}
for url in books[:4]:
page = httpx.post(f"{API}/v1/extract", headers=auth, json={"url": url, "schema": schema}, timeout=120).json()
credits = page["receipt"]["credits"] # the page, plus 5 when a model read it; learned selectors cost nothing extra
print(page["data"], f"({page['method']}, {credits} credit{'' if credits == 1 else 's'})")map: 73 URLs from the home page and sitemaps, 20 of them book pages
{'title': 'A Light in the Attic', 'price': '£51.77', 'availability': 'In stock (22 available)'} (learned_selectors, 1 credit)
{'title': 'Tipping the Velvet', 'price': '£53.74', 'availability': 'In stock (20 available)'} (learned_selectors, 1 credit)
{'title': 'Soumission', 'price': '£50.10', 'availability': 'In stock (20 available)'} (learned_selectors, 1 credit)
{'title': 'Sharp Objects', 'price': '£47.82', 'availability': 'In stock (20 available)'} (learned_selectors, 1 credit)Extraction learns the page layout the first time a model reads it, which adds 5 credits to that page; after that, pages with the same layout are extracted by the learned selectors at no extra cost. Every page above shows learned_selectors and 1 credit, because an earlier run had already learned this layout. To list a site's pages without writing code, try Find all pages on a website.
Which one you need
- Only crawling when the question is about the site itself: how many pages it has, which ones are broken, what a sitemap is missing.
- Only scraping when you already know the pages: a list of product URLs, a set of competitors' pricing pages.
- Both when you need data from pages you have not found yet: every product in a category, every article in a section. Crawl first, filter the URLs, then scrape the ones you want.
Whichever you run, go slowly, respect the site's terms and robots.txt, and avoid collecting personal data you do not need.
Questions
- What is a web crawler?
- A program that discovers web pages by following links from page to page, or by reading sitemaps. Search engines run crawlers such as Googlebot to find pages to index; data projects run them to list the pages they then scrape.
- What is the difference between a crawler and a scraper?
- A crawler finds pages and produces a list of URLs; a scraper reads pages and produces data. Crawl when you need to discover pages, scrape when you need what is on them, and do both when you need data from pages you have not found yet.
- How do you crawl a website?
- Start from its sitemap if it has one; otherwise start at the home page, collect its links, keep the ones on the same site, and repeat for each new page you have not seen, up to a limit. Pause between requests and respect robots.txt.