# Web crawling vs web scraping: what's the difference?

Web crawling finds pages by following links or sitemaps; web scraping extracts data from the pages you have. Most data projects do both: crawl to discover the URLs, then scrape each one. The example below maps a small site and then extracts the same fields from every page.

Updated 1 Oct 2026 · Tested on 1 Oct 2026 · https://skryp.dev/guides/web-crawling-vs-web-scraping

## The difference

| | Web crawling | Web scraping |
| --- | --- | --- |
| Goal | Find pages | Get data out of pages |
| Starts from | One address, a sitemap or a domain | Addresses you already have |
| Produces | A list of URLs | Records: names, prices, dates, text |
| Typical tools | Scrapy, Crawlee, a site map or crawl API, search engines' bots | BeautifulSoup, Cheerio, Playwright, an extraction API |
| What goes wrong | Loops, duplicate URLs, crawling too fast, wandering off the site | Layout changes, pages built by JavaScript, missing fields |

Search engines are crawlers: Googlebot and Bingbot follow links across the web to find pages to index. A price tracker is a scraper: it reads the same product pages again and pulls out the price. Most data projects need a bit of both.

All the code below ran as shown, against [quotes.toscrape.com](https://quotes.toscrape.com) and [books.toscrape.com](https://books.toscrape.com), sites made for practising scraping.

## Crawling: discover the pages

A crawler starts somewhere, reads the links on each page, and queues the ones it has not seen. The result is addresses, not data:

```python
# Crawling: start at one page and follow links to discover others. The output is a list of URLs, not data.
# pip install requests beautifulsoup4
import time
from collections import deque
from urllib.parse import urljoin, urlparse

import requests
from bs4 import BeautifulSoup

START = "https://quotes.toscrape.com/"
seen, queue, found = {START}, deque([START]), []

while queue and len(found) < 25:  # stop after 25 pages
    url = queue.popleft()
    soup = BeautifulSoup(requests.get(url, timeout=30).content, "html.parser")
    found.append(url)
    for a in soup.select("a[href]"):
        link = urljoin(url, a["href"]).split("#")[0]
        if urlparse(link).netloc == urlparse(START).netloc and link not in seen and "/login" not in link:
            seen.add(link)
            queue.append(link)
    time.sleep(0.5)

authors = [u for u in found if "/author/" in u]
print(f"crawled {len(found)} pages and saw {len(seen)} links on the site")
print(f"{len(authors)} of the crawled pages are author pages, e.g. {authors[:2]}")
```

Output (ran 1 Oct 2026; Python 3.12.14, beautifulsoup4 4.15.0, requests 2.34.2; 34.0 s):

```text
crawled 25 pages and saw 116 links on the site
5 of the crawled pages are author pages, e.g. ['https://quotes.toscrape.com/author/Albert-Einstein', 'https://quotes.toscrape.com/author/J-K-Rowling']
```

Three rules keep a crawler out of trouble: stay on the site you meant to crawl, never queue an address twice, and stop at a limit you chose in advance. Pause between requests, and read the site's `robots.txt` before you start: it says which paths the owner does not want crawled. Many sites also publish a sitemap that lists their pages directly, which is faster and lighter than following links.

## Scraping: get the data out

A scraper takes known pages and extracts the same fields from each. Here, three of the author pages the crawler found:

```python
# Scraping: take pages you already know and extract fields from each. The output is data.
# pip install requests beautifulsoup4
import time

import requests
from bs4 import BeautifulSoup

AUTHOR_PAGES = [
    "https://quotes.toscrape.com/author/Albert-Einstein/",
    "https://quotes.toscrape.com/author/J-K-Rowling/",
    "https://quotes.toscrape.com/author/Jane-Austen/",
]

for url in AUTHOR_PAGES:
    soup = BeautifulSoup(requests.get(url, timeout=30).content, "html.parser")
    print({
        "name": soup.select_one("h3.author-title").get_text(strip=True),
        "born": soup.select_one(".author-born-date").get_text(strip=True),
        "born_in": soup.select_one(".author-born-location").get_text(strip=True).removeprefix("in "),
    })
    time.sleep(0.5)
```

Output (ran 1 Oct 2026; Python 3.12.14, beautifulsoup4 4.15.0, requests 2.34.2; 3.8 s):

```text
{'name': 'Albert Einstein', 'born': 'March 14, 1879', 'born_in': 'Ulm, Germany'}
{'name': 'J.K. Rowling', 'born': 'July 31, 1965', 'born_in': 'Yate, South Gloucestershire, England, The United Kingdom'}
{'name': 'Jane Austen', 'born': 'December 16, 1775', 'born_in': 'Steventon Rectory, Hampshire, The United Kingdom'}
```

The selectors (`h3.author-title`, `.author-born-date`) are the fragile part: if the site changes its layout, they find nothing. That is why scrapers need monitoring that notices when a field comes back empty.

## Both together

Through an API, crawling and scraping are two calls. [Map](https://skryp.dev/product/crawl) lists a site's URLs from its sitemaps and pages in one request; [Extract](https://skryp.dev/product/extract) then pulls the same fields from each page:

```python
# Both together through an API: map a site's URLs (crawling), then extract the same fields from each (scraping).
# pip install httpx    export SKRYP_API_KEY=...
import os

import httpx

API = os.environ.get("SKRYP_API_URL", "https://api.skryp.dev")
auth = {"Authorization": f"Bearer {os.environ['SKRYP_API_KEY']}"}

site = httpx.post(f"{API}/v1/map", headers=auth, json={"url": "https://books.toscrape.com/", "limit": 200}, timeout=120).json()
urls = [link["url"] for link in site["links"]]  # each link also says where it was found: the sitemap or a page
books = [u for u in urls if "/catalogue/" in u and u.endswith("/index.html") and "/category/" not in u]
print(f"map: {len(urls)} URLs from the home page and sitemaps, {len(books)} of them book pages")

schema = {"type": "object", "properties": {"title": {"type": "string"}, "price": {"type": "string"}, "availability": {"type": "string"}}}
for url in books[:4]:
    page = httpx.post(f"{API}/v1/extract", headers=auth, json={"url": url, "schema": schema}, timeout=120).json()
    credits = page["receipt"]["credits"]  # the page, plus 5 when a model read it; learned selectors cost nothing extra
    print(page["data"], f"({page['method']}, {credits} credit{'' if credits == 1 else 's'})")
```

Output (ran 1 Oct 2026; Python 3.12.14, httpx 0.28.1; 4.6 s):

```text
map: 73 URLs from the home page and sitemaps, 20 of them book pages
{'title': 'A Light in the Attic', 'price': '£51.77', 'availability': 'In stock (22 available)'} (learned_selectors, 1 credit)
{'title': 'Tipping the Velvet', 'price': '£53.74', 'availability': 'In stock (20 available)'} (learned_selectors, 1 credit)
{'title': 'Soumission', 'price': '£50.10', 'availability': 'In stock (20 available)'} (learned_selectors, 1 credit)
{'title': 'Sharp Objects', 'price': '£47.82', 'availability': 'In stock (20 available)'} (learned_selectors, 1 credit)
```

Extraction learns the page layout the first time a model reads it, which adds 5 credits to that page; after that, pages with the same layout are extracted by the learned selectors at no extra cost. Every page above shows `learned_selectors` and 1 credit, because an earlier run had already learned this layout. To list a site's pages without writing code, try [Find all pages on a website](https://skryp.dev/tools/find-all-pages).

## Which one you need

- **Only crawling** when the question is about the site itself: how many pages it has, which ones are broken, what a sitemap is missing.
- **Only scraping** when you already know the pages: a list of product URLs, a set of competitors' pricing pages.
- **Both** when you need data from pages you have not found yet: every product in a category, every article in a section. Crawl first, filter the URLs, then scrape the ones you want.

Whichever you run, go slowly, respect the site's terms and `robots.txt`, and avoid collecting personal data you do not need.

## Questions

**What is a web crawler?**

A program that discovers web pages by following links from page to page, or by reading sitemaps. Search engines run crawlers such as Googlebot to find pages to index; data projects run them to list the pages they then scrape.

**What is the difference between a crawler and a scraper?**

A crawler finds pages and produces a list of URLs; a scraper reads pages and produces data. Crawl when you need to discover pages, scrape when you need what is on them, and do both when you need data from pages you have not found yet.

**How do you crawl a website?**

Start from its sitemap if it has one; otherwise start at the home page, collect its links, keep the ones on the same site, and repeat for each new page you have not seen, up to a limit. Pause between requests and respect robots.txt.
