# Web scraping vs an official API: which should you use?

Use a site's official API when it exposes the data you need at a price and rate you can live with; scrape when it doesn't, or when you need what users see, such as country-specific prices. Many projects use both. The decision table below covers coverage, cost, limits, terms and maintenance.

Updated 1 Oct 2026 · Tested on 1 Oct 2026 · https://skryp.dev/guides/web-scraping-vs-api

## The decision in one table

| | Official API | Scraping |
| --- | --- | --- |
| Data available | What the site chose to publish | Whatever the page shows |
| What users see (ranking, local prices, layout) | Often not included | Yes, as seen from where you fetch it |
| Format | Structured JSON with stable IDs | HTML you have to parse |
| Breaks when | The API version is retired, usually with notice | The page layout changes, without notice |
| Limits | Rate limits, quotas, sometimes paid tiers | Bot protection; your own politeness |
| Terms | Spelled out in the API's terms | Set by the site's terms and the law where you operate |
| Effort | Read the docs, get a key | Write and maintain a parser, or use a scraping service |

Use the official API when it has the data you need at a price and rate you can live with: it is more stable and its terms are clear. Scrape when it doesn't: the data is not in the API, the API costs more than the data is worth, or you need what a visitor sees, such as the prices shown in a particular country.

## The same data both ways

[Hacker News](https://news.ycombinator.com) has both: an official API and a front page. Here are the top five stories through the API, then from the page, fetched seconds apart. Both ran as shown.

```python
# Hacker News publishes an official API: the top stories as structured JSON, with stable ids and exact fields.
# pip install requests
import requests

API = "https://hacker-news.firebaseio.com/v0"
ids = requests.get(f"{API}/topstories.json", timeout=30).json()[:5]
for i in ids:
    item = requests.get(f"{API}/item/{i}.json", timeout=30).json()
    print({"id": item["id"], "title": item["title"][:60], "points": item["score"], "comments": item.get("descendants", 0), "time": item["time"]})
```

Output (ran 1 Oct 2026; Python 3.12.14, requests 2.34.2; 3.4 s):

```text
{'id': 49923692, 'title': 'Clef: our open-source decision models', 'points': 68, 'comments': 15, 'time': 1790871537}
{'id': 49923466, 'title': 'RIP, vector database', 'points': 63, 'comments': 16, 'time': 1790870516}
{'id': 49920160, 'title': 'StreetComplete on iOS is now in public beta', 'points': 390, 'comments': 82, 'time': 1790852397}
{'id': 49922515, 'title': 'RacketCon Is Saturday', 'points': 52, 'comments': 14, 'time': 1790866693}
{'id': 49920896, 'title': 'How to speed up the Rust compiler in September 2026', 'points': 166, 'comments': 80, 'time': 1790858678}
```

```python
# The same stories scraped from the front page: what a reader sees, in the order they see it.
# pip install requests beautifulsoup4
import requests
from bs4 import BeautifulSoup

soup = BeautifulSoup(requests.get("https://news.ycombinator.com/", headers={"User-Agent": "my-scraper/1.0 (you@example.com)"}, timeout=30).content, "html.parser")
for row in soup.select("tr.athing")[:5]:
    sub = row.find_next_sibling("tr")
    points = sub.select_one(".score")
    print({
        "id": int(row["id"]),
        "rank": int(row.select_one(".rank").get_text(strip=True).rstrip(".")),
        "title": row.select_one(".titleline > a").get_text(strip=True)[:60],
        "points": int(points.get_text().split()[0]) if points else None,
        "age": sub.select_one(".age").get_text(strip=True),
    })
```

Output (ran 1 Oct 2026; Python 3.12.14, beautifulsoup4 4.15.0, requests 2.34.2; 1.3 s):

```text
{'id': 49923692, 'rank': 1, 'title': 'Clef: our open-source decision models', 'points': 63, 'age': '46 minutes ago'}
{'id': 49923466, 'rank': 2, 'title': 'RIP, vector database', 'points': 63, 'age': '1 hour ago'}
{'id': 49920160, 'rank': 3, 'title': 'StreetComplete on iOS is now in public beta', 'points': 389, 'age': '6 hours ago'}
{'id': 49922515, 'rank': 4, 'title': 'RacketCon Is Saturday', 'points': 51, 'age': '2 hours ago'}
{'id': 49920896, 'rank': 5, 'title': 'How to speed up the Rust compiler in September 2026', 'points': 165, 'age': '4 hours ago'}
```

The same five stories came back in the same order, with the same IDs. The differences are the ones in the table:

- **The API gives exact fields.** Points, comment counts and a precise timestamp, as numbers. No HTML to parse, nothing to break when the design changes.
- **The page gives what a reader sees.** The rank and the relative age shown to readers are presentation, which the API leaves you to work out.
- **They disagreed.** Fetched seconds apart, the page showed a few points fewer than the API on several stories. The two are served differently, so they were not equally fresh. When you combine sources, record when and where each value came from.

## Use both

Many projects do: the API for everything it covers, scraping for the rest. Take product IDs and descriptions from a shop's API, and scrape the product page for the price shown in each country. Or use the API for the record and scrape the page to check what visitors actually see.

## A scraping API is not an official API

A scraping API, such as [Skryp](https://skryp.dev/product/scrape), is a service that fetches pages from any site for you and returns them as Markdown or JSON. It removes the work of running browsers and parsers, but it is still scraping: the site's terms still apply, and the data is still what the page shows. An official API is one the site itself publishes.

To try scraping a page without writing code, see [URL to Markdown](https://skryp.dev/tools/url-to-markdown). For the code, see [Web scraping with Python](https://skryp.dev/guides/web-scraping-python).

## Questions

**Web scraping or an API: which is better?**

An official API, when it has the data you need at an acceptable price and rate limit: it is structured, stable and its terms are clear. Scraping, when the API lacks the data or you need what visitors see, such as prices in a particular country. Many projects use both.

**How do you use a web scraping API?**

Send it the address of the page and the format you want, usually with an API key in the header; it fetches the page (with a browser or another country's connection if needed) and returns Markdown, HTML or structured JSON. Some also work as tools for AI agents over MCP.
