Skip to content
Skryp

Web scraping vs an official API: which should you use?

Updated 1 Oct 2026 · Tested on 1 Oct 2026

Use a site's official API when it exposes the data you need at a price and rate you can live with; scrape when it doesn't, or when you need what users see, such as country-specific prices. Many projects use both. The decision table below covers coverage, cost, limits, terms and maintenance.

The decision in one table

Official APIScraping
Data availableWhat the site chose to publishWhatever the page shows
What users see (ranking, local prices, layout)Often not includedYes, as seen from where you fetch it
FormatStructured JSON with stable IDsHTML you have to parse
Breaks whenThe API version is retired, usually with noticeThe page layout changes, without notice
LimitsRate limits, quotas, sometimes paid tiersBot protection; your own politeness
TermsSpelled out in the API's termsSet by the site's terms and the law where you operate
EffortRead the docs, get a keyWrite and maintain a parser, or use a scraping service

Use the official API when it has the data you need at a price and rate you can live with: it is more stable and its terms are clear. Scrape when it doesn't: the data is not in the API, the API costs more than the data is worth, or you need what a visitor sees, such as the prices shown in a particular country.

The same data both ways

Hacker News has both: an official API and a front page. Here are the top five stories through the API, then from the page, fetched seconds apart. Both ran as shown.

web-scraping-vs-api/01-official-api.py
# Hacker News publishes an official API: the top stories as structured JSON, with stable ids and exact fields.
# pip install requests
import requests

API = "https://hacker-news.firebaseio.com/v0"
ids = requests.get(f"{API}/topstories.json", timeout=30).json()[:5]
for i in ids:
    item = requests.get(f"{API}/item/{i}.json", timeout=30).json()
    print({"id": item["id"], "title": item["title"][:60], "points": item["score"], "comments": item.get("descendants", 0), "time": item["time"]})
OutputRan 1 Oct 2026 · Python 3.12.14, requests 2.34.2 · 3.4 s
{'id': 49923692, 'title': 'Clef: our open-source decision models', 'points': 68, 'comments': 15, 'time': 1790871537}
{'id': 49923466, 'title': 'RIP, vector database', 'points': 63, 'comments': 16, 'time': 1790870516}
{'id': 49920160, 'title': 'StreetComplete on iOS is now in public beta', 'points': 390, 'comments': 82, 'time': 1790852397}
{'id': 49922515, 'title': 'RacketCon Is Saturday', 'points': 52, 'comments': 14, 'time': 1790866693}
{'id': 49920896, 'title': 'How to speed up the Rust compiler in September 2026', 'points': 166, 'comments': 80, 'time': 1790858678}
web-scraping-vs-api/02-scrape-the-page.py
# The same stories scraped from the front page: what a reader sees, in the order they see it.
# pip install requests beautifulsoup4
import requests
from bs4 import BeautifulSoup

soup = BeautifulSoup(requests.get("https://news.ycombinator.com/", headers={"User-Agent": "my-scraper/1.0 (you@example.com)"}, timeout=30).content, "html.parser")
for row in soup.select("tr.athing")[:5]:
    sub = row.find_next_sibling("tr")
    points = sub.select_one(".score")
    print({
        "id": int(row["id"]),
        "rank": int(row.select_one(".rank").get_text(strip=True).rstrip(".")),
        "title": row.select_one(".titleline > a").get_text(strip=True)[:60],
        "points": int(points.get_text().split()[0]) if points else None,
        "age": sub.select_one(".age").get_text(strip=True),
    })
OutputRan 1 Oct 2026 · Python 3.12.14, beautifulsoup4 4.15.0, requests 2.34.2 · 1.3 s
{'id': 49923692, 'rank': 1, 'title': 'Clef: our open-source decision models', 'points': 63, 'age': '46 minutes ago'}
{'id': 49923466, 'rank': 2, 'title': 'RIP, vector database', 'points': 63, 'age': '1 hour ago'}
{'id': 49920160, 'rank': 3, 'title': 'StreetComplete on iOS is now in public beta', 'points': 389, 'age': '6 hours ago'}
{'id': 49922515, 'rank': 4, 'title': 'RacketCon Is Saturday', 'points': 51, 'age': '2 hours ago'}
{'id': 49920896, 'rank': 5, 'title': 'How to speed up the Rust compiler in September 2026', 'points': 165, 'age': '4 hours ago'}

The same five stories came back in the same order, with the same IDs. The differences are the ones in the table:

  • The API gives exact fields. Points, comment counts and a precise timestamp, as numbers. No HTML to parse, nothing to break when the design changes.
  • The page gives what a reader sees. The rank and the relative age shown to readers are presentation, which the API leaves you to work out.
  • They disagreed. Fetched seconds apart, the page showed a few points fewer than the API on several stories. The two are served differently, so they were not equally fresh. When you combine sources, record when and where each value came from.

Use both

Many projects do: the API for everything it covers, scraping for the rest. Take product IDs and descriptions from a shop's API, and scrape the product page for the price shown in each country. Or use the API for the record and scrape the page to check what visitors actually see.

A scraping API is not an official API

A scraping API, such as Skryp, is a service that fetches pages from any site for you and returns them as Markdown or JSON. It removes the work of running browsers and parsers, but it is still scraping: the site's terms still apply, and the data is still what the page shows. An official API is one the site itself publishes.

To try scraping a page without writing code, see URL to Markdown. For the code, see Web scraping with Python.

Questions

Web scraping or an API: which is better?
An official API, when it has the data you need at an acceptable price and rate limit: it is structured, stable and its terms are clear. Scraping, when the API lacks the data or you need what visitors see, such as prices in a particular country. Many projects use both.
How do you use a web scraping API?
Send it the address of the page and the format you want, usually with an API key in the header; it fetches the page (with a browser or another country's connection if needed) and returns Markdown, HTML or structured JSON. Some also work as tools for AI agents over MCP.