# Web scraping with JavaScript and Node.js

In Node.js, fetch static pages and parse them with Cheerio; use Puppeteer or Playwright when the page renders with JavaScript. This guide covers both with runnable code, plus pagination and the point where an API is simpler.

Updated 1 Oct 2026 · Tested on 1 Oct 2026 · https://skryp.dev/guides/web-scraping-javascript

## What you need

Node.js 18 or later, which has `fetch` built in, and two packages:

```bash
npm install cheerio puppeteer
```

Cheerio parses HTML with a jQuery-like API, without a browser. Puppeteer drives Chrome for pages that need one; installing it downloads a matching Chrome. Every example below ran as shown, against [books.toscrape.com](https://books.toscrape.com) and [quotes.toscrape.com](https://quotes.toscrape.com), sites made for practising scraping.

## 1. A static page: fetch and Cheerio

```javascript
// Fetch a static page with Node's built-in fetch and pick fields out of it with Cheerio.
// npm install cheerio   (Node.js 18 or later has fetch built in)
import * as cheerio from "cheerio";

const url = "https://books.toscrape.com/catalogue/category/books/travel_2/index.html";
const res = await fetch(url, { headers: { "User-Agent": "my-scraper/1.0 (you@example.com)" } });
if (!res.ok) throw new Error(`HTTP ${res.status} for ${url}`);
const $ = cheerio.load(await res.text());

const books = $("article.product_pod")
  .map((_, el) => ({
    title: $(el).find("h3 a").attr("title"),
    price: $(el).find(".price_color").text().trim(),
    inStock: $(el).find(".availability").text().includes("In stock"),
  }))
  .get();

console.log(`${books.length} books on the page`);
console.log(books.slice(0, 3));
```

Output (ran 1 Oct 2026; Node.js 24.16.0, cheerio 1.2.0; 1.4 s):

```text
11 books on the page
[
  { title: "It's Only the Himalayas", price: '£45.17', inStock: true },
  {
    title: 'Full Moon over Noah’s Ark: An Odyssey to Mount Ararat and Beyond',
    price: '£49.43',
    inStock: true
  },
  {
    title: 'See America: A Celebration of Our National Parks & Treasured Sites',
    price: '£48.87',
    inStock: true
  }
]
```

`$(selector)` works as in jQuery; `.map(...).get()` turns the matches into a plain array. Check `res.ok` before parsing, so a 403 or 404 page is not mistaken for data, and identify yourself in the `User-Agent`.

## 2. Many pages: follow "next" and save JSON

```javascript
// Follow "next" links across a listing, one request a second, and save every record to a JSON file.
// npm install cheerio
import { writeFile } from "node:fs/promises";
import * as cheerio from "cheerio";

let url = "https://books.toscrape.com/catalogue/category/books/mystery_3/index.html";
const rows = [];
let pages = 0;

while (url && pages < 5) {
  const $ = cheerio.load(await (await fetch(url)).text());
  pages++;
  $("article.product_pod").each((_, el) => {
    rows.push({ title: $(el).find("h3 a").attr("title"), price: $(el).find(".price_color").text().trim(), url: new URL($(el).find("h3 a").attr("href"), url).href });
  });
  const next = $("li.next a").attr("href");
  url = next ? new URL(next, url).href : null;
  await new Promise((r) => setTimeout(r, 1000));
}

await writeFile("mystery-books.json", JSON.stringify(rows, null, 2));
console.log(`${rows.length} books from ${pages} pages saved to mystery-books.json`);
console.log(rows[0]);
```

Output (ran 1 Oct 2026; Node.js 24.16.0, cheerio 1.2.0; 3.5 s):

```text
32 books from 2 pages saved to mystery-books.json
{
  title: 'Sharp Objects',
  price: '£47.82',
  url: 'https://books.toscrape.com/catalogue/sharp-objects_997/index.html'
}
```

`new URL(href, base)` resolves relative links. The page cap and the one-second pause keep a bug, or an eager loop, from hammering the site.

## 3. Pages built with JavaScript: Puppeteer

The [JavaScript version](https://quotes.toscrape.com/js/) of the quotes site sends an almost empty page and builds the quotes in the browser. `fetch` sees none of them; Puppeteer runs the page in Chrome and sees all ten.

```javascript
// A page that builds its content with JavaScript: fetch gets none of it, Puppeteer runs it in Chrome.
// npm install puppeteer
import puppeteer from "puppeteer";

const url = "https://quotes.toscrape.com/js/";
const html = await (await fetch(url)).text();
console.log(`fetch: ${(html.match(/class="quote"/g) ?? []).length} quotes in the HTML it received`);

const browser = await puppeteer.launch();
try {
  const page = await browser.newPage();
  await page.goto(url);
  await page.waitForSelector("div.quote");
  const quotes = await page.$$eval("div.quote", (els) =>
    els.map((e) => ({ text: e.querySelector(".text").textContent, author: e.querySelector(".author").textContent })),
  );
  console.log(`puppeteer: ${quotes.length} quotes`);
  console.log(quotes[0]);
} finally {
  await browser.close();
}
```

Output (ran 1 Oct 2026; Node.js 24.16.0, puppeteer 24.43.1; 4.4 s):

```text
fetch: 0 quotes in the HTML it received
puppeteer: 10 quotes
{
  text: '“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”',
  author: 'Albert Einstein'
}
```

`waitForSelector` holds until the content exists, and `$$eval` runs one function inside the page to read every match at once. Close the browser in `finally`, so an error does not leave Chrome running.

## 4. Infinite scroll

Lists that load more as you scroll need the scrolling done for them. Scroll, pause, count, and stop when you have enough or the list stops growing:

```javascript
// An infinite-scroll list: scroll until enough items have loaded or nothing new arrives.
// npm install puppeteer
import puppeteer from "puppeteer";

const browser = await puppeteer.launch();
try {
  const page = await browser.newPage();
  await page.goto("https://quotes.toscrape.com/scroll");
  await page.waitForSelector("div.quote");
  let count = 0;
  for (let i = 0; i < 20; i++) {
    const now = await page.$$eval("div.quote", (els) => els.length);
    if (now >= 50 || (i > 0 && now === count)) break; // enough, or the list stopped growing
    count = now;
    await page.mouse.wheel({ deltaY: 4000 });
    await new Promise((r) => setTimeout(r, 800)); // give the next batch time to arrive
  }
  const authors = await page.$$eval("small.author", (els) => els.map((e) => e.textContent));
  console.log(`${authors.length} quotes loaded by scrolling, ${new Set(authors).size} authors`);
} finally {
  await browser.close();
}
```

Output (ran 1 Oct 2026; Node.js 24.16.0, puppeteer 24.43.1; 3.7 s):

```text
100 quotes loaded by scrolling, 50 authors
```

The stop condition matters more than the scroll: without it, a list that never ends keeps the loop running forever. Before scrolling at all, open the browser's network tab: many infinite lists fetch each batch from a JSON endpoint you can call directly. This one loads `/api/quotes?page=2`, then page 3, and so on, which `fetch` can read without a browser.

## 5. When an API is simpler

Running Chrome yourself is fine for a few sites. For many sites, other countries, pages that block automated browsers, or jobs on a schedule, a scraping API runs the browser per request. The same page through [Skryp's extraction API](https://skryp.dev/product/extract), with the fields described as a JSON schema:

```javascript
// The same JavaScript page through a scraping API with fetch: one request, typed JSON back.
// export SKRYP_API_KEY=...
const api = process.env.SKRYP_API_URL ?? "https://api.skryp.dev";
const res = await fetch(`${api}/v1/extract`, {
  method: "POST",
  headers: { Authorization: `Bearer ${process.env.SKRYP_API_KEY}`, "Content-Type": "application/json" },
  body: JSON.stringify({
    url: "https://quotes.toscrape.com/js/",
    schema: { type: "object", properties: { quotes: { type: "array", items: { type: "object", properties: { text: { type: "string" }, author: { type: "string" } } } } } },
  }),
});
const out = await res.json();
console.log(`${out.data.quotes.length} quotes, values checked against the page: ${JSON.stringify(out.checks)}`);
console.log(out.data.quotes[0]);
console.log(`route ${out.receipt.route}, cache ${out.receipt.cache}, ${out.receipt.credits} credit${out.receipt.credits === 1 ? "" : "s"}`);
```

Output (ran 1 Oct 2026; Node.js 24.16.0; 2.7 s):

```text
10 quotes, values checked against the page: {"quotes":"on_page"}
{
  text: 'The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.',
  author: 'Albert Einstein'
}
route browser:direct, cache hit, 5 credits
```

The receipt shows the page needed a browser, and every value was checked against the page. Here the page came from Skryp's cache, so it cost nothing, and the 5 credits paid for the model that did the extracting. A fresh page costs 1 credit on Skryp's own network (5, or 10 with a browser, through a residential IP), and failed pages cost nothing. The same call is available to agents over MCP: see the [Quickstart](https://skryp.dev/docs/quickstart).

## Which approach for which page

| The page | Use | Cost |
| --- | --- | --- |
| Data in the HTML | `fetch` + Cheerio | Fast and light |
| Data loaded as JSON by the page | Call that JSON endpoint with `fetch` | Fastest; find it in the network tab |
| Built by JavaScript, or needs clicks and scrolling | Puppeteer or Playwright | A browser per job: slower, more memory |
| Many sites, countries, blocked pages, schedules | A scraping API | A price per page; no browsers to run |

To choose between browser libraries, see [Playwright vs Puppeteer vs Selenium](https://skryp.dev/guides/playwright-vs-puppeteer-vs-selenium), where we ran the same pages through all three. The same techniques in Python are in [Web scraping with Python](https://skryp.dev/guides/web-scraping-python).

## Scraping responsibly

Read the site's terms and `robots.txt` first, go slowly, store only the data you need, and cache what you fetched so a re-run of your parser does not hit the site again.

## Questions

**Can you web scrape with JavaScript?**

Yes. In Node.js, fetch downloads pages and Cheerio parses the HTML; Puppeteer or Playwright drive a real browser for pages that build their content with JavaScript. The same tools run in TypeScript.

**How do you build a web scraper in JavaScript?**

Fetch the page, load the HTML into Cheerio, select the elements you need with CSS selectors and map them to objects. Follow the next-page link with a pause between requests, save the results as JSON or CSV, and switch to Puppeteer for pages that need a browser.

**Is JavaScript or Python better for web scraping?**

Both work well; use the language your project already uses. Python has requests, BeautifulSoup, Scrapy and Playwright; Node.js has fetch, Cheerio, Puppeteer and Playwright. Browser automation behaves much the same in both.

**How do you scrape a website that uses JavaScript?**

First look in the page source and the network tab: the data is often in a script tag or a JSON request you can fetch directly. If it is not, load the page in a browser with Puppeteer or Playwright, wait for the content to appear, then read it.
