Skip to content
Skryp

Web scraping with JavaScript and Node.js

Updated 1 Oct 2026 · Tested on 1 Oct 2026

In Node.js, fetch static pages and parse them with Cheerio; use Puppeteer or Playwright when the page renders with JavaScript. This guide covers both with runnable code, plus pagination and the point where an API is simpler.

What you need

Node.js 18 or later, which has fetch built in, and two packages:

bash
npm install cheerio puppeteer

Cheerio parses HTML with a jQuery-like API, without a browser. Puppeteer drives Chrome for pages that need one; installing it downloads a matching Chrome. Every example below ran as shown, against books.toscrape.com and quotes.toscrape.com, sites made for practising scraping.

1. A static page: fetch and Cheerio

web-scraping-javascript/01-fetch-cheerio.mjs
// Fetch a static page with Node's built-in fetch and pick fields out of it with Cheerio.
// npm install cheerio   (Node.js 18 or later has fetch built in)
import * as cheerio from "cheerio";

const url = "https://books.toscrape.com/catalogue/category/books/travel_2/index.html";
const res = await fetch(url, { headers: { "User-Agent": "my-scraper/1.0 (you@example.com)" } });
if (!res.ok) throw new Error(`HTTP ${res.status} for ${url}`);
const $ = cheerio.load(await res.text());

const books = $("article.product_pod")
  .map((_, el) => ({
    title: $(el).find("h3 a").attr("title"),
    price: $(el).find(".price_color").text().trim(),
    inStock: $(el).find(".availability").text().includes("In stock"),
  }))
  .get();

console.log(`${books.length} books on the page`);
console.log(books.slice(0, 3));
OutputRan 1 Oct 2026 · Node.js 24.16.0, cheerio 1.2.0 · 1.4 s
11 books on the page
[
  { title: "It's Only the Himalayas", price: '£45.17', inStock: true },
  {
    title: 'Full Moon over Noah’s Ark: An Odyssey to Mount Ararat and Beyond',
    price: '£49.43',
    inStock: true
  },
  {
    title: 'See America: A Celebration of Our National Parks & Treasured Sites',
    price: '£48.87',
    inStock: true
  }
]

$(selector) works as in jQuery; .map(...).get() turns the matches into a plain array. Check res.ok before parsing, so a 403 or 404 page is not mistaken for data, and identify yourself in the User-Agent.

2. Many pages: follow "next" and save JSON

web-scraping-javascript/02-pagination-to-json.mjs
// Follow "next" links across a listing, one request a second, and save every record to a JSON file.
// npm install cheerio
import { writeFile } from "node:fs/promises";
import * as cheerio from "cheerio";

let url = "https://books.toscrape.com/catalogue/category/books/mystery_3/index.html";
const rows = [];
let pages = 0;

while (url && pages < 5) {
  const $ = cheerio.load(await (await fetch(url)).text());
  pages++;
  $("article.product_pod").each((_, el) => {
    rows.push({ title: $(el).find("h3 a").attr("title"), price: $(el).find(".price_color").text().trim(), url: new URL($(el).find("h3 a").attr("href"), url).href });
  });
  const next = $("li.next a").attr("href");
  url = next ? new URL(next, url).href : null;
  await new Promise((r) => setTimeout(r, 1000));
}

await writeFile("mystery-books.json", JSON.stringify(rows, null, 2));
console.log(`${rows.length} books from ${pages} pages saved to mystery-books.json`);
console.log(rows[0]);
OutputRan 1 Oct 2026 · Node.js 24.16.0, cheerio 1.2.0 · 3.5 s
32 books from 2 pages saved to mystery-books.json
{
  title: 'Sharp Objects',
  price: '£47.82',
  url: 'https://books.toscrape.com/catalogue/sharp-objects_997/index.html'
}

new URL(href, base) resolves relative links. The page cap and the one-second pause keep a bug, or an eager loop, from hammering the site.

3. Pages built with JavaScript: Puppeteer

The JavaScript version of the quotes site sends an almost empty page and builds the quotes in the browser. fetch sees none of them; Puppeteer runs the page in Chrome and sees all ten.

web-scraping-javascript/03-puppeteer-rendered-page.mjs
// A page that builds its content with JavaScript: fetch gets none of it, Puppeteer runs it in Chrome.
// npm install puppeteer
import puppeteer from "puppeteer";

const url = "https://quotes.toscrape.com/js/";
const html = await (await fetch(url)).text();
console.log(`fetch: ${(html.match(/class="quote"/g) ?? []).length} quotes in the HTML it received`);

const browser = await puppeteer.launch();
try {
  const page = await browser.newPage();
  await page.goto(url);
  await page.waitForSelector("div.quote");
  const quotes = await page.$$eval("div.quote", (els) =>
    els.map((e) => ({ text: e.querySelector(".text").textContent, author: e.querySelector(".author").textContent })),
  );
  console.log(`puppeteer: ${quotes.length} quotes`);
  console.log(quotes[0]);
} finally {
  await browser.close();
}
OutputRan 1 Oct 2026 · Node.js 24.16.0, puppeteer 24.43.1 · 4.4 s
fetch: 0 quotes in the HTML it received
puppeteer: 10 quotes
{
  text: '“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”',
  author: 'Albert Einstein'
}

waitForSelector holds until the content exists, and $$eval runs one function inside the page to read every match at once. Close the browser in finally, so an error does not leave Chrome running.

4. Infinite scroll

Lists that load more as you scroll need the scrolling done for them. Scroll, pause, count, and stop when you have enough or the list stops growing:

web-scraping-javascript/04-infinite-scroll.mjs
// An infinite-scroll list: scroll until enough items have loaded or nothing new arrives.
// npm install puppeteer
import puppeteer from "puppeteer";

const browser = await puppeteer.launch();
try {
  const page = await browser.newPage();
  await page.goto("https://quotes.toscrape.com/scroll");
  await page.waitForSelector("div.quote");
  let count = 0;
  for (let i = 0; i < 20; i++) {
    const now = await page.$$eval("div.quote", (els) => els.length);
    if (now >= 50 || (i > 0 && now === count)) break; // enough, or the list stopped growing
    count = now;
    await page.mouse.wheel({ deltaY: 4000 });
    await new Promise((r) => setTimeout(r, 800)); // give the next batch time to arrive
  }
  const authors = await page.$$eval("small.author", (els) => els.map((e) => e.textContent));
  console.log(`${authors.length} quotes loaded by scrolling, ${new Set(authors).size} authors`);
} finally {
  await browser.close();
}
OutputRan 1 Oct 2026 · Node.js 24.16.0, puppeteer 24.43.1 · 3.7 s
100 quotes loaded by scrolling, 50 authors

The stop condition matters more than the scroll: without it, a list that never ends keeps the loop running forever. Before scrolling at all, open the browser's network tab: many infinite lists fetch each batch from a JSON endpoint you can call directly. This one loads /api/quotes?page=2, then page 3, and so on, which fetch can read without a browser.

5. When an API is simpler

Running Chrome yourself is fine for a few sites. For many sites, other countries, pages that block automated browsers, or jobs on a schedule, a scraping API runs the browser per request. The same page through Skryp's extraction API, with the fields described as a JSON schema:

web-scraping-javascript/05-skryp-api.mjs
// The same JavaScript page through a scraping API with fetch: one request, typed JSON back.
// export SKRYP_API_KEY=...
const api = process.env.SKRYP_API_URL ?? "https://api.skryp.dev";
const res = await fetch(`${api}/v1/extract`, {
  method: "POST",
  headers: { Authorization: `Bearer ${process.env.SKRYP_API_KEY}`, "Content-Type": "application/json" },
  body: JSON.stringify({
    url: "https://quotes.toscrape.com/js/",
    schema: { type: "object", properties: { quotes: { type: "array", items: { type: "object", properties: { text: { type: "string" }, author: { type: "string" } } } } } },
  }),
});
const out = await res.json();
console.log(`${out.data.quotes.length} quotes, values checked against the page: ${JSON.stringify(out.checks)}`);
console.log(out.data.quotes[0]);
console.log(`route ${out.receipt.route}, cache ${out.receipt.cache}, ${out.receipt.credits} credit${out.receipt.credits === 1 ? "" : "s"}`);
OutputRan 1 Oct 2026 · Node.js 24.16.0 · 2.7 s
10 quotes, values checked against the page: {"quotes":"on_page"}
{
  text: 'The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.',
  author: 'Albert Einstein'
}
route browser:direct, cache hit, 5 credits

The receipt shows the page needed a browser, and every value was checked against the page. Here the page came from Skryp's cache, so it cost nothing, and the 5 credits paid for the model that did the extracting. A fresh page costs 1 credit on Skryp's own network (5, or 10 with a browser, through a residential IP), and failed pages cost nothing. The same call is available to agents over MCP: see the Quickstart.

Which approach for which page

The pageUseCost
Data in the HTMLfetch + CheerioFast and light
Data loaded as JSON by the pageCall that JSON endpoint with fetchFastest; find it in the network tab
Built by JavaScript, or needs clicks and scrollingPuppeteer or PlaywrightA browser per job: slower, more memory
Many sites, countries, blocked pages, schedulesA scraping APIA price per page; no browsers to run

To choose between browser libraries, see Playwright vs Puppeteer vs Selenium, where we ran the same pages through all three. The same techniques in Python are in Web scraping with Python.

Scraping responsibly

Read the site's terms and robots.txt first, go slowly, store only the data you need, and cache what you fetched so a re-run of your parser does not hit the site again.

Questions

Can you web scrape with JavaScript?
Yes. In Node.js, fetch downloads pages and Cheerio parses the HTML; Puppeteer or Playwright drive a real browser for pages that build their content with JavaScript. The same tools run in TypeScript.
How do you build a web scraper in JavaScript?
Fetch the page, load the HTML into Cheerio, select the elements you need with CSS selectors and map them to objects. Follow the next-page link with a pause between requests, save the results as JSON or CSV, and switch to Puppeteer for pages that need a browser.
Is JavaScript or Python better for web scraping?
Both work well; use the language your project already uses. Python has requests, BeautifulSoup, Scrapy and Playwright; Node.js has fetch, Cheerio, Puppeteer and Playwright. Browser automation behaves much the same in both.
How do you scrape a website that uses JavaScript?
First look in the page source and the network tab: the data is often in a script tag or a JSON request you can fetch directly. If it is not, load the page in a browser with Puppeteer or Playwright, wait for the content to appear, then read it.