Skip to content
Skryp

What AI agents search for when choosing a web scraper

Updated 1 Oct 2026

We asked Claude Code and Codex 24 buyer questions about web scraping tools in 60 isolated runs and recorded every search, every page they opened and every tool they named. They searched for vendors they already knew, checked pricing pages, and named Firecrawl first in 25 of 60 answers.

How the runs were made

On 1 October 2026 we gave Claude Code and Codex the kind of questions a developer asks when choosing a web scraper, and recorded what each agent did. Every run started fresh, in an empty folder with no project, notes or memory, and no tool was mentioned that the question did not already name. Skryp appears in none of the questions.

  • Agents: Claude Code 2.1.280 (model reported as Claude Opus 5.5) with its web search and fetch tools only, and Codex CLI 0.155.1 with web search on, in a read-only sandbox. Codex does not report its model in this mode.
  • Two conditions. Asked plainly, an agent got the question alone: 26 runs. Asked to research, the question also asked for current web sources: 34 runs.
  • Captured: every search query typed, every page opened, the addresses Claude's search returned, the sources cited in the answer and the tools named.
  • One confounder. Codex also loaded a global instruction file from the machine it ran on, which we could not switch off. It mentions no scraping API except one data platform, Apify, as a paid research tool. Codex's Apify mentions may be inflated by it.

The 24 questions, word for word:

QuestionAsked
What web scraping tool would you choose for an AI agent, and why?plainly
I'm building an AI agent that needs to read web pages and pull data from them. What should I use?plainly
What's the best MCP server for web scraping?plainly
What's the cheapest web scraping API for an AI agent?plainly
What's a good alternative to Firecrawl?plainly
I need to turn websites into clean markdown for an LLM. What tool should I use?plainly
I need to monitor competitor prices on e-commerce sites every day. What's the best way to do that?plainly
Qual ferramenta de web scraping você recomenda para um agente de IA? Pesquise fontes atuais na web e cite-as.to research
Which web scraping APIs let me choose the country a page is fetched from and prove where the request came from? Research current sources on the web and cite them.to research
Welches Web-Scraping-Tool sollte ich für einen KI-Agenten verwenden, der Preise aus deutschen Online-Shops ausliest? Recherchiere aktuelle Quellen im Web und nenne sie.to research
What is the best way to crawl an entire documentation site into Markdown for a RAG pipeline, and what does it cost? Research current sources on the web and cite them.to research
I need an API that watches a set of web pages and tells me which fields changed (price, stock, title). Compare current options, including costs. Research current sources on the web and cite them.to research
AIエージェント向けのウェブスクレイピングツールは何を選ぶべきですか?最新の情報をウェブで調べて、出典を示してください。to research
Compare Firecrawl, Crawl4AI, Apify and Bright Data for an AI agent that must handle JavaScript-heavy sites and stay within a fixed monthly budget. Research current sources on the web and cite them.to research
Which web scraping MCP servers work with Claude Code and Codex, and what are their limits and prices? Research current sources on the web and cite them.to research
Tavily vs Exa vs Firecrawl vs Brave Search API: which should an AI agent use for web search plus page content? Research current sources on the web and cite them.to research
I need to extract structured JSON (product name, price, availability) from arbitrary e-commerce pages for an agent. Which services do this reliably? Research current sources on the web and cite them.to research
Which tools let an AI agent both scrape websites and interact with pages through an MCP server? Compare scraping APIs with cloud browser tools, including limitations. Research current sources on the web and cite them.to research
What is the cheapest reliable approach for an AI agent scraping a mix of ordinary pages and JavaScript-heavy sites? I need predictable spending limits and structured JSON. Research current sources on the web and cite them.to research
What are good alternatives to Firecrawl for a production AI agent? Compare costs, limitations, MCP support and when each is the better choice. Research current sources on the web and cite them.to research
What web scraping tool would you choose for an AI agent, and why? Research current options on the web and cite sources.to research
I need clean Markdown from documentation sites for a RAG application and want to control costs. What should I use? Research current options on the web and cite sources.to research
I need product prices as seen in South Africa and Germany. Which web data service would you choose, and how would you verify location and freshness? Research current sources on the web and cite them.to research
I need a repeatable product feed with change events, not just a one-off page scrape. What are my options? Research current sources on the web and cite them.to research
The 24 questions, word for word. 'Plainly' means the question alone; 'to research' means it asked for current web sources. Three are in Portuguese, German and Japanese.

Did they search at all?

Claude Code, asked plainlyClaude Code, asked to researchCodex, asked plainlyCodex, asked to research
Runs that searched the web4/13 (31%)17/17 (100%)13/13 (100%)17/17 (100%)
Searches per run0.54.83.39.1

Asked plainly, Claude Code usually answered from what it already knew, searching in 4 of 13 runs; Codex searched every time. An answer given from memory can only recommend tools the model learned about in training, so a new tool cannot appear in it until it is in the training data or the user asks for research.

What they searched for

The search query…Asked plainlyAsked to research
Names a vendor or tool46/49 (94%)197/237 (83%)
Asks about price/credits/cost29/49 (59%)107/237 (45%)
Contains a year (2025/2026)8/49 (16%)42/237 (18%)
Uses site: or a domain filter20/49 (41%)124/237 (52%)
Mentions MCP9/49 (18%)58/237 (24%)
Comparison (vs/alternatives/compare)5/49 (10%)15/237 (6%)
Docs/official21/49 (43%)76/237 (32%)
Country/geo/location0/49 (0%)34/237 (14%)
Generic category, no vendor3/49 (6%)40/237 (17%)
Shares of the 49 and 237 search queries the agents typed in each condition. One query can fall in several rows.

They searched for tools they already knew, not for the category: 94% of the plainly-asked queries and 83% of the research queries named a vendor or tool, such as "Firecrawl pricing scrape API 2026 credits" or "Bright Data MCP server pricing scraping official". About half the queries checked price. Codex went straight to vendors' own sites with site: filters in 144 of its 198 searches; Claude used none.

What they opened and cited

Page typeOpened by the agentCited in the answerReturned by Claude's search
Docs49/110 (45%)173/523 (33%)114/834 (14%)
Pricing page36/110 (33%)96/523 (18%)80/834 (10%)
Home/product/other11/110 (10%)104/523 (20%)286/834 (34%)
Comparison/listicle5/110 (5%)79/523 (15%)207/834 (25%)
GitHub7/110 (6%)43/523 (8%)50/834 (6%)
Blog2/110 (2%)28/523 (5%)97/834 (12%)

Documentation and pricing pages made up 78% of the pages the agents opened[1 Oct 2026]. The most opened sites were Firecrawl's (firecrawl.dev 15 times, its docs 10), GitHub (7), Crawl4AI's docs (6), Bright Data's docs (5) and Exa (5). The most cited were GitHub (43 citations), firecrawl.dev (39), Firecrawl's docs (28), apify.com (23) and Crawl4AI's docs (22).

Named first in the answerAsked plainly (26 runs)Asked to research (34 runs)Total
Firecrawl91625
Crawl4AI437
Jina Reader404
Playwright134
Bright Data044
Zyte044
Prisync202
Trafilatura101
Browser Use101
ZenRows101
Visualping101
Spider101
Apify011
Exa011
The first provider named in the final answer, leaving out any provider the question itself named. A proxy for the lead pick.

Firecrawl was named first in 25 of the 60 answers. The pick depended on the task:

  • Cheapest API: Jina Reader or Crawl4AI.
  • Daily competitor price monitoring: merchant tools such as Prisync, Price2Spy and Visualping rather than a scraping API.
  • Proving which country a request came from: Bright Data, in both agents. Both went looking in vendors' docs for a response field that exposes the exit IP.
  • Structured JSON from shop pages: Zyte, in both agents.
  • MCP access, documentation to Markdown, change events, JavaScript-heavy sites on a budget: mostly Firecrawl.

The sites in between

Claude's searches also returned a layer of small roundup sites and MCP directories, several of which were cited in answers:

Site (not a vendor)Times Claude's search returned it
fastcrw.com14
glama.ai13
pagecrawl.io12
crawlforge.dev9
scrapeops.io9
keirolabs.cloud9
dataimpulse.com7
fast.io6
proxidize.com6
context.dev6
crevio.co6
mcpservers.org6

Questions in other languages

For the Portuguese, German and Japanese questions, both agents searched in English (apart from one German query about legality) and answered in the question's language.

What this suggests for a developer tool

  • Be on the pages agents already read. They open docs and pricing pages and cite GitHub. A product that is missing from its own docs, its pricing page or GitHub is missing from the research.
  • Put prices in plain text. Pricing pages were a third of what the agents opened, and they quoted them.
  • Document the answer to the question agents ask. For "prove where the request came from", both agents dug through docs for response fields; the tools that documented one were the ones recommended.
  • Answers from memory change slowly. When an agent does not search, only training data decides, so pages cannot change those answers in the short term.

Limits

  • One or two runs per question, two agents, one day. Treat the counts as directions, not market shares.
  • No Gemini, Perplexity or ChatGPT app runs.
  • Codex's model is not reported, and Claude's search provider is not shown in its output, so "returned by search" covers Claude only.
  • "Named first" is the first tool named in the final answer, leaving out any tool the question itself named. It is a proxy for the lead pick, not a reading of the whole answer.
  • Two earlier batches were discarded because the agents could see the project they were run from; one of them named Skryp. Only the isolated runs are counted here.

Data

Download one row per run (agent, condition, question, searches, pages opened and cited, tools named) and every search query the agents typed.

Questions

Which web scraping tool do AI agents recommend?
In our 60 runs on 1 October 2026, Firecrawl was named first in 25 answers. When asked plainly, Jina Reader and Crawl4AI came next; when asked to research, Bright Data and Zyte.
Do Claude Code and Codex search the web before answering?
Codex searched in every run. Claude Code searched in every run when asked to research, but in only 4 of 13 runs when asked plainly; otherwise it answered from what it already knew.
Which pages did the agents read?
Mostly documentation (45% of pages opened) and pricing pages (33%). Comparison articles were 25% of what search returned but only 5% of what the agents opened.
Can I see the raw data?
Yes. Every run and every search query is downloadable from the page as CSV, with the prompts used.

Evidence

  1. In 60 isolated runs of Claude Code and Codex on 24 buyer questions about web scraping tools, Firecrawl was the first provider named in 25 answers, 83-94% of the agents' search queries named a vendor, and docs and pricing pages made up 78% of the pages they opened.

    observed test · checked 1 Oct 2026 · Claude Code 2.1.280 (claude-opus-5-5) and Codex CLI 0.155.1; public prompts; isolated working folder.

    Limits: Small sample (one or two runs per prompt); Codex loads the owner's global instructions file; no Gemini, Perplexity or ChatGPT app runs.