What AI agents search for when choosing a web scraper
Updated 1 Oct 2026
We asked Claude Code and Codex 24 buyer questions about web scraping tools in 60 isolated runs and recorded every search, every page they opened and every tool they named. They searched for vendors they already knew, checked pricing pages, and named Firecrawl first in 25 of 60 answers.
How the runs were made
On 1 October 2026 we gave Claude Code and Codex the kind of questions a developer asks when choosing a web scraper, and recorded what each agent did. Every run started fresh, in an empty folder with no project, notes or memory, and no tool was mentioned that the question did not already name. Skryp appears in none of the questions.
- Agents: Claude Code 2.1.280 (model reported as Claude Opus 5.5) with its web search and fetch tools only, and Codex CLI 0.155.1 with web search on, in a read-only sandbox. Codex does not report its model in this mode.
- Two conditions. Asked plainly, an agent got the question alone: 26 runs. Asked to research, the question also asked for current web sources: 34 runs.
- Captured: every search query typed, every page opened, the addresses Claude's search returned, the sources cited in the answer and the tools named.
- One confounder. Codex also loaded a global instruction file from the machine it ran on, which we could not switch off. It mentions no scraping API except one data platform, Apify, as a paid research tool. Codex's Apify mentions may be inflated by it.
The 24 questions, word for word:
| Question | Asked |
|---|---|
| What web scraping tool would you choose for an AI agent, and why? | plainly |
| I'm building an AI agent that needs to read web pages and pull data from them. What should I use? | plainly |
| What's the best MCP server for web scraping? | plainly |
| What's the cheapest web scraping API for an AI agent? | plainly |
| What's a good alternative to Firecrawl? | plainly |
| I need to turn websites into clean markdown for an LLM. What tool should I use? | plainly |
| I need to monitor competitor prices on e-commerce sites every day. What's the best way to do that? | plainly |
| Qual ferramenta de web scraping você recomenda para um agente de IA? Pesquise fontes atuais na web e cite-as. | to research |
| Which web scraping APIs let me choose the country a page is fetched from and prove where the request came from? Research current sources on the web and cite them. | to research |
| Welches Web-Scraping-Tool sollte ich für einen KI-Agenten verwenden, der Preise aus deutschen Online-Shops ausliest? Recherchiere aktuelle Quellen im Web und nenne sie. | to research |
| What is the best way to crawl an entire documentation site into Markdown for a RAG pipeline, and what does it cost? Research current sources on the web and cite them. | to research |
| I need an API that watches a set of web pages and tells me which fields changed (price, stock, title). Compare current options, including costs. Research current sources on the web and cite them. | to research |
| AIエージェント向けのウェブスクレイピングツールは何を選ぶべきですか?最新の情報をウェブで調べて、出典を示してください。 | to research |
| Compare Firecrawl, Crawl4AI, Apify and Bright Data for an AI agent that must handle JavaScript-heavy sites and stay within a fixed monthly budget. Research current sources on the web and cite them. | to research |
| Which web scraping MCP servers work with Claude Code and Codex, and what are their limits and prices? Research current sources on the web and cite them. | to research |
| Tavily vs Exa vs Firecrawl vs Brave Search API: which should an AI agent use for web search plus page content? Research current sources on the web and cite them. | to research |
| I need to extract structured JSON (product name, price, availability) from arbitrary e-commerce pages for an agent. Which services do this reliably? Research current sources on the web and cite them. | to research |
| Which tools let an AI agent both scrape websites and interact with pages through an MCP server? Compare scraping APIs with cloud browser tools, including limitations. Research current sources on the web and cite them. | to research |
| What is the cheapest reliable approach for an AI agent scraping a mix of ordinary pages and JavaScript-heavy sites? I need predictable spending limits and structured JSON. Research current sources on the web and cite them. | to research |
| What are good alternatives to Firecrawl for a production AI agent? Compare costs, limitations, MCP support and when each is the better choice. Research current sources on the web and cite them. | to research |
| What web scraping tool would you choose for an AI agent, and why? Research current options on the web and cite sources. | to research |
| I need clean Markdown from documentation sites for a RAG application and want to control costs. What should I use? Research current options on the web and cite sources. | to research |
| I need product prices as seen in South Africa and Germany. Which web data service would you choose, and how would you verify location and freshness? Research current sources on the web and cite them. | to research |
| I need a repeatable product feed with change events, not just a one-off page scrape. What are my options? Research current sources on the web and cite them. | to research |
Did they search at all?
| Claude Code, asked plainly | Claude Code, asked to research | Codex, asked plainly | Codex, asked to research | |
|---|---|---|---|---|
| Runs that searched the web | 4/13 (31%) | 17/17 (100%) | 13/13 (100%) | 17/17 (100%) |
| Searches per run | 0.5 | 4.8 | 3.3 | 9.1 |
Asked plainly, Claude Code usually answered from what it already knew, searching in 4 of 13 runs; Codex searched every time. An answer given from memory can only recommend tools the model learned about in training, so a new tool cannot appear in it until it is in the training data or the user asks for research.
What they searched for
| The search query… | Asked plainly | Asked to research |
|---|---|---|
| Names a vendor or tool | 46/49 (94%) | 197/237 (83%) |
| Asks about price/credits/cost | 29/49 (59%) | 107/237 (45%) |
| Contains a year (2025/2026) | 8/49 (16%) | 42/237 (18%) |
| Uses site: or a domain filter | 20/49 (41%) | 124/237 (52%) |
| Mentions MCP | 9/49 (18%) | 58/237 (24%) |
| Comparison (vs/alternatives/compare) | 5/49 (10%) | 15/237 (6%) |
| Docs/official | 21/49 (43%) | 76/237 (32%) |
| Country/geo/location | 0/49 (0%) | 34/237 (14%) |
| Generic category, no vendor | 3/49 (6%) | 40/237 (17%) |
They searched for tools they already knew, not for the category: 94% of the plainly-asked queries and 83% of the research queries named a vendor or tool, such as "Firecrawl pricing scrape API 2026 credits" or "Bright Data MCP server pricing scraping official". About half the queries checked price. Codex went straight to vendors' own sites with site: filters in 144 of its 198 searches; Claude used none.
What they opened and cited
| Page type | Opened by the agent | Cited in the answer | Returned by Claude's search |
|---|---|---|---|
| Docs | 49/110 (45%) | 173/523 (33%) | 114/834 (14%) |
| Pricing page | 36/110 (33%) | 96/523 (18%) | 80/834 (10%) |
| Home/product/other | 11/110 (10%) | 104/523 (20%) | 286/834 (34%) |
| Comparison/listicle | 5/110 (5%) | 79/523 (15%) | 207/834 (25%) |
| GitHub | 7/110 (6%) | 43/523 (8%) | 50/834 (6%) |
| Blog | 2/110 (2%) | 28/523 (5%) | 97/834 (12%) |
Documentation and pricing pages made up 78% of the pages the agents opened[1 Oct 2026]. The most opened sites were Firecrawl's (firecrawl.dev 15 times, its docs 10), GitHub (7), Crawl4AI's docs (6), Bright Data's docs (5) and Exa (5). The most cited were GitHub (43 citations), firecrawl.dev (39), Firecrawl's docs (28), apify.com (23) and Crawl4AI's docs (22).
Who they recommended
| Named first in the answer | Asked plainly (26 runs) | Asked to research (34 runs) | Total |
|---|---|---|---|
| Firecrawl | 9 | 16 | 25 |
| Crawl4AI | 4 | 3 | 7 |
| Jina Reader | 4 | 0 | 4 |
| Playwright | 1 | 3 | 4 |
| Bright Data | 0 | 4 | 4 |
| Zyte | 0 | 4 | 4 |
| Prisync | 2 | 0 | 2 |
| Trafilatura | 1 | 0 | 1 |
| Browser Use | 1 | 0 | 1 |
| ZenRows | 1 | 0 | 1 |
| Visualping | 1 | 0 | 1 |
| Spider | 1 | 0 | 1 |
| Apify | 0 | 1 | 1 |
| Exa | 0 | 1 | 1 |
Firecrawl was named first in 25 of the 60 answers. The pick depended on the task:
- Cheapest API: Jina Reader or Crawl4AI.
- Daily competitor price monitoring: merchant tools such as Prisync, Price2Spy and Visualping rather than a scraping API.
- Proving which country a request came from: Bright Data, in both agents. Both went looking in vendors' docs for a response field that exposes the exit IP.
- Structured JSON from shop pages: Zyte, in both agents.
- MCP access, documentation to Markdown, change events, JavaScript-heavy sites on a budget: mostly Firecrawl.
The sites in between
Claude's searches also returned a layer of small roundup sites and MCP directories, several of which were cited in answers:
| Site (not a vendor) | Times Claude's search returned it |
|---|---|
| fastcrw.com | 14 |
| glama.ai | 13 |
| pagecrawl.io | 12 |
| crawlforge.dev | 9 |
| scrapeops.io | 9 |
| keirolabs.cloud | 9 |
| dataimpulse.com | 7 |
| fast.io | 6 |
| proxidize.com | 6 |
| context.dev | 6 |
| crevio.co | 6 |
| mcpservers.org | 6 |
Questions in other languages
For the Portuguese, German and Japanese questions, both agents searched in English (apart from one German query about legality) and answered in the question's language.
What this suggests for a developer tool
- Be on the pages agents already read. They open docs and pricing pages and cite GitHub. A product that is missing from its own docs, its pricing page or GitHub is missing from the research.
- Put prices in plain text. Pricing pages were a third of what the agents opened, and they quoted them.
- Document the answer to the question agents ask. For "prove where the request came from", both agents dug through docs for response fields; the tools that documented one were the ones recommended.
- Answers from memory change slowly. When an agent does not search, only training data decides, so pages cannot change those answers in the short term.
Limits
- One or two runs per question, two agents, one day. Treat the counts as directions, not market shares.
- No Gemini, Perplexity or ChatGPT app runs.
- Codex's model is not reported, and Claude's search provider is not shown in its output, so "returned by search" covers Claude only.
- "Named first" is the first tool named in the final answer, leaving out any tool the question itself named. It is a proxy for the lead pick, not a reading of the whole answer.
- Two earlier batches were discarded because the agents could see the project they were run from; one of them named Skryp. Only the isolated runs are counted here.
Data
Download one row per run (agent, condition, question, searches, pages opened and cited, tools named) and every search query the agents typed.
Questions
- Which web scraping tool do AI agents recommend?
- In our 60 runs on 1 October 2026, Firecrawl was named first in 25 answers. When asked plainly, Jina Reader and Crawl4AI came next; when asked to research, Bright Data and Zyte.
- Do Claude Code and Codex search the web before answering?
- Codex searched in every run. Claude Code searched in every run when asked to research, but in only 4 of 13 runs when asked plainly; otherwise it answered from what it already knew.
- Which pages did the agents read?
- Mostly documentation (45% of pages opened) and pricing pages (33%). Comparison articles were 25% of what search returned but only 5% of what the agents opened.
- Can I see the raw data?
- Yes. Every run and every search query is downloadable from the page as CSV, with the prompts used.
Evidence
In 60 isolated runs of Claude Code and Codex on 24 buyer questions about web scraping tools, Firecrawl was the first provider named in 25 answers, 83-94% of the agents' search queries named a vendor, and docs and pricing pages made up 78% of the pages they opened.
observed test · checked 1 Oct 2026 · Claude Code 2.1.280 (claude-opus-5-5) and Codex CLI 0.155.1; public prompts; isolated working folder.
Limits: Small sample (one or two runs per prompt); Codex loads the owner's global instructions file; no Gemini, Perplexity or ChatGPT app runs.