Web scraping for AI agents
Updated 1 Oct 2026 · Tested on 1 Oct 2026
Give your agent one tool to read, search and collect from the web. It picks the cheapest route that works, returns lean Markdown or typed JSON, and reports what each request cost.
An agent with web access, recorded
This script connects Claude Code to Skryp over MCP and asks a question that needs prices from two countries. The agent chose the right tool on its own. It ran as shown:
# Give Claude Code web access through Skryp's MCP server, then ask a question that needs two countries' prices.
# Needs Claude Code and SKRYP_API_KEY. Runs in a scratch folder so the server config stays out of your project.
# timeout: 400
set -euo pipefail
cd "$(mktemp -d)"
claude mcp add --scope project --transport http skryp "${SKRYP_API_URL:-https://api.skryp.dev}/mcp" \
--header "Authorization: Bearer $SKRYP_API_KEY" > /dev/null
claude -p --mcp-config .mcp.json --strict-mcp-config --allowedTools mcp__skryp --output-format stream-json --verbose \
"What does Portal 2 cost on Steam in the United Kingdom and in South Africa, in US dollars? Use the skryp tools and answer in two short lines." \
< /dev/null > session.jsonl
# What the agent did, what Skryp charged, and the answer
python3 - <<'PY'
import json
for line in open("session.jsonl"):
e = json.loads(line)
if e.get("type") == "system" and e.get("subtype") == "init":
tools = [t for t in e["tools"] if t.startswith("mcp__skryp__")]
print(f"connected: {[s['name'] + ' ' + s['status'] for s in e['mcp_servers']]}, {len(tools)} Skryp tools")
if e.get("type") == "assistant":
for c in e["message"]["content"]:
if c.get("type") == "tool_use" and c["name"].startswith("mcp__skryp__"):
print(f"tool call: {c['name'].removeprefix('mcp__skryp__')}({json.dumps(c['input'])[:150]})")
if e.get("type") == "user":
for c in e["message"]["content"]:
if c.get("type") == "tool_result":
parts = c["content"] if isinstance(c["content"], list) else [{"text": c["content"]}]
for part in parts:
try:
receipt = json.loads(part.get("text", "")).get("receipt", {})
except (ValueError, AttributeError):
continue
if "credits" in receipt:
print(f" Skryp charged {receipt['credits']} credits for it")
if e.get("type") == "result":
print(f"\nanswer:\n{e['result']}\n\nagent: {e['num_turns']} turns, {e['duration_ms'] / 1000:.0f} s")
PYconnected: ['skryp connected'], 17 Skryp tools
tool call: compare({"countries": ["GB", "ZA"], "url": "https://store.steampowered.com/app/620/Portal_2/", "base_currency": "USD"})
Skryp charged 4 credits for it
answer:
UK: £1.70 ≈ **$2.26**. South Africa: R20 ≈ **$1.22**. Both prices were read from Steam today through local connections in Salford and Johannesburg.
Steam normally charges more than that in both countries, so this is probably a sale price. The price was pulled out of the page by a model, not read from Steam's own data, so treat it as close but not exact.
agent: 3 turns, 37 sOne tool call fetched the store page from both countries, checked each exit, extracted the prices and converted them; the agent turned that into two lines and passed on the caveat the result carried. The receipt reported the call's cost back to the agent: 4 credits under the earlier prices in force when this was recorded. At the current prices a page through a residential IP costs 5 credits and a model extraction 5, so the same call costs more; see Pricing.
What your agent gets
Seventeen tools, so it can use the cheapest one that does the job:
| The job | Tools |
|---|---|
| Read a page as Markdown, HTML or the page's own data | scrape, and read for more of a long page |
| Search the web, optionally from a country | search |
| Get named fields as JSON, checked against the page | extract |
| Find a site's pages, or crawl it in the background | map, crawl, job |
| Keep a list of records current, with change events | collect, dataset, recipe, monitor |
| Compare one page across countries | compare |
| Click, type and scroll, including signed in to your own accounts | browser, workflow |
| Read a page as it was on a past date | archive |
| See the balance and what is available | account |
The MCP tools reference lists every argument. Skryp tries the cheapest route that works for each page: a plain request first, then residential IPs, a browser, a real Chrome window and finally a third-party unblocker, and it remembers what worked for each site.
Connect in one command
Claude Code:
claude mcp add --transport http skryp https://api.skryp.dev/mcp \
--header "Authorization: Bearer $SKRYP_API_KEY"Codex, in ~/.codex/config.toml:
[mcp_servers.skryp]
url = "https://api.skryp.dev/mcp"
bearer_token_env_var = "SKRYP_API_KEY"Both were tested on 1 October 2026: each connected, listed the tools and called scrape[1 Oct 2026]. Cursor, Claude Desktop and VS Code take the same server as a standard MCP configuration; see Integrations.
Lean output leaves room in the context window
Raw HTML is enormous for a model: on 170 public pages the median page was 65,378 tokens as served, and 1,800 tokens as Skryp's Markdown[1 Oct 2026]. On the 153 pages both services read, Skryp's output was a median 40% shorter than Firecrawl's[1 Oct 2026], at a slightly lower completeness, as the token study shows. Long pages stay on Skryp's side behind a reference the agent reads in sections, and focus returns only the parts of a page about a topic.
Every result has a receipt
Each result tells the agent how it was fetched and what it cost: the route, every attempt, the exit's country and whether it was verified, whether it came from cache, and the credits charged. An agent can use that to judge a result rather than trust a status code, and you can see the same receipts in your dashboard's Activity. See Receipts.
Controls that keep costs predictable
- Failed and blocked pages cost nothing[1 Oct 2026], and a cached result costs nothing.
- A maximum per page:
max_creditsstops Skryp before it tries a route that costs more. - Spending limits: daily and monthly limits on each API key and on the workspace. A request that would pass one is refused before it runs[2 Oct 2026].
- Caps on schedules: every monitor has its own credit cap and pauses when it reaches it.
- A balance and limits an agent can check with the
accounttool before a large job.
Every workspace starts with 1,000 free credits a month. See Pricing.
Limits
- Pages behind a sign-in need a saved login: your own account, connected once in the dashboard, which the agent then uses by name. See Live browser.
- Some sites block automated readers on every route at times; the result says so and costs nothing.
- In our benchmark Skryp missed three of 23 documentation pages that Firecrawl returned; check a few pages before indexing a whole docs site.
Questions
- What is AI web scraping?
- Web scraping done by or for an AI system: an agent or model reads web pages through a tool, gets back clean text or structured data, and uses it to answer or act. The scraping still needs a fetcher that handles JavaScript, blocked pages and countries; the AI decides what to read and what to keep.
- Which AI is best for web scraping?
- The model matters less than the web tool it has. Current models from Anthropic, OpenAI and Google can all choose pages and pull out data; what limits them is whether their tool can render JavaScript, get past blocks, return lean output and report what each request cost. Give your agent a tool that does those things.
Evidence
Skryp's HTTP MCP endpoint worked from Claude Code 2.1.280 (added with claude mcp add --transport http, all 17 tools listed, scrape called and answered correctly) and from Codex CLI 0.159.0 (mcp_servers.skryp with url and bearer_token_env_var, scrape called and answered correctly) on 1 October 2026.
observed test · checked 1 Oct 2026 · Claude Code and Codex CLI against the preview service; one task each.
Limits: Cursor, Claude Desktop and VS Code were not tested. Codex CLI 0.155.1 could not run the account's configured model, so the run used the 0.159.0 CLI bundled with the ChatGPT app.
Fetched with a plain HTTP GET on 1 October 2026, the median benchmark page with an HTML count was 65,378 tokens of HTML; on the 122 pages where Skryp also returned usable Markdown, the HTML was a median 34 times larger.
observed test · checked 1 Oct 2026 · 126 of the benchmark's 170 pages had an HTML count; the others were blocked or errors (30), failed to load (9) or were PDFs (5).
Limits: HTML as served, before JavaScript runs: for JavaScript apps it is a small shell without the content. Fetched hours after the Markdown benchmark, so some pages changed in between. Country variants of one URL share a single fetch.
On the same 170 public pages, the median output was 1,800 tokens per page from Skryp and 3,520 from Firecrawl.
observed test · checked 1 Oct 2026 · Same benchmark run; Markdown output.
Limits: Fewer tokens is not automatically better: completeness was slightly lower for Skryp (median 0.98 vs 1.00). PDFs were the exception: on the five PDFs both services read, they returned about the same length (Skryp shorter on three, longer on two; median ratio 0.95).
Blocked and failed pages cost 0 credits.
observed test · checked 1 Oct 2026 · All page reads through the API, MCP and dashboard.
Limits: Applies to page reads; a model extraction that runs is charged when it runs.
On the 153 public pages both services returned as usable on 1 October 2026, Skryp's Markdown was a median 40% shorter than Firecrawl's, and shorter on 135 of them.
observed test · checked 1 Oct 2026 · Same benchmark run; Markdown output; pages where both services returned usable content.
Limits: Fewer tokens is not automatically better: Skryp kept slightly less of each page's reference text (median completeness 0.98 vs 1.00). PDFs came out about the same length. One run on 170 pages we chose.
Each Skryp API key can have a daily limit, a monthly limit and a maximum per page, and each workspace a daily and a monthly limit; a request that would pass one is refused with HTTP 402 before it runs and costs nothing. On 2 October 2026 a key limited to 1 credit a day read one page for 1 credit, and its next request was refused.
observed test · checked 2 Oct 2026 · Skryp's API and MCP server; limits set in the dashboard by a workspace owner or admin.
Limits: Days and months are UTC. Requests in progress count at their reserved cost, so a limit can refuse a request that would have cost less. An API key cannot change its own limits.