Netscraper

Ship the scraper as a manifest, not an image, a package or a source tree. A Netscraper actor is one JSON file that runs on any browser driver, in Node or in a browser tab. It is tested continuously, refactored in moments for under a dollar, and deployed by serving the file. This one collected 100 of 100 records in 64 seconds.

github.com/drksci/netscraper
The idea
A scraper should be a file you can serve.

Scrapers break. The site ships a redesign, a modal appears, an endpoint moves, and the fix has to travel through an image build, a package release or a source checkout before anything runs again. Netscraper takes that weight out. The whole actor is a manifest: plain JSON that describes how to read the pages, how to move through the site and what each record looks like. A generic runtime executes it. So a fix is a new file, checked against what the page actually shows, written by an agent in a few minutes for less than a dollar of model time. Serve it over HTTP, attach it to an email or put it on a floppy disk, Goodbye harness and DevOps: when the site changes, you move quickly.

Production run
100/100
records in 64 s
Parity
5/5
targets matched field for field
Manifest
≈ 24 KB
about 4 KB gzipped
Drivers
3 + WASM
one file, no changes
Compute per 1M
≈ $21
AWS Lambda; no proxies

Watch it work

The Studio session of 6 October 2026, replayed. Time is compressed. Page imagery is mosaicked, and creator identities (names, handles, captions, links, images) are replaced with placeholders; counts and timings are as recorded. It plays while on screen and pauses when you scroll away.
JSON
XState + A2UI. That’s it.

The whole actor is one JSON file: an XState machine, plus A2UI views and JSONata routes. No code ships with it.

Target
TypeScript · WASM. Runs anywhere.

One runtime for Node and for a browser tab, where it runs inside QuickJS compiled to WebAssembly.

BYOD · bring your own driver
Any driver, same records.
  • Playwright
  • Puppeteer
  • raw Chrome DevTools Protocol
  • attach to a running browser
  • CloakBrowser as the browser

In the replay, an AI agent (Claude, working only through Netscraper’s own tools) is asked for TikTok’s explore feed, logged out. It reads how open-source TikTok scrapers work, then browses like a visitor. When a login prompt covers the page it closes it, checks that it has gone, and carries on scrolling until it can show that paging really fired. Then it writes the manifest, tests it on 20 items, matches 5/5 target records field for field, and runs it at scale: 100 of 100 videos in 64 seconds, in 17 tool calls.

…but why?

A conventional scraper is a program. The site’s structure is buried in its code, the browser library is wired through it, and the program ships however you ship programs. When the site changes, which it will, you are back in the build-and-deploy loop to change one selector.

Netscraper splits the scraper in two. The runtime is generic and never changes per site. Everything site-specific lives in the manifest, and the manifest is data. The properties below follow from that split.

Portable
The same file runs on Playwright, Puppeteer and a raw Chrome DevTools Protocol socket, with byte-identical output in the project’s end-to-end test. It also runs in an ordinary browser tab as WebAssembly. Nothing is installed per site.
Cheap to change
A fix is a new manifest. An agent in the Studio re-reads the page, patches the file and proves it against what it saw. The recorded session took 17 tool calls from a blank page to a production run.
Continuously testable
Parity compares a run’s output with records captured from the live page, field by field. Each view carries a structural fingerprint, so a redesign is detected on the next run rather than discovered weeks later in bad data.
Failures are declared
Login prompts, cookie banners, captchas, missing accounts and empty pages are named states in the file, not exceptions in code. Work is split into idempotent units, so a crashed run restarts without duplicating records.
Small
The recorded manifest is about 24 KB as formatted JSON, 4 KB gzipped. Gzipped, the WebAssembly runtime and engine add about 505 KB, so the whole stack fits on a 1.44 MB floppy with room to spare.

Anatomy of a manifest

The design rests on a standard in the middle. A2UI is Google’s open protocol for agents to describe user interfaces as data: components such as text, lists and buttons, bound to a data model, sent as a stream of small JSON messages. Netscraper uses it the other way round. It describes an existing web page as if it were an A2UI interface. The rest of the system never sees the DOM, only a typed surface it can read and press buttons on.

Architecture
The state machine from the demo above (tiktok.explore)
Views
Each kind of page is a view: a match rule, A2UI components and a data model. The model is filled from the DOM, or better, from the page’s own network responses: a net pointer names a response by URL and selects from it with JSONata, a small query language for JSON. Capture is passive. Netscraper reads what the page already received and never replays or forges a request.
Machine
Navigation is an XState v5 state machine written as JSON. Its states navigate, wait for a view, or press an A2UI button; its guards read the data model. A selector never appears in it. Interrupt views, such as a login prompt, sit over the page as their own surface: the machine dismisses them and a deep history state resumes exactly where it was.
Routes
A route turns a surface into records: a key for de-duplication, a JSONata expression, a JSON Schema every record is checked against. Field names follow the most-used commercial actors, so the output drops in where theirs went.
The recorded manifest, abridged: a view fed by the page’s own responses, an interrupt, a route, the machine
{
"a2flow": "0.2",
"id": "tiktok.explore",
"inputs": {
"properties": {
"maxItems": { "type": "integer", "default": 100, "maximum": 500 }
}
},
"views": {
"explore": {
"match": { "path": "^/explore/?$", "selector": "#explore-item-list" },
"model": {
"/videos": {
"net": {
"url": "/api/explore/item_list/",
"when": "$exists(itemList)",
"key": "id",
"mode": "append",
"select": "itemList.{ \"id\": id, \"text\": desc, \"playCount\": stats.playCount, \"authorMeta\": { \"name\": author.uniqueId } }"
}
}
},
"actions": {
"loadMore": {
"op": "scroll",
"by": "page",
"intent": "Scroll the explore grid down so more videos load"
}
}
},
"loginModal": {
"interrupt": true,
"match": { "selector": "#loginModalContentContainer" },
"actions": {
"dismiss": {
"op": "click",
"anchor": "dismiss",
"intent": "Close the login prompt with its X button"
}
}
},
"captcha": {
"outcome": true,
"match": {
"selector": "#captcha-verify-container, [id*=\"captcha_container\"], iframe[src*=\"captcha\"]"
}
}
},
"routes": {
"tiktok.video": {
"key": "id",
"from": {
"view": "explore",
"on": "item",
"each": "videos[[0..($surfaces.control.run.maxItems - 1)]]"
},
"extract": "$item"
}
},
"machine": {
"on": { "INTERRUPT": { "target": ".interrupted" } },
"states": {
"interrupted": {
"invoke": {
"src": "a2ui.act",
"input": { "on": "dismiss" },
"onDone": { "target": "resume" }
}
},
"resume": { "type": "history", "history": "deep" },
"loadFeed": {
"always": [
{
"guard": {
"type": "cond",
"params": {
"or": [
{ "gte": [{ "count": "/videos" }, "{{ maxItems }}"] },
{ "stalled": 3 }
]
}
},
"target": "emit"
},
{ "target": "loadingMore" }
]
},
"loadingMore": {
"invoke": {
"src": "a2ui.act",
"input": { "on": "load_more", "action": "loadMore" },
"onDone": "loadFeed"
}
}
}
}
}
What one record must look like (route schema, abridged)
{
"type": "object",
"properties": {
"id": { "type": "string" },
"text": { "type": "string" },
"createTime": { "type": "number" },
"authorMeta": { "type": "object" },
"playCount": { "type": "number" },
"diggCount": { "type": "number" },
"webVideoUrl": { "type": "string" },
"hashtags": { "type": "array" }
}
}

Those three lines under machine.on and resume are the whole interrupt story. Whatever the machine is doing when a login prompt appears, it jumps to interrupted, presses the prompt’s own close button through an A2UI action, and returns to the exact nested state it left. The interrupted step runs again, which is safe because every unit of work is idempotent.

Bring your own driver

Everything the runtime does to a browser goes through one small interface, and it evaluates expressions as strings, never closures. That is what lets one engine drive a browser from another process, another machine or a WebAssembly sandbox.

The browser surface a manifest needs (abridged)
interface BrowserAdapter {
readonly kind: "playwright" | "puppeteer" | "cdp"
goto(url: string): Promise<{ status?: number }>
evaluate<T>(expression: string): Promise<T> // a string, never a closure
click(selector: string): Promise<void>
wheel(dy: number): Promise<void>
screenshot(): Promise<Buffer>
cdp<T>(method: string, params?: object): Promise<T> // raw DevTools command
onCdp(event: string, fn: (params: any) => void): void
}
Same manifest, different driver: one flag
npx tsx bin/a2flow.ts run tiktok.explore.a2flow.json --adapter puppeteer \
--input '{"maxItems":100}'

In a browser tab, the runtime is bundled into QuickJS, a small JavaScript engine compiled to WebAssembly, running in a Web Worker. It drives a sandboxed browser on the server over a page-scoped DevTools socket. In the recorded session it started in 202 ms and collected 20 of 20 explore records in 50.1 seconds.

Drivers and what each has shown

Test results are against the project’s deterministic look-alike site
Driver or hostTestLive site
Playwrightreference records100 of 100 in 64 s
Puppeteeridentical to Playwrightnot separately measured
Raw DevTools socketidentical to Playwrightnot separately measured
WebAssembly in a tabsame records as Node20 of 20 in 50.1 s
browser_oxide (Rust)not runpassed the bot check; feed never loaded
Identical means the record sets compare equal when one manifest runs on all three Node drivers in turn. An earlier, cold WebAssembly run took 109 s for the same 20 records.

The run

On the live site, logged out, from Australia: 100 of 100 videos in 64 seconds, about 1.6 a second, after matching 5/5 target records field for field on a 20-item test.

Wall time per record
641 ms
page loads and scrolling included
CPU per record
287 ms
Node plus browser renderer
Cycles per record
316M
at the host’s clock

An earlier run the same day was four times faster per record (158 ms) and stopped at 66 of 100. Its browsing had reported success on a page that never actually scrolled, so the manifest was designed against a feed that never loaded its second screen. Scrolling is now verified step by step, and the agent may not write a manifest until it can show the page moved, nothing covers it and paging fired. The run is slower as a result, but it completes.

What it costs

Estimated cost per 1 million records

log scale · compute only, list prices · Apify prices include everything
$0.1$1$10$100$1k$10k$0.47WebAssembly on a VM · extraction only$0.82Cloudflare Workers · extraction only$16.02Cloudflare Browser Rendering$17.80Browserbase$21.37AWS Lambda, 2 GB$33.97Google Cloud Run, 2 vCPU$300Apify tiktok-search-scraper$500Apify tiktok-scraper, best tier$3.0kApify free-tiktok-scraper$3.7kApify tiktok-scraper, free tier

The same estimates, with their basis

Excludes proxies, egress, maintenance and blocking risk on the Netscraper rows
WherePer 1MBasis
AWS Lambda, 2 GB$21.37$0.0000166667 per GB-second, wall-clock
Google Cloud Run, 2 vCPU$33.97per vCPU-second and GiB-second
Apify tiktok-scraper, free tier≈ $3,700$0.0037 per result, plus a start fee per run
Apify tiktok-scraper, best tier≈ $500$0.0005 per result at the top volume tier
Cloudflare Browser Rendering$16.02$0.09 per browser-hour
Browserbase$17.80about $0.10 per browser-hour
Cloudflare Workers, extraction only$0.82$0.02 per million CPU-ms
WebAssembly on a VM, extraction only$0.47one vCPU at about $0.04 an hour
Apify free-tiktok-scraper≈ $3,000$0.003 per result
Apify tiktok-search-scraper≈ $300$0.0003 per post
Netscraper rows: 641 ms wall and 287 ms CPU per record from the recorded run, at the list prices in the Studio’s pricing table. Apify rows: published pay-per-result prices.

These figures are not like for like. A commercial actor’s price buys residential proxies, someone to fix it when the site changes, and support. What they show is that running the engine is cheap, and that fixing it, the part that used to need an engineer and a release, is now a supervised agent session measured in minutes.

Netscraper is a d/rksci labs prototype. The runtime, the manifest format and the Studio are working code; they are not offered as a hosted service.