Writing a scraper usually means inspecting the page, finding CSS selectors, handling edge cases, and then fixing it all again when the site changes its layout. For a lot of jobs that's overkill: you just want these five facts from these pages.
So I built an extractor that works the other way around. You describe the data, an LLM reads the page, and you get JSON back.
The whole input
{
"startUrls": [{ "url": "https://github.com/apify/crawlee" }],
"fields": {
"name": "string",
"description": "string",
"license": "string",
"primary_language": "string"
}
}
The output (a real run)
{
"url": "https://github.com/apify/crawlee",
"success": true,
"data": {
"name": "crawlee",
"description": "Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers...",
"license": "Apache License 2.0",
"primary_language": "JavaScript"
}
}
No selectors, no code, and the same input works on any other repository page, or on a completely different site if you change the fields.
Things I cared about
-
No invented data. The model is instructed to use only what's on the page and return
nullotherwise. Missing is better than made up. -
Types you choose.
string,number,integer,booleanorarrayper field, or a full JSON Schema for nested results like "all jobs on this page with title, location and remote flag". - Plain-language instructions. "Prices in USD as numbers" or "only the first listing" steer the result without code.
- Works on stubborn sites. Pages are fetched over plain HTTP first; if a site blocks that or needs JavaScript, a real browser takes over automatically.
- Pay only for results. Blocked, failed and empty pages are free. It's $0.01 per extracted page, about a third of what comparable AI scrapers on the same platform charge.
Good fits
- Product pages → name, price, availability
- Job posts → title, company, location, salary, remote
- Company websites → what they do, industry, headquarters, contact page
- Articles → headline, author, date, summary
- Event pages → date, venue, price
When not to use it
If you need millions of pages from one site with a fixed layout, a classic selector-based scraper is cheaper. This tool shines when pages vary, layouts change, or you only need a handful of fields from many different sites.
Try it
AI Web Data Extractor on Apify. New Apify accounts get free monthly credit, so you can try it on your own URLs for free. It can also be called from code or used as a tool by AI agents through the Apify MCP server.
Feedback welcome, especially examples where it gets something wrong.
Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.