DEV Community

Jeffrey Turov
Jeffrey Turov

Posted on

I built 11 pay-per-use data APIs that AI agents can call directly (Google Maps, TikTok, LinkedIn, review monitoring) — here's what I learned

A few months ago I published a set of Actors on Apify Store. Today they're all AI-agent ready: any LLM agent (Claude, GPT, Cursor, LangChain, n8n) can discover and call them through the Apify MCP server — no custom integration code needed.

This post is the full playbook: what the tools do, how pay-per-event monetization works (including the pricing trap that made me lose money on every sale), the bugs I hit, and how AI agents actually consume the tools.

The toolbox

Actor What it extracts Price (pay-per-event)
Google Maps Business Scraper Names, phones, websites, ratings, reviews, GPS $1.00 / run + $0.03 / business
TikTok Profile & Video Scraper Followers, likes, bio, per-video stats $0.01 / profile + $0.002 / video
Instagram Profile Scraper Followers, bio, verified, engagement $0.01 / profile
YouTube Video & Channel Scraper Views, likes, subscribers, search results $0.002 / video
LinkedIn Profile Scraper Headlines, companies, skills, experience $0.02 / profile
RAG Web Browser Clean Markdown from any URL + Google search $0.003 / page
Fuel Prices France API Real-time prices, 9,800 stations, GPS $0.20 / run + $0.01 / 1k stations
Hotel Rate Monitoring Competitor rates, parity checks fractions of a cent per item
API Breaking-Change Radar Diffs OpenAPI specs, classifies changelogs, alerts $0.50 / run + $0.50 / breaking change
Review Radar New Google reviews for a business, Slack alerts $0.25 / run + $0.01 / new review
Review Pitch Generator Worst reviews → ready-to-send sales report $0.25 / run + $0.10 / pitch

The last three are a different breed: not scrapers but monitors — they keep state between runs and only bill when they find something.

Why "AI-agent ready" changes everything

The old model: a human finds your scraper on the store, reads the docs, clicks buttons.

The new model: an AI agent gets a task ("find me 50 plumbers in Austin with their phone numbers"), searches the Apify Store via MCP, reads the actor's README and input schema, and calls it — end to end, no human.

For that to work, three things must be true:

  1. Your README is written for an LLM, not just humans. Mine now all start with a "Use this tool when..." section — that's what the agent pattern-matches against the user's request.
  2. Your input schema has a description on every field. The agent constructs the JSON input from those descriptions. No description = hallucinated parameters = failed runs = no revenue.
  3. Your output is documented field by field. The agent needs to know what it gets back to reason over it.

Here's the actual flow with the Apify MCP server (https://mcp.apify.com — add it to Claude Desktop or Cursor in 30 seconds):

User: "Get me the follower counts of these 5 TikTok creators"
Agent: → search-actors("tiktok profile")
       → fetch-actor-details (reads README + input schema)
       → call-actor(travelmonitorlab/tiktok-scraper,
                    {"profiles": [...], "maxVideosPerProfile": 0})
       → returns structured JSON
Enter fullscreen mode Exit fullscreen mode

Or skip MCP entirely — every actor is a single synchronous HTTP call:

curl -X POST "https://api.apify.com/v2/acts/travelmonitorlab~google-maps-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"queries": ["plumbers Austin TX"], "maxResults": 50}'
Enter fullscreen mode Exit fullscreen mode

Monetization: pay-per-event (and the trap that cost me real money)

Apify offers several pricing models. For new actors, PRICE_PER_DATASET_ITEM is rejected — you must use PAY_PER_EVENT. The model is better anyway: you define events and charge explicitly in code:

await Actor.charge(event_name="business-scraped")
await Actor.push_data(item)
Enter fullscreen mode Exit fullscreen mode

Trap #1 — the vanishing charge: put the charge AFTER crawler.run() and your event loop may already be closed — the charge silently vanishes, you deliver data for free. Charge inside the handler, right before pushing data. Always verify with chargedEventCounts in the run object after a test run.

Trap #2 — pricing below your compute cost (this one actually bit me): on Apify's free plan, the actor owner pays the user's compute. My Google Maps actor was priced at $0.005/business. A customer scraped 16 businesses → I earned $0.08, and paid $1.23 in Playwright + residential proxy costs. Margin: −1436%. Every sale lost money. The structural fix: a flat run-started event that covers fixed costs (browser launch, proxy session) plus a per-result event with real margin:

await Actor.charge(event_name="run-started")   # first line of main()
Enter fullscreen mode Exit fullscreen mode

That same 16-business run now bills $1.48 instead of $0.08. If you build browser-based actors, do this from day one.

Setting pricing is pure API — and pricingInfos is append-only: you can't edit a tier in place, you append a new one:

PUT /v2/acts/{actorId}
{"pricingInfos": [existing_tier_verbatim, {
  "pricingModel": "PAY_PER_EVENT",
  "reasonForChange": "Flat run fee to cover compute",
  "pricingPerEvent": {"actorChargeEvents": {
    "run-started": {"eventTitle": "Run started", "eventDescription": "...", "eventPriceUsd": 1.0},
    "business-scraped": {"eventTitle": "Business scraped", "eventDescription": "...", "eventPriceUsd": 0.03}
  }}}]}
Enter fullscreen mode Exit fullscreen mode

Apify takes 20%. Compute is paid by the user (on paid plans); you pocket the event fees.

Trap #3 — self-billing: running your own PPE actor bills your own account. Build a dryRun input flag that wraps every Actor.charge() and skips it — you can then integration-test the full pipeline for free and only flip billing on for real runs.

The QA gauntlet (or: your actor must survive an empty input)

Apify automatically runs every store actor with the schema's prefilled input and expects success within 5 minutes — three failures and your actor gets flagged "Under maintenance", then deprecated. Two lessons learned the hard way:

  • Never raise on empty input. If the user (or the QA bot) provides no target, fall back to a small built-in demo (2 results, dry-run forced) and exit 0. An actor that raises RuntimeError("no target") is an actor on the deprecation list.
  • Keep the prefill tiny. The QA run is baked into your build: a heavy prefill (20 Playwright results through residential proxies) blows the 5-minute budget and flags you. Prefill = 2 results max, dry-run on.

Battle scars (so you don't get them)

  • Crawlee 1.8 breaking changes: purge_on_start and navigation_timeout_secs are no longer valid kwargs — use page.set_default_navigation_timeout() in a pre_navigation_hook.
  • Google Maps never fires load: analytics keep streaming forever, so navigation always times out. Fix: abort images/fonts/media via page.route() (the handler must be a coroutine, not a lambda).
  • Residential proxies are mandatory for Google Maps, and you must pass actor_proxy_input= as a named argument to Actor.create_proxy_configuration().
  • French number formats will crash your floats: "4,8" → replace comma; "1 234" reviews can use \xa0 or \u202f as thousand separator.
  • A "SUCCEEDED" run can contain zero useful data. Always check itemCount + sample the dataset + read the end of the log.
  • Don't run 7 queries × 25 results in one run. Split into parallel runs of ≤5 queries × 15 results; retry failed ones sequentially (residential proxy tunnels occasionally die).
  • Pin your dependencies. apify>=3.2,<4 + crawlee>=1.7,<2 is a proven combo; open ranges break within weeks as PyPI drifts.

The distribution strategy

Publishing on the store is step 0. What actually moves the needle:

  1. README written for LLMs — agents choose tools whose docs they can parse. Clear "use when", typed inputs, example I/O. The MCP server's own system prompt tells agents to read your README before calling — it's your storefront.
  2. Niche SEO titles — "Google Maps Scraper" is saturated (the official one has 500k+ users); "Fuel Prices France API" has zero competition on the store.
  3. Monitors > scrapers — a scraper competes with everyone; a stateful monitor (breaking changes, new reviews) with Slack alerts answers a recurring pain and justifies a per-run flat fee.
  4. Dogfooding — I use my own Google Maps actor to build lead lists I sell elsewhere. Every sale is also a demo.

Try them

All 11 actors are live on Apify Store. If you build agents, add https://mcp.apify.com to your MCP client and just ask for the data — the agent will find the tools.

Feedback, bugs, feature requests: open an issue on any actor page, I answer fast.

Top comments (0)