<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Jeffrey Turov</title>
    <description>The latest articles on DEV Community by Jeffrey Turov (@jeffreyturov).</description>
    <link>https://dev.to/jeffreyturov</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4050944%2Fabfb8306-248c-4a94-8471-d954f41366c3.png</url>
      <title>DEV Community: Jeffrey Turov</title>
      <link>https://dev.to/jeffreyturov</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/jeffreyturov"/>
    <language>en</language>
    <item>
      <title>Nature Detective — the offline AI that sends my kids back outside</title>
      <dc:creator>Jeffrey Turov</dc:creator>
      <pubDate>Tue, 06 Oct 2026 09:10:54 +0000</pubDate>
      <link>https://dev.to/jeffreyturov/nature-detective-the-offline-ai-that-sends-my-kids-back-outside-5e5g</link>
      <guid>https://dev.to/jeffreyturov/nature-detective-the-offline-ai-that-sends-my-kids-back-outside-5e5g</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/hacktoberfest-week1-2026-10-05"&gt;Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Last Sunday I watched my kids walk past a hundred fascinating things without seeing a single one.&lt;/p&gt;

&lt;p&gt;We live in Mersch, in the middle of Luxembourg — fifteen minutes from forests that have more species per square meter than my phone has apps. And yet the walk was negotiations: five more minutes of screen, then we'll go. The screen is the destination. The forest is the commute.&lt;/p&gt;

&lt;p&gt;So for this challenge I built the opposite deal: &lt;strong&gt;the screen becomes the reason to go into the forest — and the forest stays the point.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Nature Detective&lt;/strong&gt; is an offline-first nature companion for kids aged 5-10.&lt;/p&gt;

&lt;p&gt;A child photographs something alive — a plant, a beetle, a mushroom. An &lt;strong&gt;open-weight vision model (Gemma 3 4B, running entirely on local hardware)&lt;/strong&gt; identifies the find, tells one true fun fact a seven-year-old would love, sets a safety rule, and then does the thing no engagement-optimized app would ever do: &lt;strong&gt;it sends the kid back outside.&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;🎯 &lt;em&gt;"Count how many different colored flowers you can spot before the next corner!"&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Every identification comes with a mission — count, find, compare, observe. Missions never involve the screen. The screen is the shortest part of the experience, by design.&lt;/p&gt;

&lt;p&gt;Every discovery lands in a &lt;strong&gt;field journal&lt;/strong&gt; on the device, with a detective rank that grows with real outdoor finds (🐣 → 🐾 → 🦊 → 🦉). No account, no ads, no streak anxiety — just a collection of afternoons.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://main-michael-stanford-myrtle.trycloudflare.com" rel="noopener noreferrer"&gt;Try the live demo&lt;/a&gt;&lt;/strong&gt; — it runs on my own machine, on Gemma 3, right now.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fum3qhout3dvhrby56zxp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fum3qhout3dvhrby56zxp.png" alt="The result screen: a bee identified, a safety badge, the next outdoor mission" width="800" height="2089"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;The full flow, no signup, works from a phone browser:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Open the &lt;a href="https://main-michael-stanford-myrtle.trycloudflare.com" rel="noopener noreferrer"&gt;live demo&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Snap a discovery (or upload any plant/insect photo)&lt;/li&gt;
&lt;li&gt;Get the identification, the fact, the safety badge — and your mission&lt;/li&gt;
&lt;li&gt;Check the field journal&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Fair warning: inference on a 24-core CPU takes ~15-60 seconds depending on load. Slow enough to look at the real thing while you wait — which, honestly, is also the point.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/jeffreyturov-dev" rel="noopener noreferrer"&gt;
        jeffreyturov-dev
      &lt;/a&gt; / &lt;a href="https://github.com/jeffreyturov-dev/nature-detective" rel="noopener noreferrer"&gt;
        nature-detective
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;🔍 Nature Detective&lt;/h1&gt;
&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;An offline-first nature companion that gets kids off the screen and into the woods.&lt;/strong&gt;
Built for the Hacktoberfest 2026 DEV Challenge — Week 1: &lt;em&gt;Touch Grass&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Live demo:&lt;/strong&gt; &lt;a href="https://main-michael-stanford-myrtle.trycloudflare.com" rel="nofollow noopener noreferrer"&gt;https://main-michael-stanford-myrtle.trycloudflare.com&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;A child on a hike photographs a plant, a bug, a mushroom. An &lt;strong&gt;open-weight vision
model (Gemma 3 4B) running locally&lt;/strong&gt; identifies the find, tells one true fun fact
sets a safety rule, and hands out the &lt;strong&gt;next outdoor mission&lt;/strong&gt; — count, find
compare. Every discovery lands in a field journal on the device.&lt;/p&gt;
&lt;p&gt;No account. No cloud. No photo, no GPS coordinate, no child's data ever leaves
your own hardware. It runs where the forest has no bars — that's the whole point.&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;Why open matters here&lt;/h2&gt;

&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No internet, no problem.&lt;/strong&gt; Inference runs on a local machine (a laptop, or a
Raspberry Pi in a backpack acting as a Wi-Fi hotspot). The forest has no 5G…&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/jeffreyturov-dev/nature-detective" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;p&gt;The whole stack, MIT licensed:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Piece&lt;/th&gt;
&lt;th&gt;Choice&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Vision + reasoning&lt;/td&gt;
&lt;td&gt;Gemma 3 4B via Ollama (Q4_K_M)&lt;/td&gt;
&lt;td&gt;open-weight, multimodal, runs on CPU&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Backend&lt;/td&gt;
&lt;td&gt;FastAPI + SQLite&lt;/td&gt;
&lt;td&gt;one file, zero infra&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Frontend&lt;/td&gt;
&lt;td&gt;Vanilla JS PWA, zero external assets&lt;/td&gt;
&lt;td&gt;the offline claim has to be true&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost&lt;/td&gt;
&lt;td&gt;€0&lt;/td&gt;
&lt;td&gt;no API key, no per-call fee, ever&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;One model call turns a photo into a strict JSON object: identification, kid fact, safety level, next mission, quiz question. Everything the screen shows, in a single pass — because a children's app in a forest cannot afford round trips, and mine literally has none to afford: it runs where there is no signal.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;📱 phone camera ──► FastAPI ──► Gemma 3 (local vision) ──► strict JSON
                                                              │
        field journal (SQLite) ◄── safety net (code) ◄── verifier pass
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The interesting part is not the pipeline. It's what the pipeline refuses to do.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part I'm most proud of: what it refuses to do
&lt;/h2&gt;

&lt;p&gt;During testing, I fed the app an abstract watercolor — green and yellow blur, no living thing in it. A single-pass version of the app answered, with high confidence: &lt;strong&gt;"Spiderweb, Araneae."&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;It was a confident, detailed, completely fabricated answer. The most dangerous kind of wrong — the kind a child would believe, remember, and repeat at school.&lt;/p&gt;

&lt;p&gt;An app that teaches nature to kids has one job that outranks every other: &lt;strong&gt;never teach a lie.&lt;/strong&gt; So now every identification goes through a second pass — a verifier, running on the same local model, whose only job is to doubt the first answer. The trick that makes it actually work: the verifier &lt;strong&gt;never sees the claim first&lt;/strong&gt;. It must describe what it objectively sees — &lt;em&gt;"The image is an abstract painting, not a living organism."&lt;/em&gt; — and only then decide whether the claim matches its own independent observation. Ask it to judge the claim directly and it politely agrees with everything; make it commit to its own eyes first, and it catches the lie.&lt;/p&gt;

&lt;p&gt;The abstract painting now gets the honest answer: &lt;em&gt;"A first guess said this might be a spiderweb, but my double-check disagreed — and a good detective never teaches a maybe as a fact. Ask a grown-up or a field guide!"&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Teaching a child that "I don't know" is a respectable answer might be the most valuable feature in the whole app.&lt;/p&gt;

&lt;p&gt;Safety works the same way — enforced twice, because prompt rules are suggestions and &lt;strong&gt;code rules are guarantees&lt;/strong&gt;. The model is instructed that mushrooms and berries are always &lt;em&gt;look, don't touch&lt;/em&gt;. And then the backend overrides it in code anyway, every time, no matter what the model says. When I tested it with a fly agaric (&lt;em&gt;Amanita muscaria&lt;/em&gt; — the red one with white dots, beautiful and toxic), it identified it correctly &lt;em&gt;and&lt;/em&gt; refused any interaction with it, twice over.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Does Open Innovation Matter?
&lt;/h2&gt;

&lt;p&gt;This project doesn't just &lt;em&gt;use&lt;/em&gt; open weights. &lt;strong&gt;It only exists because of them.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The forest has no bars.&lt;/strong&gt; The whole point is a place with no signal. A closed API is a product that stops working exactly where mine starts working. Local inference isn't an optimization here — it's the product.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;I will never upload my kids' photos to someone else's server.&lt;/strong&gt; Not their faces, not their location, not their finds. With Gemma running on my own hardware, nothing ever leaves the device. That's not a privacy policy; it's physics.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;€0, forever.&lt;/strong&gt; No API key, no subscription, no per-call cost — so it can run in a school, a scout group, a nature club with no budget, anywhere in the world.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ownable and swappable.&lt;/strong&gt; Any open-weight vision model Ollama can serve drops in with a one-line change. When a better small multimodal model ships next month, every Nature Detective gets smarter for free. Try getting that guarantee from a deprecated API version.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A closed model would have made this easier to demo and impossible to believe in.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prize Categories
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Best Use of Gemma&lt;/strong&gt; — Gemma 3 4B is the entire brain of the project: species identification, kid-level explanations, mission generation, and the self-verification pass, all running locally through Ollama.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;This Sunday we're going back to the forest. My kids already asked if the detective is coming.&lt;/p&gt;

&lt;p&gt;He's in my pocket. He knows nothing about engagement metrics. And his favorite sentence, by far, is: &lt;em&gt;"Now go look for yourself."&lt;/em&gt; 🌿&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>hf26challenge</category>
      <category>gemma</category>
      <category>opensource</category>
    </item>
    <item>
      <title>I Put 6,522 Birds in My Backpack: Offline Bird Call ID on a Laptop, Zero Internet</title>
      <dc:creator>Jeffrey Turov</dc:creator>
      <pubDate>Tue, 06 Oct 2026 08:11:43 +0000</pubDate>
      <link>https://dev.to/jeffreyturov/i-put-6522-birds-in-my-backpack-offline-bird-call-id-on-a-laptop-zero-internet-4pda</link>
      <guid>https://dev.to/jeffreyturov/i-put-6522-birds-in-my-backpack-offline-bird-call-id-on-a-laptop-zero-internet-4pda</guid>
      <description>&lt;p&gt;&lt;em&gt;This is my submission for the &lt;a href="https://dev.to/challenges/hacktoberfest-week1-2026-10-05"&gt;Hacktoberfest Open-Source AI Challenge: Week 1 — Touch Grass&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem with birding apps
&lt;/h2&gt;

&lt;p&gt;The best birding happens exactly where your phone becomes a brick: deep forest, mountain trails, that marsh 40 minutes from the nearest cell tower. The moment a call you don't recognize echoes through the trees, the usual flow is: record it, &lt;em&gt;hope you remember it later&lt;/em&gt;, upload it when you get home, wait for the cloud.&lt;/p&gt;

&lt;p&gt;I wanted the whole loop to close &lt;strong&gt;on the trail&lt;/strong&gt;. So I built &lt;code&gt;trail-bird-id&lt;/code&gt;: point it at an audio file, get the species. No internet. No account. No API key. No per-call cost. Just a laptop (or a Raspberry Pi) and an open-weight model.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;python identify.py trail_recording_07h42.mp3

Analyzing trail_recording_07h42.mp3 &lt;span class="o"&gt;(&lt;/span&gt;fully offline, BirdNET V2.4, 6,522 species&lt;span class="o"&gt;)&lt;/span&gt;...

  1. Black-capped Sparrow &lt;span class="o"&gt;(&lt;/span&gt;Arremon abeillei&lt;span class="o"&gt;)&lt;/span&gt;      &lt;span class="nv"&gt;conf&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;0.99  @ 0.0-3.0s
  2. Streaked Saltator &lt;span class="o"&gt;(&lt;/span&gt;Saltator striatipectus&lt;span class="o"&gt;)&lt;/span&gt;   &lt;span class="nv"&gt;conf&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;0.86  @ 3.0-6.0s
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The build: embarrassingly little code
&lt;/h2&gt;

&lt;p&gt;The heavy lifting is &lt;a href="https://github.com/birdnet-team/BirdNET-Analyzer" rel="noopener noreferrer"&gt;BirdNET&lt;/a&gt; — an open-source audio classifier from the Cornell Lab of Ornithology and Chemnitz University of Technology (CC BY-SA 4.0), trained on 6,522 species. It ships as a ~50 MB TFLite/TensorFlow model &lt;em&gt;inside the pip package&lt;/em&gt;, which is the whole trick: installing the library means installing the brain.&lt;/p&gt;

&lt;p&gt;My contribution is a ~60-line wrapper, &lt;a href="https://github.com/jeffreyturov-dev/trail-bird-id/blob/master/identify.py" rel="noopener noreferrer"&gt;&lt;code&gt;identify.py&lt;/code&gt;&lt;/a&gt;, that runs the model in 3-second windows over any wav/mp3/ogg file and aggregates to the best-confidence detection per species:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;uv venv &lt;span class="nt"&gt;--python&lt;/span&gt; 3.11 .venv
uv pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;--python&lt;/span&gt; .venv/bin/python birdnet-analyzer
.venv/bin/python identify.py your_recording.mp3 &lt;span class="nt"&gt;--top&lt;/span&gt; 5
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's the entire setup. No Docker, no GPU, no cloud credentials.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does it actually work? Verified against ground truth
&lt;/h2&gt;

&lt;p&gt;I don't trust demos that only show one cherry-picked clip, so I tested against four field recordings from Wikimedia Commons where the species is known from the recording metadata (xeno-canto / iNaturalist sourced):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Recording&lt;/th&gt;
&lt;th&gt;Expected species&lt;/th&gt;
&lt;th&gt;Top-1 prediction&lt;/th&gt;
&lt;th&gt;Confidence&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Black-capped Sparrow XC250490 (Niels Krabbe, CC BY-SA)&lt;/td&gt;
&lt;td&gt;&lt;em&gt;Arremon abeillei&lt;/em&gt;&lt;/td&gt;
&lt;td&gt;✅ &lt;strong&gt;Black-capped Sparrow&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;0.99&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Australian Magpie song (CC BY-SA)&lt;/td&gt;
&lt;td&gt;&lt;em&gt;Gymnorhina tibicen&lt;/em&gt;&lt;/td&gt;
&lt;td&gt;✅ &lt;strong&gt;Australian Magpie&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;0.99&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Brown Hawk-Owl, South Bengal (CC BY)&lt;/td&gt;
&lt;td&gt;&lt;em&gt;Ninox scutulata&lt;/em&gt;&lt;/td&gt;
&lt;td&gt;✅ &lt;strong&gt;Brown Boobook&lt;/strong&gt; &lt;em&gt;(same bird, name updated by taxonomists — the model is more current than the file title)&lt;/em&gt;
&lt;/td&gt;
&lt;td&gt;1.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Guira Cuckoo (CC0)&lt;/td&gt;
&lt;td&gt;&lt;em&gt;Guira guira&lt;/em&gt;&lt;/td&gt;
&lt;td&gt;✅ &lt;strong&gt;Guira Cuckoo&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;0.95&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;4/4 top-1 correct&lt;/strong&gt;, across four continents and four very different vocalizations (sparrow song, magpie caroling, an owl's hoot, cuckoo chatter). Runtime: ~6.7 seconds wall-clock for a 40-second recording, on a plain CPU.&lt;/p&gt;

&lt;h2&gt;
  
  
  Proving "offline" isn't marketing
&lt;/h2&gt;

&lt;p&gt;"Works offline" is easy to claim. So I re-ran identification with &lt;strong&gt;all egress blocked&lt;/strong&gt; — every proxy environment variable pointed at a dead localhost port, so any HTTP call from the Python stack would fail instantly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;env &lt;/span&gt;&lt;span class="nv"&gt;HTTPS_PROXY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;http://127.0.0.1:9 &lt;span class="nv"&gt;HTTP_PROXY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;http://127.0.0.1:9 &lt;span class="se"&gt;\&lt;/span&gt;
      &lt;span class="nv"&gt;ALL_PROXY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;socks5://127.0.0.1:9 &lt;span class="nv"&gt;NO_PROXY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
      .venv/bin/python identify.py audio/australian_magpie.ogg

  1. Australian Magpie &lt;span class="o"&gt;(&lt;/span&gt;Gymnorhina tibicen&lt;span class="o"&gt;)&lt;/span&gt;   &lt;span class="nv"&gt;conf&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;0.99  @ 9.0-12.0s
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Still works, because there is no network code to fail. The weights live in &lt;code&gt;site-packages/birdnet_analyzer/checkpoints/&lt;/code&gt;. The trail is the deployment target, and the trail has no SLA.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why open innovation matters here
&lt;/h2&gt;

&lt;p&gt;This is where the open approach doesn't just match the closed one — it &lt;em&gt;wins&lt;/em&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;It works where the closed apps can't.&lt;/strong&gt; Zero-connectivity is the core use case, not an edge case. A cloud API is a non-starter on a ridgeline.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Nothing leaves your device.&lt;/strong&gt; Field recordings often capture more than birds — your voice, your location, your hiking companions. With a local model, that data never touches a server you don't control.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You can actually own it.&lt;/strong&gt; Pin the model version, tune the confidence threshold for your region, swap in a species list for your local patch (&lt;code&gt;--lat/--lon&lt;/code&gt; filters to species plausible for your coordinates), or fine-tune on your own recordings. Try doing that with a closed mobile app.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost at scale is zero.&lt;/strong&gt; A citizen-science project processing 10,000 hours of soundscape pays nothing per call — because there are no calls.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Honest limitations
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;BirdNET needs a reasonably clean recording; heavy wind or overlapping dawn-chorus can drop confidence. The 3-second windowing means very short calls can be missed.&lt;/li&gt;
&lt;li&gt;My test set is 4 recordings, not 4,000 — it's a smoke test with known ground truth, not a benchmark. (BirdNET's own peer-reviewed evaluations cover the rigorous part.)&lt;/li&gt;
&lt;li&gt;I built and validated this on real field recordings made by others; the full "take it outside" bonus round happens this weekend on my local trails — I'll update the repo with what it hears.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Repo:&lt;/strong&gt; &lt;a href="https://github.com/jeffreyturov-dev/trail-bird-id" rel="noopener noreferrer"&gt;https://github.com/jeffreyturov-dev/trail-bird-id&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model:&lt;/strong&gt; &lt;a href="https://github.com/birdnet-team/BirdNET-Analyzer" rel="noopener noreferrer"&gt;BirdNET-Analyzer&lt;/a&gt; (CC BY-SA 4.0)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The screen is the shortest part of this experience: record outside, identify anywhere, get back to listening. 🌲&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>hf26challenge</category>
      <category>touchgrass</category>
      <category>opensource</category>
    </item>
    <item>
      <title>A Voz dos Meus — my father waited with flowers. I built a memory that never forgets</title>
      <dc:creator>Jeffrey Turov</dc:creator>
      <pubDate>Sun, 04 Oct 2026 19:16:01 +0000</pubDate>
      <link>https://dev.to/jeffreyturov/a-voz-dos-meus-my-father-waited-with-flowers-i-built-a-memory-that-never-forgets-1h97</link>
      <guid>https://dev.to/jeffreyturov/a-voz-dos-meus-my-father-waited-with-flowers-i-built-a-memory-that-never-forgets-1h97</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/hacktoberfest-weekend-2026-10-01"&gt;Hacktoberfest Weekend Challenge: Build for a Friend&lt;/a&gt; — and for the **Best Use of Gemma&lt;/em&gt;* category.*&lt;/p&gt;




&lt;p&gt;In 1996, a six-year-old boy got on a bus in Amares, near Braga, for a long ride north with other Portuguese families. He doesn't remember any other children on that bus — emigration was mostly parents back then. When the bus finally reached Luxembourg, his father was standing there, waiting for them, holding a bunch of flowers for his mother.&lt;/p&gt;

&lt;p&gt;That boy was me. The man with the flowers is my father.&lt;/p&gt;

&lt;p&gt;And today I'm a father myself — two kids who know Luxembourg as home, who speak French at school, whose Portuguese is a work in progress, and who have never seen their grandfather wait for anyone with flowers. We film our children growing up. Nobody films our parents growing old. My father carried his whole life to Luxembourg and lets the story out only in pieces, at the table, when something reminds him — and if you ask directly, he shrugs: &lt;em&gt;"Não é nada."&lt;/em&gt; It's nothing.&lt;/p&gt;

&lt;p&gt;It's not nothing. So for this challenge I built the person I love most a memory the family can keep talking to — and the very first voice I put in it was my own, telling the story where he's the hero. Because my memories belong to my two kids now, too: one day they'll want to know who their father was at six, and their grandfather with the flowers.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Voz dos Meus — The Voice of the Ones We Love
&lt;/h2&gt;

&lt;p&gt;The concept fits in one sentence, and it isn't really about my father:&lt;br&gt;
&lt;strong&gt;the people we love most are the ones we record least — so capture their voices, and let the family ask their memory questions. Forever.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Father, grandmother, aunt, the friend who's like family — &lt;em&gt;pai, avó, tia&lt;/em&gt; — the concept is the same, which is why the app is called &lt;em&gt;A Voz dos Meus&lt;/em&gt;, "the voice of my people". My father is simply the reason it exists.&lt;/p&gt;

&lt;p&gt;Three screens, because he's not a user, he's my father:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7yeqxca8pv5dklvc7048.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7yeqxca8pv5dklvc7048.png" alt="Asking the memory a question"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Contar história&lt;/strong&gt; — one big red button. Press it, talk as long as you want, press it again. That's the entire interface. (There is no step 2.)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Histórias&lt;/strong&gt; — every story transcribed, with the original audio underneath, so the grandchildren hear &lt;em&gt;his&lt;/em&gt; voice, not a synthesizer's.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Perguntar&lt;/strong&gt; — the family asks: &lt;em&gt;"Pai, como é que nós viemos para o Luxemburgo?"&lt;/em&gt; — and the memory answers in warm, simple European Portuguese. Read it instantly with &lt;strong&gt;Ouvir 🔊&lt;/strong&gt;, translate it for the grandkids with &lt;strong&gt;Traduire 🇫🇷&lt;/strong&gt; — or press the button that still gives me goosebumps:&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Na voz dele 🧡&lt;/strong&gt; — the answer is synthesized &lt;em&gt;in his own cloned voice&lt;/em&gt;, by an open-source model running in my home.&lt;/p&gt;

&lt;p&gt;Full honesty: today it comes out with a slight &lt;strong&gt;Brazilian&lt;/strong&gt; accent — the open model's Portuguese was trained mostly on Brazilian data, and his is from Amares. A closed service would hide that kind of rough edge behind a polished voice you can't inspect. Here, we can see the limitation, name it, and swap the component the day a better European-Portuguese voice model ships — without asking anyone's permission. That's the point of building on open pieces: the roadmap belongs to us.&lt;/p&gt;

&lt;p&gt;Here's the full flow in motion — question, grounded answer, the story library, the one-button recorder:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcgm4wposnkchqtgxbu6j.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcgm4wposnkchqtgxbu6j.gif" alt="Demo of the full flow"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Live demo:&lt;/strong&gt; &lt;a href="https://tariff-ventures-income-insights.trycloudflare.com" rel="noopener noreferrer"&gt;https://tariff-ventures-income-insights.trycloudflare.com&lt;/a&gt; — a temporary tunnel to the same home box my family uses; it's a kitchen appliance, not a datacenter, so be patient with it (voice cloning on a CPU takes ~2 minutes — real time for real love).&lt;br&gt;
&lt;strong&gt;Source:&lt;/strong&gt; &lt;a href="https://github.com/jeffreyturov-dev/avozdosmeus" rel="noopener noreferrer"&gt;github.com/jeffreyturov-dev/avozdosmeus&lt;/a&gt; (MIT — take it, build it for someone you love).&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4c3ni3r0j6s5q2n7so7o.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4c3ni3r0j6s5q2n7so7o.png" alt="The first story in the archive"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The part that matters most: what it refuses to do
&lt;/h2&gt;

&lt;p&gt;The first version of the "ask" feature did something that looked like a success and was actually a betrayal.&lt;/p&gt;

&lt;p&gt;I asked it: &lt;em&gt;"Pai, qual era o teu prato preferido quando eras pequeno?"&lt;/em&gt; — what was your favorite dish as a kid? I (the &lt;em&gt;Pai&lt;/em&gt; of this archive) have never told that story. The model, eager to please, &lt;strong&gt;invented one&lt;/strong&gt; — grilled sardines with potatoes and kale, warm and plausible and completely false. A memory machine that hallucinates isn't a memory machine. It's a fiction machine wearing a father's voice.&lt;/p&gt;

&lt;p&gt;For a product, that's a bug. For a family archive, it's a moral failure. If my children ask this thing a question in twenty years, I need to trust the answer the way I'd trust their grandfather.&lt;/p&gt;

&lt;p&gt;So I rebuilt the answer path around a refusal, in three layers:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Retrieval&lt;/strong&gt; — the question is embedded and matched against story chunks (open embeddings, local SQLite).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A strict classifier&lt;/strong&gt; — before &lt;em&gt;any&lt;/em&gt; prose is generated, a separate pass asks: &lt;em&gt;"Do these stories explicitly contain the answer? Reply only SIM or NÃO."&lt;/em&gt; If NÃO, the generative model &lt;strong&gt;is never called&lt;/strong&gt;. A model cannot hallucinate what it never sees.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No-invention rules&lt;/strong&gt; — when the stories do contain the answer, the answering prompt runs under absolute rules: no invented people, places, dates or details. Not even to complete a sentence nicely.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Today, that same trap question gets this answer — every time, deterministically:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"O Pai ainda não contou essa história. Pergunta-lhe na próxima visita — e se ele contar, grava-a aqui. 🧡"&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;"Your father hasn't told that story yet. Ask him on the next visit — and if he tells it, record it here."&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Forb1va1rjhee4g7ht3bu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Forb1va1rjhee4g7ht3bu.png" alt="The honest refusal"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It admits the gap, and — my favorite detail — &lt;strong&gt;it sends the family back to him&lt;/strong&gt;. The app knows its job is not to replace my father. Its job is to make sure he gets asked while he can still answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it works (everything open, everything local)
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt; his voice ─▶ faster-whisper ─▶ story text ─▶ chunks ─▶ embeddings ─▶ SQLite
                                                         │
 a question ─▶ embed ─▶ retrieve ─▶ strict classifier ─▶ Gemma 3 (local,
              "never invent" rules) ─▶ answer ─▶ read aloud ─▶ or cloned
              in HIS voice by XTTS v2 — all of it inside my home
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;faster-whisper&lt;/strong&gt; (&lt;code&gt;large-v3-turbo&lt;/code&gt;, int8) transcribes the stories on a plain CPU — no GPU anywhere in this project.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;nomic-embed-text&lt;/strong&gt; vectors in a humble SQLite file. His whole memory is a folder you can copy to a USB stick.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gemma 3 (4B)&lt;/strong&gt; via Ollama is the brain: the honesty classifier, the grounded answering, story titling, and the Portuguese→French translation for the grandchildren.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;XTTS v2&lt;/strong&gt; (open-source voice cloning) reads answers &lt;em&gt;in his voice&lt;/em&gt; — synthesized locally, from one minute of his real audio.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Browser SpeechSynthesis&lt;/strong&gt; handles the instant read-aloud — so even that works with the router unplugged.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The backend is a single Flask file. The frontend is one HTML page. No Docker, no accounts, no API keys, no build step. An old laptop on the home Wi-Fi is the entire infrastructure — which is precisely the point.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why open-source wasn't a choice here — it was the requirement
&lt;/h2&gt;

&lt;p&gt;Every "why open matters" section talks philosophy. Mine is simpler: &lt;strong&gt;I was not going to put my father's voice on a stranger's server.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;His stories are among the most private data my family owns.&lt;/strong&gt; Where he came from, who he waited for with flowers. A closed API would mean shipping that intimacy to a datacenter I can't see, under terms I can't negotiate. With open-weight models running in my home, his voice — and its digital double — physically never leaves the house. The threat model is my Wi-Fi password.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Voice cloning makes this non-negotiable.&lt;/strong&gt; I'm creating a model of a real person's voice. If that capability lived in a cloud account, the consent conversation gets murky fast. Local and open means &lt;em&gt;his voiceprint belongs to him, full stop&lt;/em&gt; — I can show him exactly where it lives, and delete it by deleting a folder.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It has to work in a kitchen, not in a demo.&lt;/strong&gt; No internet? Everything above still runs. Closed models turn into a blank screen the moment the connection drops — his memory doesn't.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It has to cost €0 forever.&lt;/strong&gt; A subscription is a thing that gets cancelled. This is meant to outlive my father, and ideally me. Open models don't send invoices to my children.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It has to be ownable by the next family.&lt;/strong&gt; Every component — Whisper, the embeddings, Gemma, XTTS, the UI — is open and swappable. When a better open Portuguese model appears, his memory upgrades without asking anyone's permission.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That's what "open innovation" means when the person is someone you love: not a license badge, but a promise you can actually keep — &lt;em&gt;your voice stays yours&lt;/em&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What happens next Sunday
&lt;/h2&gt;

&lt;p&gt;The roadmap writes itself, and it isn't mine anymore — it's the family's: one archive per person (&lt;em&gt;Pai&lt;/em&gt; today, &lt;em&gt;Avó&lt;/em&gt; next, the tia who tells the scandalous ones after that), WhatsApp voice-message import, a printed family book for Christmas.&lt;/p&gt;

&lt;p&gt;But none of that is the point. The point is that next Sunday, at the table, when he lets out one of those pieces and shrugs &lt;em&gt;"não é nada"&lt;/em&gt; — I'll press a red button. And one day, when my two kids ask this memory who their father was at six, and who their grandfather was with the flowers, it will answer — carefully, honestly, in our voice.&lt;/p&gt;

&lt;p&gt;It will be something. Forever. In his own voice.&lt;/p&gt;




&lt;h2&gt;
  
  
  Built with
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://ollama.com/library/gemma3" rel="noopener noreferrer"&gt;Gemma 3&lt;/a&gt;&lt;/strong&gt; (open-weight, via &lt;a href="https://ollama.com" rel="noopener noreferrer"&gt;Ollama&lt;/a&gt;) — the brain: honesty classifier, grounded answering, titling, translation. &lt;em&gt;Entered in **Best Use of Gemma&lt;/em&gt;&lt;em&gt;.&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://github.com/SYSTRAN/faster-whisper" rel="noopener noreferrer"&gt;faster-whisper&lt;/a&gt;&lt;/strong&gt; — open-source speech-to-text, CPU-only.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://github.com/idiap/coqui-ai-TTS" rel="noopener noreferrer"&gt;XTTS v2&lt;/a&gt;&lt;/strong&gt; — open-source voice cloning (CPML, non-commercial) so answers can sound like &lt;em&gt;him&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://ollama.com/library/nomic-embed-text" rel="noopener noreferrer"&gt;nomic-embed-text&lt;/a&gt;&lt;/strong&gt; — open embeddings for memory retrieval.&lt;/li&gt;
&lt;li&gt;Python + Flask + SQLite + one HTML page. Repo: &lt;strong&gt;&lt;a href="https://github.com/jeffreyturov-dev/avozdosmeus" rel="noopener noreferrer"&gt;github.com/jeffreyturov-dev/avozdosmeus&lt;/a&gt;&lt;/strong&gt; (MIT).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;If you have someone whose voice you'd miss — fork it. This weekend is long enough to save it.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>hf26challenge</category>
      <category>gemma</category>
    </item>
    <item>
      <title>I turned my mother-in-law's voice memos into a recipe book — 100% local open-source AI</title>
      <dc:creator>Jeffrey Turov</dc:creator>
      <pubDate>Sun, 04 Oct 2026 08:34:18 +0000</pubDate>
      <link>https://dev.to/jeffreyturov/i-turned-my-mother-in-laws-voice-memos-into-a-recipe-book-100-local-open-source-ai-46m</link>
      <guid>https://dev.to/jeffreyturov/i-turned-my-mother-in-laws-voice-memos-into-a-recipe-book-100-local-open-source-ai-46m</guid>
      <description>&lt;p&gt;&lt;em&gt;Posted during Hacktoberfest 2026 — a weekend build celebrating open-source AI.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Who it's for
&lt;/h2&gt;

&lt;p&gt;Odilia is my mother-in-law, and her cooking lives in voice memos. Decades of family recipes — including her Luxembourgish &lt;em&gt;Gromperekichelcher&lt;/em&gt; (crispy potato fritters) — exist only as rambling WhatsApp audios she sends when one of us asks "how do you make that again?". Nobody ever scrolls back to find them, and nobody writes them down.&lt;/p&gt;

&lt;p&gt;So this weekend I built her &lt;strong&gt;voice-recipe-book&lt;/strong&gt;: drop a voice memo in, get a clean recipe card out. One real person, one real problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it does
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;voice memo (.mp3/.m4a/.wav)
   │
   ├─ transcribe.py        faster-whisper (open-weight Whisper, int8 CPU) → raw text
   └─ structure_recipe.py  Qwen2.5-3B-Instruct (open-weight LLM, local CPU) → Markdown card
                              title · servings/times · ingredients · steps · "the secret"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Everything runs locally. No cloud, no API key, €0 per recipe.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;👉 &lt;strong&gt;Repo: &lt;a href="https://github.com/jeffreyturov-dev/voice-recipe-book" rel="noopener noreferrer"&gt;https://github.com/jeffreyturov-dev/voice-recipe-book&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo (real run)
&lt;/h2&gt;

&lt;p&gt;The demo memo is TTS-synthesized because I won't publish a real family recording — the pipeline is identical either way. Whisper (small, CPU) transcribed the 40-second French memo at confidence 1.00. One honest failure: &lt;em&gt;Gromperekichelcher&lt;/em&gt; came out as "grands-pèreux qui chèrent" — regional dish names are exactly the words a speech model has never seen, so the app accepts an explicit title (a feature, not a bug).&lt;/p&gt;

&lt;p&gt;Then Qwen2.5-3B-Instruct, running on the same CPU, structured the raw transcript — hesitations and all — into this card (actual output, unedited except the title):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# Gromperekichelcher d'Odilia&lt;/span&gt;
&lt;span class="gs"&gt;**Portions :**&lt;/span&gt; 8 | &lt;span class="gs"&gt;**Préparation :**&lt;/span&gt; 30 min | &lt;span class="gs"&gt;**Cuisson :**&lt;/span&gt; 15 min

&lt;span class="gu"&gt;## Ingrédients&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; 1 kg de pommes de terre farineuses
&lt;span class="p"&gt;-&lt;/span&gt; 1 oignon
&lt;span class="p"&gt;-&lt;/span&gt; 2 œufs
&lt;span class="p"&gt;-&lt;/span&gt; 2 cuillères à soupe de farine
&lt;span class="p"&gt;-&lt;/span&gt; Sel, poivre, noix de muscade
&lt;span class="p"&gt;-&lt;/span&gt; Persil haché (non précisé)

&lt;span class="gu"&gt;## Étapes&lt;/span&gt;
&lt;span class="p"&gt;1.&lt;/span&gt; Râpez les pommes de terre farineuses et essorez-les bien dans un torchon.
&lt;span class="p"&gt;2.&lt;/span&gt; Mélangez avec l'oignon râpé, les œufs, la farine, le sel, le poivre et la muscade.
&lt;span class="p"&gt;3.&lt;/span&gt; Chauffez de l'huile dans une poêle à feu moyen.
&lt;span class="p"&gt;4.&lt;/span&gt; Formez des galettes et aplatissez-les.
&lt;span class="p"&gt;5.&lt;/span&gt; Faites frire 3 à 4 minutes de chaque côté jusqu'à ce qu'elles soient croustillantes.
&lt;span class="p"&gt;6.&lt;/span&gt; Servez chaud avec une compote de pommes.

&lt;span class="gu"&gt;## Le secret&lt;/span&gt;
&lt;span class="gt"&gt;&amp;gt; Essorer bien les pommes de terre pour éviter qu'elles détrempe.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It even extracted "the secret" — squeeze the potatoes dry — the one line of grandma-wisdom that makes the dish.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why open innovation matters here
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Privacy&lt;/strong&gt;: family voice memos never leave the laptop. No third-party server ever hears Odilia's voice. With a cloud speech API, that's simply not true.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost&lt;/strong&gt;: €0 per recipe, forever, vs ~€0.02–0.10 per memo with a cloud speech+LLM stack. A whole family book costs nothing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Offline&lt;/strong&gt;: after a one-time model download, it works in a kitchen with no internet.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hackable&lt;/strong&gt;: every layer is an open weight — swap Whisper sizes, swap the LLM, fine-tune on her dialect, change the card template. A closed API gives you none of that.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The stack: &lt;strong&gt;Whisper&lt;/strong&gt; (MIT, via faster-whisper/CTranslate2) + &lt;strong&gt;Qwen2.5-3B-Instruct&lt;/strong&gt; (Apache 2.0) on plain CPU. Total new code: ~150 lines of Python, written this weekend — the open models do the heavy lifting.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/jeffreyturov-dev/voice-recipe-book
&lt;span class="nb"&gt;cd &lt;/span&gt;voice-recipe-book &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; requirements.txt
python3 app.py memo.m4a &lt;span class="s2"&gt;"Grandma's apple pie"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Next for Odilia: a weekly cron that pulls her memos and prints the growing book for Christmas. She doesn't need to know what an LLM is. She just needs her recipes to outlive the chat history.&lt;/p&gt;

</description>
      <category>opensource</category>
    </item>
    <item>
      <title>A 3B model that beats a 7B: failure-driven orchestration on uncontaminated knowledge</title>
      <dc:creator>Jeffrey Turov</dc:creator>
      <pubDate>Thu, 01 Oct 2026 19:53:32 +0000</pubDate>
      <link>https://dev.to/jeffreyturov/a-3b-model-that-beats-a-7b-failure-driven-orchestration-on-uncontaminated-knowledge-1gfh</link>
      <guid>https://dev.to/jeffreyturov/a-3b-model-that-beats-a-7b-failure-driven-orchestration-on-uncontaminated-knowledge-1gfh</guid>
      <description>&lt;p&gt;&lt;strong&gt;How a fully sovereign QA stack — local Wikipedia index, distilled Qwen2.5-3B reader, and crutches built only from measured failures — went from 33% to 52% on facts no LLM can have memorized, and matched a zero-shot 7B more than twice its size.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Configuration&lt;/th&gt;
&lt;th&gt;Post-cutoff-150 accuracy&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;3B naked (no retrieval)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;0.0%&lt;/strong&gt; [0–2.5]&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;7B naked&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;0.0%&lt;/strong&gt; (by construction — see below)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3B + sovereign chain (start, Sep 29)&lt;/td&gt;
&lt;td&gt;33.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3B + sovereign chain (final, Oct 1)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;52.0%&lt;/strong&gt; [44.1–59.8]&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;7B zero-shot, same chain&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;50.0%&lt;/strong&gt; [42.1–57.9]&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Every number comes from a deterministic gate with Wilson 95% CIs. Every rollback is published below, not just the wins. The benchmark is public on Kaggle (post-cutoff-150).&lt;/p&gt;




&lt;h2&gt;
  
  
  1. Why "post-cutoff": most retrieval benchmarks measure memory, not retrieval
&lt;/h2&gt;

&lt;p&gt;HotpotQA (2018) sits inside the training data of every 2024 model. We measured it: our naked 3B scores 30.5% on HotpotQA — from pure memorization. On such a benchmark, injecting retrieved evidence can only &lt;em&gt;hurt&lt;/em&gt;: the "does retrieval help?" question is rigged from the start.&lt;/p&gt;

&lt;p&gt;So we built &lt;strong&gt;post-cutoff-150&lt;/strong&gt;: 150 questions generated from local Wikipedia pages whose lead paragraphs cite 2025–2026 facts, with answers that are exact spans of the source text. The contamination filter is deterministic and brutal: &lt;strong&gt;if the naked 3B can answer a question without context, the question is rejected&lt;/strong&gt;. Final control: naked model = 0.0% [0–2.5]. Whatever the pipeline scores from there is &lt;em&gt;net retrieval value&lt;/em&gt; — nothing else.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. The sovereign stack
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Index&lt;/strong&gt;: 7.2M Wikipedia pages + 11.4M redirects, FTS5, 49 GB, fully local (crash-proof indexer, resumable to infinity).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reader&lt;/strong&gt;: Qwen2.5-3B-Instruct + LoRA stack trained inside the harness (SFT on noisy evidence, then DPO — the only technique that ever broke a plateau, twice).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Orchestrator&lt;/strong&gt;: deterministic arms chained as a superset — anchor extraction → lead fetch → gated joint reading → deterministic bridge hop (n-grams certified against the 7.2M titles) → full-text search retry → Wikipedia &lt;em&gt;tables&lt;/em&gt; arm (wikitext parsing with rowspan inheritance) → semantic lens → live full-text search.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gate discipline&lt;/strong&gt;: a challenger is PROMOTED only if it beats the champion on the same 150 questions, through the same matcher, with Wilson CIs reported. Otherwise ROLLBACK — and a rollback is a success of the system, not a failure to hide.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Two infrastructure rules that cost us bugs before they became rules: &lt;strong&gt;never evaluate a model in memory right after training&lt;/strong&gt; (post-train state measures 0% on a perfectly healthy model — save, reload from disk, then gate), and &lt;strong&gt;all extraction LLM calls at temperature 0&lt;/strong&gt; (temp 0.1 anchoring injected ±4–6 points of run-to-run variance into every gate we ever ran).&lt;/p&gt;

&lt;h2&gt;
  
  
  3. The campaign: 33.3% → 52.0% in four days
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Step&lt;/th&gt;
&lt;th&gt;Mechanism&lt;/th&gt;
&lt;th&gt;Score&lt;/th&gt;
&lt;th&gt;Verdict&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Baseline pipeline&lt;/td&gt;
&lt;td&gt;anchors → lens → joint read&lt;/td&gt;
&lt;td&gt;33.3%&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;+ fullpool&lt;/td&gt;
&lt;td&gt;show the reader the whole pool, not just top sentences&lt;/td&gt;
&lt;td&gt;34.7%&lt;/td&gt;
&lt;td&gt;kept (reserve)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;+ FTS-retry&lt;/td&gt;
&lt;td&gt;full-text enrichment when reading fails&lt;/td&gt;
&lt;td&gt;39.3%&lt;/td&gt;
&lt;td&gt;kept&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;+ entropy-gated routing&lt;/td&gt;
&lt;td&gt;accept a read only if token entropy ≤ 0.6&lt;/td&gt;
&lt;td&gt;42.0%&lt;/td&gt;
&lt;td&gt;PROMOTE&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;+ SLOTS&lt;/td&gt;
&lt;td&gt;2-slot "preuve → réponse" prompt with deterministic verification&lt;/td&gt;
&lt;td&gt;43.3%&lt;/td&gt;
&lt;td&gt;PROMOTE&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;+ deterministic anchors fix, echo-retry, FR meta-strip&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;46.0%&lt;/td&gt;
&lt;td&gt;kept (reserve)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;+ FR→EN translation arm&lt;/td&gt;
&lt;td&gt;French meta-questions translated once, read in English&lt;/td&gt;
&lt;td&gt;48.0%&lt;/td&gt;
&lt;td&gt;PROMOTE&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;+ Wikipedia tables arm&lt;/td&gt;
&lt;td&gt;wikitext tables → "header: cell" lines → lens&lt;/td&gt;
&lt;td&gt;48.7%&lt;/td&gt;
&lt;td&gt;PROMOTE&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;+ &lt;strong&gt;semantic lens&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;12B judges sentence relevance instead of token overlap&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;52.0%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;PROMOTE&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;What did &lt;strong&gt;not&lt;/strong&gt; work — measured, gated, published:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Self-consistency ×5&lt;/strong&gt;: null. The literature assumes random errors that voting smooths; ours are &lt;em&gt;systematic&lt;/em&gt; (wrong page → same wrong answer 5 times).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Best-of-N by confidence, inter-path consensus&lt;/strong&gt;: null, same reason — independent paths converge to the &lt;em&gt;same&lt;/em&gt; errors.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;IRCoT&lt;/strong&gt; (academic interleaved retrieval-CoT): 10% naive, 25% hybrid — our sequential pipeline beats it by 17+ points. For a 3B, anchoring by exact titles beats conversational retrieval.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Query planner&lt;/strong&gt; (typed decomposition): rollback both directions — deterministic decomposition doesn't convert hard failures; they're hard for rules too.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SpanLift&lt;/strong&gt; (generation-free QA: enumerate spans, judge by evidence lift): 1.3% — catastrophic. The candidate enumerator produced sentence fragments, and the scorer missed even when the truth was a candidate. Paradigm abandoned on 3B.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GRPO&lt;/strong&gt; (both reader and GSM8K): gradient desert is structural (&lt;code&gt;frac_reward_zero_std = 0.8&lt;/code&gt;). We fixed the mechanism (dense rewards → 0.2, reward flowing) — and transfer to the target distribution was &lt;em&gt;still null&lt;/em&gt;. Two independent adapters scored &lt;em&gt;identically&lt;/em&gt; to the champion.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deep plaintext retrieval&lt;/strong&gt; (full pages beyond the lead) and &lt;strong&gt;live Wikipedia full-text search&lt;/strong&gt;: both net zero — every win overlapped an existing arm's rescue.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  4. The capstone: 3B trained + orchestrated vs 7B zero-shot
&lt;/h2&gt;

&lt;p&gt;Same chain, same questions, same matcher, same deterministic gates. The &lt;em&gt;only&lt;/em&gt; variable swapped: the reader (Qwen2.5-7B-Instruct, 4-bit, no LoRA, no tuning).&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Arm wins&lt;/th&gt;
&lt;th&gt;3B (78/150 = 52.0%)&lt;/th&gt;
&lt;th&gt;7B (75/150 = 50.0%)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Primary reader arm&lt;/td&gt;
&lt;td&gt;50&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;66&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Crutch arms (semantic lens, tables, fallbacks)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;+28&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;+9&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Raw capacity wins the primary arm (+16 for the 7B). But the crutches — each one built on a &lt;em&gt;measured&lt;/em&gt; 3B failure mode (echo killed by SLOTS, entropy calibrated at 0.05–0.21 for correct vs 1.0+ for guessing, deterministic proof verification) — add 28 points to the 3B and only 9 to the 7B. The 7B has no echo to fix and its entropy sits near 0.002, making the calibrated gate useless. &lt;strong&gt;The crutches don't transfer because they were never generic — they were prosthetics fitted to a specific patient's limp.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Honest caveats: the confidence intervals overlap (we claim "ties or beats", not "beats"); the 7B is 4-bit quantized and zero-shot — a &lt;em&gt;tuned&lt;/em&gt; 7B would do better, and that's a fine next experiment.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. What we actually learned (the publishable part)
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Contamination decides what a benchmark measures.&lt;/strong&gt; On memorized knowledge, retrieval adds nothing; on post-cutoff knowledge, a sovereign pipeline turns 0% into 52%.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Small-model errors are systematic, not random.&lt;/strong&gt; Everything built on smoothing randomness (voting, consensus, best-of-n) failed four independent times. Everything built on &lt;em&gt;diagnosing the systematic class&lt;/em&gt; worked.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Entropy is a calibrated confidence signal nobody was reading.&lt;/strong&gt; Correct answers: 0.05–0.21. Echoes/guesses: 1.0–2.27. One threshold, zero training, +2 points.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SLOTS = externalized working memory.&lt;/strong&gt; Force the model to first &lt;em&gt;copy the proof sentence&lt;/em&gt;, then &lt;em&gt;copy the answer from inside the proof&lt;/em&gt;, and verify both deterministically at receipt. The meta-question echo class (24 questions) died (3 left).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context quality beats context coverage.&lt;/strong&gt; The semantic lens won 8 net questions — 4× its coverage diagnosis — because a relevance-filtered context &lt;em&gt;reads better&lt;/em&gt;, not just &lt;em&gt;finds more&lt;/em&gt;. Measure arms in-chain, not by detection rate.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deterministic gates catch the bugs that fake results.&lt;/strong&gt; In-memory post-training evals, a &lt;code&gt;urllib&lt;/code&gt; import swallowed by &lt;code&gt;try/except&lt;/code&gt; (the "reader" we benchmarked for days was the fallback), a wedged CUDA process serving empty answers, ±6-point anchoring noise — every one of these would have become a false claim without the discipline.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DPO is the only training move that ever broke a plateau&lt;/strong&gt; (GSM8K 60→67, reader 42→45, twice confirmed). SFT saturates; GRPO's gradient desert is structural.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  6. Reproducibility
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Benchmark: &lt;strong&gt;post-cutoff-150&lt;/strong&gt;, public on Kaggle, with the contamination filter (naked = 0% verified).&lt;/li&gt;
&lt;li&gt;All scores from deterministic gates (Wilson 95% CI), same matcher everywhere (a lower bound, consistently applied — comparisons valid).&lt;/li&gt;
&lt;li&gt;6 promotions, 6 rollbacks — all archived with per-question outputs.&lt;/li&gt;
&lt;li&gt;Hardware: one RTX 5090. No API calls to any frontier model for the &lt;em&gt;answers&lt;/em&gt; — the 12B orchestration LLM and the 3B reader are both fully local.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  7. Limitations
&lt;/h2&gt;

&lt;p&gt;150 questions is honest but small; we report CIs and refuse to claim deltas inside the noise. The matcher (containment + token subset) undercounts morphological variants — scores are lower bounds. The 7B comparison is zero-shot/4-bit by design (we isolated the &lt;em&gt;orchestration&lt;/em&gt; variable, not the best possible 7B). The residual 72 failures split between evidence absent from every source we can reach (obscure entities, join-queries) and pure 3B reading capacity — that ceiling is documented, not hidden.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Built by an autonomous harness that diagnoses its own failures, trains only through deterministic gates, and publishes its rollbacks. The rollbacks taught us more than the wins.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>nlp</category>
      <category>research</category>
    </item>
    <item>
      <title>I benchmarked what frontier models actually know about 2026 — most of it, they don't</title>
      <dc:creator>Jeffrey Turov</dc:creator>
      <pubDate>Thu, 01 Oct 2026 14:03:21 +0000</pubDate>
      <link>https://dev.to/jeffreyturov/i-benchmarked-what-frontier-models-actually-know-about-2026-most-of-it-they-dont-nm5</link>
      <guid>https://dev.to/jeffreyturov/i-benchmarked-what-frontier-models-actually-know-about-2026-most-of-it-they-dont-nm5</guid>
      <description>&lt;p&gt;&lt;strong&gt;Benchmark:&lt;/strong&gt; &lt;a href="https://www.kaggle.com/benchmarks/tasks/jeffreyturov/post-cutoff-150" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/tasks/jeffreyturov/post-cutoff-150&lt;/a&gt;&lt;br&gt;
&lt;strong&gt;Dataset:&lt;/strong&gt; &lt;a href="https://www.kaggle.com/datasets/jeffreyturov/post-cutoff-knowledge-150" rel="noopener noreferrer"&gt;https://www.kaggle.com/datasets/jeffreyturov/post-cutoff-knowledge-150&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The itch
&lt;/h2&gt;

&lt;p&gt;Every LLM leaderboard tells you how models score on knowledge from their training data.&lt;br&gt;
I wanted the opposite: &lt;strong&gt;what do models know about things that happened AFTER their&lt;br&gt;
training cutoff?&lt;/strong&gt; Not retrieval-augmented answers — raw parametric knowledge of&lt;br&gt;
2025-2026 facts.&lt;/p&gt;

&lt;h2&gt;
  
  
  The benchmark: Post-Cutoff Knowledge 150
&lt;/h2&gt;

&lt;p&gt;150 factual questions built from a 7.2M-page web index, restricted to pages citing&lt;br&gt;
2025-2026 events. Two construction guarantees make contamination impossible:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Span-verified&lt;/strong&gt;: each gold answer is an EXACT span of the source page's lead
(deterministic string check, no LLM judge).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Contamination-filtered&lt;/strong&gt;: every candidate question is asked to a 2024-cutoff
3B model with NO context. If the small model can answer it from memory, the
question is REJECTED (~4% of candidates were). Kept questions are provably
outside 2024 training data: the 3B control scores &lt;strong&gt;0/150&lt;/strong&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Scoring: case-insensitive, token-boundary containment of the gold span — a lower&lt;br&gt;
bound (morphological variants of a correct answer can be missed). Prompts demand&lt;br&gt;
a bare fact, no sentence.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I ran it against
&lt;/h2&gt;

&lt;p&gt;9 frontier models, 150 questions each, single run, no retries on the scored attempt&lt;br&gt;
(prompt: "answer with just the requested fact, no sentence"). k = correct answers&lt;br&gt;
out of 150; CI = Wilson 95%.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Accuracy&lt;/th&gt;
&lt;th&gt;k/150&lt;/th&gt;
&lt;th&gt;Wilson 95% CI&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.7 Flash&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;24.0%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;36&lt;/td&gt;
&lt;td&gt;[17.9%, 31.4%]&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek-R1 (0528)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;20.0%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;30&lt;/td&gt;
&lt;td&gt;[14.4%, 27.1%]&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 2.5 Pro&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;16.0%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;24&lt;/td&gt;
&lt;td&gt;[11.1%, 22.4%]&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.4&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;11.3%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;17&lt;/td&gt;
&lt;td&gt;[7.2%, 17.4%]&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.5 Flash-Lite&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;8.7%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;13&lt;/td&gt;
&lt;td&gt;[5.1%, 14.3%]&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 2.5 Flash&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;7.3%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;11&lt;/td&gt;
&lt;td&gt;[4.1%, 12.7%]&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemma 4 31B (open weights)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;6.0%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;[3.2%, 11.0%]&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Haiku 4.5&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;5.3%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;[2.7%, 10.2%]&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.4 nano&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;3.3%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;[1.4%, 7.6%]&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Claude Sonnet 4.5 was scheduled three times and never started (infrastructure-side,&lt;br&gt;
not a scoring failure) — it is excluded rather than reported as zero. Gemma 4 31B&lt;br&gt;
hit per-question timeouts on 5 of 150 items (slow open-weights serving); the&lt;br&gt;
unanswered items are counted as wrong, standard practice — 9/150 = 6.0%.&lt;/p&gt;

&lt;h2&gt;
  
  
  What surprised me
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The best model still fails 3 questions out of 4.&lt;/strong&gt; Gemini 3.7 Flash leads at&lt;br&gt;
24% — meaning even the freshest frontier model has no parametric trace of most&lt;br&gt;
2025-2026 facts. Anyone building on "the model probably knows" is wrong 76% of&lt;br&gt;
the time.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Reasoning does not rescue knowledge.&lt;/strong&gt; DeepSeek-R1 (20.0%) scores second and&lt;br&gt;
beats several newer generalist models — but its long chains cannot invent a fact&lt;br&gt;
that was never in training. Reasoning moves the needle on problems, not on&lt;br&gt;
missing data.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Size and price do not order the ranking.&lt;/strong&gt; GPT-5.4 (11.3%) sits below&lt;br&gt;
DeepSeek-R1 and far below Gemini 3.7 Flash; Gemini 2.5 Pro (16.0%) beats its own&lt;br&gt;
family's newer Flash-Lite (8.7%). Freshness of training data matters more than&lt;br&gt;
benchmark muscle.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The honest-control design works&lt;/strong&gt;: the 2024-cutoff 3B control scores exactly&lt;br&gt;
0/150 — by construction — which makes every point above zero a genuine&lt;br&gt;
post-cutoff signal, not contamination.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Honest limits
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Questions are single-hop by construction (generated from one page lead). This
benchmark does not test multi-hop reasoning.&lt;/li&gt;
&lt;li&gt;The "3B fails" filter makes the control ≈0 BY DESIGN — the number that matters
is each frontier model's absolute score, not the gap to control.&lt;/li&gt;
&lt;li&gt;Span containment is a lower bound on true accuracy.&lt;/li&gt;
&lt;li&gt;150 items → Wilson CI ≈ ±8 points. Direction is reliable, fine ranking is not.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Extend to 500 questions to tighten CIs.&lt;/li&gt;
&lt;li&gt;Multi-hop variant: questions whose answer requires joining TWO 2025-2026 pages.&lt;/li&gt;
&lt;li&gt;RAG arm: same questions WITH the source page as context, to measure the
retrieval uplift per model (preliminary local data on a 3B pipeline: 0% → 33.3%).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Built with the &lt;code&gt;kaggle-benchmarks&lt;/code&gt; library. Task source is public on the&lt;br&gt;
benchmark page — fork it and run your own lineup.&lt;/p&gt;

</description>
      <category>kagglechallenge</category>
      <category>ai</category>
      <category>benchmarking</category>
      <category>llm</category>
    </item>
    <item>
      <title>I built 11 pay-per-use data APIs that AI agents can call directly (Google Maps, TikTok, LinkedIn, review monitoring) — here's what I learned</title>
      <dc:creator>Jeffrey Turov</dc:creator>
      <pubDate>Tue, 29 Sep 2026 15:13:34 +0000</pubDate>
      <link>https://dev.to/jeffreyturov/i-built-11-pay-per-use-data-apis-that-ai-agents-can-call-directly-google-maps-tiktok-linkedin-24b9</link>
      <guid>https://dev.to/jeffreyturov/i-built-11-pay-per-use-data-apis-that-ai-agents-can-call-directly-google-maps-tiktok-linkedin-24b9</guid>
      <description>&lt;p&gt;A few months ago I published a set of Actors on Apify Store. Today they're all &lt;strong&gt;AI-agent ready&lt;/strong&gt;: any LLM agent (Claude, GPT, Cursor, LangChain, n8n) can discover and call them through the Apify MCP server — no custom integration code needed.&lt;/p&gt;

&lt;p&gt;This post is the full playbook: what the tools do, how pay-per-event monetization works (including the pricing trap that made me lose money on every sale), the bugs I hit, and how AI agents actually consume the tools.&lt;/p&gt;

&lt;h2&gt;
  
  
  The toolbox
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Actor&lt;/th&gt;
&lt;th&gt;What it extracts&lt;/th&gt;
&lt;th&gt;Price (pay-per-event)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://apify.com/travelmonitorlab/google-maps-scraper" rel="noopener noreferrer"&gt;Google Maps Business Scraper&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Names, phones, websites, ratings, reviews, GPS&lt;/td&gt;
&lt;td&gt;$1.00 / run + $0.03 / business&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://apify.com/travelmonitorlab/tiktok-scraper" rel="noopener noreferrer"&gt;TikTok Profile &amp;amp; Video Scraper&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Followers, likes, bio, per-video stats&lt;/td&gt;
&lt;td&gt;$0.01 / profile + $0.002 / video&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://apify.com/travelmonitorlab/instagram-scraper" rel="noopener noreferrer"&gt;Instagram Profile Scraper&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Followers, bio, verified, engagement&lt;/td&gt;
&lt;td&gt;$0.01 / profile&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://apify.com/travelmonitorlab/youtube-scraper" rel="noopener noreferrer"&gt;YouTube Video &amp;amp; Channel Scraper&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Views, likes, subscribers, search results&lt;/td&gt;
&lt;td&gt;$0.002 / video&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://apify.com/travelmonitorlab/linkedin-profile-scraper" rel="noopener noreferrer"&gt;LinkedIn Profile Scraper&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Headlines, companies, skills, experience&lt;/td&gt;
&lt;td&gt;$0.02 / profile&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://apify.com/travelmonitorlab/rag-web-browser" rel="noopener noreferrer"&gt;RAG Web Browser&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Clean Markdown from any URL + Google search&lt;/td&gt;
&lt;td&gt;$0.003 / page&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://apify.com/travelmonitorlab/hermes-revenu-api" rel="noopener noreferrer"&gt;Fuel Prices France API&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Real-time prices, 9,800 stations, GPS&lt;/td&gt;
&lt;td&gt;$0.20 / run + $0.01 / 1k stations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://apify.com/travelmonitorlab/travel-monitor-launch" rel="noopener noreferrer"&gt;Hotel Rate Monitoring&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Competitor rates, parity checks&lt;/td&gt;
&lt;td&gt;fractions of a cent per item&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://apify.com/travelmonitorlab/api-breaking-change-radar" rel="noopener noreferrer"&gt;API Breaking-Change Radar&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Diffs OpenAPI specs, classifies changelogs, alerts&lt;/td&gt;
&lt;td&gt;$0.50 / run + $0.50 / breaking change&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://apify.com/travelmonitorlab/review-radar" rel="noopener noreferrer"&gt;Review Radar&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;New Google reviews for a business, Slack alerts&lt;/td&gt;
&lt;td&gt;$0.25 / run + $0.01 / new review&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://apify.com/travelmonitorlab/review-pitch-generator" rel="noopener noreferrer"&gt;Review Pitch Generator&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Worst reviews → ready-to-send sales report&lt;/td&gt;
&lt;td&gt;$0.25 / run + $0.10 / pitch&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The last three are a different breed: not scrapers but &lt;strong&gt;monitors&lt;/strong&gt; — they keep state between runs and only bill when they find something.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why "AI-agent ready" changes everything
&lt;/h2&gt;

&lt;p&gt;The old model: a human finds your scraper on the store, reads the docs, clicks buttons.&lt;/p&gt;

&lt;p&gt;The new model: an AI agent gets a task ("find me 50 plumbers in Austin with their phone numbers"), searches the Apify Store via MCP, reads the actor's README and input schema, and calls it — end to end, no human.&lt;/p&gt;

&lt;p&gt;For that to work, three things must be true:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Your README is written for an LLM, not just humans.&lt;/strong&gt; Mine now all start with a "Use this tool when..." section — that's what the agent pattern-matches against the user's request.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your input schema has a description on every field.&lt;/strong&gt; The agent constructs the JSON input from those descriptions. No description = hallucinated parameters = failed runs = no revenue.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your output is documented field by field.&lt;/strong&gt; The agent needs to know what it gets back to reason over it.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Here's the actual flow with the Apify MCP server (&lt;code&gt;https://mcp.apify.com&lt;/code&gt; — add it to Claude Desktop or Cursor in 30 seconds):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User: "Get me the follower counts of these 5 TikTok creators"
Agent: → search-actors("tiktok profile")
       → fetch-actor-details (reads README + input schema)
       → call-actor(travelmonitorlab/tiktok-scraper,
                    {"profiles": [...], "maxVideosPerProfile": 0})
       → returns structured JSON
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or skip MCP entirely — every actor is a single synchronous HTTP call:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-X&lt;/span&gt; POST &lt;span class="s2"&gt;"https://api.apify.com/v2/acts/travelmonitorlab~google-maps-scraper/run-sync-get-dataset-items?token=&lt;/span&gt;&lt;span class="nv"&gt;$APIFY_TOKEN&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"queries": ["plumbers Austin TX"], "maxResults": 50}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Monetization: pay-per-event (and the trap that cost me real money)
&lt;/h2&gt;

&lt;p&gt;Apify offers several pricing models. For new actors, &lt;code&gt;PRICE_PER_DATASET_ITEM&lt;/code&gt; is &lt;strong&gt;rejected&lt;/strong&gt; — you must use &lt;code&gt;PAY_PER_EVENT&lt;/code&gt;. The model is better anyway: you define events and charge explicitly in code:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;Actor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;charge&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;event_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;business-scraped&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;Actor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;push_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Trap #1 — the vanishing charge:&lt;/strong&gt; put the charge AFTER &lt;code&gt;crawler.run()&lt;/code&gt; and your event loop may already be closed — the charge silently vanishes, you deliver data for free. Charge inside the handler, right before pushing data. Always verify with &lt;code&gt;chargedEventCounts&lt;/code&gt; in the run object after a test run.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Trap #2 — pricing below your compute cost (this one actually bit me):&lt;/strong&gt; on Apify's free plan, the &lt;em&gt;actor owner&lt;/em&gt; pays the user's compute. My Google Maps actor was priced at $0.005/business. A customer scraped 16 businesses → I earned $0.08, and paid $1.23 in Playwright + residential proxy costs. Margin: −1436%. Every sale lost money. The structural fix: &lt;strong&gt;a flat &lt;code&gt;run-started&lt;/code&gt; event that covers fixed costs&lt;/strong&gt; (browser launch, proxy session) plus a per-result event with real margin:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;Actor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;charge&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;event_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;run-started&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# first line of main()
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That same 16-business run now bills $1.48 instead of $0.08. If you build browser-based actors, do this from day one.&lt;/p&gt;

&lt;p&gt;Setting pricing is pure API — and &lt;code&gt;pricingInfos&lt;/code&gt; is &lt;strong&gt;append-only&lt;/strong&gt;: you can't edit a tier in place, you append a new one:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;PUT&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;v2&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;acts&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;actorId&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pricingInfos&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;existing_tier_verbatim&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pricingModel&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;PAY_PER_EVENT&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;reasonForChange&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Flat run fee to cover compute&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pricingPerEvent&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;actorChargeEvents&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;run-started&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;eventTitle&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Run started&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;eventDescription&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;eventPriceUsd&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;1.0&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;business-scraped&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;eventTitle&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Business scraped&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;eventDescription&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;eventPriceUsd&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;0.03&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;}}}]}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Apify takes 20%. Compute is paid by the user (on paid plans); you pocket the event fees.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Trap #3 — self-billing:&lt;/strong&gt; running your own PPE actor bills your own account. Build a &lt;code&gt;dryRun&lt;/code&gt; input flag that wraps every &lt;code&gt;Actor.charge()&lt;/code&gt; and skips it — you can then integration-test the full pipeline for free and only flip billing on for real runs.&lt;/p&gt;

&lt;h2&gt;
  
  
  The QA gauntlet (or: your actor must survive an empty input)
&lt;/h2&gt;

&lt;p&gt;Apify automatically runs every store actor with the schema's prefilled input and expects success within 5 minutes — three failures and your actor gets flagged "Under maintenance", then deprecated. Two lessons learned the hard way:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Never raise on empty input.&lt;/strong&gt; If the user (or the QA bot) provides no target, fall back to a small built-in demo (2 results, dry-run forced) and exit 0. An actor that raises &lt;code&gt;RuntimeError("no target")&lt;/code&gt; is an actor on the deprecation list.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep the prefill tiny.&lt;/strong&gt; The QA run is baked into your build: a heavy prefill (20 Playwright results through residential proxies) blows the 5-minute budget and flags you. Prefill = 2 results max, dry-run on.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Battle scars (so you don't get them)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Crawlee 1.8 breaking changes:&lt;/strong&gt; &lt;code&gt;purge_on_start&lt;/code&gt; and &lt;code&gt;navigation_timeout_secs&lt;/code&gt; are no longer valid kwargs — use &lt;code&gt;page.set_default_navigation_timeout()&lt;/code&gt; in a &lt;code&gt;pre_navigation_hook&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Google Maps never fires &lt;code&gt;load&lt;/code&gt;:&lt;/strong&gt; analytics keep streaming forever, so navigation always times out. Fix: abort images/fonts/media via &lt;code&gt;page.route()&lt;/code&gt; (the handler must be a coroutine, not a lambda).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Residential proxies are mandatory for Google Maps&lt;/strong&gt;, and you must pass &lt;code&gt;actor_proxy_input=&lt;/code&gt; as a &lt;em&gt;named&lt;/em&gt; argument to &lt;code&gt;Actor.create_proxy_configuration()&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;French number formats will crash your floats:&lt;/strong&gt; &lt;code&gt;"4,8"&lt;/code&gt; → replace comma; &lt;code&gt;"1 234"&lt;/code&gt; reviews can use &lt;code&gt;\xa0&lt;/code&gt; &lt;em&gt;or&lt;/em&gt; &lt;code&gt;\u202f&lt;/code&gt; as thousand separator.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A "SUCCEEDED" run can contain zero useful data.&lt;/strong&gt; Always check &lt;code&gt;itemCount&lt;/code&gt; + sample the dataset + read the end of the log.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Don't run 7 queries × 25 results in one run.&lt;/strong&gt; Split into parallel runs of ≤5 queries × 15 results; retry failed ones sequentially (residential proxy tunnels occasionally die).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pin your dependencies.&lt;/strong&gt; &lt;code&gt;apify&amp;gt;=3.2,&amp;lt;4&lt;/code&gt; + &lt;code&gt;crawlee&amp;gt;=1.7,&amp;lt;2&lt;/code&gt; is a proven combo; open ranges break within weeks as PyPI drifts.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The distribution strategy
&lt;/h2&gt;

&lt;p&gt;Publishing on the store is step 0. What actually moves the needle:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;README written for LLMs&lt;/strong&gt; — agents choose tools whose docs they can parse. Clear "use when", typed inputs, example I/O. The MCP server's own system prompt tells agents to read your README before calling — it's your storefront.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Niche SEO titles&lt;/strong&gt; — "Google Maps Scraper" is saturated (the official one has 500k+ users); "Fuel Prices France API" has zero competition on the store.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Monitors &amp;gt; scrapers&lt;/strong&gt; — a scraper competes with everyone; a stateful monitor (breaking changes, new reviews) with Slack alerts answers a recurring pain and justifies a per-run flat fee.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dogfooding&lt;/strong&gt; — I use my own Google Maps actor to build lead lists I sell elsewhere. Every sale is also a demo.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Try them
&lt;/h2&gt;

&lt;p&gt;All 11 actors are live on &lt;a href="https://apify.com/travelmonitorlab" rel="noopener noreferrer"&gt;Apify Store&lt;/a&gt;. If you build agents, add &lt;code&gt;https://mcp.apify.com&lt;/code&gt; to your MCP client and just ask for the data — the agent will find the tools.&lt;/p&gt;

&lt;p&gt;Feedback, bugs, feature requests: open an issue on any actor page, I answer fast.&lt;/p&gt;

</description>
      <category>apify</category>
      <category>webscraping</category>
      <category>mcp</category>
      <category>ai</category>
    </item>
    <item>
      <title>Lux Stay Agent — a hotel travel agent that only works because its content is structured</title>
      <dc:creator>Jeffrey Turov</dc:creator>
      <pubDate>Fri, 18 Sep 2026 21:24:36 +0000</pubDate>
      <link>https://dev.to/jeffreyturov/lux-stay-agent-a-hotel-travel-agent-that-only-works-because-its-content-is-structured-5ff6</link>
      <guid>https://dev.to/jeffreyturov/lux-stay-agent-a-hotel-travel-agent-that-only-works-because-its-content-is-structured-5ff6</guid>
      <description>&lt;p&gt;I run a real 3-star hotel next to Luxembourg's central station. For the Sanity Challenge, instead of building a demo on fake data, I pointed a production AI agent at a Sanity Knowledge Base filled with the hotel's &lt;strong&gt;actual operational content&lt;/strong&gt; — live room prices, the real 54-dish snack menu, transport facts, multilingual FAQs. The result: &lt;strong&gt;Lux Stay Agent&lt;/strong&gt;, a travel assistant that answers travelers' questions (FR/EN/DE) with exact figures it cannot afford to get wrong, and that would be completely useless without structured content.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Live demo (real recorded sessions):&lt;/strong&gt; &lt;a href="https://yasha.phoenix--ai.com/agent/" rel="noopener noreferrer"&gt;https://yasha.phoenix--ai.com/agent/&lt;/a&gt;&lt;br&gt;
&lt;strong&gt;Sanity project ID:&lt;/strong&gt; &lt;code&gt;vxozr96i&lt;/code&gt; (dataset &lt;code&gt;production&lt;/code&gt;)&lt;/p&gt;
&lt;h2&gt;
  
  
  The problem: answers a hotel can't get wrong
&lt;/h2&gt;

&lt;p&gt;"How much is a double room?" sounds trivial until you realize the price changes daily (our rates sync with Booking.com every night). "How do I get from the airport?" has one correct answer (bus 29, free — Luxembourg made all public transport free in 2020). "What's on the menu?" is 54 items with individual prices living in the hotel's production ordering system.&lt;/p&gt;

&lt;p&gt;A keyword search over a website gets you approximate, stale, or wrong answers. An LLM without grounding gets you confident nonsense. What works is an agent that &lt;em&gt;queries&lt;/em&gt; structured content — exact fields, exact numbers — through a scoped, read-only window.&lt;/p&gt;
&lt;h2&gt;
  
  
  The Knowledge Base: 78 real documents
&lt;/h2&gt;

&lt;p&gt;I modeled the hotel's world as six Sanity document types:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Type&lt;/th&gt;
&lt;th&gt;Content&lt;/th&gt;
&lt;th&gt;Why it must be structured&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;hotel&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Address, phone, check-in/out times, distances (100 m to station, 400 m to center)&lt;/td&gt;
&lt;td&gt;Exact facts, zero tolerance for approximation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;room&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;3 types, capacity, bed, size, &lt;strong&gt;basePriceEur&lt;/strong&gt; + price-sync note&lt;/td&gt;
&lt;td&gt;A price is a number field, not prose&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;menuItem&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;54 dishes, category, &lt;strong&gt;priceEur&lt;/strong&gt;, availability, ordering note&lt;/td&gt;
&lt;td&gt;Pulled from the live ordering system (RoomEats)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;guide&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Airport/station/transport guides with &lt;code&gt;facts[]&lt;/code&gt; arrays&lt;/td&gt;
&lt;td&gt;Bus lines, durations, the free-transport rule&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;faq&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;FR/EN/DE questions &amp;amp; answers&lt;/td&gt;
&lt;td&gt;The agent answers in the user's language&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;attraction&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Sights with &lt;strong&gt;walkingMinutesFromHotel&lt;/strong&gt;, UNESCO flags&lt;/td&gt;
&lt;td&gt;"What can I visit on foot?" is a numeric query&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Every document traces back to a real system: the booking engine I built for the hotel (rates synced nightly from Booking.com), the production QR-ordering database for the menu, and verified local transport facts.&lt;/p&gt;
&lt;h2&gt;
  
  
  Architecture: agent → Sanity Context MCP → Content Lake
&lt;/h2&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Traveler question (FR/EN/DE)
        │
        ▼
Lux Stay Agent (Python, OpenAI-compatible LLM, function calling)
        │  1. fetches /initial-context over HTTP → schema-aware system prompt
        ▼
Sanity Context MCP endpoint
https://api.sanity.io/v2026-03-03/context/mcp/vxozr96i/production?embeddings=true
        │  tools: schema_explorer, groq_query, array_field_reader
        ▼
Sanity Content Lake — 78 real documents, embeddings enabled
        (semantic search via text::semanticSimilarity)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The agent loop is deliberately boring — that's the point. The intelligence lives in the content model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;run_agent&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;tools&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;get_tools&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;  &lt;span class="c1"&gt;# MCP tools/list -&amp;gt; OpenAI function schemas
&lt;/span&gt;    &lt;span class="n"&gt;messages&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;system&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;SYSTEM&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;{ctx}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;initial_context_http&lt;/span&gt;&lt;span class="p"&gt;())},&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;MAX_ROUNDS&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;msg&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;llm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;tools&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tool_calls&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;           &lt;span class="c1"&gt;# grounded final answer
&lt;/span&gt;        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;call&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;msg&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tool_calls&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
            &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;mcp_tool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;call&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;function&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
                              &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;loads&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;call&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;function&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;arguments&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]))&lt;/span&gt;
            &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tool&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tool_call_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;call&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
                             &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The system prompt is strict: exact figures only via &lt;code&gt;groq_query&lt;/code&gt;, never from memory; cite the source document type; if the base doesn't know, say so and hand over to the hotel's phone/email; answer in the user's language.&lt;/p&gt;

&lt;h2&gt;
  
  
  It only works because the content is structured
&lt;/h2&gt;

&lt;p&gt;Three moments from real sessions:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Combien coûte une nuit en chambre double et à quelle distance de la gare ?"&lt;/strong&gt;&lt;br&gt;
The agent fires one &lt;code&gt;groq_query&lt;/code&gt; on &lt;code&gt;room&lt;/code&gt; (&lt;code&gt;basePriceEur: 95&lt;/code&gt;) and one on &lt;code&gt;hotel&lt;/code&gt; (&lt;code&gt;distanceToStationM: 100&lt;/code&gt;). A keyword search would find a page mentioning "double room" and "station" — it would not reliably bind 95 € to &lt;em&gt;this&lt;/em&gt; room type on &lt;em&gt;this&lt;/em&gt; date.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Quels plats avec du kebab, et à quel prix ?"&lt;/strong&gt;&lt;br&gt;
The menu is 54 &lt;code&gt;menuItem&lt;/code&gt; documents with &lt;code&gt;name&lt;/code&gt; and &lt;code&gt;priceEur&lt;/code&gt; as queryable fields. The agent answers with the exact list and prices — and surfaces the pricing contradiction honestly: each item carries the Wolt delivery price &lt;em&gt;and&lt;/em&gt; the note that ordering in-room via QR is 15–25% cheaper. Both claims, with their sources, side by side — exactly what structured content enables.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Was kostet ein Einzelzimmer und wie komme ich vom Flughafen zum Hotel?"&lt;/strong&gt;&lt;br&gt;
German question → German answer, because the FAQ documents are tagged &lt;code&gt;language: "de"&lt;/code&gt;, while the room price still comes from the language-neutral numeric field. Translation happens at the LLM layer; facts stay exact at the content layer.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxmjolxy0ne1gu4g9k9ax.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxmjolxy0ne1gu4g9k9ax.png" alt="The agent replaying a real session — user question, MCP tool calls, grounded answer" width="800" height="933"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Rigor: I tested the base before trusting the agent
&lt;/h2&gt;

&lt;p&gt;Before writing the agent, I ran a 15-assertion QA battery against the Context MCP endpoint itself: schema visibility for all 6 types, exact price queries, reference resolution (&lt;code&gt;room.hotel-&amp;gt;name&lt;/code&gt;), document counts (54 menu items), guide facts, semantic search with embeddings, and multilingual retrieval. &lt;strong&gt;15/15.&lt;/strong&gt; Then a 5-question battery against the full agent loop (German, menu prices, UNESCO sights, English, free transport) asserting exact figures in the answers and actual tool usage. &lt;strong&gt;5/5.&lt;/strong&gt; The demo page replays these unedited sessions.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I learned
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;/initial-context&lt;/code&gt; over HTTP is a real optimization&lt;/strong&gt;: injecting the compressed schema into the system prompt saved a tool call per conversation and visibly improved the agent's GROQ accuracy on the first try.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Embeddings were one flag (&lt;code&gt;embeddings=true&lt;/code&gt;) but needed enabling per-dataset&lt;/strong&gt; (&lt;code&gt;sanity datasets embeddings enable&lt;/code&gt;) — the error message was clear, the fix took a minute with the CLI once I had a token with the right grant (&lt;code&gt;deploy-studio&lt;/code&gt; role for hosting, &lt;code&gt;developer&lt;/code&gt; for embeddings).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scope by default&lt;/strong&gt;: the read-only token plus the Context endpoint's design meant I could point an agent at production hospitality data without ever risking a write. For a business dataset, that's the difference between "fun demo" and "deployable".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The hard part was never the agent&lt;/strong&gt; — it was deciding what deserved to be a field. &lt;code&gt;walkingMinutesFromHotel&lt;/code&gt; as a number turns "what can I visit?" into a sortable query. That modeling decision is the whole game.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Try it, read it, reuse it
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Live demo with real session replays:&lt;/strong&gt; &lt;a href="https://yasha.phoenix--ai.com/agent/" rel="noopener noreferrer"&gt;https://yasha.phoenix--ai.com/agent/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The hotel's real booking site&lt;/strong&gt; (the content source): &lt;a href="https://yasha.phoenix--ai.com/" rel="noopener noreferrer"&gt;https://yasha.phoenix--ai.com/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sanity Studio:&lt;/strong&gt; &lt;a href="https://lux-stay-agent.sanity.studio/" rel="noopener noreferrer"&gt;https://lux-stay-agent.sanity.studio/&lt;/a&gt; (project &lt;code&gt;vxozr96i&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sanity Context docs:&lt;/strong&gt; &lt;a href="https://www.sanity.io/docs/ai/sanity-context" rel="noopener noreferrer"&gt;https://www.sanity.io/docs/ai/sanity-context&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you have content an agent can't afford to get wrong — prices, inventory, schedules, errata — structure it, scope it, and let the MCP endpoint do the rest. The agent is the easy part.&lt;/p&gt;

</description>
      <category>sanitychallenge</category>
      <category>ai</category>
      <category>mcp</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Your Actor should never see my credentials: the security case for MCP connectors</title>
      <dc:creator>Jeffrey Turov</dc:creator>
      <pubDate>Thu, 17 Sep 2026 22:34:36 +0000</pubDate>
      <link>https://dev.to/apify/your-actor-should-never-see-my-credentials-the-security-case-for-mcp-connectors-3ief</link>
      <guid>https://dev.to/apify/your-actor-should-never-see-my-credentials-the-security-case-for-mcp-connectors-3ief</guid>
      <description>&lt;p&gt;Every automation platform has the same dirty secret: the easiest way to connect a third-party service is to paste a token into an input field. It works on the first try. It is also how credentials end up in places nobody intended.&lt;/p&gt;

&lt;p&gt;I learned this building lead-generation Actors on Apify. My first working version took a GitHub token as an Actor input. It shipped fast, and it was wrong. This piece is about the threat model behind that mistake, and how MCP connectors invert it: the Actor I run never sees a credential at all.&lt;/p&gt;

&lt;p&gt;Everything below comes from real builds and real run logs, including the failures.&lt;/p&gt;

&lt;h2&gt;
  
  
  The threat model nobody writes down
&lt;/h2&gt;

&lt;p&gt;Passing a service token as an Actor input creates three leak paths, and all three are boring:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Run logs.&lt;/strong&gt; Inputs are echoed, logged, and retained. Anyone with whom you share a run (a teammate, a support ticket, a screenshot in a bug report) potentially shares the token with it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Shared input JSON.&lt;/strong&gt; Actors get cloned, forked, and re-run from saved inputs. A token inside an input object travels with every copy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Third-party Actor code.&lt;/strong&gt; The moment you run code you did not write (a Store Actor, a fork, a colleague's experiment), any credential you hand it is only as safe as that code's worst &lt;code&gt;console.log&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Notice what these have in common: none of them require an attacker. The leak vector is the architecture itself. The token is simply in more places than it needs to be.&lt;/p&gt;

&lt;h2&gt;
  
  
  The inversion: credentials live with the connector, not the code
&lt;/h2&gt;

&lt;p&gt;Apify's &lt;a href="https://docs.apify.com/integrations/mcp-connectors" rel="noopener noreferrer"&gt;MCP connectors&lt;/a&gt; restructure where the secret sits:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You authorize the third-party service (GitHub, Notion, Slack) &lt;strong&gt;once&lt;/strong&gt;, in Apify Console under Settings, Integrations. The OAuth credential is stored with the connector, on the platform side.&lt;/li&gt;
&lt;li&gt;At run time, the Actor receives a &lt;strong&gt;connector ID&lt;/strong&gt;: an opaque string like &lt;code&gt;ebw4ThD4cQbEKzC2l&lt;/code&gt;. It is not a token, and it is useless outside the platform.&lt;/li&gt;
&lt;li&gt;The Actor talks to the &lt;strong&gt;Apify MCP proxy&lt;/strong&gt; at &lt;code&gt;${ACTOR_MCP_CONNECTOR_BASE_URL}/&amp;lt;connectorId&amp;gt;&lt;/code&gt;, authenticating with the run's own &lt;code&gt;APIFY_TOKEN&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;The proxy enforces the tool permissions the Actor declared in its input schema.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The practical consequence: you can publish the Actor's source, share its runs, and paste its input JSON into a forum. There is nothing in any of those artifacts worth stealing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Least privilege is declared, not hoped for
&lt;/h2&gt;

&lt;p&gt;The second half of the model is the &lt;code&gt;mcpServers&lt;/code&gt; rule in the input schema. This is where the Actor states which tools it needs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="nl"&gt;"githubConnector"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"title"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"GitHub connector"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"string"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"resourceType"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"mcpConnector"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"mcpServers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"*"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"tools"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"required"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"create_*"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"push_*"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"update_*"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"write_*"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"commit_*"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"get_*"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"readOnly"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two layers constrain what the Actor can do: what the connector exposed at authorization time, and what the schema declares. The intersection is enforced by the proxy at run time, not by documentation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verified in a real run log.&lt;/strong&gt; My Actor's schema used the declaration above. When it called &lt;code&gt;list_tools()&lt;/code&gt; through the proxy, it saw exactly 7 tools: &lt;code&gt;create_branch&lt;/code&gt;, &lt;code&gt;create_or_update_file&lt;/code&gt;, &lt;code&gt;create_pull_request&lt;/code&gt;, &lt;code&gt;create_repository&lt;/code&gt;, &lt;code&gt;push_files&lt;/code&gt;, &lt;code&gt;update_pull_request&lt;/code&gt;, &lt;code&gt;update_pull_request_branch&lt;/code&gt;. Nothing else existed as far as the Actor could tell. The constraint layer is not a promise on a marketing page; it is observable behavior.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest limits
&lt;/h2&gt;

&lt;p&gt;A security argument that only lists strengths is marketing. Three limits I hit or verified:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. The connector's tool set is frozen at authorization time.&lt;/strong&gt; When you authorize a connector, the platform discovers the available tools once. If the upstream server later adds tools you want, you re-authorize. The set does not refresh itself. I lost time to this before reading it properly in the docs: it is "Layer 1" of their two-layer model, and it is by design (a silent expansion of your Actor's powers would be a security hole, not a feature).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. The run still authenticates with a token: the starter's &lt;code&gt;APIFY_TOKEN&lt;/code&gt;.&lt;/strong&gt; The proxy trusts whoever started the run. That means access control moves to &lt;em&gt;who can start your Actor with your connector&lt;/em&gt;, which is an Apify permissions question, not a code question. For Actors you share or publish, treat connector-enabled inputs as capability grants and review who can run what.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Connectors do not sanitize your payloads.&lt;/strong&gt; The proxy enforces &lt;em&gt;which tools&lt;/em&gt; the Actor calls, not &lt;em&gt;what data&lt;/em&gt; flows through them. An Actor scraping personal data and pushing it to your repo is still your compliance problem. The connector solves credential custody, not data governance.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pitfalls from real runs
&lt;/h2&gt;

&lt;p&gt;Three failures I actually hit while building on this model, because they cost me runs:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SDK drift.&lt;/strong&gt; The MCP Python SDK docs show &lt;code&gt;streamable_http_client&lt;/code&gt; unpacking into three values (&lt;code&gt;read, write, _&lt;/code&gt;). The version my Docker image installed yields two. &lt;code&gt;ValueError: not enough values to unpack&lt;/code&gt;. Indexing the tuple (&lt;code&gt;streams[0], streams[1]&lt;/code&gt;) works across versions. Same class of issue: the tool result error flag is &lt;code&gt;result.is_error&lt;/code&gt; (snake_case), not the &lt;code&gt;isError&lt;/code&gt; casing the TypeScript-flavored docs suggest.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Default branch assumptions.&lt;/strong&gt; &lt;code&gt;create_or_update_file&lt;/code&gt; failed with &lt;code&gt;Branch main not found&lt;/code&gt;: my repo's default is &lt;code&gt;master&lt;/code&gt;. The tool does not fall back to the repository default; it fails loudly, which is the correct behavior for a security boundary. Detect and retry with the other common name, or pass the branch explicitly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Write-only connectors and the SHA problem.&lt;/strong&gt; Updating an existing file on GitHub requires its current blob SHA. My write-scoped connector exposed no &lt;code&gt;get_file_contents&lt;/code&gt; to fetch it. Rather than widening permissions, I switched the pipeline to immutable, timestamped snapshots: each run writes a new dated file, and Git itself becomes the history. The least-privilege constraint pushed me to a better design, which is what least privilege is supposed to do.&lt;/p&gt;

&lt;h2&gt;
  
  
  When to use what
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Situation&lt;/th&gt;
&lt;th&gt;Token as input&lt;/th&gt;
&lt;th&gt;MCP connector&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Local throwaway script, one run, you watch it&lt;/td&gt;
&lt;td&gt;Fine&lt;/td&gt;
&lt;td&gt;Overkill&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Actor you will run on a schedule for weeks&lt;/td&gt;
&lt;td&gt;Leak path&lt;/td&gt;
&lt;td&gt;Right answer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Actor you publish or share&lt;/td&gt;
&lt;td&gt;Irresponsible&lt;/td&gt;
&lt;td&gt;Right answer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Running someone else's Actor with your services&lt;/td&gt;
&lt;td&gt;Never&lt;/td&gt;
&lt;td&gt;The only sane option&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The last row is the one that convinced me. If you consume third-party Actors, connectors are the only way to grant access to your services where a sloppy or malicious &lt;code&gt;print()&lt;/code&gt; cannot exfiltrate your credentials.&lt;/p&gt;

&lt;h2&gt;
  
  
  The takeaway
&lt;/h2&gt;

&lt;p&gt;Security models for automation are usually about adding vigilance: rotate tokens, restrict scopes, audit logs. MCP connectors remove the attack surface instead: the credential never enters the Actor's world, so it cannot leak from it. You trade a small amount of convenience (one authorization step in Console, a frozen tool set) for the ability to share runs, publish source, and run untrusted code without a knot in your stomach.&lt;/p&gt;

&lt;p&gt;The working example this article is drawn from is public: &lt;a href="https://github.com/jeffreyturov-dev/apify-scraping-toolbox" rel="noopener noreferrer"&gt;maps-to-stack on GitHub&lt;/a&gt;, and the build walkthrough with the scraping pitfalls is &lt;a href="https://dev.to/apify/from-scraper-to-stack-pushing-google-maps-leads-straight-into-github-with-mcp-connectors-43f3"&gt;on dev.to&lt;/a&gt;. Connector documentation: &lt;a href="https://docs.apify.com/integrations/mcp-connectors" rel="noopener noreferrer"&gt;docs.apify.com/integrations/mcp-connectors&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Build log details verified against run logs from August 2026: tool filtering (7 tools), branch fallback, SDK unpacking, and the immutable-snapshot pattern all occurred in real runs of the maps-to-stack Actor.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>apify</category>
      <category>mcp</category>
      <category>security</category>
      <category>automation</category>
    </item>
    <item>
      <title>Your scraper forgets everything between runs. Here's a review monitor that doesn't.</title>
      <dc:creator>Jeffrey Turov</dc:creator>
      <pubDate>Tue, 01 Sep 2026 17:55:54 +0000</pubDate>
      <link>https://dev.to/apify/your-scraper-forgets-everything-between-runs-heres-a-review-monitor-that-doesnt-4ddb</link>
      <guid>https://dev.to/apify/your-scraper-forgets-everything-between-runs-heres-a-review-monitor-that-doesnt-4ddb</guid>
      <description>&lt;p&gt;Your scraper forgets everything between runs. Here's a review monitor that doesn't.&lt;/p&gt;

&lt;p&gt;A scraper that pulls Google Maps reviews once is a toy. What businesses actually pay for is a monitor: "tell me the moment a bad review lands." The difference between the two is one unglamorous feature — remembering what you've already seen.&lt;/p&gt;

&lt;p&gt;I built Review Radar, an Apify Actor that watches a list of Google Maps places, scrapes recent reviews on a schedule, and posts a Slack alert the moment a new review at or below your star threshold appears. The Slack side uses Apify's MCP connectors, so the Actor never touches a token. Everything below comes from real runs — including the two state bugs that almost shipped.&lt;/p&gt;

&lt;p&gt;What we're building&lt;/p&gt;

&lt;p&gt;Review Radar:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;takes a list of Google Maps place URLs (or a search query),&lt;/li&gt;
&lt;li&gt;scrapes the latest reviews of each place with Playwright,&lt;/li&gt;
&lt;li&gt;diffs them against a persistent store of already-seen review IDs,&lt;/li&gt;
&lt;li&gt;pushes only genuinely new reviews to the dataset, flagged &lt;code&gt;isNew: true&lt;/code&gt;, and&lt;/li&gt;
&lt;li&gt;posts a Slack alert for each new review at or below your threshold — via an MCP connector, so no Slack token ever enters the Actor's code.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Why Google Maps reviews? Because for hotels, restaurants, and local agencies, a 1-star review answered within an hour is recoverable; the same review discovered three weeks later is a lost customer. Review monitoring is a product businesses already pay monthly for — and the official Google Business API only covers businesses you own, not your competitors or your clients' portfolios.&lt;/p&gt;

&lt;p&gt;The three parts that actually matter&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Scraping the reviews panel&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;On a Maps place page, reviews live behind the "Avis"/"Reviews" tab. The extraction selectors that survived contact with production:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Review blocks: &lt;code&gt;div[data-review-id]&lt;/code&gt; — with a trap I detail below&lt;/li&gt;
&lt;li&gt;Author: &lt;code&gt;div.d4r55&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Rating: &lt;code&gt;span.kvMYJc[role="img"]&lt;/code&gt; — parse the number from the &lt;code&gt;aria-label&lt;/code&gt; ("4 étoiles" / "4 stars")&lt;/li&gt;
&lt;li&gt;Text: &lt;code&gt;span.wiI7pd&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Relative date: &lt;code&gt;span.rsqaWe&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If no review block is visible on load, click the tab first: &lt;code&gt;button[aria-label*="Avis"]&lt;/code&gt; or &lt;code&gt;button[aria-label*="Reviews"]&lt;/code&gt;. Then scroll the panel (&lt;code&gt;div.m6QErb.DxyBCb.kA9KIf.dS8AEf&lt;/code&gt;) until you have enough blocks. Google Maps requires a residential proxy — datacenter IPs get consent-walled or blocked outright.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;State: the difference between a scraper and a monitor&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The Actor keeps a named key-value store (&lt;code&gt;review-radar-state&lt;/code&gt;) with one record per place: the set of review IDs already seen. Each run:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;hashes each review to a stable ID (the native &lt;code&gt;data-review-id&lt;/code&gt; when present, otherwise a SHA-1 of author+date+text),&lt;/li&gt;
&lt;li&gt;flags &lt;code&gt;isNew: true&lt;/code&gt; only for IDs not in the store,&lt;/li&gt;
&lt;li&gt;persists the updated set at the end.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Proof from two consecutive real runs on the same restaurant. First run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Cindy P. | 5★ | isNew=True
Raquel G. Urbano | 4★ | isNew=True
Chri Cou1967 | 5★ | isNew=True
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Second run, minutes later, identical reviews:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Cindy P. | 5★ | isNew=False
Raquel G. Urbano | 4★ | isNew=False
Chri Cou1967 | 5★ | isNew=False
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That &lt;code&gt;False&lt;/code&gt; is the entire product. A scheduled run now only surfaces what changed.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Slack alerts without a token in sight&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Same connector model as my previous build: the user authorizes Slack once in Apify Console → Settings → Integrations. The Actor declares the connector in its input schema and receives a connector ID at runtime — never a token:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="nl"&gt;"slackConnector"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"title"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Slack connector (optional)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"string"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"resourceType"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"mcpConnector"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"mcpServers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"*"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"tools"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"required"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"*message*"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"*chat*"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"*post*"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"*send*"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"readOnly"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;mcpServers&lt;/code&gt; declaration does double duty: it filters which connectors the picker offers, and the proxy refuses any tool call outside that list. At runtime the Actor lists the connector's actual tools and picks the first one matching &lt;code&gt;post/chat/message/send&lt;/code&gt;, then formats the alert:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;🚨 Nouvel avis 2★ sur *La Maison Lefèvre* (J. Dupont, il y a 2 jours)
Service décevant, attente de 40 minutes...
https://google.com/maps/place/...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Making the connector input optional (&lt;code&gt;"required": []&lt;/code&gt;) was deliberate: without it, the Actor still produces the full dataset — which also makes the Actor testable without touching your Slack workspace.&lt;/p&gt;

&lt;p&gt;The two bugs that almost shipped&lt;/p&gt;

&lt;p&gt;Bug 1 — key-value store key charset. I keyed place records by the Maps feature ID (&lt;code&gt;0x45d3f1...:0x...&lt;/code&gt;). Apify record keys only allow &lt;code&gt;a-zA-Z0-9!-_.'()&lt;/code&gt; — the colon is illegal. The Actor scraped everything perfectly, then crashed on the very last line of the run. One regex fixed it, but it's exactly the class of bug that only appears after a full successful scrape:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;place_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sub&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;[^a-zA-Z0-9!\-_.&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;()]&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;_&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;place_key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Bug 2 — duplicated review blocks. Querying both &lt;code&gt;div[data-review-id]&lt;/code&gt; and the legacy &lt;code&gt;div.jftiEf.fontBodyMedium&lt;/code&gt; selector returns overlapping containers — the same review twice. My first test dataset had Cindy P. duplicated. Fix: dedupe by review ID inside the scrape loop, before anything reaches the dataset.&lt;/p&gt;

&lt;p&gt;What I'd do differently at scale&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Relative dates ("il y a 2 semaines") don't sort. For precise alerting windows, resolve them against the run date.&lt;/li&gt;
&lt;li&gt;The Slack tool argument names vary by connector (&lt;code&gt;text&lt;/code&gt;, &lt;code&gt;message&lt;/code&gt;, &lt;code&gt;content&lt;/code&gt;). I inspect the tool's input schema and map fields — brittle but workable until connector schemas stabilize.&lt;/li&gt;
&lt;li&gt;Photos in reviews aren't extracted yet — for hospitality clients, a photo of a dirty room matters more than the text.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Try it&lt;/p&gt;

&lt;p&gt;The Actor is review-radar on my Apify account. Inputs: place URLs (or a search query), reviews per place, star threshold, optional Slack connector. First run reports the existing batch as new; every run after that reports only what changed. Point a daily schedule at it and you have a review monitoring product for the cost of a few compute units.&lt;/p&gt;

&lt;p&gt;A note on terms: Google Maps scraping sits uneasily with Google's ToS. Keep volumes polite (a handful of places, daily cadence), use a residential proxy, weigh the risk for production use, and prefer official sources where they cover your need — the Business Profile API works fine for businesses you own, it just can't watch anyone else's.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>tutorial</category>
      <category>javascript</category>
      <category>scraping</category>
    </item>
    <item>
      <title>From scraper to stack: pushing Google Maps leads straight into GitHub with MCP connectors</title>
      <dc:creator>Jeffrey Turov</dc:creator>
      <pubDate>Mon, 31 Aug 2026 23:39:38 +0000</pubDate>
      <link>https://dev.to/apify/from-scraper-to-stack-pushing-google-maps-leads-straight-into-github-with-mcp-connectors-43f3</link>
      <guid>https://dev.to/apify/from-scraper-to-stack-pushing-google-maps-leads-straight-into-github-with-mcp-connectors-43f3</guid>
      <description>&lt;p&gt;From scraper to stack: pushing Google Maps leads straight into GitHub with MCP connectors&lt;br&gt;
Every lead-generation pipeline I build ends the same way: a dataset full of freshly scraped businesses… that I then have to export, transform, and commit somewhere by hand. The scrape is automated; the last mile never is.&lt;br&gt;
When Apify launched MCP connectors - a new kind of Actor input that lets Actors securely call third-party services like GitHub, Notion, or Slack during a run - I rebuilt that last mile. This article walks through a working Actor that scrapes businesses from Google Maps and commits the results straight into a GitHub repository, without the Actor code ever seeing a credential.&lt;br&gt;
Everything below comes from a real build: the pitfalls are ones I actually hit, and the fixes are what actually solved them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Full source code:&lt;/strong&gt; &lt;a href="https://codeberg.org/veysel-devia/apify-scraping-toolbox" rel="noopener noreferrer"&gt;Codeberg repository&lt;/a&gt; (canonical mirror, readable without an account). Also available as a &lt;a href="https://drive.google.com/file/d/1_j3DTZgmCRjM70LA9tSUTcHi2vlk5QGS/view" rel="noopener noreferrer"&gt;source ZIP&lt;/a&gt;. The GitHub original at github.com/jeffreyturov-dev/apify-scraping-toolbox is currently unreachable for logged-out viewers (account-level restriction, under review with GitHub Trust &amp;amp; Safety) - use the mirror.&lt;br&gt;
What we're building&lt;br&gt;
Maps to Stack: an Actor that&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;takes a Google Maps search query (e.g. restaurants Esch-sur-Alzette),&lt;/li&gt;
&lt;li&gt;scrapes each place page with Playwright (name, address, phone, website, rating, GPS),&lt;/li&gt;
&lt;li&gt;pushes the results to the dataset, and&lt;/li&gt;
&lt;li&gt;writes a JSON snapshot into your GitHub repo via an MCP connector - one commit per run, versioned by Git.
Why GitHub as the destination? Because for lead-gen pipelines, a repo is a free CRM: versioned, diffable, and already wired into everything else (CI, dashboards, git-based CMSs). The same pattern works identically for Notion or Slack - only the connector changes.
Prerequisites&lt;/li&gt;
&lt;li&gt;An Apify account (the free plan works for testing this)&lt;/li&gt;
&lt;li&gt;A GitHub account with a repository to write into, authorized as an MCP connector in Apify Console → Settings → Integrations (one OAuth flow, two clicks)&lt;/li&gt;
&lt;li&gt;A residential proxy for the Google Maps scraping half - Maps blocks datacenter IPs outright. A note on terms: routing around Google's consent wall and IP blocks sits uneasily with Google Maps' Terms of Service. For production use you should weigh that risk, keep request volumes polite, and consider official sources such as the Places API where they cover your need; the residential proxy here is what makes the unoffical path technically reliable, not legally bulletproof.&lt;/li&gt;
&lt;li&gt;Five minutes to read the Actor source top to bottom; it's intentionally small
The security model that makes this interesting
The classic way to do this is to pass a GitHub token as an Actor input. That means every run log, every shared input JSON, and every fork of your Actor is one leak away from a compromised token.
MCP connectors invert this:&lt;/li&gt;
&lt;li&gt;You authorize GitHub once in Apify Console → Settings → Integrations. The credential lives with the connector, not with your code.&lt;/li&gt;
&lt;li&gt;At run time, the Actor receives a connector ID (a string like ebw4ThD4cQbEKzC2l) - not a token.&lt;/li&gt;
&lt;li&gt;The Actor talks to the Apify MCP proxy at ${ACTOR_MCP_CONNECTOR_BASE_URL}/, authenticating with the run's own APIFY_TOKEN.&lt;/li&gt;
&lt;li&gt;The proxy enforces the tool permissions your Actor declared in its input schema. The Actor physically cannot call tools outside its declaration.
Declaring the connector input
In .actor/INPUT_SCHEMA.json, set resourceType: "mcpConnector". The mcpServers rule list both filters which connectors the picker offers and caps which tools the proxy will let the Actor call:
"githubConnector": {
"title": "GitHub connector",
"description": "MCP connector to your GitHub account. The Actor only sees a connector ID.",
"type": "string",
"resourceType": "mcpConnector",
"mcpServers": [
{
 "url": "&lt;em&gt;",
 "tools": {
   "required": ["create_&lt;/em&gt;", "push_&lt;em&gt;", "update_&lt;/em&gt;", "write_&lt;em&gt;", "commit_&lt;/em&gt;", "get_*"],
   "readOnly": false
 }
}
]
}&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;At run time the input value is just the connector ID string.&lt;br&gt;
Verified in the run log: the proxy filtered the connector's tool list down to exactly what matched my declaration. My run saw 7 tools - create_branch, create_or_update_file, create_pull_request, create_repository, push_files, update_pull_request, update_pull_request_branch - and nothing else. The constraint layer works.&lt;br&gt;
Connecting from the Actor (Python)&lt;br&gt;
Two environment variables are injected into every run: ACTOR_MCP_CONNECTOR_BASE_URL (the proxy) and APIFY_TOKEN (the starter's token). You use the standard MCP Python SDK - nothing Apify-specific:&lt;br&gt;
import os&lt;br&gt;
import httpx&lt;br&gt;
from mcp import ClientSession&lt;br&gt;
from mcp.client.streamable_http import streamable_http_client&lt;/p&gt;

&lt;p&gt;base_url = os.environ["ACTOR_MCP_CONNECTOR_BASE_URL"]&lt;br&gt;
token = os.environ["APIFY_TOKEN"]&lt;br&gt;
proxy_url = f"{base_url}/{connector_id}"&lt;/p&gt;

&lt;p&gt;async with httpx.AsyncClient(&lt;br&gt;
   headers={"Authorization": f"Bearer {token}"}&lt;br&gt;
) as http:&lt;br&gt;
   async with streamable_http_client(proxy_url, http_client=http) as streams:&lt;br&gt;
       read, write = streams[0], streams[1]&lt;br&gt;
       async with ClientSession(read, write) as session:&lt;br&gt;
           await session.initialize()&lt;br&gt;
           tools = (await session.list_tools()).tools&lt;/p&gt;

&lt;p&gt;(Import verified against the installed SDK: mcp.client.streamable_http exposes both streamable_http_client and the alias streamablehttp_client, so the line above runs as pasted; the docs showcase the alias, which is worth knowing given the drift discussed below.)&lt;br&gt;
Pitfall #1 - SDK drift. The docs unpack streamable_http_client into three values (read, write, _), but the mcp version my image installed yields two. ValueError: not enough values to unpack (expected 3, got 2). Indexing the tuple (streams[0], streams[1]) works across versions. Same story for the result object: the error flag is result.is_error (snake_case), not isError - the docs' TypeScript casing leaks into expectations.&lt;br&gt;
Writing the file - and surviving reality&lt;br&gt;
With the session up, pick the file-writing tool from whatever the connector actually exposes (don't hard-code - the tool list is filtered by your schema &lt;em&gt;and&lt;/em&gt; by what the server advertised at authorization time):&lt;br&gt;
patterns = ["create_or_update_file", "push_files", "create_file", "update_file", "write_file"]&lt;br&gt;
tool = next((t for p in patterns for t in tools if p in t.name), None)&lt;/p&gt;

&lt;p&gt;owner, _, repo_name = actor_input["repo"].partition("/")&lt;br&gt;
args = {&lt;br&gt;
   "owner": owner,&lt;br&gt;
   "repo": repo_name,&lt;br&gt;
   "path": actor_input.get("path", "leads/output.json"),&lt;br&gt;
   "content": payload_json,&lt;br&gt;
   "message": f"leads: {query} ({date.today()})",&lt;br&gt;
   "branch": actor_input.get("branch", "main"),&lt;br&gt;
}&lt;br&gt;
result = await session.call_tool(tool.name, arguments=args)&lt;/p&gt;

&lt;p&gt;Two failures hit me immediately in real runs - both easy to fix once you see them:&lt;br&gt;
Pitfall #2 - Branch main not found. My repo's default branch is master, not main. The MCP tool doesn't fall back to the repo's default; it fails loudly. Cheap robust fix: detect the failure and retry once with the other common branch name.&lt;br&gt;
Pitfall #3 - File already exists… provide the current file's SHA. The second run collided with the first run's file. The GitHub API wants the blob SHA for updates - but the connector I authorized only exposes write tools (no get_file_contents to fetch the SHA). Instead of fighting it, I leaned into a better pattern for pipelines: immutable, timestamped snapshots. Every run writes a new dated file; Git itself becomes the history. If a read tool &lt;em&gt;is&lt;/em&gt; available, the code uses the SHA path instead:&lt;br&gt;
text = "".join(getattr(c, "text", "") for c in (result.content or []))&lt;br&gt;
if "already exists" in text.lower():&lt;br&gt;
   get_tool = next((t for t in tools if "get_file_contents" in t.name), None)&lt;br&gt;
   if get_tool:&lt;br&gt;
       ...  # fetch sha, retry with args["sha"] = sha&lt;br&gt;
   else:&lt;br&gt;
       stamp = datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%SZ")&lt;br&gt;
       args["path"] = re.sub(r".json$", f"-{stamp}.json", args["path"])&lt;br&gt;
   result = await session.call_tool(tool.name, arguments=args)&lt;/p&gt;

&lt;p&gt;Worth knowing: the connector's tool set is fixed at authorization time ("Layer 1" in the docs). If you authorize a connector and later wish it exposed more tools, re-authorize it - the discovered set doesn't refresh on its own.&lt;br&gt;
The scraping side, briefly&lt;br&gt;
The Maps half of the Actor uses Crawlee's PlaywrightCrawler with a residential proxy (Google blocks datacenter IPs outright). Three details that make it reliable, all learned the hard way on earlier Actors:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Kill the consent wall first. consent.google.com intercepts the first navigation; click "Accept" before waiting for any results selector.&lt;/li&gt;
&lt;li&gt;Block heavy resources in a pre-navigation hook. The load event never fires on Maps (continuous analytics), so route-abort images/fonts/media, or every navigation times out.&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Don't trust a green run. A SUCCEEDED run can still hold zero useful items - always assert on dataset item count and log per-item progress.&lt;br&gt;
The first two are a few lines each, and they are the difference between an Actor that works and one that silently dies on every third navigation:&lt;br&gt;
&lt;a class="mentioned-user" href="https://dev.to/crawler"&gt;@crawler&lt;/a&gt;.pre_navigation_hook&lt;br&gt;
async def optimize(context: PlaywrightCrawlingContext):&lt;br&gt;
page = context.page&lt;br&gt;
page.set_default_navigation_timeout(120_000)&lt;/p&gt;

&lt;p&gt;async def _abort(route):&lt;br&gt;
    await route.abort()&lt;/p&gt;
&lt;h1&gt;
  
  
  The load event never fires on Maps (continuous analytics).
&lt;/h1&gt;
&lt;h1&gt;
  
  
  Without this, every navigation hits the navigation timeout.
&lt;/h1&gt;

&lt;p&gt;await page.route(&lt;br&gt;
    "*&lt;em&gt;/&lt;/em&gt;.{png,jpg,jpeg,gif,webp,svg,ico,woff,woff2,ttf,mp4,webm,avi}",&lt;br&gt;
    _abort,&lt;br&gt;
)&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a class="mentioned-user" href="https://dev.to/crawler"&gt;@crawler&lt;/a&gt;.router.default_handler&lt;br&gt;
async def search_handler(context: PlaywrightCrawlingContext):&lt;br&gt;
    page = context.page&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# consent.google.com intercepts the first navigation.
# Click "Accept" before waiting for any results selector.
if "consent.google" in page.url:
    for sel in ('button[aria-label*="Accept"]',
                'button[aria-label*="Tout accepter"]',
                'button[aria-label*="Accepter"]',
                'form[action*="consent"] button'):
        btn = await page.query_selector(sel)
        if btn:
            await btn.click()
            await page.wait_for_timeout(2000)
            break

await page.wait_for_selector('div[role="feed"]', timeout=15000)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;The full scraper is ~150 lines; the dataset holds one item per business with name, category, address, phone, website, rating, coordinates, and the Maps URL.&lt;br&gt;
What a real run looks like&lt;br&gt;
Input:&lt;br&gt;
{&lt;br&gt;
 "query": "restaurants Esch-sur-Alzette",&lt;br&gt;
 "maxResults": 3,&lt;br&gt;
 "githubConnector": "ebw4ThD4cQbEKzC2l",&lt;br&gt;
 "repo": "jeffreyturov-dev/apify-scraping-toolbox",&lt;br&gt;
 "path": "leads/esch-restaurants.json"&lt;br&gt;
}&lt;/p&gt;

&lt;p&gt;Run log (abridged):&lt;br&gt;
Scraped 3 businesses&lt;br&gt;
Connector exposes 7 tools: ['create_branch', 'create_or_update_file', ...]&lt;br&gt;
Using tool: create_or_update_file&lt;br&gt;
Branch 'main' not found - retrying 'master'&lt;br&gt;
Timestamped snapshot: leads/esch-restaurants-20260903T220554Z.json&lt;br&gt;
Done - create_or_update_file wrote 3 leads&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdrive.google.com%2Fuc%3Fid%3D19mE12lo1DIsu6GEznRC6wU6d0M_ETxF2" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdrive.google.com%2Fuc%3Fid%3D19mE12lo1DIsu6GEznRC6wU6d0M_ETxF2" alt="Run log of a real Maps-to-Stack run in Apify Console" width="1440" height="900"&gt;&lt;/a&gt;&lt;br&gt;
And the commit lands in GitHub with the full JSON: 3 businesses, ratings, phone numbers, coordinates - queryable, diffable, and already where the rest of my tooling lives.&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdrive.google.com%2Fuc%3Fid%3D1i8DhkCjxyAeRgE3_2vWuNw51dQWjFp5p" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdrive.google.com%2Fuc%3Fid%3D1i8DhkCjxyAeRgE3_2vWuNw51dQWjFp5p" alt="The timestamped snapshot commit landing in the GitHub repo" width="1440" height="1500"&gt;&lt;/a&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdrive.google.com%2Fuc%3Fid%3D1RALZ2iRAtRr0S7u0WWCmCwfgFYRS7D6g" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdrive.google.com%2Fuc%3Fid%3D1RALZ2iRAtRr0S7u0WWCmCwfgFYRS7D6g" alt="Apify Console Integrations section with the MCP connectors card" width="1265" height="1382"&gt;&lt;/a&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdrive.google.com%2Fuc%3Fid%3D1HSLcZVg4_gv-kWizDvRgWc9Fff2Lr_MO" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdrive.google.com%2Fuc%3Fid%3D1HSLcZVg4_gv-kWizDvRgWc9Fff2Lr_MO" alt="Actor source: INPUT_SCHEMA.json and the connector session code" width="1265" height="1202"&gt;&lt;/a&gt;&lt;br&gt;
Where this pattern earns its keep&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Lead-gen pipelines: scrape on a schedule, each run commits a dated snapshot; a git log &lt;em&gt;is&lt;/em&gt; your change history of who opened/closed in a neighborhood.&lt;/li&gt;
&lt;li&gt;Any-to-any delivery: swap the connector for Notion (a database row per business) or Slack (a summary message) - the Actor code changes by a dozen lines; the security model stays identical.&lt;/li&gt;
&lt;li&gt;Untrusted Actor code: if you consume third-party Actors, connectors are the only sane way to give them access to &lt;em&gt;your&lt;/em&gt; services - the schema's tool constraints and the proxy's enforcement mean a malicious or sloppy Actor can't exceed its brief.
Why the connector model is the only sane way to share access
That last point deserves its own paragraph, because it generalizes beyond Apify. Every integration platform faces the same dilemma: users want automations that touch their GitHub, their CRM, their Slack - but handing a raw token to third-party code is an unacceptable blast radius. The connector pattern resolves it with three properties that are hard to get simultaneously any other way. First, mediation: every call passes through a proxy that authenticates the caller and authorizes the tool, so the credential is never exposed to the code that uses it. Second, declared least privilege: the input schema caps what the Actor may do, and the cap is enforced outside the Actor, where the Actor cannot rewrite it. Third, auditability: because calls flow through one chokepoint, every tool invocation is attributable to a specific run. Tokens in code give you none of these; connectors give you all three for the price of one OAuth flow. If you build or consume automations that touch shared services, this is the shape the access layer should have.
Try it
The Actor source (schema, scraper, connector logic) is intentionally small - read it top to bottom in five minutes. The moving parts that matter:&lt;/li&gt;
&lt;li&gt;resourceType: "mcpConnector" in the input schema, with mcpServers tool constraints.&lt;/li&gt;
&lt;li&gt;${ACTOR_MCP_CONNECTOR_BASE_URL}/ + APIFY_TOKEN bearer, via the stock MCP SDK.&lt;/li&gt;
&lt;li&gt;Defensive write logic: branch fallback, SHA-if-available, timestamped snapshot otherwise.
The connector model removes the part of pipeline building I liked least - sprinkling credentials through code - and replaces it with one authorization, one ID, and a proxy that keeps everyone honest.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Useful links: the &lt;a href="https://docs.apify.com/platform/actors/development/actor-definition/input-schema" rel="noopener noreferrer"&gt;Apify input schema reference&lt;/a&gt;, the &lt;a href="https://docs.apify.com/platform/actors/publishing/monetize" rel="noopener noreferrer"&gt;pay-per-event monetization docs&lt;/a&gt;, the &lt;a href="https://docs.apify.com/platform/integrations/mcp" rel="noopener noreferrer"&gt;Apify MCP integration docs&lt;/a&gt;, and the &lt;a href="https://apify.com/mcp" rel="noopener noreferrer"&gt;Apify MCP server page&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>tutorial</category>
      <category>javascript</category>
      <category>scraping</category>
    </item>
    <item>
      <title>I built 8 pay-per-use scraping APIs that AI agents can call directly (Google Maps, TikTok, Instagram, YouTube, LinkedIn) — here's what I learned</title>
      <dc:creator>Jeffrey Turov</dc:creator>
      <pubDate>Tue, 28 Jul 2026 09:28:47 +0000</pubDate>
      <link>https://dev.to/jeffreyturov/i-built-8-pay-per-use-scraping-apis-that-ai-agents-can-call-directly-google-maps-tiktok-o61</link>
      <guid>https://dev.to/jeffreyturov/i-built-8-pay-per-use-scraping-apis-that-ai-agents-can-call-directly-google-maps-tiktok-o61</guid>
      <description>&lt;p&gt;A few weeks ago I published a set of Actors on Apify Store. Today they're all &lt;strong&gt;AI-agent ready&lt;/strong&gt;: any LLM agent (Claude, GPT, Cursor, LangChain, n8n) can discover and call them through the Apify MCP server — no custom integration code needed.&lt;/p&gt;

&lt;p&gt;This post is the full playbook: what the tools do, how the pay-per-event monetization works, the bugs I hit, and how AI agents actually consume them.&lt;/p&gt;

&lt;h2&gt;
  
  
  The toolbox
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Actor&lt;/th&gt;
&lt;th&gt;What it extracts&lt;/th&gt;
&lt;th&gt;Price (pay-per-event)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://apify.com/travelmonitorlab/google-maps-scraper" rel="noopener noreferrer"&gt;Google Maps Business Scraper&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Names, phones, websites, ratings, reviews, GPS&lt;/td&gt;
&lt;td&gt;$0.005 / business&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://apify.com/travelmonitorlab/tiktok-scraper" rel="noopener noreferrer"&gt;TikTok Profile &amp;amp; Video Scraper&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Followers, likes, bio, per-video stats&lt;/td&gt;
&lt;td&gt;$0.01 / profile&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://apify.com/travelmonitorlab/instagram-scraper" rel="noopener noreferrer"&gt;Instagram Profile Scraper&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Followers, bio, verified, engagement&lt;/td&gt;
&lt;td&gt;$0.01 / profile&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://apify.com/travelmonitorlab/youtube-scraper" rel="noopener noreferrer"&gt;YouTube Video &amp;amp; Channel Scraper&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Views, likes, subscribers, search results&lt;/td&gt;
&lt;td&gt;$0.002 / video&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://apify.com/travelmonitorlab/linkedin-profile-scraper" rel="noopener noreferrer"&gt;LinkedIn Profile Scraper&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Headlines, companies, skills, experience&lt;/td&gt;
&lt;td&gt;$0.02 / profile&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://apify.com/travelmonitorlab/rag-web-browser" rel="noopener noreferrer"&gt;RAG Web Browser&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Clean Markdown from any URL + Google search&lt;/td&gt;
&lt;td&gt;$0.003 / page&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://apify.com/travelmonitorlab/hermes-revenu-api" rel="noopener noreferrer"&gt;Fuel Prices France API&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Real-time prices, 9,800 stations, GPS&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;FREE&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://apify.com/travelmonitorlab/travel-monitor-launch" rel="noopener noreferrer"&gt;Hotel Rate Monitoring&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Competitor rates, parity checks&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;FREE&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The two free ones are deliberate lead magnets — more on that below.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why "AI-agent ready" changes everything
&lt;/h2&gt;

&lt;p&gt;The old model: a human finds your scraper on the store, reads the docs, clicks buttons.&lt;/p&gt;

&lt;p&gt;The new model: an AI agent gets a task ("find me 50 plumbers in Austin with their phone numbers"), searches the Apify Store via MCP, reads the actor's README and input schema, and calls it — end to end, no human.&lt;/p&gt;

&lt;p&gt;For that to work, three things must be true:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Your README is written for an LLM, not just humans.&lt;/strong&gt; Mine now all start with a "Use this tool when..." section — that's what the agent pattern-matches against the user's request.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your input schema has a description on every field.&lt;/strong&gt; The agent constructs the JSON input from those descriptions. No description = hallucinated parameters = failed runs = no revenue.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your output is documented field by field.&lt;/strong&gt; The agent needs to know what it gets back to reason over it.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Here's the actual flow with the Apify MCP server (&lt;code&gt;https://mcp.apify.com&lt;/code&gt; — add it to Claude Desktop or Cursor in 30 seconds):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User: "Get me the follower counts of these 5 TikTok creators"
Agent: → search-actors("tiktok profile")
       → fetch-actor-details (reads README + input schema)
       → call-actor(travelmonitorlab/tiktok-scraper,
                    {"profiles": [...], "maxVideosPerProfile": 0})
       → returns structured JSON
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or skip MCP entirely — every actor is a single synchronous HTTP call:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-X&lt;/span&gt; POST &lt;span class="s2"&gt;"https://api.apify.com/v2/acts/travelmonitorlab~google-maps-scraper/run-sync-get-dataset-items?token=&lt;/span&gt;&lt;span class="nv"&gt;$APIFY_TOKEN&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"queries": ["plumbers Austin TX"], "maxResults": 50}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Monetization: pay-per-event (and the trap that cost me hours)
&lt;/h2&gt;

&lt;p&gt;Apify offers several pricing models. For new actors, &lt;code&gt;PRICE_PER_DATASET_ITEM&lt;/code&gt; is &lt;strong&gt;rejected&lt;/strong&gt; — you must use &lt;code&gt;PAY_PER_EVENT&lt;/code&gt;. The model is better anyway: you define events (e.g. &lt;code&gt;business-scraped&lt;/code&gt; at $0.005) and charge explicitly in code:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;Actor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;charge&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;event_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;business-scraped&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;Actor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;push_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;The trap:&lt;/strong&gt; put the charge AFTER &lt;code&gt;crawler.run()&lt;/code&gt; and your event loop may already be closed — the charge silently vanishes, you deliver data for free. Charge inside the handler, right before pushing data. Always verify with &lt;code&gt;chargedEventCounts&lt;/code&gt; in the run object after a test run.&lt;/p&gt;

&lt;p&gt;Setting pricing is pure API:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;PUT&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;v2&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;acts&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;actorId&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pricingInfos&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt;
  &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pricingModel&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;PAY_PER_EVENT&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;reasonForChange&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Launch pricing&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pricingPerEvent&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;actorChargeEvents&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;business-scraped&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;eventTitle&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Business scraped&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;eventDescription&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;One Google Maps business record&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;eventPriceUsd&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;0.005&lt;/span&gt;
    &lt;span class="p"&gt;}}}}&lt;/span&gt;
&lt;span class="p"&gt;]}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Apify takes 20%. Compute is paid by the user; you pocket the event fees.&lt;/p&gt;

&lt;h2&gt;
  
  
  Battle scars (so you don't get them)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Crawlee 1.8 breaking changes:&lt;/strong&gt; &lt;code&gt;purge_on_start&lt;/code&gt; and &lt;code&gt;navigation_timeout_secs&lt;/code&gt; are no longer valid kwargs — use &lt;code&gt;page.set_default_navigation_timeout()&lt;/code&gt; in a &lt;code&gt;pre_navigation_hook&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Google Maps never fires &lt;code&gt;load&lt;/code&gt;:&lt;/strong&gt; analytics keep streaming forever, so navigation always times out. Fix: abort images/fonts/media via &lt;code&gt;page.route()&lt;/code&gt; (the handler must be a coroutine, not a lambda).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Residential proxies are mandatory for Google Maps&lt;/strong&gt;, and you must pass &lt;code&gt;actor_proxy_input=&lt;/code&gt; as a &lt;em&gt;named&lt;/em&gt; argument to &lt;code&gt;Actor.create_proxy_configuration()&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;French number formats will crash your floats:&lt;/strong&gt; &lt;code&gt;"4,8"&lt;/code&gt; → replace comma; &lt;code&gt;"1 234"&lt;/code&gt; reviews can use &lt;code&gt;\xa0&lt;/code&gt; &lt;em&gt;or&lt;/em&gt; &lt;code&gt;\u202f&lt;/code&gt; as thousand separator.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A "SUCCEEDED" run can contain zero useful data.&lt;/strong&gt; Always check &lt;code&gt;itemCount&lt;/code&gt; + sample the dataset + read the end of the log.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Don't run 7 queries × 25 results in one run.&lt;/strong&gt; Split into parallel runs of ≤5 queries × 15 results; retry failed ones sequentially (residential proxy tunnels occasionally die).&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The distribution strategy
&lt;/h2&gt;

&lt;p&gt;Publishing on the store is step 0. What actually moves the needle:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Free lead magnets&lt;/strong&gt; — the fuel-price and hotel-rate actors are 100% free. Free tools get users, ratings, and store ranking; ranked actors surface in MCP search results; MCP visibility drives paying users to the paid actors.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;README written for LLMs&lt;/strong&gt; — agents choose tools whose docs they can parse. Clear "use when", typed inputs, example I/O.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Niche SEO titles&lt;/strong&gt; — "Google Maps Scraper" is saturated; "Fuel Prices France API" has zero competition on the store.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dogfooding&lt;/strong&gt; — I use my own Google Maps actor to build lead lists I sell elsewhere. Every sale is also a demo.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Try them
&lt;/h2&gt;

&lt;p&gt;All 8 actors are live on &lt;a href="https://apify.com/travelmonitorlab" rel="noopener noreferrer"&gt;Apify Store&lt;/a&gt;. If you build agents, add &lt;code&gt;https://mcp.apify.com&lt;/code&gt; to your MCP client and just ask for the data — the agent will find the tools.&lt;/p&gt;

&lt;p&gt;Feedback, bugs, feature requests: open an issue on any actor page, I answer fast.&lt;/p&gt;

</description>
      <category>apify</category>
      <category>webscraping</category>
      <category>mcp</category>
      <category>ai</category>
    </item>
  </channel>
</rss>
