<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Jay Pokale</title>
    <description>The latest articles on DEV Community by Jay Pokale (@jaypokale).</description>
    <link>https://dev.to/jaypokale</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4145200%2Ffaac088a-cfd1-4bfd-b1e7-f155cf2dd81e.jpg</url>
      <title>DEV Community: Jay Pokale</title>
      <link>https://dev.to/jaypokale</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/jaypokale"/>
    <language>en</language>
    <item>
      <title>Caveman vs Ponytail vs Chisle: I benchmarked the Claude Code token-saving plugins on 20 tasks</title>
      <dc:creator>Jay Pokale</dc:creator>
      <pubDate>Tue, 06 Oct 2026 20:14:43 +0000</pubDate>
      <link>https://dev.to/jaypokale/caveman-vs-ponytail-vs-chisle-i-benchmarked-the-claude-code-token-saving-plugins-on-20-tasks-bg8</link>
      <guid>https://dev.to/jaypokale/caveman-vs-ponytail-vs-chisle-i-benchmarked-the-claude-code-token-saving-plugins-on-20-tasks-bg8</guid>
      <description>&lt;p&gt;If you use &lt;strong&gt;Claude Code&lt;/strong&gt;, you've probably seen the two popular plugins that promise to cut your token bill: &lt;strong&gt;&lt;a href="https://github.com/JuliusBrussee/caveman" rel="noopener noreferrer"&gt;caveman&lt;/a&gt;&lt;/strong&gt;, which makes Claude talk like a caveman, and &lt;strong&gt;&lt;a href="https://github.com/dietrichgebert/ponytail" rel="noopener noreferrer"&gt;ponytail&lt;/a&gt;&lt;/strong&gt;, which pushes it to write less code. I built a third one, &lt;strong&gt;&lt;a href="https://chisle.jaypokale.me" rel="noopener noreferrer"&gt;Chisle&lt;/a&gt;&lt;/strong&gt;, and benchmarked all three against Claude with &lt;em&gt;no plugin at all&lt;/em&gt;: &lt;strong&gt;20 live tasks, 59+ model runs&lt;/strong&gt; on Haiku and Sonnet, with the arms differing only in the injected ruleset.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; caveman cut the bill to &lt;strong&gt;80%&lt;/strong&gt;, ponytail to &lt;strong&gt;68%&lt;/strong&gt;, and &lt;strong&gt;Chisle to 52%&lt;/strong&gt;, nearly half. Chisle also had the smallest worst day (&lt;strong&gt;173%&lt;/strong&gt; vs &lt;strong&gt;424%&lt;/strong&gt; and &lt;strong&gt;227%&lt;/strong&gt;) and backfired once in 20 tasks, where caveman backfired 6 times and ponytail 8.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzgihasih6mgt9oo3pbu8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzgihasih6mgt9oo3pbu8.png" alt="Total billed output across 20 tasks as percent of the no-plugin baseline: caveman 80% (worst day 424%), ponytail 68% (worst day 227%), Chisle 52% (worst day 173%)" width="799" height="316"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;20 live tasks, billed output vs no plugin&lt;/th&gt;
&lt;th&gt;total bill&lt;/th&gt;
&lt;th&gt;average task&lt;/th&gt;
&lt;th&gt;worst case&lt;/th&gt;
&lt;th&gt;backfires&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;no plugin (baseline)&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;caveman&lt;/td&gt;
&lt;td&gt;80%&lt;/td&gt;
&lt;td&gt;98%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;424%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;6 / 20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ponytail&lt;/td&gt;
&lt;td&gt;68%&lt;/td&gt;
&lt;td&gt;91%&lt;/td&gt;
&lt;td&gt;227%&lt;/td&gt;
&lt;td&gt;8 / 20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Chisle&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;52%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;69%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;173%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1 / 20&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;In the July re-verification run, &lt;strong&gt;every answer from every arm was graded correct&lt;/strong&gt;. None of these tools buys its savings with wrong answers. Every number comes from committed raw transcripts in the &lt;a href="https://github.com/JayPokale/Chisle/tree/main/benchmarks/results" rel="noopener noreferrer"&gt;Chisle repo&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the savings come from: coding vs explanation
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;tasks&lt;/th&gt;
&lt;th&gt;caveman&lt;/th&gt;
&lt;th&gt;ponytail&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Chisle&lt;/strong&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;coding&lt;/strong&gt; (wants working code)&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;74%&lt;/td&gt;
&lt;td&gt;59%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;44%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;explanation&lt;/strong&gt; (wants prose)&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;103%&lt;/td&gt;
&lt;td&gt;104%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;87%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;On coding, Chisle bills &lt;strong&gt;44%&lt;/strong&gt; of a bare model, a third less than ponytail, the closest thing to a dedicated "lazy code" tool. On explanation prompts both specialists go &lt;strong&gt;above 100%&lt;/strong&gt;: tools built to write less made Claude write &lt;em&gt;more&lt;/em&gt; than using nothing. Chisle is the only one that stays under.&lt;/p&gt;

&lt;p&gt;Split by answer length, the gap widens: on &lt;strong&gt;long answers Chisle bills 45%&lt;/strong&gt;, caveman 79%, ponytail 59%.&lt;/p&gt;

&lt;h2&gt;
  
  
  Caveman alternative: great prose compressor, no engineering judgment
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;caveman is genuinely good at compressing prose&lt;/strong&gt;, and on some short prose prompts it's a hair leaner than Chisle. But it has no judgment about &lt;em&gt;what&lt;/em&gt; to build. Asked to &lt;em&gt;"add caching"&lt;/em&gt;, it produced three implementations (&lt;strong&gt;330 tokens&lt;/strong&gt;). Chisle gave one &lt;code&gt;@cache&lt;/code&gt; decorator and a one-line upgrade path (&lt;strong&gt;151 tokens&lt;/strong&gt;). Its worst day cost &lt;strong&gt;4.2×&lt;/strong&gt; a bare model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Ponytail alternative: right instinct on code, padded prose
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;ponytail has the right instinct&lt;/strong&gt;: smallest thing that works. But it pads prose so much that it backfires: on a "retry logic" prompt it ran &lt;strong&gt;227%&lt;/strong&gt; of baseline. A tool whose whole job is writing less wrote more than twice as much.&lt;/p&gt;

&lt;p&gt;Installing &lt;strong&gt;both&lt;/strong&gt; to cover both axes doesn't fix it either: they fight over prose style and double per-session overhead. On one task the pair did &lt;em&gt;worse&lt;/em&gt; (605 tokens) than Chisle alone (595).&lt;/p&gt;

&lt;h2&gt;
  
  
  Chisle compresses three things; caveman and ponytail compress one each
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;prose&lt;/th&gt;
&lt;th&gt;code judgment&lt;/th&gt;
&lt;th&gt;input / context&lt;/th&gt;
&lt;th&gt;publishes failures&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;caveman&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;td&gt;❌&lt;/td&gt;
&lt;td&gt;❌&lt;/td&gt;
&lt;td&gt;❌&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ponytail&lt;/td&gt;
&lt;td&gt;❌&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;td&gt;❌&lt;/td&gt;
&lt;td&gt;❌&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Chisle&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Output prose:&lt;/strong&gt; no filler, no hedging, no manufactured structure.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Output code:&lt;/strong&gt; a YAGNI "efficiency ladder" (use what exists, ask instead of guessing, skip speculative abstractions). Safety, error handling, validation and accessibility are never cut.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Input context, the axis neither rival touches.&lt;/strong&gt; Across 171 real Claude Code sessions, &lt;strong&gt;tool output was 67.5% of the context window&lt;/strong&gt;, re-billed on every later request. Chisle's &lt;code&gt;PostToolUse&lt;/code&gt; hook trims oversized tool output by &lt;strong&gt;~46%&lt;/strong&gt; before it re-enters context (&lt;code&gt;Read&lt;/code&gt;/&lt;code&gt;Edit&lt;/code&gt;/&lt;code&gt;Write&lt;/code&gt; are never touched).&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Same prompt, same model
&lt;/h3&gt;

&lt;p&gt;&lt;em&gt;"Add debounce to a search input that currently fires an API call on every keystroke."&lt;/em&gt; Verbatim committed output:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No plugin: 142 lines, 1,506 tokens.&lt;/strong&gt; A generic &lt;code&gt;useDebounce&amp;lt;T&amp;gt;&lt;/code&gt; hook in its own file, then Option 2, Option 3, a comparison table and caveats.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Chisle: 35 lines, 602 tokens.&lt;/strong&gt; &lt;code&gt;setTimeout&lt;/code&gt; in the effect you already have, two lines on why, and &lt;em&gt;"use &lt;code&gt;lodash.debounce&lt;/code&gt; if it's already installed."&lt;/em&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight jsx"&gt;&lt;code&gt;&lt;span class="nf"&gt;useEffect&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;timer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;setTimeout&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;query&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="cm"&gt;/* fetch */&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="mi"&gt;300&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;return &lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;clearTimeout&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;timer&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;query&lt;/span&gt;&lt;span class="p"&gt;]);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Not golfed, just boring: one less file, one less abstraction, same behaviour.&lt;/p&gt;
&lt;h2&gt;
  
  
  The fine print (because the repo publishes it)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The 20-task total includes one &lt;code&gt;cache&lt;/code&gt; task where the bare model wrote an unusually long answer (4,910 tokens, against 375–813 in later runs), and a few cells that an older ruleset example may have primed. Without those cells, Chisle's 20-task total is &lt;strong&gt;70%&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;A later, smaller rerun (13 prompts × 2 seeds on the newest plugin versions) again put &lt;strong&gt;Chisle lowest: 83%, vs caveman 102% and ponytail 105%&lt;/strong&gt;. It was the only one of the three below a bare model.&lt;/li&gt;
&lt;li&gt;On &lt;strong&gt;short answers&lt;/strong&gt;, every tool breaks even or worse; there's little to cut in a three-line reply.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Chisle is the only tool in this class that publishes the runs where it lost: &lt;a href="https://github.com/JayPokale/Chisle/tree/main/benchmarks/results" rel="noopener noreferrer"&gt;benchmarks/results&lt;/a&gt;.&lt;/p&gt;
&lt;h2&gt;
  
  
  Install Chisle in Claude Code
&lt;/h2&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx chisle             &lt;span class="c"&gt;# installs for Claude Code and any other agents it finds&lt;/span&gt;
npx chisle &lt;span class="nt"&gt;--dry-run&lt;/span&gt;   &lt;span class="c"&gt;# preview first&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Zero dependencies, MIT licensed. The same ruleset ships to &lt;strong&gt;Cursor, Codex, Gemini CLI, GitHub Copilot, Windsurf, Cline, OpenCode, Kiro, Antigravity, Hermes and Pi&lt;/strong&gt;. Claude Code and Pi also get the input-side compressor and live modes (&lt;code&gt;lite&lt;/code&gt;, &lt;code&gt;full&lt;/code&gt;, &lt;code&gt;ultra&lt;/code&gt;). &lt;code&gt;/chisle-audit&lt;/code&gt; flags over-engineered code &lt;em&gt;and&lt;/em&gt; bloated prose/docs in one ranked report; ponytail's audit is code-only and caveman has none.&lt;/p&gt;
&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Is Chisle better than caveman?&lt;/strong&gt;&lt;br&gt;
Across 20 tasks: &lt;strong&gt;52% vs 80%&lt;/strong&gt; of the bare-model bill, worst case &lt;strong&gt;173% vs 424%&lt;/strong&gt;, &lt;strong&gt;1 vs 6&lt;/strong&gt; backfires. On coding prompts &lt;strong&gt;44% vs 74%&lt;/strong&gt;. caveman is a little leaner on some short prose prompts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is Chisle better than ponytail?&lt;/strong&gt;&lt;br&gt;
Across 20 tasks: &lt;strong&gt;52% vs 68%&lt;/strong&gt;, worst case &lt;strong&gt;173% vs 227%&lt;/strong&gt;, &lt;strong&gt;1 vs 8&lt;/strong&gt; backfires. On explanation prompts ponytail goes above 100% (104%); Chisle stays at 87%.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I install caveman and ponytail together instead?&lt;/strong&gt;&lt;br&gt;
You can, but they fight over prose style and double the overhead. On one task the pair did worse (605 tokens) than Chisle alone (595).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does it make answers wrong?&lt;/strong&gt;&lt;br&gt;
In the July re-verification every answer from every arm graded correct. The rule is &lt;em&gt;necessary&lt;/em&gt;, not &lt;em&gt;fewest characters&lt;/em&gt;, and safety is never cut.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where's the raw data?&lt;/strong&gt;&lt;br&gt;
All transcripts are committed: &lt;a href="https://github.com/JayPokale/Chisle/tree/main/benchmarks/results" rel="noopener noreferrer"&gt;benchmarks/results&lt;/a&gt;. Full tables: &lt;a href="https://github.com/JayPokale/Chisle/blob/main/docs/benchmarks.md" rel="noopener noreferrer"&gt;docs/benchmarks.md&lt;/a&gt;.&lt;/p&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/JayPokale" rel="noopener noreferrer"&gt;
        JayPokale
      &lt;/a&gt; / &lt;a href="https://github.com/JayPokale/Chisle" rel="noopener noreferrer"&gt;
        Chisle
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Cut your AI coding agent's token bill on three axes: terse prose, YAGNI-first code, and tool-output compression. Claude Code, Pi, Cursor, Codex, Gemini + 4 more. Zero deps, published benchmarks including the runs it loses.
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;p&gt;
  &lt;a rel="noopener noreferrer" href="https://github.com/JayPokale/Chisle/assets/logo.png"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2FJayPokale%2FChisle%2FHEAD%2Fassets%2Flogo.png" width="120" alt="Chisle"&gt;&lt;/a&gt;
&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;Chisle&lt;/h1&gt;
&lt;/div&gt;

&lt;p&gt;
  &lt;em&gt;Your AI talks less, builds less, reads less, and says more. Like a senior dev who bills by the syllable.&lt;/em&gt;
&lt;/p&gt;

&lt;p&gt;
  &lt;em&gt;The only tool in this class that publishes the runs where it lost. &lt;a href="https://jaypokale.me/writing/chisle-benchmarks-it-loses" rel="nofollow noopener noreferrer"&gt;Here's why.&lt;/a&gt;&lt;/em&gt;
&lt;/p&gt;

&lt;p&gt;
  &lt;a href="https://www.npmjs.com/package/chisle" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/945719fdc04e4dd512464a1c4b8c272d8f922b663714fbc2df04726d89302f13/68747470733a2f2f696d672e736869656c64732e696f2f6e706d2f762f636869736c653f7374796c653d666c61742d73717561726526636f6c6f723d643738613363" alt="npm version"&gt;&lt;/a&gt;
  &lt;a rel="noopener noreferrer nofollow" href="https://camo.githubusercontent.com/0d43d2e3a0ca52af29704db00064b0ac23b7a84d2f67d4d65a61c9a1f08681b1/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f776f726b73253230776974682d31322532306167656e74732d6437386133633f7374796c653d666c61742d737175617265"&gt;&lt;img src="https://camo.githubusercontent.com/0d43d2e3a0ca52af29704db00064b0ac23b7a84d2f67d4d65a61c9a1f08681b1/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f776f726b73253230776974682d31322532306167656e74732d6437386133633f7374796c653d666c61742d737175617265" alt="Works with 12 agents"&gt;&lt;/a&gt;
  &lt;a href="https://github.com/JayPokale/Chisle/actions/workflows/test.yml" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/cb31b8cb9b49c154085465ab33384d14cbcab51f40305f1103042cf71faa9b43/68747470733a2f2f696d672e736869656c64732e696f2f6769746875622f616374696f6e732f776f726b666c6f772f7374617475732f4a6179506f6b616c652f436869736c652f746573742e796d6c3f7374796c653d666c61742d737175617265266c6162656c3d4349" alt="CI"&gt;&lt;/a&gt;
  &lt;a rel="noopener noreferrer nofollow" href="https://camo.githubusercontent.com/dff1f6228a3ea413e99e8089b1bddc580ab3174dc8493ab36170a5b64df9311a/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f646570732d302d3264613434653f7374796c653d666c61742d737175617265"&gt;&lt;img src="https://camo.githubusercontent.com/dff1f6228a3ea413e99e8089b1bddc580ab3174dc8493ab36170a5b64df9311a/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f646570732d302d3264613434653f7374796c653d666c61742d737175617265" alt="Zero deps"&gt;&lt;/a&gt;
  &lt;a rel="noopener noreferrer nofollow" href="https://camo.githubusercontent.com/72a0e4d42cb8f86d5dab6fda16879e90e6ce7ef278183d655fca9a04aef575c5/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f6c6963656e73652d4d49542d6437386133633f7374796c653d666c61742d737175617265"&gt;&lt;img src="https://camo.githubusercontent.com/72a0e4d42cb8f86d5dab6fda16879e90e6ce7ef278183d655fca9a04aef575c5/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f6c6963656e73652d4d49542d6437386133633f7374796c653d666c61742d737175617265" alt="MIT"&gt;&lt;/a&gt;
  &lt;a href="https://github.com/JayPokale/Chisle/stargazers" rel="noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/9171d8224b2072eaac59ca83f1a903b66385b07a6a55dc565872cbba7460d782/68747470733a2f2f696d672e736869656c64732e696f2f6769746875622f73746172732f4a6179506f6b616c652f436869736c653f7374796c653d736f6369616c" alt="Star Chisle on GitHub"&gt;&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;
  &lt;strong&gt;Built for Claude Code: coding answers come back 33% shorter and 24% cheaper, while caveman and ponytail make them longer · 12 agents · zero dependencies · one command&lt;/strong&gt;
&lt;/p&gt;

&lt;p&gt;Chisle is a Claude Code plugin that makes Claude cheaper to run without making it dumber. It cuts what Claude writes: no filler, no hedging, no speculative abstractions, just the smallest code that works. It also cuts what Claude reads: a &lt;code&gt;PostToolUse&lt;/code&gt; hook trims oversized tool output by ~46% before it re-enters the context window, where it would be re-billed on every later request. In agent-loop tests Claude with Chisle passes the same tasks as Claude without it. One &lt;code&gt;npx chisle&lt;/code&gt; installs it, and the…&lt;/p&gt;&lt;/div&gt;


&lt;/div&gt;
&lt;br&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/JayPokale/Chisle" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;br&gt;
&lt;/div&gt;
&lt;br&gt;


&lt;p&gt;Site: &lt;strong&gt;&lt;a href="https://chisle.jaypokale.me" rel="noopener noreferrer"&gt;chisle.jaypokale.me&lt;/a&gt;&lt;/strong&gt;. If you run your own comparison against caveman or ponytail, I'd like to see it, especially where Chisle loses.&lt;/p&gt;

</description>
      <category>claudecode</category>
      <category>ai</category>
      <category>productivity</category>
      <category>opensource</category>
    </item>
    <item>
      <title>My speech-flaw detector flagged 41 false alarms a minute. One line of math fixed it.</title>
      <dc:creator>Jay Pokale</dc:creator>
      <pubDate>Tue, 06 Oct 2026 20:04:50 +0000</pubDate>
      <link>https://dev.to/jaypokale/my-speech-flaw-detector-flagged-41-false-alarms-a-minute-one-line-of-math-fixed-it-1id1</link>
      <guid>https://dev.to/jaypokale/my-speech-flaw-detector-flagged-41-false-alarms-a-minute-one-line-of-math-fixed-it-1id1</guid>
      <description>&lt;p&gt;I built &lt;strong&gt;Podium&lt;/strong&gt;, a tool that compares your reading of a speech with a great delivery of the &lt;em&gt;same text&lt;/em&gt; (JFK, Reagan) and tells you exactly where you drift and why: "14.3 syllables/s here vs 7.7 in the reference (+85%)", pinned to a time range.&lt;/p&gt;

&lt;p&gt;The first version worked beautifully on my test set. Then I gave it a different speaker, and it flagged &lt;strong&gt;41 "flaws" per minute&lt;/strong&gt; on a perfectly clean recording.&lt;/p&gt;

&lt;p&gt;This post covers why that happened, the one-line fix, and the evaluation set-up that caught it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvvt909v4wwvxdpaqa4m6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvvt909v4wwvxdpaqa4m6.png" alt="Podium dashboard" width="800" height="1486"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: put both recordings on the same word grid
&lt;/h2&gt;

&lt;p&gt;Comparing two deliveries frame by frame is hopeless: they never line up. Comparing them &lt;em&gt;word by word&lt;/em&gt; is easy, as long as both are force-aligned to the same transcript. I used torchaudio's &lt;code&gt;MMS_FA&lt;/code&gt; (a wav2vec2 CTC aligner):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;bundle&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;torchaudio&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;pipelines&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;MMS_FA&lt;/span&gt;
&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;tokenizer&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;aligner&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;bundle&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;with_star&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;bundle&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_tokenizer&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="n"&gt;bundle&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_aligner&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;emission&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;wav&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;spans&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;aligner&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;emission&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="nf"&gt;tokenizer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;words&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;   &lt;span class="c1"&gt;# one span list per word
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Now word 17 of your reading is word 17 of JFK's. For every word I compute speaker-normalised features with Praat (via Parselmouth) and librosa:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;pitch in &lt;strong&gt;semitones relative to your own median&lt;/strong&gt; (a deep and a high voice with the same intonation give the same curve);&lt;/li&gt;
&lt;li&gt;loudness in &lt;strong&gt;dB relative to your own level&lt;/strong&gt;;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;articulation rate&lt;/strong&gt;, syllables per second of actual speech, pauses excluded, over a 5-word window;&lt;/li&gt;
&lt;li&gt;the pause before the word, and &lt;strong&gt;voiced sound inside that pause&lt;/strong&gt; (more on that below);&lt;/li&gt;
&lt;li&gt;high-frequency energy above 2.5 kHz, which drops when consonants get swallowed.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;
  
  
  Step 2: build a dataset where the answers are exact
&lt;/h2&gt;

&lt;p&gt;There's no public dataset that pairs a good delivery with bad deliveries of the same words. So I made one: &lt;strong&gt;207 recordings&lt;/strong&gt;, starting from three public-domain speeches.&lt;/p&gt;

&lt;p&gt;The trick is that I &lt;em&gt;inject&lt;/em&gt; the flaws myself, so every label is sample-accurate:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Flaw&lt;/th&gt;
&lt;th&gt;How it's injected&lt;/th&gt;
&lt;th&gt;Severity 1 / 2 / 3&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;rushed / dragged&lt;/td&gt;
&lt;td&gt;phase-vocoder time-scale of 4–9 words&lt;/td&gt;
&lt;td&gt;×1.3/1.6/2.0 · ×0.8/0.65/0.5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;monotone&lt;/td&gt;
&lt;td&gt;WORLD vocoder, pitch contour squashed toward its mean&lt;/td&gt;
&lt;td&gt;50/75/95% removed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;mumbled&lt;/td&gt;
&lt;td&gt;gain down + low-pass&lt;/td&gt;
&lt;td&gt;−6 dB @ 3 kHz … −16 dB @ 1.1 kHz&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;awkward pause&lt;/td&gt;
&lt;td&gt;room-tone silence inserted mid-phrase&lt;/td&gt;
&lt;td&gt;0.7 / 1.3 / 2.2 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;missing breath&lt;/td&gt;
&lt;td&gt;a natural pause squeezed out&lt;/td&gt;
&lt;td&gt;50 / 20 / 0% kept&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;stutter&lt;/td&gt;
&lt;td&gt;word onset repeated&lt;/td&gt;
&lt;td&gt;1 / 2 / 3 repeats&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Edits are joined with 5 ms fades that &lt;strong&gt;preserve length&lt;/strong&gt;, so the label times never drift. My first version used 15 ms crossfades, which quietly shifted every later label by 15 ms per edit.&lt;/p&gt;

&lt;p&gt;I also had the same texts read by open &lt;strong&gt;Piper&lt;/strong&gt; TTS voices, with flaws injected into those too. That's the "different speaker" test.&lt;/p&gt;

&lt;p&gt;And I split it honestly: thresholds are tuned &lt;strong&gt;only on JFK&lt;/strong&gt;, then frozen and tested on &lt;strong&gt;Reagan&lt;/strong&gt; plus an unseen voice.&lt;/p&gt;
&lt;h2&gt;
  
  
  Step 3: the bug that wasn't a bug
&lt;/h2&gt;

&lt;p&gt;Version 1 scored each word by how far it departed from the reference, as a robust z-score (median/MAD) after removing your overall offset. On Reagan's own audio with injected flaws it got F1 0.64 with &lt;strong&gt;zero&lt;/strong&gt; false alarms.&lt;/p&gt;

&lt;p&gt;On a TTS voice reading the same text:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;F1&lt;/th&gt;
&lt;th&gt;False alarms on clean audio&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Same speaker&lt;/td&gt;
&lt;td&gt;0.64&lt;/td&gt;
&lt;td&gt;0 / min&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Different speaker&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.03&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;41 / min&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A synthetic voice doesn't breathe where Reagan breathes or lift the words he lifts. Measured against Reagan, &lt;em&gt;everything&lt;/em&gt; it does is a deviation. Technically correct, but useless as feedback: nobody wants to hear that every sentence is wrong because they aren't Reagan.&lt;/p&gt;

&lt;p&gt;The fix was to separate &lt;strong&gt;style&lt;/strong&gt; from &lt;strong&gt;flaws&lt;/strong&gt;. A flaw is something that's off compared with the reference &lt;em&gt;and&lt;/em&gt; off compared with &lt;strong&gt;your own&lt;/strong&gt; delivery. So every word gets two scores:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;z_ref&lt;/span&gt;  &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;robust_z&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;participant&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;reference&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# departs from the reference, beyond your overall style
&lt;/span&gt;&lt;span class="n"&gt;z_self&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;robust_z&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;participant_feature&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;       &lt;span class="c1"&gt;# stands out within your own recording
&lt;/span&gt;&lt;span class="n"&gt;flaw&lt;/span&gt;   &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fmin&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;z_ref&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;z_self&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;              &lt;span class="c1"&gt;# a soft AND
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;A consistently different voice moves &lt;code&gt;z_ref&lt;/code&gt; everywhere but &lt;code&gt;z_self&lt;/code&gt; almost nowhere, so it's reported once as &lt;em&gt;style&lt;/em&gt; ("33% faster and flatter than the reference") instead of 40 times as flaws.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;False alarms on clean different-speaker recordings went from 41/min to 3/min.&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Step 4: a stutter that looked like a pause
&lt;/h2&gt;

&lt;p&gt;Injected stutters kept being reported as "awkward pause". The cause: the aligner usually places the word at its &lt;em&gt;last&lt;/em&gt; restart, so "w- w- we" becomes a long gap followed by a normal "we".&lt;/p&gt;

&lt;p&gt;The fix was one feature: &lt;strong&gt;seconds of voiced audio inside the gap&lt;/strong&gt;. A real pause is silent; a gap full of voiced sound is a restart.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;gap_voiced&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;~&lt;/span&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;isnan&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;f0&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;prev_end&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="n"&gt;word_start&lt;/span&gt;&lt;span class="p"&gt;]))&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;HOP&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Long pauses now require &lt;code&gt;gap_voiced &amp;lt; 0.15 s&lt;/code&gt;, and stutters can trigger on &lt;code&gt;gap_voiced &amp;gt; 0.12 s&lt;/code&gt;. Stutter recall went from 0.11 to 0.44 on the held-out set.&lt;/p&gt;
&lt;h2&gt;
  
  
  Results on the held-out set (Reagan + an unseen voice)
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;detected regions overlap the true flaw with &lt;strong&gt;mean IoU 0.81&lt;/strong&gt;; overall &lt;strong&gt;F1 0.51&lt;/strong&gt;;&lt;/li&gt;
&lt;li&gt;awkward pauses, missing breaths, mumbling and stutters are located within &lt;strong&gt;0.00–0.06 s&lt;/strong&gt;;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;0 false alarms per minute&lt;/strong&gt; on untouched audio of the reference speaker;&lt;/li&gt;
&lt;li&gt;the rubric score falls monotonically from L0 to L4 on &lt;strong&gt;every&lt;/strong&gt; excerpt (Spearman ρ = −1.00).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What doesn't work yet:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Pacing is weakest (F1 0.29–0.35).&lt;/strong&gt; Slowing a phrase also stretches its pauses, so "dragged" is often reported as "awkward pause"; the confusion matrix shows exactly that.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Different speakers are still hard (F1 0.21).&lt;/strong&gt; Fixed noise floors are the next thing to replace with learned ones.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;
  
  
  Three things I'd tell anyone building an evaluator
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Make your labels exact before tuning anything.&lt;/strong&gt; Injecting the flaws yourself turns an argument about subjective judgement into a number.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hold out a speaker, not just files.&lt;/strong&gt; My calibration F1 (0.62) and held-out F1 (0.51) are close &lt;em&gt;because&lt;/em&gt; the split was honest. Same-speaker numbers alone would have hidden the 41-per-minute problem.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Decide what &lt;em&gt;shouldn't&lt;/em&gt; count before deciding what should.&lt;/strong&gt; The hardest part wasn't detecting deviations, it was deciding which deviations matter.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Code, dataset (207 labelled recordings) and the full evaluation are open source:&lt;/p&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/JayPokale" rel="noopener noreferrer"&gt;
        JayPokale
      &lt;/a&gt; / &lt;a href="https://github.com/JayPokale/podium" rel="noopener noreferrer"&gt;
        podium
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Contrastive speech delivery analytics: find exactly where a delivery drifts from a great speech, and why. Multimodal AI Hackathon 2026, Track C.
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;🎙️ Podium: contrastive speech delivery analytics with temporal flaw grounding&lt;/h1&gt;
&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Read a great speech. Podium shows exactly where your delivery drifts from it, to the tenth of a second
and explains each flaw with the numbers behind it.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Built for the Multimodal AI Hackathon 2026, &lt;strong&gt;Track C: Contrastive Speech Analytics &amp;amp; Temporal Flaw Grounding&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;&lt;a rel="noopener noreferrer" href="https://github.com/JayPokale/podium/docs/img/dashboard_full.png"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2FJayPokale%2Fpodium%2FHEAD%2Fdocs%2Fimg%2Fdashboard_full.png" alt="Podium dashboard: rubric, pitch/loudness/pace overlays with flaw regions, explained flaw cards"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;What it does&lt;/h2&gt;
&lt;/div&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Pick a reference&lt;/strong&gt;: an excerpt of a great public-domain speech (JFK, Reagan).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Give your delivery&lt;/strong&gt;: record yourself reading the same text in the browser, upload a file, or try one of
the 207 dataset samples.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Get grounded feedback&lt;/strong&gt;
&lt;ul&gt;
&lt;li&gt;a deterministic &lt;strong&gt;0-10 rubric&lt;/strong&gt; for pacing, pauses, pitch, volume and fluency;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;flaw regions&lt;/strong&gt; with exact start/end times, shaded on time-series overlays of &lt;em&gt;your&lt;/em&gt; pitch, loudness and
pace against the reference (time-warped onto your timeline word by word);&lt;/li&gt;
&lt;li&gt;a &lt;strong&gt;causal explanation&lt;/strong&gt; for every region, built from the measured delta ("14.3 syllables/s here…&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/JayPokale/podium" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;p&gt;&lt;em&gt;Built for the Multimodal AI Hackathon 2026 (Track C), with heavy help from Claude Code, an AI coding agent, for implementation and evaluation runs.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>machinelearning</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Outside Quest: I asked a local Gemma to name birds. It couldn't, so it sends you on a scavenger hunt instead</title>
      <dc:creator>Jay Pokale</dc:creator>
      <pubDate>Tue, 06 Oct 2026 06:05:08 +0000</pubDate>
      <link>https://dev.to/jaypokale/outside-quest-i-asked-a-local-gemma-to-name-birds-it-couldnt-so-it-sends-you-on-a-scavenger-l2a</link>
      <guid>https://dev.to/jaypokale/outside-quest-i-asked-a-local-gemma-to-name-birds-it-couldnt-so-it-sends-you-on-a-scavenger-l2a</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/hacktoberfest-week1-2026-10-05"&gt;Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Outside Quest&lt;/strong&gt; is a photo scavenger hunt that runs on Google's open-weight &lt;strong&gt;Gemma 4&lt;/strong&gt;, entirely on a laptop with Ollama. It's designed so the screen is the shortest part of the experience:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Before the walk (about 2 minutes):&lt;/strong&gt; type where you're going. Gemma writes a six-item quest card that fits the real place and season. Print it, or glance at it once. Then a full-screen &lt;strong&gt;"Phone away 🌳"&lt;/strong&gt; card takes over.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;On the walk (an hour):&lt;/strong&gt; the phone stays in your pocket. You only take it out to photograph a find.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Back home (a few minutes):&lt;/strong&gt; drop in your photos. Gemma checks which quests each photo completes and says what it saw. You get points, badges and a walk journal: a timeline from the photos' timestamps and a small route map drawn from their GPS tags, with no map tiles and no internet.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;It's for anyone who says "I should go outside more" and then doesn't: families with kids, friends who need an excuse for a walk, or me after a day of staring at code.&lt;/p&gt;

&lt;p&gt;The quest card adapts to where you are. For &lt;em&gt;Ambazari lake garden, Nagpur, India&lt;/em&gt; in October it asked for "a seed pod or fruit on the ground" and "moss or damp earth" (post-monsoon tropical). For &lt;em&gt;Central Park, New York&lt;/em&gt; it asked for "crimson leaf under a sturdy oak" and "a squirrel busy gathering nuts". That only happened after I told the prompt to think about the real local climate; the first version cheerfully asked Nagpur for "autumn colour change".&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5rk493o785hzxn3ml0wt.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5rk493o785hzxn3ml0wt.jpg" alt="Quest card for Ambazari lake garden, Nagpur, generated locally by Gemma 4" width="799" height="416"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyrdimvpgpawk4ryl2nb5.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyrdimvpgpawk4ryl2nb5.jpg" alt="Phone away screen" width="799" height="416"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq8g17k2jm4xzqsn1cl3k.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq8g17k2jm4xzqsn1cl3k.jpg" alt="Results board after checking six photos" width="799" height="416"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9ow8bygrgl8940njbvdp.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9ow8bygrgl8940njbvdp.jpg" alt="Walk journal with timeline and offline route" width="799" height="416"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The results above use six Wikimedia Commons photos (credited in the repo) with made-up timestamps and GPS points, so I could test the journal and map.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/JayPokale" rel="noopener noreferrer"&gt;
        JayPokale
      &lt;/a&gt; / &lt;a href="https://github.com/JayPokale/outside-quest" rel="noopener noreferrer"&gt;
        outside-quest
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Two minutes on a screen, an hour outside: an offline photo scavenger hunt powered by Gemma 4 + Ollama. Hacktoberfest 2026 Touch Grass.
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;Outside Quest 🌿&lt;/h1&gt;
&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;Two minutes on a screen, an hour outside.&lt;/strong&gt; An offline photo scavenger hunt powered by Google's open-weight
&lt;strong&gt;Gemma 4&lt;/strong&gt;, running on your own computer with &lt;a href="https://ollama.com" rel="nofollow noopener noreferrer"&gt;Ollama&lt;/a&gt;.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Before you go:&lt;/strong&gt; tell it where you're walking. Gemma writes a 6-item quest card that fits the real local
season ("a seed pod on the ground" in tropical Nagpur in October, "crimson leaves under an oak" in New York)
Print it or glance at it once.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;On the walk:&lt;/strong&gt; phone in your pocket. You only take it out to snap a photo of each find.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Back home:&lt;/strong&gt; drop in your photos. Gemma checks which quests each photo completes, with the visible evidence
and builds a walk journal: a timeline from the photos' timestamps and a little route map from their GPS tags.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;No account, no upload, no internet needed. Your photos (and their location data) never leave the computer.&lt;/p&gt;
&lt;p&gt;…&lt;/p&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/JayPokale/outside-quest" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;h3&gt;
  
  
  Try it
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ollama pull gemma4:e2b
git clone https://github.com/JayPokale/outside-quest &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cd &lt;/span&gt;outside-quest
pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; requirements.txt &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; python3 server.py   &lt;span class="c"&gt;# http://127.0.0.1:8777&lt;/span&gt;
&lt;span class="c"&gt;# optional, for durable checking:&lt;/span&gt;
temporal server start-dev &lt;span class="nt"&gt;--db-filename&lt;/span&gt; ~/.outside-quest/temporal.db
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No account, no API key. There's also a &lt;strong&gt;Quick card&lt;/strong&gt; button that works with no model at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Stack:&lt;/strong&gt; &lt;code&gt;gemma4:e2b&lt;/code&gt; through Ollama on a 6 GB GTX 1660 Ti laptop GPU · a Python standard-library server plus Pillow (for EXIF) · an optional local Temporal server for durable checking · one HTML file, no framework. A quest card takes ~30 seconds, checking a photo ~30 seconds.&lt;/p&gt;

&lt;h3&gt;
  
  
  Plan A failed, and that's the most useful thing I learned
&lt;/h3&gt;

&lt;p&gt;My first idea was the obvious one: a bird and plant identifier that works on the trail with no signal. Before building the UI, I tested the model on six photos: a Common Myna, a neem tree, a Monarch butterfly, a fly agaric mushroom, a peacock and a storm cloud.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Photo&lt;/th&gt;
&lt;th&gt;gemma4:e2b said&lt;/th&gt;
&lt;th&gt;gemma4:e4b said&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Common Myna&lt;/td&gt;
&lt;td&gt;"Jackdaw", confidence &lt;strong&gt;high&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;"Weaver Bird", confidence &lt;strong&gt;high&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Monarch butterfly&lt;/td&gt;
&lt;td&gt;"butterfly, undetermined"&lt;/td&gt;
&lt;td&gt;"Heliconius"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Neem leaves&lt;/td&gt;
&lt;td&gt;"deciduous trees, undetermined"&lt;/td&gt;
&lt;td&gt;"trees / forest canopy"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Peacock&lt;/td&gt;
&lt;td&gt;Peacock ✓&lt;/td&gt;
&lt;td&gt;Peacock ✓&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Asking for a top-3 list didn't rescue it: the Myna never appeared. Gemma 3 4B did worse and broke its JSON. A tool that names the wrong bird &lt;em&gt;with high confidence&lt;/em&gt; is worse than no tool, especially outdoors, where people make decisions about mushrooms and berries.&lt;/p&gt;

&lt;p&gt;But in every run, the &lt;strong&gt;coarse&lt;/strong&gt; answer was right: &lt;em&gt;there's a bird, a butterfly on a flower, red-capped mushrooms, a storm cloud, a canopy of leaves&lt;/em&gt;. So I flipped the design: the model only answers the question small open models are good at, "does this photo show a bird?", and the naming is left to you and a field guide. A scavenger hunt turned out to be a better way to get people outside anyway.&lt;/p&gt;

&lt;h3&gt;
  
  
  Making the judge strict
&lt;/h3&gt;

&lt;p&gt;The first judge was too generous: it gave "a leaf" points for the butterfly photo because there were leaves in the background. Word-matching heuristics made it worse (it then rejected the peacock for "a bird"). What worked was asking Gemma, through a JSON schema, two explicit yes/no questions for &lt;strong&gt;every&lt;/strong&gt; quest on &lt;strong&gt;every&lt;/strong&gt; photo:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"q2"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"evidence"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"A butterfly (an insect) is clearly resting on the flower."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"is_main_subject"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"completed"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A quest only counts if both are true. On a test card (a bird, an insect on a flower, a mushroom, a stormy cloud, leaves up close, something red) and the six test photos, that gave &lt;strong&gt;5 correct matches and no false ones&lt;/strong&gt;. The miss was the storm cloud, which it read as a seascape. It isn't perfect: in the run shown above it gave "a leaf bigger than your hand" to a canopy photo, where nobody can judge leaf size. But it errs strict far more often than generous, and for a game that's the right way round.&lt;/p&gt;

&lt;h3&gt;
  
  
  Safety lives in code, not in the prompt
&lt;/h3&gt;

&lt;p&gt;Quests the model writes pass through plain-code filters before anyone sees them. Anything about climbing, entering water, touching, picking, eating, feeding or chasing animals, roads, railway tracks, private land or night walks is dropped and replaced from a safe built-in pool. Mushrooms, berries, nests and eggs always get "(photo only, don't touch)". If the model is down or returns junk, the card fills from the same pool, and there's an instant "Quick card" button that skips the model entirely.&lt;/p&gt;

&lt;h3&gt;
  
  
  Making the photo check durable with Temporal
&lt;/h3&gt;

&lt;p&gt;Checking is the slow part: about 30 seconds per photo, so a 40-photo walk is 20 minutes of a laptop GPU working. The first version looped over photos in the browser. Close the tab, let the laptop sleep, or have Ollama run out of memory on photo 23, and the rest of the batch was simply gone.&lt;/p&gt;

&lt;p&gt;So the Gemma judge now runs inside a &lt;a href="https://temporal.io" rel="noopener noreferrer"&gt;Temporal&lt;/a&gt; workflow, on the open-source Temporal dev server on the same laptop. Each walk is a &lt;code&gt;CheckWalk&lt;/code&gt; workflow and each photo is a &lt;code&gt;check_one&lt;/code&gt; activity:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Photos are saved to disk first, so you can close the tab as soon as the upload finishes. Reopen the page later and it picks the walk back up.&lt;/li&gt;
&lt;li&gt;If Ollama is down, out of memory, or returns broken JSON, the activity is retried with backoff (5 seconds up to 2 minutes, 30 attempts) instead of being skipped.&lt;/li&gt;
&lt;li&gt;If the server dies, the workflow resumes from the next unchecked photo. Each answer is cached next to its photo, so nothing is checked twice.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;To test it, I stopped Ollama in the middle of a six-photo walk and then killed the app server too. The red bar is the photo that kept failing while Ollama was down. Attempt 4 succeeded once it came back, and the walk finished with all six photos checked:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdx40wq1gqwpzwv5tgbtn.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdx40wq1gqwpzwv5tgbtn.jpg" alt="Temporal timeline: a photo check retried while Ollama was down, then the walk finished" width="800" height="148"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If Temporal isn't running, the app works exactly as before. Durable mode is an extra, not a requirement.&lt;/p&gt;

&lt;h3&gt;
  
  
  What a Reddit reader caught
&lt;/h3&gt;

&lt;p&gt;After I shared Outside Quest on Reddit, a reader pointed out something I hadn't tested: WhatsApp strips EXIF data, and iPhone sharing can leave location out. So for many people the route map and timeline would just be empty, with no error. The journal now says so ("no location in 2 of 5 photos") and suggests copying the originals over USB or AirDrop. Fixing that also turned up a bug: the walk length mixed camera times with file dates, which made one test walk "3175 minutes" long. It now uses camera times only.&lt;/p&gt;

&lt;p&gt;They also asked whether the card changes with the season. It does. Same place, Central Park, three months: October asked for "a fallen maple leaf", January for "ice crystals on a puddle", and April for "a daffodil or early tulip bloom".&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Does Open Innovation Matter?
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Your photos are a map of where you were.&lt;/strong&gt; Walk photos carry GPS coordinates and timestamps. Sending them to a cloud API would hand someone a log of when and where you walk. Here the model and the photos stay on your laptop.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It works where the walk is.&lt;/strong&gt; No API key, no signal needed, no per-photo cost, so a family can check a hundred photos after a holiday without a bill.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;I could test it honestly and change course.&lt;/strong&gt; Because the model runs locally and is free to call, I could benchmark e2b against e4b and Gemma 3 on the same photos in an evening, see exactly where it fails, and design around that instead of trusting a demo.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Where closed would win:&lt;/strong&gt; a frontier cloud model would almost certainly name the Myna. For this app that's the wrong trade: I'd rather have a private, free, offline game that's honest about what it knows.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  My Agent Session
&lt;/h2&gt;

&lt;p&gt;Built with heavy help from Claude Code (an AI coding agent): it ran the model benchmarks above, wrote most of the code, and drove the UI in a browser to test it. The design decisions (dropping species ID, strict judging, code-level safety filters) came out of those test runs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prize Categories
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Best Use of Gemma&lt;/li&gt;
&lt;li&gt;Best Use of Temporal&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>devchallenge</category>
      <category>hf26challenge</category>
      <category>gemma</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
