<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Nitsuj</title>
    <description>The latest articles on DEV Community by Nitsuj (@nitsuj).</description>
    <link>https://dev.to/nitsuj</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3929911%2Fc45fcf27-1e29-4234-a374-4226cbb85497.jpg</url>
      <title>DEV Community: Nitsuj</title>
      <link>https://dev.to/nitsuj</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/nitsuj"/>
    <language>en</language>
    <item>
      <title>Touch Grass Bingo: an offline AI walk card that never leaves your browser</title>
      <dc:creator>Nitsuj</dc:creator>
      <pubDate>Wed, 07 Oct 2026 14:03:47 +0000</pubDate>
      <link>https://dev.to/nitsuj/touch-grass-bingo-an-offline-ai-walk-card-that-never-leaves-your-browser-9kj</link>
      <guid>https://dev.to/nitsuj/touch-grass-bingo-an-offline-ai-walk-card-that-never-leaves-your-browser-9kj</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/hacktoberfest-week1-2026-10-05"&gt;Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;Every "AI feature" in a web app today is a round trip: collect input, send it to a hosted&lt;br&gt;
model, stream tokens back. For a game whose whole point is &lt;strong&gt;getting you off the screen&lt;/strong&gt;, a&lt;br&gt;
server would have been both a contradiction and a liability — someone else would know where&lt;br&gt;
you walked.&lt;/p&gt;

&lt;p&gt;So &lt;strong&gt;Touch Grass Bingo&lt;/strong&gt; runs Gemma 3 1B &lt;em&gt;inside the browser&lt;/em&gt;, via&lt;br&gt;
&lt;a href="https://github.com/mlc-ai/web-llm" rel="noopener noreferrer"&gt;WebLLM&lt;/a&gt; and WebGPU. Pick a habitat, a time of day, and an&lt;br&gt;
optional focus ("birds, street art, dogs"). The model writes a 5×5 bingo card of things you&lt;br&gt;
can actually spot in the next 15 minutes. You tap squares on your walk — with zero signal —&lt;br&gt;
and afterwards the same model writes a short field recap of your outing.&lt;/p&gt;

&lt;p&gt;No accounts, no API keys, no analytics, no server. After the first visit it works in&lt;br&gt;
airplane mode.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it asks for, and what you get
&lt;/h2&gt;



&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    subgraph Browser["Browser (static files, any CDN)"]
        UI["UI shell&amp;lt;br/&amp;gt;13 KB app JS"]
        SW["Service Worker&amp;lt;br/&amp;gt;precaches app + engine"]
        subgraph Worker["Web Worker"]
            H["WebWorkerMLCEngineHandler"]
            E["MLCEngine / TVM"]
        end
        Cache["Browser weight cache"]
        Store["localStorage&amp;lt;br/&amp;gt;current card"]
    end
    HF[("HuggingFace CDN&amp;lt;br/&amp;gt;gemma3-1b-it-q4f16_1-MLC&amp;lt;br/&amp;gt;~570 MB, once")]

    UI -- "dynamic import()" --&amp;gt; E
    UI --- Store
    UI --- SW
    SW -. precache .- E
    E -- "one-time fetch" --&amp;gt; HF
    E --- Cache
    H -- "chat.completions.create()" --&amp;gt; E
    E -- "24 squares / recap" --&amp;gt; UI&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Generation runs in a &lt;strong&gt;Web Worker&lt;/strong&gt; so the UI never janks while the model prefills, and the&lt;br&gt;
6 MB engine chunk is &lt;strong&gt;lazy-loaded on first Generate&lt;/strong&gt; — the initial page is 13.8 KB of app&lt;br&gt;
code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why open innovation is the whole project, not a footnote
&lt;/h2&gt;

&lt;p&gt;The challenge asks why the open piece matters. For this app it &lt;em&gt;is&lt;/em&gt; the app:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;It runs on a phone in the backcountry with no internet.&lt;/strong&gt; Weights are cached after the
first load; generation, checking, and recaps are all local. A closed API can't promise
that, and I'd have nothing to demo at mile three of a trail.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your location data stays off a server you don't control.&lt;/strong&gt; The app never sees where you
are. There is no backend to leak.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It costs nothing to run.&lt;/strong&gt; Static hosting, no inference bill, no rate limits.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The model is swappable in one line.&lt;/strong&gt; WebLLM ships Llama, Qwen, Phi and more in its
prebuilt config — change &lt;code&gt;MODEL_ID&lt;/code&gt; in &lt;code&gt;src/engine.ts&lt;/code&gt; and the card's personality changes.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The interesting engineering problem: 1B models are flaky
&lt;/h2&gt;

&lt;p&gt;The fun part wasn't calling an API, it was making a small model reliable enough for a game&lt;br&gt;
that must always produce exactly 24 squares. Defenses, in order:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Failure mode&lt;/th&gt;
&lt;th&gt;Defense&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;JSON wrapped in prose or markdown fences&lt;/td&gt;
&lt;td&gt;bracket-hunting parser&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model answers with a numbered list instead of JSON&lt;/td&gt;
&lt;td&gt;line-based fallback parser&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;JSON arrays with trailing commas (Gemma's signature)&lt;/td&gt;
&lt;td&gt;salvage quoted strings when strict &lt;code&gt;JSON.parse&lt;/code&gt; fails&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Repeated squares across attempts&lt;/td&gt;
&lt;td&gt;case-insensitive dedupe&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Single-word lazy squares (&lt;code&gt;dog&lt;/code&gt;, &lt;code&gt;leaf&lt;/code&gt;, &lt;code&gt;flash&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;bare-noun gate drops them; card tops up from the pad pool, or regenerates&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Recap falls into a repetition loop (&lt;code&gt;a *another* dog…&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;repetition_penalty&lt;/code&gt; + degenerate-output detector + colder retry&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fewer than 24 items returned&lt;/td&gt;
&lt;td&gt;top-up from a themed pad pool&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Genuinely bad response&lt;/td&gt;
&lt;td&gt;retry with a stricter prompt, then a built-in preset deck&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;No WebGPU / failed worker&lt;/td&gt;
&lt;td&gt;preset decks + an honest status chip, never a dead end&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Every path terminates in a playable card. The parser and defense tests live in &lt;code&gt;tests/extract.test.ts&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A detail worth stealing:&lt;/strong&gt; before picking the model I checked WebLLM's prebuilt config for&lt;br&gt;
&lt;code&gt;required_features&lt;/code&gt;. Plenty of popular &lt;code&gt;q4f16_1&lt;/code&gt; builds declare&lt;br&gt;
&lt;code&gt;required_features: ["shader-f16"]&lt;/code&gt; and throw &lt;code&gt;ShaderF16SupportError&lt;/code&gt; on GPUs without that&lt;br&gt;
feature. &lt;code&gt;gemma3-1b-it-q4f16_1-MLC&lt;/code&gt; (711 MB VRAM, &lt;code&gt;low_resource_required: true&lt;/code&gt;) does not —&lt;br&gt;
which makes it the right default for a game anyone should be able to open. Another landmine:&lt;br&gt;
Gemma 3's config ships both &lt;code&gt;context_window_size&lt;/code&gt; and &lt;code&gt;sliding_window_size&lt;/code&gt; positive, which&lt;br&gt;
WebLLM refuses (&lt;code&gt;WindowSizeConfigurationError&lt;/code&gt;) — unless you disable one at load time via&lt;br&gt;
&lt;code&gt;chatOpts: { sliding_window_size: -1 }&lt;/code&gt;. Feature-detecting &lt;code&gt;navigator.gpu&lt;/code&gt; up front and&lt;br&gt;
falling back to preset decks keeps the promise: the app never dead-ends.&lt;/p&gt;

&lt;h2&gt;
  
  
  Getting it running
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/hey-nitsuj/Touch-Grass-Bingo
npm &lt;span class="nb"&gt;install
&lt;/span&gt;npm run dev        &lt;span class="c"&gt;# local dev&lt;/span&gt;
npm run &lt;span class="nb"&gt;test&lt;/span&gt;       &lt;span class="c"&gt;# parser tests&lt;/span&gt;
npm run build      &lt;span class="c"&gt;# typecheck + production build&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;WebGPU needs a secure context — &lt;code&gt;https://&lt;/code&gt; or &lt;code&gt;localhost&lt;/code&gt;, nowhere else.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd do differently
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Streaming.&lt;/strong&gt; Card generation currently resolves as one blob; streaming the 24 squares in
as they decode would make the wait feel shorter. WebLLM supports it; I chose simplicity
first.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A service worker engine host.&lt;/strong&gt; WebLLM can host the engine in a service worker so the
model survives reloads without a worker respawn. Worth it for a v2.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Try it, then go outside
&lt;/h2&gt;

&lt;p&gt;The point of the game is that the screen is the shortest part of the experience. Generate a&lt;br&gt;
card, put the phone in your pocket, walk, and come back for the recap.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Play it live: &lt;a href="https://hey-nitsuj.github.io/Touch-Grass-Bingo/" rel="noopener noreferrer"&gt;https://hey-nitsuj.github.io/Touch-Grass-Bingo/&lt;/a&gt;&lt;/strong&gt; — or read the source at&lt;br&gt;
&lt;a href="https://github.com/hey-nitsuj/Touch-Grass-Bingo" rel="noopener noreferrer"&gt;github.com/hey-nitsuj/Touch-Grass-Bingo&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Built with WebLLM + Gemma 3 1B (open weights), Vite, and no backend at all. MIT licensed.&lt;/p&gt;

</description>
      <category>hf26challenge</category>
      <category>gemma</category>
      <category>webgpu</category>
      <category>devchallenge</category>
    </item>
  </channel>
</rss>
