<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Shishir Shrivastava</title>
    <description>The latest articles on DEV Community by Shishir Shrivastava (@shishir_shrivastava_86e7f).</description>
    <link>https://dev.to/shishir_shrivastava_86e7f</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1905855%2F93d4035c-4761-41cb-84f2-257a445e4feb.png</url>
      <title>DEV Community: Shishir Shrivastava</title>
      <link>https://dev.to/shishir_shrivastava_86e7f</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/shishir_shrivastava_86e7f"/>
    <language>en</language>
    <item>
      <title>Tuft: a scavenger hunt that checks your photos on your phone, offline</title>
      <dc:creator>Shishir Shrivastava</dc:creator>
      <pubDate>Wed, 07 Oct 2026 18:17:18 +0000</pubDate>
      <link>https://dev.to/shishir_shrivastava_86e7f/tuft-a-scavenger-hunt-that-checks-your-photos-on-your-phone-offline-2c6k</link>
      <guid>https://dev.to/shishir_shrivastava_86e7f/tuft-a-scavenger-hunt-that-checks-your-photos-on-your-phone-offline-2c6k</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/hacktoberfest-week1-2026-10-05"&gt;Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;Tuft is a scavenger hunt you mostly listen to. You pick a hunt ("Neighbourhood Basics", "Autumn Walk"), put the phone in your pocket, and a spoken clue tells you what to find: &lt;em&gt;"Find a tree. Put your hand on the bark, then take a photo of the trunk."&lt;/em&gt; When you find it, you take one photo. An open-weight vision model running on your phone decides whether the photo really shows what the clue asked for. Then you get the next clue.&lt;/p&gt;

&lt;p&gt;The problem: "go outside" is easy advice, and a game gives a walk a purpose. But a game that trusts everyone has nothing to prove, and one that verifies photos in the cloud sends pictures of your street, your kids and your dog to someone else's server. Tuft verifies locally, so it can be honest and private at once. The screen is the short part: a clue is a sentence you hear, a photo takes a few seconds, and the rest is walking.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Deployed app: &lt;a href="https://shishir2405.github.io/tuft/" rel="noopener noreferrer"&gt;https://shishir2405.github.io/tuft/&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/Shishir2405" rel="noopener noreferrer"&gt;
        Shishir2405
      &lt;/a&gt; / &lt;a href="https://github.com/Shishir2405/tuft" rel="noopener noreferrer"&gt;
        tuft
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      On-device scavenger hunt: an open-weight vision model checks your photos on your phone. No server, works offline.
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;Tuft&lt;/h1&gt;
&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;A scavenger hunt that talks to you, checks your photos on your phone, and works with no signal.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;You pick a hunt, pocket the phone, and listen. Each clue is spoken ("Find a tree. Put your hand on the bark…"). When you find the thing, you take one photo; an open-weight vision model running &lt;strong&gt;on the device&lt;/strong&gt; decides whether it really is that thing. No account, no server, no photo ever uploaded.&lt;/p&gt;
&lt;p&gt;The screen is the shortest part of the walk: a clue is a sentence you hear, a photo is a few seconds, and the rest is outdoors.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Built for the Hacktober Open-Source AI Challenge, Week 1: "Touch Grass".&lt;/p&gt;
&lt;/blockquote&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;Why it exists&lt;/h2&gt;
&lt;/div&gt;
&lt;p&gt;"Go outside" is easy advice and hard to follow. A scavenger hunt gives a walk a purpose, but the usual app versions either trust you (no fun, nothing checked) or send every photo to a cloud…&lt;/p&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/Shishir2405/tuft" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;p&gt;Apache-2.0.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Stack.&lt;/strong&gt; React + TypeScript + Vite, a service-worker PWA, no backend at all. Inference uses transformers.js and ONNX Runtime Web in a Web Worker.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Model.&lt;/strong&gt; SigLIP base (&lt;code&gt;google/siglip-base-patch16-224&lt;/code&gt;, Apache-2.0), 8-bit ONNX weights. A hunt target is just a sentence, so there is no label set to train: the photo and the text "a photo of a mushroom" are embedded into the same space and compared.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Split encoders.&lt;/strong&gt; The full model is about 200 MB of image tower plus text tower. I run them as separate graphs and precompute text embeddings for the built-in hunts at build time (&lt;code&gt;npm run embed&lt;/code&gt;, 35 vectors, 144 KB). Players only download the ~100 MB image tower. The text tower is fetched lazily if someone imports a hunt with new labels. The &lt;code&gt;scale&lt;/code&gt; and &lt;code&gt;bias&lt;/code&gt; that SigLIP applies after the towers are read from the upstream checkpoint, since a split export drops them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Deciding "found".&lt;/strong&gt; My first rule, an absolute probability threshold, rejected every correct photo: SigLIP's per-pair probabilities for ordinary photos are tiny (a clear dog photo scored 3e-4). The working rule has three parts. The target must be the best of all candidate prompts. It must hold at least half the softmax share against its look-alikes ("a cat" for "a dog") and a fixed list of cheat scenes (a screen, a room, a hand). And it needs a very small absolute probability floor, because without it several wrong pairs reached a high share by default when nothing matched.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Measured, honestly.&lt;/strong&gt; On 19 Wikipedia lead images (15 true, 4 unrelated), 13/15 true photos counted as finds, the other 2 came back "close, try again", and 0 of 175 wrong photo/target pairs were accepted. The thresholds were tuned on those same images and they are encyclopedic rather than phone photos, so treat the result as a smoke test. Image embedding took 143 to 240 ms in Node on an M1 across two runs. I have not tested outdoors or on real phones.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Engineering decisions.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Model behind an &lt;code&gt;InferenceEngine&lt;/code&gt; interface; UI and core never import the ML library. A deterministic hash-embedding engine makes CI fast, and a build-time &lt;code&gt;e2e&lt;/code&gt; mode swaps it in for browser tests.&lt;/li&gt;
&lt;li&gt;Untrusted input (imported hunt files, localStorage) is validated; corrupt state is dropped, not trusted.&lt;/li&gt;
&lt;li&gt;ONNX runtime files are served from our own origin. The library defaults to a CDN, which would have been a hidden online dependency. A test records that no CDN request happens.&lt;/li&gt;
&lt;li&gt;The native camera opens through a file input instead of &lt;code&gt;getUserMedia&lt;/code&gt;: simpler and more accessible.&lt;/li&gt;
&lt;li&gt;WebGPU failed in headless Chromium ("Failed to get GPU adapter") and killed model loading, so WASM is the default and WebGPU is opt-in.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What did not work first time.&lt;/strong&gt; The absolute threshold above; WebGPU auto-detection; and my first offline test, which only passed once I stopped waiting for a "model ready" label that the hunt screen deliberately hides.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tests.&lt;/strong&gt; 58 unit and component tests, 2 mock-engine browser tests in CI, and one real-model browser test that downloads the model, goes offline, reloads, and verifies a real photo (passed once, headless Chromium, first verification about 3.1 s on WASM).&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Does Open Innovation Matter?
&lt;/h2&gt;

&lt;p&gt;What the open approach actually bought, as implemented:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Privacy by construction.&lt;/strong&gt; The photos never leave the device because there is nowhere to send them. A closed vision API would have received every photo.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Offline.&lt;/strong&gt; Once the model is cached, a hunt works in a field with no signal. That is verified by the offline test, in Chromium only.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Open vocabulary without a server.&lt;/strong&gt; Anyone can write a hunt in JSON and the model scores it. No retraining, no per-request fee.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Control over the pipeline.&lt;/strong&gt; Splitting the encoders to save a 112 MB download, reading the calibration constants, precomputing embeddings and tagging them by model id are only possible with the weights in hand.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Replaceable.&lt;/strong&gt; A different CLIP-style model needs one interface implementation and a re-run of the embed script.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What it did not buy: better accuracy than a large hosted model (it almost certainly isn't), a small first download (about 100 MB), or fine-tuning, which I did not do.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Agent Session
&lt;/h2&gt;

&lt;p&gt;No DevRelay session was saved for this project, so I'm not including one.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>hf26challenge</category>
      <category>ai</category>
      <category>webdev</category>
    </item>
  </channel>
</rss>
