<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Mohammed Aftab Ahamed</title>
    <description>The latest articles on DEV Community by Mohammed Aftab Ahamed (@mohammed_aftabahamed_2).</description>
    <link>https://dev.to/mohammed_aftabahamed_2</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4172226%2F86175137-0e87-44ed-8909-2f9e87991a3c.png</url>
      <title>DEV Community: Mohammed Aftab Ahamed</title>
      <link>https://dev.to/mohammed_aftabahamed_2</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/mohammed_aftabahamed_2"/>
    <language>en</language>
    <item>
      <title>SideQuest AI: a game that shows you less screen so you can see more</title>
      <dc:creator>Mohammed Aftab Ahamed</dc:creator>
      <pubDate>Sat, 10 Oct 2026 22:25:53 +0000</pubDate>
      <link>https://dev.to/mohammed_aftabahamed_2/sidequest-ai-a-game-that-shows-you-less-screen-so-you-can-see-more-ak8</link>
      <guid>https://dev.to/mohammed_aftabahamed_2/sidequest-ai-a-game-that-shows-you-less-screen-so-you-can-see-more-ak8</guid>
      <description>&lt;h2&gt;
  
  
  What I built
&lt;/h2&gt;

&lt;p&gt;SideQuest AI is a small outdoor exploration game with a screen-time budget built into its architecture. You tell it how long you have and what you like. It gives you one mission. You put the phone in your pocket and go outside. When you come back you take one photo, and it tells you — honestly — whether you actually did the thing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why an app about getting off your phone needs almost no phone
&lt;/h2&gt;

&lt;p&gt;The obvious way to build this is wrong. You add a streak counter, a daily challenge feed, a social leaderboard, achievement notifications. Now you have built another thing people use indoors.&lt;/p&gt;

&lt;p&gt;So I treated the screen budget as a hard constraint rather than a feature:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Moment&lt;/th&gt;
&lt;th&gt;Screen interaction&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Choosing a mission&lt;/td&gt;
&lt;td&gt;~30 seconds of form input&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Doing the mission&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Zero&lt;/strong&gt; — the app is closed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Coming back&lt;/td&gt;
&lt;td&gt;One photo, one submit&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;There is no feed, no chat, no streak reminder, no notification designed to pull you back in. The result screen shows your XP and gets out of the way.&lt;/p&gt;

&lt;h2&gt;
  
  
  The interesting problem wasn't the model — it was the lying
&lt;/h2&gt;

&lt;p&gt;The AI part of this app is a vision model looking at one photo and deciding whether it matches the mission. That is exactly the task where a small model fails in the most embarrassing way: ask a 4B model "did this photo show a flower?" and it will cheerfully say yes about a photograph of a car park.&lt;/p&gt;

&lt;p&gt;Most apps in this space paper over that with a green tick. This one is built so the failure is impossible to hide:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Verdict&lt;/th&gt;
&lt;th&gt;Meaning&lt;/th&gt;
&lt;th&gt;XP?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;verified&lt;/code&gt; / &lt;code&gt;likely&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;A model ran, looked, and judged it&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;uncertain&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Photo unclear, or no model reachable&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;No&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;heuristic&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;No model ran at all&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;No&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;code&gt;is_ai_verified&lt;/code&gt; is set in exactly one function in the codebase, and only when a model actually ran, didn't ask for a better photo, and returned a positive verdict. No button, no heuristic, and no fallback can set it. If the model returns a verdict label the app doesn't recognise, it defaults to &lt;code&gt;uncertain&lt;/code&gt; — never to "verified".&lt;/p&gt;

&lt;h2&gt;
  
  
  Gemma, and the deterministic fallback
&lt;/h2&gt;

&lt;p&gt;The model is &lt;a href="https://ai.google.dev/gemma" rel="noopener noreferrer"&gt;Gemma 3&lt;/a&gt; via HuggingFace transformers, running locally. Every AI call goes through one interface (&lt;code&gt;InferenceEngine&lt;/code&gt;), and there are two implementations behind it: &lt;code&gt;GemmaEngine&lt;/code&gt; (real weights) and &lt;code&gt;UnavailableEngine&lt;/code&gt; (no weights, and honest about it).&lt;/p&gt;

&lt;p&gt;Because the fallback sits &lt;em&gt;behind the same interface&lt;/em&gt;, the code path a user gets with no model installed is the exact path the tests exercise. "It degrades gracefully" isn't a hope here — it's measured, in four failure-injection scenarios that break the engine in specific ways and check the app still answers without claiming AI verification. All four survive.&lt;/p&gt;

&lt;h2&gt;
  
  
  The safety gate the model can't argue with
&lt;/h2&gt;

&lt;p&gt;Model output is treated as untrusted input. Every mission passes a deterministic gate before it's returned: 6 blocking rules (open water, height, traffic, trespass, storms, extreme heat) and 6 warnings (heat load, darkness, being alone, cold water, wildlife, litter). A blocked mission is thrown away and regenerated deterministically. &lt;code&gt;plan_with_retry&lt;/code&gt; retries on &lt;em&gt;parse&lt;/em&gt; failures but never on safety blocks — retrying an unsafe generation is precisely how an unsafe mission eventually gets through.&lt;/p&gt;

&lt;h2&gt;
  
  
  Open innovation: why local and open-weight matters here
&lt;/h2&gt;

&lt;p&gt;"The app looks at your photo" has usually meant a multibillion-dollar API call — an outbound request carrying a user's image to someone else's datacenter, metered per call. That is structurally hostile to an app whose thesis is &lt;em&gt;less&lt;/em&gt; screen. A 4B open-weight model changes the trade: the weights are small enough to run on hardware people already own, which makes a different category of app possible — the model is a local component rather than a dependency, the photo doesn't have to go anywhere, and the interesting engineering becomes the &lt;em&gt;policy&lt;/em&gt; around the model rather than surviving an API.&lt;/p&gt;

&lt;h2&gt;
  
  
  Privacy
&lt;/h2&gt;

&lt;p&gt;Progress is a local JSON file (&lt;code&gt;store.py&lt;/code&gt;, atomic writes). Nothing leaves the device except a photo you explicitly submit. There's no GPS anywhere — &lt;code&gt;location_label&lt;/code&gt; is a free-text label you may choose to fill in, never coordinates. No analytics, no third-party scripts, no fonts from a CDN (&lt;code&gt;sw.js&lt;/code&gt; never caches &lt;code&gt;/api/*&lt;/code&gt;).&lt;/p&gt;

&lt;h2&gt;
  
  
  Credits
&lt;/h2&gt;

&lt;p&gt;Built with open weights and open libraries: &lt;a href="https://huggingface.co/google/gemma-3-4b-it" rel="noopener noreferrer"&gt;google/gemma-3-4b-it&lt;/a&gt; (gated HF repo, evaluated here); &lt;a href="https://github.com/huggingface/transformers" rel="noopener noreferrer"&gt;transformers&lt;/a&gt; (local inference); &lt;a href="https://python-pillow.org/" rel="noopener noreferrer"&gt;Pillow&lt;/a&gt; / &lt;a href="https://pillow.readthedocs.io/" rel="noopener noreferrer"&gt;PIL&lt;/a&gt; (synthetic verification images); &lt;a href="https://vitejs.dev" rel="noopener noreferrer"&gt;vite&lt;/a&gt; / vanilla JS (front-end); standard Python (&lt;code&gt;json&lt;/code&gt;, &lt;code&gt;hashlib&lt;/code&gt;, &lt;code&gt;re&lt;/code&gt;, &lt;code&gt;os&lt;/code&gt;, &lt;code&gt;pathlib&lt;/code&gt;).&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo and repository
&lt;/h2&gt;

&lt;p&gt;Repository (verified, 3 commits): &lt;a href="https://github.com/MoAftab-01/SideQuest-AI" rel="noopener noreferrer"&gt;https://github.com/MoAftab-01/SideQuest-AI&lt;/a&gt;&lt;br&gt;
Live demo URL: &lt;strong&gt;PENDING&lt;/strong&gt; — Render free-tier deploy (&lt;code&gt;render.yaml&lt;/code&gt;, CPU-only, deterministic mode, no secrets in source) was interrupted before authorization completed. The config is committed; the instance is not live. This gap is stated plainly rather than hidden.&lt;/p&gt;

&lt;h2&gt;
  
  
  Measured results (truthful, 2026-10-11)
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Value (verified)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Planning engine used&lt;/td&gt;
&lt;td&gt;deterministic (100%) — model unavailable 401&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Safety gate pass rate&lt;/td&gt;
&lt;td&gt;100.0% (8/8)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Safety gate block count&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Robustness: 4 injection scenarios&lt;/td&gt;
&lt;td&gt;all survived (&lt;code&gt;survived: true&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;False &lt;code&gt;AI verified&lt;/code&gt; results&lt;/td&gt;
&lt;td&gt;0 (pass criterion = 0; not tested with real images yet)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vision verification accuracy&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;not measured&lt;/strong&gt; (synthetic 8-image smoke test only; no labelled &lt;code&gt;--cases&lt;/code&gt; set)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Full offline with loaded weights&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;not claimed&lt;/strong&gt; — deterministic offline verified; model-first-load needs network, unverified&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tests passing&lt;/td&gt;
&lt;td&gt;853/853 (7 test files)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Template coverage&lt;/td&gt;
&lt;td&gt;30 templates, zero contradictions across 12,096 constraint-combination missions&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;What is measured is saved in the repo (&lt;code&gt;eval/results/eval_results.json&lt;/code&gt;, &lt;code&gt;tests/&lt;/code&gt; regression suite covering the two &lt;code&gt;BUG_*&lt;/code&gt; fixes: library-poisoning + gate self-block). What is NOT measured is stated plainly rather than hidden. No fabricated numbers; no "verified" label applied where only a button click or hardcoded rule could set it; no claim of full offline operation until inference has been tested without network access after model setup.&lt;/p&gt;

&lt;p&gt;This submission is part of &lt;a href="https://dev.to/challenges/hacktoberfest-week1-2026-10-05"&gt;Hacktoberfest 2026 — Week 1 "Touch Grass"&lt;/a&gt; (#devchallenge #hf26challenge).&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>hf26challenge</category>
    </item>
    <item>
      <title>Hello</title>
      <dc:creator>Mohammed Aftab Ahamed</dc:creator>
      <pubDate>Thu, 08 Oct 2026 23:21:45 +0000</pubDate>
      <link>https://dev.to/mohammed_aftabahamed_2/hello-2g58</link>
      <guid>https://dev.to/mohammed_aftabahamed_2/hello-2g58</guid>
      <description></description>
    </item>
  </channel>
</rss>
