DEV Community

Mohammed Aftab Ahamed
Mohammed Aftab Ahamed

Posted on

SideQuest AI: a game that shows you less screen so you can see more

Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass Submission 🌿

What I built

SideQuest AI is a small outdoor exploration game with a screen-time budget built into its architecture. You tell it how long you have and what you like. It gives you one mission. You put the phone in your pocket and go outside. When you come back you take one photo, and it tells you — honestly — whether you actually did the thing.

Why an app about getting off your phone needs almost no phone

The obvious way to build this is wrong. You add a streak counter, a daily challenge feed, a social leaderboard, achievement notifications. Now you have built another thing people use indoors.

So I treated the screen budget as a hard constraint rather than a feature:

Moment Screen interaction
Choosing a mission ~30 seconds of form input
Doing the mission Zero — the app is closed
Coming back One photo, one submit

There is no feed, no chat, no streak reminder, no notification designed to pull you back in. The result screen shows your XP and gets out of the way.

The interesting problem wasn't the model — it was the lying

The AI part of this app is a vision model looking at one photo and deciding whether it matches the mission. That is exactly the task where a small model fails in the most embarrassing way: ask a 4B model "did this photo show a flower?" and it will cheerfully say yes about a photograph of a car park.

Most apps in this space paper over that with a green tick. This one is built so the failure is impossible to hide:

Verdict Meaning XP?
verified / likely A model ran, looked, and judged it Yes
uncertain Photo unclear, or no model reachable No
heuristic No model ran at all No

is_ai_verified is set in exactly one function in the codebase, and only when a model actually ran, didn't ask for a better photo, and returned a positive verdict. No button, no heuristic, and no fallback can set it. If the model returns a verdict label the app doesn't recognise, it defaults to uncertain — never to "verified".

Gemma, and the deterministic fallback

The model is Gemma 3 via HuggingFace transformers, running locally. Every AI call goes through one interface (InferenceEngine), and there are two implementations behind it: GemmaEngine (real weights) and UnavailableEngine (no weights, and honest about it).

Because the fallback sits behind the same interface, the code path a user gets with no model installed is the exact path the tests exercise. "It degrades gracefully" isn't a hope here — it's measured, in four failure-injection scenarios that break the engine in specific ways and check the app still answers without claiming AI verification. All four survive.

The safety gate the model can't argue with

Model output is treated as untrusted input. Every mission passes a deterministic gate before it's returned: 6 blocking rules (open water, height, traffic, trespass, storms, extreme heat) and 6 warnings (heat load, darkness, being alone, cold water, wildlife, litter). A blocked mission is thrown away and regenerated deterministically. plan_with_retry retries on parse failures but never on safety blocks — retrying an unsafe generation is precisely how an unsafe mission eventually gets through.

Open innovation: why local and open-weight matters here

"The app looks at your photo" has usually meant a multibillion-dollar API call — an outbound request carrying a user's image to someone else's datacenter, metered per call. That is structurally hostile to an app whose thesis is less screen. A 4B open-weight model changes the trade: the weights are small enough to run on hardware people already own, which makes a different category of app possible — the model is a local component rather than a dependency, the photo doesn't have to go anywhere, and the interesting engineering becomes the policy around the model rather than surviving an API.

Privacy

Progress is a local JSON file (store.py, atomic writes). Nothing leaves the device except a photo you explicitly submit. There's no GPS anywhere — location_label is a free-text label you may choose to fill in, never coordinates. No analytics, no third-party scripts, no fonts from a CDN (sw.js never caches /api/*).

Credits

Built with open weights and open libraries: google/gemma-3-4b-it (gated HF repo, evaluated here); transformers (local inference); Pillow / PIL (synthetic verification images); vite / vanilla JS (front-end); standard Python (json, hashlib, re, os, pathlib).

Demo and repository

Repository (verified, 3 commits): https://github.com/MoAftab-01/SideQuest-AI
Live demo URL: PENDING — Render free-tier deploy (render.yaml, CPU-only, deterministic mode, no secrets in source) was interrupted before authorization completed. The config is committed; the instance is not live. This gap is stated plainly rather than hidden.

Measured results (truthful, 2026-10-11)

Metric Value (verified)
Planning engine used deterministic (100%) — model unavailable 401
Safety gate pass rate 100.0% (8/8)
Safety gate block count 0
Robustness: 4 injection scenarios all survived (survived: true)
False AI verified results 0 (pass criterion = 0; not tested with real images yet)
Vision verification accuracy not measured (synthetic 8-image smoke test only; no labelled --cases set)
Full offline with loaded weights not claimed — deterministic offline verified; model-first-load needs network, unverified
Tests passing 853/853 (7 test files)
Template coverage 30 templates, zero contradictions across 12,096 constraint-combination missions

What is measured is saved in the repo (eval/results/eval_results.json, tests/ regression suite covering the two BUG_* fixes: library-poisoning + gate self-block). What is NOT measured is stated plainly rather than hidden. No fabricated numbers; no "verified" label applied where only a button click or hardcoded rule could set it; no claim of full offline operation until inference has been tested without network access after model setup.

This submission is part of Hacktoberfest 2026 — Week 1 "Touch Grass" (#devchallenge #hf26challenge).

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •
You need to verify your account.
Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to