DEV Community

Harsh Shah
Harsh Shah

Posted on AI-assisted

I Built an AI That Wants You to Close the App

Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass Submission 🌿

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass

What I Built

I built Offscreen β€” an AI-powered outdoor discovery journal.

The idea is simple: instead of giving people another reason to stay on their phones, the app gives them a reason to put their phones down.

Every session starts with a small outdoor challenge:

Capture the shadow of a person walking past you.

Or:

Spot a piece of faded, hand-painted advertising that has almost vanished into a building's walls.

The challenge might take five minutes or twenty-five.

Then comes the important part:

Put the phone down. Go outside.

When you come back, you bring a photograph of what you discovered. Gemma looks at the photograph, interprets what you found, and scores the discovery.

The app deliberately has no timer, no feed, and nothing to scroll. The shortest part of the experience is supposed to be the time spent in the app.

The goal isn't to make people spend more time with an AI.

It's to give them a reason to spend less time looking at a screen.

Demo

Live demo: https://offscreen-one.vercel.app/

Judge instructions: create an account with any email address and a password of at least 8 characters. There is no email verification step, so you can be signed in within seconds.

Demo video (63s, the real app end to end):

The complete experience is:

Get a challenge β†’ Put the phone down β†’ Go outside β†’ Take a photo β†’ Come back β†’ Let Gemma evaluate your discovery.

Code

GitHub repository:

Offscreen

A photo walk, one challenge at a time.

Outside, Not Online β€” an AI-powered outdoor discovery journal built for the Hacktoberfest 2026 Open-Source AI Challenge: "Touch Grass".

You get one small, interesting real-world challenge. You leave the app, find something, photograph it, and come back. Gemma looks at the photo, judges whether it satisfies the challenge, writes a short reflection, and awards discovery points.

The design goal is that you leave the app. There is no feed, no streak pressure no notifications, and no infinite scroll.

Open app  β†’  Today's challenge  β†’  Close the app, go outside  β†’  Photograph
          β†’  Gemma looks  β†’  Score + reflection  β†’  Journal  β†’  Come back tomorrow

Watch it

The actual demo β€” the real app, real Gemma, real photograph, no cuts or narration:

The 21-second launch cut β€” a recreation of the product's screens, rendered from HTML:

The demo is the honest…

The project is an AI-powered web application with the AI workflow separated from the user interface, orchestrated as two explicit state graphs rather than a single opaque agent loop.

How I Built It

The core of the project is Gemma, Google's open-weight model.

Gemma writes every challenge, looks at every photograph, and decides every score. Nothing in the app works without it.

The application uses LangGraph to orchestrate the workflow. There are two graphs, and both are deliberately boring:

challenge:  START β†’ load_context β†’ generate_challenge β†’ END

discovery:  START β†’ load_challenge β†’ analyze_photo β†’ evaluate_discovery
                 β†’ generate_feedback β†’ save_discovery β†’ update_profile β†’ END
Enter fullscreen mode Exit fullscreen mode

There is no autonomous loop and no self-reflection. Each node runs once. The whole system has exactly one conditional edge, and it is an error guard: if the challenge can't be found, the graph jumps straight to END, because there is no point paying for a vision call to judge a photograph against nothing.

Three decisions were made against the obvious implementation.

1. One multimodal call, not four. The textbook pipeline is identify β†’ describe β†’ evaluate β†’ feedback. That's four model calls, and it lets the model contradict its own description. Instead, a single request returns the description and the evaluation together. The two-stage path is still in the codebase behind a flag so the two can be compared.

2. Gemma ignores the function schema. My first live call used the provider's structured-output helper. Gemma ignored the JSON schema I declared and returned its own key names β€” asked for title, it gave me challenge_title. So the schema is now written into the prompt as an annotated JSON skeleton, the reply is parsed and validated locally with Pydantic, and a validation failure buys the model exactly one chance to fix itself. Nothing the model returns is trusted.

3. Personalization you can measure. "Favour what the user likes" left to the model is a suggestion it may ignore, and you cannot check it afterwards. So the category is chosen before generation β€” about 60% weighted toward the user's favourites and 40% uniform exploration β€” and handed to the model as a constraint it cannot override. A test draws 2,000 seeds and asserts the ratio. A score isn't a vibe here; it's a number somebody checked.

The backend is FastAPI, Pydantic, and SQLite through SQLAlchemy; the frontend is Next.js. There are 176 tests, all hermetic β€” no test makes a network call, none touches the real database, and the model is stubbed everywhere. The suite runs in about eighteen seconds and costs nothing, which meant I could run it after every single change.

Why Does Open Innovation Matter?

The honest version of this answer starts with a limitation: this MVP runs against Google's hosted AI Studio API, so photographs do leave the machine. I'm not claiming local inference or privacy properties I haven't demonstrated.

What open AI actually bought me was three concrete things.

1. It made the project possible in the time I had. A multimodal model I could call with an API key, in an afternoon, for free at this scale, is what turned "an idea about putting the phone down" into something I could iterate on. The interesting work went into prompt design, evaluation, and agent structure rather than into access. For a project whose entire premise is an AI loop, that's the difference between shipping and not.

2. Gemma's actual behaviour shaped the architecture. Asked for a JSON schema through the provider's function calling, it returned its own field names. A closed API would have given me that same result β€” but with an open model I could inspect the failure, decide to stop trusting the tooling, and build the fix into my own prompt and validation layer. Every subsequent discovery followed the same pattern: see what the model does, then design around it. I now know it hedges with "this appears to be" when it's unsure, which is exactly the behaviour the prompts demand, and I know it doesn't reliably β€” both of which I learned by testing, not by reading a model card.

3. The model is a configuration value, not an architecture. The model id lives in an environment variable. That made one optimisation obvious: challenge writing is mostly text generation and barely uses vision, so splitting a small model for challenges and the larger one only for judging became a straightforward next step rather than a rewrite. It's only obvious because swapping models isn't a rewrite.

LangGraph matters for the same reason β€” the reason this agent is explainable is that its control flow is an open graph I can read. When the graph made a mistake, I found it because the whole path was six lines long and visible. That determinism is a property of choosing an open harness, not something I wrote.

What Actually Happened When I Used It

This is the part I found most interesting, and most of it is wrong.

The judge is genuinely strict. I submitted a photograph of a walking shadow against a challenge asking for a dated plaque on a building. It came back 0/10, completed: false, with: "This is a great capture of city life and textures, but I couldn't spot a plaque or date in this frame!" No score inflation, no encouragement to try again β€” it just told me the truth. That mattered more to me than a high score would have, because the obvious failure mode for an AI judge is being agreeable.

It invented an object. On the same photograph it correctly described the low angle, the mid-stride legs, and the warm afternoon light β€” and then titled the discovery "Streetlight Silhouettes." There is no streetlight in that frame. My prompts explicitly forbid inventing things that aren't visible, and it did anyway. This is a real hallucination that shipped, and it's the clearest example of why an AI evaluator needs a human who actually goes outside and checks.

Its confidence is meaningless. Both submissions returned confidence: 1.0 β€” for the correct verdict and for the wrong one. A field that's always 1.0 carries no information at all. It's either a prompt problem or a calibration problem, and right now I can't tell which. The fix is a small labelled set of my own photographs with scores I'd defend, and measuring the model against them.

Real usage found bugs that 169 tests missed. The profile showed a streak of 0 and 0 completed challenges forever. Two separate bugs, in the same function: SQLite returns naive datetimes even from timezone-aware columns, so a timezone comparison always failed β€” meaning the streak would have worked perfectly on PostgreSQL, and the field I'd have migrated to later. And a counter was declared and returned by the API but never incremented anywhere in the codebase. Both only appeared once a real photograph went through the system. The tests had been asserting on in-memory objects; the new ones round-trip through the database, which is exactly the gap that let them through.

The latency is real. Generating a challenge takes about 40 seconds on the free tier. That's why the app pre-generates tomorrow's challenge while you're reading today's one β€” the wait happens once, quietly, instead of in front of you every morning.

My Agent Session

I built Offscreen from the initial idea through implementation and submission planning with an agentic coding workflow β€” twenty-two commits, phase by phase, from a written specification.

You can see the development process here:

Building Offscreen: an outdoor discovery journal from spec to submission
You

hi, read @spec.md and then start bilding the app, phase by phase, ensuring proper commits for each feature built

Agent

I'll start by reading the spec.

Agent

No git repo yet. Starting Phase 0 (repo scaffolding).

Agent

Now the backend skeleton.

Agent

Now the frontend.

Agent

No API key in the environment β€” I'll build Phase 1 fully with an injectable client + tests, then need a key from you for the live check.

Agent

Now the prompt modules and the test endpoint.

Agent

Now schemas + the /api/test/gemma endpoint.

Agent

Now tests for the Gemma service with a fake client.

You

git's author must be me

Agent

Author fixed β€” the Phase 0 commit is now yours (Harsh-0986 <[REDACTED]>), and it's set in local repo config so all future commits are too. (I won't pass -c user.name=... again.)

Now back to the 3 failing Phase 1 tests.

Agent

Now let me verify the live endpoint behaves correctly with no key, then commit Phase 1.

Agent

Phase 0 and 1 are committed. One blocker before I call Phase 1 done:

I need a GEMINI_API_KEY (AI Studio). There's none in the environment and no .env yet.

aistudio.google.com/apikey β†’ Create API key

Put it in .env at the repo root (git-ignored):

GEMINI_API_KEY=your_key_here
GEMMA_MODEL=gemma-4-26b-a4b-it

Then I'll run the acceptance check:

cd backend && .venv/bin/uvicorn app.main:app --reload
curl -X POST localhost:8000/api/test/gemma
# β†’ {"success":true, "structured_ok":true, ...}

Status:

Phase State
0 β€” Repo setup βœ… 821ffad β€” both apps run, /health ok, landing page serves
1 β€” Gemma connection ⚠️ code + 17 tests pass; live call needs your key
2–10 not started

Two notes on Phase 1:
- The model id gemma-4-26b-a4b-it from the spec is a placeholder β€” the live call will tell us if that exact id resolves. If it 404s I'll swap it via GEMMA_MODEL only, as the spec instructs.
- I added GEMMA_RETRY_BASE_DELAY as an env-tunable (default 1.5s) so tests don't actually sleep.

Two things worth your call before I continue:
1. Do you want to paste the key now, or should I press on with Phase 2 (challenge generator β€” fully unit-testable offline) and verify Phase 1 when you have the key?
2. GEMMA_MODEL β€” do you have a specific Gemma 4 model id from AI Studio you want pinned, or should I take whatever the key's project exposes?

You

added key, test

You

don't read env, just start

Agent

Fair β€” I'll make the config pick up backend/.env on its own instead of me reading it.

Agent

Gemma ignores the function-calling schema and invents its own keys. I'll switch the default to instructed-JSON with explicit schema rendering + one repair retry.

Agent

The repair loop caught a bad category and fixed it. Now the live endpoint:

Agent

Let me also remove the duplicated repair loop in the two call paths.

Agent

Now re-verify live (the enum now reaches the model) and commit.

Agent

Phase 1 acceptance met β€” live Gemma call works.

{"success":true,"model":"gemma-4-26b-a4b-it","structured_ok":true,
 "sample":"Take a deep breath and enjoy the fresh air."}

Three real problems surfaced that only a live call could have caught:

  1. .env location β€” your key was in backend/.env, config only read the repo root. Now reads both.
  2. CORS_ORIGINS crashed the app β€” pydantic-settings tries to JSON-decode dotenv values, so http://localhost:3000 was a hard SettingsError. Now parsed as a comma string.
  3. Gemma ignores the function schema β€” this was the big one. with_structured_output returned {"challenge_title": "Nature...", ...}, not my schema. Fixed by making the default path instructed-JSON (schema rendered into the prompt as an annotated skeleton β†’ parsed β†’ Pydantic-validated) plus one self-repair attempt. The repair loop earned its keep on the first live run: the model returned category: "Exploration", failed validation, and corrected itself to a valid category. Old behaviour available via GEMMA_STRUCTURED_OUTPUT_MODE=tooling.

gemma-4-26b-a4b-it resolves fine β€” no model swap needed.

Committed as 03a07b0. 22 tests pass, ruff clean.

Next: Phase 2 β€” Challenge Generator. That means the LangGraph state graph (START β†’ LoadContext β†’ GenerateChallenge β†’ END), build_challenge_graph(), POST /api/challenges/generate returning a validated challenge, and a dependency-injected model so the tests stay offline.

Two decisions I'd like your input on before I build it, since both shape the DB work in Phase 5:

  1. Single-user MVP or multi-user? The spec's schema has user_id on challenges and discoveries, but there's no auth in the MVP (spec Β§3 lists it as a non-goal). I'd default to a single implicit user (created on first request) so /journal and /profile work today without auth, while keeping user_id columns so multi-user is a later migration. Agreed?
  2. Does "today's challenge" expire? /api/challenges/today implies one challenge per day. I'd generate on first request and reuse until the next UTC day, so a refresh doesn't burn a new challenge. Yes?
You

multi user, yes

Agent

Multi-user without auth it is β€” anonymous per-browser identity via a UUID header, real user_id on everything, auth deferred.

Agent

Now the API schemas and routes for challenges.

Agent

Now tests for the graph and endpoint (offline, fake service).

Agent

Adding the global error handler so no stack trace ever reaches the client (SPEC Β§26):

Agent

Now verify Phase 2 live against the real model:

Agent

Live works, but 35.9s against a <10s target (SPEC Β§43). Let me find out why.

You

bud leave it, it can consume time since it's a free api

Agent

Fair enough β€” dropping the latency optimization. Committing Phase 2.

Agent

Phase 2 committed (89b5627) β€” 36 tests, verified live.

Two things I noticed in the live run that you should know about, because they'll bite in Phase 5:

  1. The duplicate guard can't work yet. User A's 2nd challenge was "a single leaf with a hole", the 3rd from user B was "a single leaf turned yellow" β€” all three leaves. The guard only compares against persisted history, and there is none until Phase 5. So the spec's "two consecutive challenges must not be identical" acceptance criterion is not actually met yet β€” it's coded and unit-tested against injected history, but not true end-to-end. I'll re-verify it after the DB lands.

  2. 35.9s per challenge, dropped at your instruction. Worth remembering that this eats your demo's 2-minute script budget if the frontend chains calls.

Next up: Phase 3 β€” Photo Analysis (upload + validation + multimodal Gemma, no scoring yet). That means the image upload endpoint, MIME/size validation with sanitized UUID filenames, and image resizing before it goes to the model.

Two questions on that:

  1. /submit needs a challenge_id, but nothing is persisted yet. Options: (a) generate an ephemeral in-memory challenge handle now and make it a real FK in Phase 5, or (b) pull the DB forward into Phase 3 so photo submission has something real to reference. The spec puts DB in Phase 5, but (a) means building a throwaway id layer. I lean (b) β€” pull the DB forward, since every remaining phase touches it. OK?
  2. Test photos β€” do you have a few real outdoor photos on this machine I can use as fixtures for the AI test cases in SPEC Β§41 (obvious success, obvious failure, ambiguous, unrelated, poor quality)? That matters a lot for Phase 4 accuracy, and SPEC Β§37/Β§38 require real photos for the demo anyway.
You

pull db, and leave the testing, i'll do it at the end of the app

Agent

Pulling the DB forward. I'll generate synthetic fixture images programmatically for plumbing tests β€” real photo testing is yours at the end.

You

no need to create images, that we'll see later

Agent

Understood β€” no fixtures. Continuing the DB layer.

Agent

Now wire the graph to the DB and make /generate and /today persist.

Agent

Now repository tests for the DB layer itself:

Agent

Found it β€” a real coupling bug: preferences only update if a user row exists. Preferences should always record.

Agent

Now live-verify persistence + the challenge-reuse behaviour, including the dedup criterion that failed before:

You

hey, let's move testing to the end the last phase

Agent

Understood β€” no more live calls from me until the final phase. Committing the DB work on unit tests alone.

Agent

Now Phase 3 β€” Photo Analysis: upload, validation, and multimodal Gemma. No scoring yet, no live calls.

Agent

Now the photo schemas, analysis service, and upload endpoint:

Agent

Now tests for image validation and the upload endpoint β€” using in-memory images, no fixture files on disk:

Agent

Found a real trap: SQLite in-memory uses a per-thread connection, so the TestClient's thread got a different empty database.

Agent

Phase 3 committed (c385caf) β€” 85 tests, no live calls made.

Notable finds along the way:
- StaticPool bug in tests β€” in-memory SQLite gives each thread its own connection, so TestClient's thread was hitting a different, empty database than the test session. Silent and would have made ownership tests lie.
- Ownership is enforced at the photo layer too β€” passing another user's challenge_id yields "no challenge context" rather than leaking their prompt.

Next: Phase 4 β€” Challenge Evaluation. This is the heart of the app: the discovery graph

START β†’ load_challenge β†’ analyze_photo β†’ evaluate_discovery β†’ generate_feedback β†’ save_discovery β†’ update_profile β†’ END

All model calls stay injectable and offline-tested. Scoring, the self-repair loop, and the "appears to be" hedging rule (SPEC Β§14) get exercised by fakes now; real-photo accuracy waits for your final phase.

Two decisions I want your call on first, because they're visible in the product:

  1. One multimodal call or two? SPEC Β§43 says prefer one call, but Β§12–13 describe analysis and evaluation as separate prompts. Combining them is faster and avoids the model contradicting its own description, at the cost of a less interesting prompt. I'd do one call that returns description + evaluation together, and keep the two-stage path behind a flag for comparison in your real testing. OK?

  2. Should a failed challenge be saved to the journal? SPEC Β§15 says "every completed challenge creates a journal entry", which reads as successes only β€” but a score of 2/10 with kind feedback is motivating, and hiding misses feels dishonest. I'd save everything, and let the journal show completed: false with a muted treatment. Yes?

You

prefer 1, ok

Agent

Both confirmed. Building Phase 4 with the single-call pipeline.

Agent

Now the discovery graph nodes:

Agent

Now the discovery graph:

Agent

image_bytes in the state violates SPEC Β§9's "state should be serializable". Moving it into the graph closure instead:

Agent

Now the discovery/journal/profile endpoints:

Agent

Now the offline test suite for the evaluation graph and endpoints:

Agent

Three real bugs found. Fixing: duplicate field placement, and the graph kept calling the model after a missing challenge, and a wrong error type on failure.

Agent

Phase 4 committed (efb28fa) β€” 106 tests, still no live calls.

The most important bug this phase caught: the graph kept calling the model after load_challenge failed. A user submitting against a nonexistent challenge would have burned a paid vision call and asked Gemma to judge a photo against nothing. A conditional edge now short-circuits to END β€” the only branching in the graph, and it's an error guard, not autonomy.

Backend MVP is now feature-complete against SPEC Β§19. Remaining:

Phase Scope State
6 Journal frontend not started
7 Personalization (favour top categories ~60%) plumbing exists, not tuned
8 Frontend polish + animations not started
9–10 Your real outdoor testing, demo data, empty/error states yours

Next I'll build Phase 6 + 7 together, since the journal UI is small and personalization is a small backend change β€” one commit each.

For Phase 7 there's a judgement call I want your input on. The spec says roughly 60% familiar / 40% exploration, but my bump_preference just adds +1 per discovery, which makes the ratio emergent rather than enforced. Two ways to do it:

  1. Prompt-level (what the spec implies): tell the model your favourite categories and let it choose. Zero code, but the model may ignore the ratio β€” and every generation is a coin flip.
  2. Deterministic category selection: pick the category ourselves (60% weighted pick from your top categories, 40% uniform from all 11), then constrain the model to it. The 60/40 split becomes measurable, and generations get cheaper because we skip the repair retry. But challenges may feel samey, since a category can only go so many ways.

I'd go with a hybrid: deterministic category pick, but only ~50% of the time β€” otherwise force a full-choice generation. Best of both, still measurable. Which do you want?

Also, one thing I noticed and should flag: the app currently has no way to change display_name, and no endpoint exists for it. The profile is anonymous. Do you want a tiny PATCH /api/profile for a name, or leave it anonymous for the hackathon?

You

second, yes

Agent

Option 2 it is β€” deterministic 60/40 category selection, plus the profile name endpoint.

Agent

Now wire it into the generate_challenge node:

Agent

Now the PATCH /api/profile name endpoint and personalization tests:

Agent

Now tests that personalization actually reaches the prompt and that consecutive challenges differ (SPEC Β§35):

Agent

The invalid-category path raised a raw ValidationError out of the graph. Making it degrade safely instead:

Agent

Adding a test for the profile-name endpoint, then committing:

Agent

Phase 7 committed (e92916f) β€” 128 tests, still no live calls.

Your deterministic option paid off immediately. Two bugs it exposed:

  • build_challenge_graph silently ignored select_category β€” the constraint was never reaching the prompt at all. Caught only because the tests inject a fake picker.
  • An invalid category override raised a raw ValidationError out of the graph, which would have surfaced as an unhandled 500. Now refused, with the model's own category kept.

The 60/40 split is now measurable, not aspirational β€” a test draws 2000 seeds and asserts familiar lands in 55–65%.

Next: Phase 6 β€” the frontend. This is the first time I'll write real UI, and it's the bulk of what makes the demo land. Before I start, three decisions that shape every screen:

  1. Identity in the browser. The backend needs X-User-Id. I'll generate a UUID on first visit, keep it in localStorage, and attach it in the API client β€” no login screen, matching the "anti-social" tone. Agreed?

  2. Fonts. SPEC Β§21 says large typography, photography, whitespace. I'd pair a heavy grotesque for the big statements (Inter Tight or Archivo) with a quiet serif for reflections β€” serif for the reflection text specifically sells "journal" over "app". Do you have a preference, or shall I pick?

  3. Colour. SPEC Β§21 doesn't state a palette. I want to avoid the generic AI default of purple gradients on dark. My instinct: warm paper white background, near-black text, one mossy-green accent, and the photographs carry all the colour. It's on-brief for "anti-social media" β€” the app recedes, the photo dominates. Yes, or do you want something bolder?

Also worth deciding: /today takes ~35s to generate on a cold cache. That means the /today screen needs a genuinely good loading state ("Finding something interesting..." per SPEC Β§38), and the challenge should probably be generated in advance rather than on first view. I'd add a "tomorrow's challenge" pre-generation in the background β€” cheap to build now, and it removes the wait from the demo. Want that?

You
  1. login, 2. your wish, 3. your wish, bonus yes
Agent

Login it is β€” real accounts. That changes the backend's identity model, so I'll do that first as its own commit.

Agent

Now the repositories and auth routes:

Agent

Now replace the anonymous identity dependency with real auth:

Agent

Now update the routes that used the anonymous header:

Agent

My string replacements got sloppy. Rewriting both route files cleanly:

Agent

Now update the tests to the account model:

Agent

Now the endpoint tests β€” swapping X-User-Id headers for real bearer tokens:

Agent

The regex pass is getting fragile. Let me rewrite the endpoint section of that file directly.

Agent

This file has too many patch scars. Rewriting it cleanly:

Agent

Now auth tests:

Agent

Auth committed (2447240) β€” 162 tests. Anonymous header identity is fully gone; every route requires a bearer token now.

Security choices worth naming:
- PBKDF2-HMAC-SHA256 (600k rounds) from stdlib β€” no native bcrypt dependency to break on deploys.
- Stateless HMAC tokens β€” no session table, so nothing to clean up; rotate SECRET_KEY to log everyone out.
- Login gives one message for unknown-email and wrong-password alike, so the endpoint can't be used to enumerate accounts. There's a test asserting both produce identical text.
- Signed token bodies β€” a user can't edit their own id out of the token.

One migration note: users gained email/password_hash as non-null columns and create_all won't alter an existing table. Since you've never had real data, I deleted the dev DB. Worth knowing before your testing phase β€” if you have an old backend/outside.db, delete it or signups will fail on the missing columns.

Now the frontend (Phase 6 + 8), with your three decisions:
1. Login/signup screen, real accounts.
2. My call on type: heavy grotesque for statements, quiet serif for reflections.
3. Warm paper white, near-black, one mossy accent.
4. Bonus: pre-generate tomorrow's challenge in the background.

I'll build it as: (a) design tokens + shell/nav, (b) auth screens, (c) /today + mission mode, (d) /submit + /result, (e) /journal + /profile, (f) pre-generation + animations.

One thing I should ask rather than assume: should / (landing) stay public, or require login? I'd keep it public as the demo entry point with a START EXPLORING CTA that routes to /login if signed out β€” otherwise your 2-minute demo script opens straight into a login form, which is a bad first five seconds. Agreed?

You

/ is public

Agent

Building the frontend now, starting with the design system, API client, and public landing page.

Agent

Builds clean. Committing the design system + landing, then the auth screens.

Agent

Now the login/signup screen and the app shell:

Agent

useSearchParams needs a Suspense boundary for static prerender:

Agent

Now the core loop: /today with mission mode.

Agent

Now /result with the score animation, then the journal and profile.

Agent

Three legitimate React 19 lint errors. Fixing them properly rather than suppressing:

Agent

That lint warning is pointing at a real bug β€” getSnapshot returned a fresh object each call, which would loop forever. Caching it:

Agent

Lint and build both clean. Committing the full loop:

Agent

App is feature-complete. Final pass on config docs and README before handing over to your testing phase.

Agent

The app is built and committed β€” 14 commits, 162 backend tests, 8 frontend routes building and linting clean.

What's done

Every phase except the ones that need you:

Phase State
0 Repo Β· 1 Gemma Β· 2 Challenges Β· 3 Photo Β· 4 Evaluation Β· 5 DB Β· 6 Journal Β· 7 Personalization Β· 8 Frontend βœ…
9 Real outdoor test · 10 Demo data ⏳ yours

The full loop works: public landing β†’ login β†’ /today challenge β†’ /mission (close the app) β†’ /submit photo β†’ /result scored by Gemma β†’ /journal β†’ /profile. Empty, loading, and error states are in.

Bugs the offline tests caught

Worth knowing, since several would have been invisible in a demo:

  • The graph kept calling Gemma after a failed challenge lookup β€” burning a paid vision call to judge a photo against nothing. Now short-circuits.
  • getSnapshot returned a fresh object each call in the result hand-off; useSyncExternalStore compares with Object.is and would have re-rendered forever.
  • The in-memory test DB gave each thread its own connection, so TestClient's thread saw a different, empty database than the test.
  • Gemma ignores the provider's function schema and invents its own key names β€” the single biggest one, found only by testing live.

Two things to watch

  1. ~/.env lives in backend/.env, and it's git-ignored. If you move it, keep SECRET_KEY set or every session token breaks.
  2. Gemma has never seen a real photograph. Every model interaction so far was text-only or stubbed. Scoring quality, the hedging behaviour in SPEC Β§14, and whether completed=false comes back sensibly are all unverified.

Your testing checklist

  1. cd backend && .venv/bin/uvicorn app.main:app --reload, then frontend && pnpm dev.
  2. curl -X POST localhost:8000/api/test/gemma β†’ expect success: true.
  3. Delete any old backend/outside.db β€” users gained non-null email/password_hash columns and create_all won't migrate them.
  4. Complete 3 real challenges. Watch for: over-generous scores, hedged language ("this appears to be" β€” correct per SPEC Β§14), and whether a bad photo still gets completed: true.
  5. If scoring feels wrong, try DISCOVERY_PIPELINE=two_stage in .env and compare β€” that's the flag's whole purpose.

Want me to pre-write the DEV submission scaffold from SPEC Β§47 while you're out testing, so you only have to paste in the photos?

You

did you not use uv ?

Agent

Fair challenge. I used uv venv + uv pip install, but not properly β€” there's no lockfile, and dependencies are duplicated across pyproject.toml and requirements.txt. Let me fix that properly.

You

just use uv

You

us sync, uv add, uv run

You

i don't want venv, or pip

Agent

Got it β€” uv only, no pip/venv instructions anywhere.

Agent

Committed c4ff53d. Now uv-only end to end:

cd backend
uv sync                                        # .venv + deps from pyproject.toml/uv.lock
uv run uvicorn app.main:app --reload
uv run pytest
uv run ruff check app tests

What changed:

  • Deleted requirements.txt β€” dependencies were declared twice, in pyproject.toml and requirements.txt, which is exactly how they drift.
  • Added backend/uv.lock β€” there was none, so installs weren't reproducible. Now pinned across 69 packages.
  • Dev tools moved to a PEP 735 [dependency-groups] block via uv add --dev, not [project.optional-dependencies], so they're never installed by a production sync.
  • Dropped the pythonpath = ["."] hack from the pytest config. uv sync installs the project itself, so import app now works from any directory instead of only when the cwd happens to be backend/.
  • README has no pip or venv instructions anymore.

One thing to flag honestly: you still have a .venv directory. That's not something I can avoid β€” it's where uv puts the environment, and it's git-ignored (backend/.gitignore:3). If what you actually meant was "no environment directory in the repo at all," the alternative is uv run --isolated, which builds a throwaway env per command β€” correct, but it re-resolves dependencies on every run and would be painfully slow for a test suite. Tell me if that's what you want and I'll switch.

Verified after the change: uv run pytest β†’ 162 passed, uv run ruff check app tests β†’ clean.

You

GET /variant1 404 in 936ms (next.js: 719ms, application-code: 217ms)
GET / 200 in 72ms (next.js: 22ms, application-code: 50ms)
[browser] β¨― unhandledRejection: AbortError: signal is aborted without reason
at AuthProvider.useEffect (src/lib/auth-context.tsx:52:29)
50 | })();
51 |

52 | return () => controller.abort();
| ^
53 | }, []);
54 |
55 | const adopt = useCallback((response: { token: string; user: AuthUser }) => {
GET /login 200 in 287ms (next.js: 272ms, application-code: 16ms)
GET /today 200 in 159ms (next.js: 138ms, application-code: 21ms)
[browser] β¨― unhandledRejection: AbortError: signal is aborted without reason
at TodayPage.useEffect (src/app/today/page.tsx:43:29)
41 | });
42 |
43 | return () => controller.abort();
| ^
44 | }, [ready]);
45 |
46 | // Keep tomorrow's challenge warm so the morning wait is short. the loader at /today is shit,

Agent

Two separate problems. Let me look at the actual code before guessing.

Agent

Found the real bug. Look at the abort call sites β€” the rejection escapes because:

  1. In auth-context the request runs in a floating async IIFE ((async () => {...})()) whose promise nobody holds.
  2. AbortError is never checked as a type, only via signal.aborted β€” and the abort carries no reason, hence "aborted without reason".

And the loader is genuinely broken β€” all three lines render at once instead of rotating:

Agent

Now the same fix in the three pages that fetch:

Agent

Now the loader β€” it renders all three lines at once instead of rotating them. Rebuilding it:

Agent

The submit screen has the same 15s wait with no real loading state. Wiring it up too:

You

please don't build again and again

Agent

Understood β€” no more repeated builds. Committing what exists:

You

//localhost:3000
- Network: http://[REDACTED]:3000
βœ“ Ready in 431ms
βœ“ Running next.config.ts took 120ms

GET /today 200 in 802ms (next.js: 575ms, application-code: 226ms)
[browser] "β¨― unhandledRejection:" unmounted

Agent

My abort reason string itself is now leaking as an unhandled rejection, which means a rejection path escapes my try. Hardening it:

Agent

Committed 13474d3. Two things, both now fixed:

The unhandled rejection

I'd misdiagnosed it the first time. The real cause: my try block only wrapped the fetch() call. Aborting between response headers and the body rejects at await response.json() β€” outside the guard β€” so the rejection escaped with whatever reason I'd attached. That's why my "fix" simply changed the message from "aborted without reason" to "unmounted" instead of eliminating it.

Now:
- request() guards the body read too, and everything funnels through one normaliseFailure().
- abortOnUnmount() aborts with a real DOMException("...", "AbortError") instead of a bare string, so even a future leak is recognisable.
- Callers ignore aborts by error type, not by poking signal.aborted.

The loader

You were right that it was bad β€” and it was worse than a styling problem. All three status lines rendered simultaneously as separate spans with staggered fade delays, so the intended rotation never happened; you got three lines stacked and overlapping.

Rebuilt: one line at a time on a 4s crossfade, an elapsed timer, and a progress rule that caps at 92% rather than animating to 100% β€” a bar that completes and then sits there while you wait is worse than no bar. The copy now says the wait is real and that you can close the app. Also wired into /submit, which had a 15s wait behind a static button label.

I skipped the rebuild as you asked β€” pnpm lint is clean, and pnpm build was green on this same code before the last two edits touched only imports and abort plumbing.

One loose end from your earlier log worth naming: GET /variant1 404 is not a bug in your app. That's Next.js 16's dev-overlay asset probing a route your app doesn't have. Harmless; it'll disappear in a production build.

You

the app is calling /api/me rapidly, also in the your journal page, images aren't loading

Agent

Both are real bugs. Let me confirm the image one first:

Agent

Confirmed β€” POST /api/discoveries returns /uploads/x.jpg (line 137) but the journal returns uploads/x.jpg with no leading slash (line 147). The frontend concatenates, producing http://localhost:8000uploads/... β†’ 404. Fixing at the source so both paths agree:

Agent

That's the bug β€” /uploads concatenated onto the base with no separator. The backend fix is in; now let me make the frontend immune to it and add regression tests.

Agent

Now the /api/me flooding β€” dedupe and short-cache the session check:

Agent

Now backend regression tests so this URL bug can't come back:

Agent

Committed 6573718. Both bugs fixed.

http://localhost:8000uploads/... β€” missing separator

Two code paths returned the same discovery but built the URL differently:

  • POST /api/discoveries β†’ "/uploads/x.jpg" βœ…
  • GET /api/journal β†’ "uploads/x.jpg" ❌

The frontend concatenates ${API_BASE}${image_url}, so the journal produced localhost:8000uploads/... β€” note the missing /. Fixed at the source: image_url now lives in to_public_dict() as the single source of truth, always slash-normalised, and neither endpoint builds it by hand.

I also made the client defensive with imageSrc(), which joins with a guaranteed separator and passes through absolute/blob URLs. One malformed field from the backend can't 404 an entire journal again.

Three regression tests now assert the submit and journal URLs are byte-identical.

/api/auth/me firing repeatedly

Real cause: dev StrictMode double-mounts the provider, and the session check also re-runs on navigation. Rather than fight it with a ref, I made the call idempotent β€” concurrent requests collapse into one in-flight promise, and a result is reused for 15 seconds.

Two safety details worth naming:
- The cache key includes the bearer token, so switching accounts can't serve the previous user's session.
- clearRequestCache() runs on sign-in and sign-out.

Applied to me() only β€” challenge generation and submission must never be cached.

165 backend tests pass, frontend lint clean. I didn't rebuild the frontend, per your instruction β€” worth a pnpm build when you're next in there.

One thing I should be straight about: I can't verify the images actually render from here, since that needs a real browser and a real uploaded file. The unit tests prove the URL is now well-formed and consistent between endpoints, but visual confirmation is still yours. If an image still fails after this, check the Network tab β€” if it's now localhost:8000/uploads/... and still 404s, the file isn't on disk and the problem moves to UPLOAD_DIR resolution.

You

File "[REDACTED]/WebstormProjects/hackathons/hacktoberfest-2026-week1/backend/app/db/session.py", line 45, in _set_sqlite_pragma
cursor.execute("PRAGMA journal_mode=WAL")
sqlalchemy.exc.OperationalError: (sqlite3.OperationalError) disk I/O error

Agent

Disk and directory are fine (33 GiB free, local APFS), so this is a WAL/pragma failure, not capacity. Two likely triggers: the stale -wal/-shm files I deleted earlier while a server held the DB, and uvicorn --reload running two processes against one file.

Either way the pragma shouldn't be able to take the app down:

Agent

Now clearing the possibly-corrupt WAL state and verifying it recovers:

Agent

DB is actually healthy (integrity: ok, all four tables, WAL active) β€” the error was transient. Let me verify the new fallback path works:

Agent

That's the real diagnosis: "database is locked", not disk. Your uvicorn --reload server holds the DB, and my fallback made it worse β€” PRAGMA journal_mode=DELETE also needs an exclusive lock, so the retry guaranteed a crash.

The bug in my fix: journal_mode is a persistent DB property, not a per-connection setting. You shouldn't be setting it on every connect at all.

Agent

Now it connects cleanly while your server still holds the file. Adding a regression test for exactly this scenario:

Agent

Two test bugs of mine: SQLAlchemy connection isolation (the first connection sees a pre-insert snapshot, which is correct behaviour), and sqlite3.Connection is immutable so it can't be spied on. Rewriting both:

Agent

Committed 61516d5. Your database is fine β€” I checked:

journal_mode: wal
integrity:    ok
tables:       users, challenges, user_preferences, discoveries

What actually happened

The disk I/O error was misleading. When I reproduced it, the real message was database is locked β€” your uvicorn --reload server was holding the file.

The root cause is a design mistake in my code: I was running PRAGMA journal_mode=WAL on every new connection. But journal_mode is a property of the database, not the connection, and changing it requires an exclusive lock. So:

  • Any second process β€” your reloader, a stale -wal/-shm from earlier, a stray script β€” makes it fail.
  • My first fix made it strictly worse: I "fell back" to journal_mode=DELETE, which needs the same exclusive lock, so the fallback guaranteed a crash.

The fix

  • The journal mode is attempted once per process, not per connection.
  • Failure is logged and swallowed β€” the app boots and uses whatever mode the database already has. WAL is a performance nicety, not a correctness requirement.
  • busy_timeout is now set before any locking pragma, so concurrent writers wait instead of erroring.

Verified against your live DB while your server held it: connects fine, journal_mode: wal, busy_timeout: 5000.

Four tests added, including one that grabs an EXCLUSIVE lock on the database and asserts the app still connects.

Two things worth flagging:

  1. I contributed to this. Earlier I ran rm -f backend/outside.db-wal outside.db-shm while your server was live. Deleting WAL sidecar files out from under a running SQLite connection is exactly what produces these errors. My apologies β€” I should have checked for a running process first.
  2. If it recurs, the fix is to stop all processes touching the DB, then delete outside.db-wal and outside.db-shm β€” they're regenerated. Never delete them while a server is running. SQLITE_JOURNAL_MODE=DELETE in .env also sidesteps WAL entirely, at some cost in concurrent-write speed.

169 tests pass. I haven't rebuilt the frontend.

You

fine is it commited?

Agent

Yes β€” committed, working tree clean.

61516d5  Make SQLite pragmas survivable instead of fatal
6573718  Fix journal images and stop /api/auth/me request bursts
13474d3  Stop cancelled requests escaping as unhandled rejections
c4ff53d  Use uv as the only Python toolchain

18 commits total. Everything is under your identity (Harsh-0986 <[REDACTED]>), and the key isn't in any of them β€” .env is git-ignored.

One thing to double-check on your side: git status is clean because backend/outside.db, outside.db-wal, and outside.db-shm are git-ignored. Your uploaded photos in backend/uploads/ are ignored too, which is right β€” but it means demo discoveries are not in git. If you photograph real discoveries for the DEV submission (SPEC Β§37), commit those images deliberately rather than relying on them being tracked.

Ready when you are for the outdoor testing phase.

You

what shall be the name of the app? suggest

Agent

Here's my honest take, in tiers.

Tier 1 β€” keep what you have

Outside, Not Online

It's the strongest statement of the three, and it's memorable precisely because it reads as an argument rather than a product name. For a DEV post title it's excellent. The downside is length and that it describes a philosophy instead of an activity β€” someone reading the repo has to work out it's about going for walks.

Tier 2 β€” my recommendation

Offscreen

Short, one word, .dev/.app shaped, and it's the positive framing of the same idea. "Outside, Not Online" is what the app argues; "Offscreen" is what it does. It also dodges the trap of Touch Grass, which is the challenge's own name and already has a competing browser extension β€” using it invites an avoidable conversation about squatting someone else's name.

Runner-up: Notice. Single word, means both "pay attention" and "a discovery," and it's about the behaviour rather than the location.

Tier 3 β€” the rest, honestly ranked

Name For Against
Unscroll Attacks the feed directly; reads as an act Slightly tech-jargon
Small Quest Conveys low-stakes adventure, sets expectations "Quest" is gamified; SPEC says avoid
Trailhead Concrete, outdoorsy, implies a start point Bland
Go Outside Blunt, funny, unmistakable Reads as a command, not a brand
Fieldnote Journal feel, alludes to field research Generic; many existing apps

Names I'd steer away from: Glimpse, Scout, Wanderlust, Look Up, and Not Online β€” all crowded, several actively confusing (Look Up especially, next to Google).

How I'd structure it

Use both, in a specific way:

Offscreen β€” Outside, Not Online

Product name short, manifesto preserved as the tagline. That gives you the punchy repo name (offscreen) without discarding the phrase you've already written into the SPEC and the demo closing scene.

So: go with Offscreen, or tell me the vibe you're after β€” dry/serious, warm/playful, or wilderness-y β€” and I'll push harder in that direction.

You

photo walk

Agent

It's a decent instinct β€” walking is more concrete than "outside" β€” but I'd push back on it as the product name, for two reasons.

1. It's a photography-community genre, not your product. "Photo walk" already means something specific: a group that walks out with cameras and shoots together. Search it and you get camera clubs, Nikon/Zeiss guides, and a genre of organised outings. Your app isn't that β€” you're not inviting anyone, you're sending one person out alone with one assignment and one photo. The name oversells the social half and undersells the challenge half.

2. It's crowded and generic. There are existing apps and services called Photo Walk / Photowalk. For a hackathon that's not fatal, but it means you spend your first line of the DEV post explaining what you are, and your repo name doesn't differentiate you from anything.

The good part β€” it does capture "go outside with a camera," which Offscreen doesn't. So I'd split the difference:

Offscreen β€” a photo walk, one challenge at a time

That keeps your instinct, fixes the collision, and reads accurately. It also gives the demo script a line: "It's a photo walk. Except you're alone, and instead of a route someone gives you a thing to find."

If you'd rather the name itself lead with walking, these are better than plain "Photo Walk":

Name Why
One Walk Pairs with "one photo, one discovery" in your own tagline; ownable, calm
Field Walk Outdoorsy, less generic than Photo Walk, not a known genre term
Twenty-Minute Walk Concrete and funny; matches your "~20 minutes" challenges

My ranking of what we just discussed: Offscreen > One Walk > Photo Walk > Field Walk.

Which do you want β€” and do you want me to do the rename across the repo (SPEC title, README, package.json, frontend metadata, page copy), or keep "Outside, Not Online" as the product and only use a short name for the repo directory?

You

update the readme to use offscreen, and generate all the docs

Agent

Writing the docs now. Starting with the README as the front door:

You

build failes

Agent

That's my fault β€” I refactored api.ts and added lib/image.ts but only linted, never rebuilt. Let me see it:

You

frontend

Agent

lib/image.ts uses API_BASE without importing it β€” eslint's config isn't type-aware, so it passed lint but failed tsc.

Agent

Fixed β€” all 8 routes build. Now applying the Offscreen branding to the frontend:

Agent

Build and lint clean. Writing the last doc, then committing:

Agent

Committed 96ea417.

The build failure was mine

lib/image.ts used API_BASE without importing it. Lint passed β€” the eslint config isn't type-aware β€” and I only ran lint because you told me to stop rebuilding. That's the trade-off biting: I traded the only check that would have caught it. Worth noting for the rest of the session: pnpm build is the real gate here, not pnpm lint.

Docs written

File Contents
README.md Rewritten as the front door, with a doc index and honest status table
docs/ARCHITECTURE.md Both graphs, the single branch, state design, the Gemma workaround, one-call pipeline, layer boundaries
docs/API.md Every endpoint with request/response examples and status codes
docs/TESTING.md Offline strategy, coverage, the bugs tests caught, and the AI cases still unwritten
docs/SECURITY.md Threat model per surface, with limits stated rather than implied
docs/DEV_SUBMISSION.md Draft write-up, every gap marked [REQUIRED]
CONTRIBUTING.md Workflow and the rules that keep the suite offline

Branding

Offscreen is the product name, Outside, Not Online is kept as the manifesto it argues, and "a photo walk, one challenge at a time" is the tagline. Applied to the landing headline, nav wordmark, page metadata, login back-link, and package.json. All 8 routes build, lint clean.

Two things I want to flag rather than let you discover later:

  1. DEV_SUBMISSION.md is a skeleton, not a submission. The real-world test section is [REQUIRED] with an empty checklist, and I wrote an "honest limits" section stating that photos leave the machine and that local inference is unimplemented. That's deliberate β€” the spec warns against unsupported privacy claims, and the challenge asks you to show real work. Fill it in with your actual trips.

  2. SPEC.md is untouched. It still says "Outside, Not Online" throughout, including the phase table. I didn't rewrite your spec β€” it's the historical document the build was derived from. If you'd rather it match the new naming, say so and I'll do it as a separate pass.

You

fine push ,and did you fix the frontend build issue?

Agent

Yes β€” fixed and verified. lib/image.ts now imports API_BASE, build compiles clean.

Agent

Only 1 commit ahead of origin/main? Let me check what's actually on the remote before pushing:

Agent

18 commits are already on origin β€” only the docs commit is unpushed. Let me verify no secrets are tracked before pushing:

Agent

Pushed. main is in sync with origin/main at 96ea417.

The build issue β€” yes, fixed

import { API_BASE } from "./api";   // was missing entirely

lib/image.ts referenced API_BASE without importing it. It only surfaced now because pnpm lint passed while tsc failed β€” the eslint config has no type-aware rules, so lint is not a substitute for build in this project. My earlier "verified, lint clean" was overclaiming; the build is what counts. Verified just now: βœ“ Compiled successfully, all 8 routes.

On the push

Two things worth knowing:

18 of 19 commits were already on the remote. I hadn't run git push at any point this session, so something else pushed them β€” an IDE, a hook, or you. Not a problem, but I'd rather flag it than let you assume I'd been pushing silently.

Secret check before pushing, since this was the first push I did:
- Only .env.example is tracked β€” your real .env is ignored.
- No GEMINI_API_KEY= value anywhere in the commit history.
- No outside.db, -wal, -shm, or uploads/ files tracked.

That last one has a consequence for you: your photographs aren't in the repo. When you take the real discovery photos for the DEV submission, add them deliberately with git add -f backend/uploads/... or move them somewhere tracked β€” otherwise the demo imagery exists only on your laptop.

Also worth noting: the remote is Harsh-0986/OffScreen β€” camelCase, which GitHub normalises to Offscreen in URLs. If you want the README links and DEV post to match exactly, you may want to rename the repo to offscreen.

You

how do i add api url to frontend?

Agent

Important catch: the root .env is invisible to Next.js. Next only reads env files from its own directory (frontend/), so that NEXT_PUBLIC_API_BASE_URL line in the root .env has been doing nothing β€” your frontend has been silently falling back to the hardcoded http://localhost:8000.

Setting it up properly:

Agent

Done and pushed (39844e6).

How to set it

The frontend needs its own env file β€” Next.js only reads env files from its own directory, never the repo root. So:

cd frontend
cp .env.example .env.local

Edit frontend/.env.local:

NEXT_PUBLIC_API_BASE_URL=http://localhost:8000

I've already created that file for you with the localhost default, so it works as-is for development.

The bug this uncovered

Your root .env had NEXT_PUBLIC_API_BASE_URL in it, and it was doing nothing. Next never reads the repo root, so the frontend was silently using the hardcoded http://localhost:8000 fallback. It happened to be the same URL, which is why nothing looked broken β€” but the setting was inert, and it would have pointed at localhost in a real deployment without any warning.

Two guards added:
- A one-time console.warn in a production build when the variable is missing.
- The root .env.example now points at frontend/.env.example instead of implying the variable lives there.

Gotcha worth memorising

NEXT_PUBLIC_* is inlined at build time, not read at runtime. So changing it requires restarting pnpm dev or re-running pnpm build β€” editing the file while the dev server runs can appear to do nothing. (Same reason GEMINI_API_KEY can never be a NEXT_PUBLIC_ variable: anything with that prefix ships to the browser.)

Confirmed the value is inlined into the client bundle. Build and lint clean.

For a real deployment

Hardcoding a backend URL is fine locally, but for the hackathon demo or any deploy you'll want either a full URL in .env.local, or better β€” a Next rewrite proxy so the browser calls same-origin /api and CORS disappears entirely:

// next.config.ts
async rewrites() {
  return [{ source: "/api/:path*", destination: `${process.env.BACKEND_URL}/api/:path*` },
          { source: "/uploads/:path*", destination: `${process.env.BACKEND_URL}/uploads/:path*` }];
}

Want me to add that? It removes CORS config, removes the env var from the browser bundle, and makes the app deployable behind one origin.

Prize Categories

I'm entering the overall prize, Best Use of Gemma and Best Use of Render the FastAPI backend and its Gemma calls run as a web service on Render.

I'm not entering the other partner categories. Nothing in this project uses Arduino, TabPFN, or Tinker, and a category is a claim that a technology did real work β€” not a list of things I could have used.

Final Thought

There are already millions of apps competing for our attention.

I wanted to build one that does the opposite.

Offscreen isn't trying to keep you here.

It's trying to give you a reason to leave.

One photo. One discovery. One reason to go outside.

Top comments (0)