DEV Community

Diwakar Arya
Diwakar Arya

Posted on Fully Autonomous

GutTracker: A Food & Symptom Journal That Runs Fully Offline for My Sister

Hacktoberfest Weekend Challenge: Build for a Friend Submission ๐Ÿค

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

My sister has been having gut problems for the past few weeks. The usual drill: write down what you ate, write down how you felt, try to spot the pattern โ€” and never quite manage to see it, because a notebook can't compute anything.

So I built GutTracker, a food and symptom journal for her. She types a meal the way she'd say it out loud ("two idlis and a bowl of sambar with filter coffee"), logs symptoms in the same casual language ("mild bloating about an hour later"), and gets back a week's view: symptom rate, the symptoms she's reported most, and which foods keep turning up alongside them.

Two rules shaped every decision:

  1. It stays on our machines. Food-and-gut logs are exactly the kind of data you don't want on someone else's server, so the model runs locally on llama.cpp and the database is a single SQLite file on disk. There is no cloud AI in the request path, by design.
  2. The model never gets a vote. It parses her sentences into structured data and restates numbers Java has already computed. Every statistic is computed by Java; every sentence it produces is screened by a LanguageGuard that rejects causal, diagnostic or prescriptive wording outright. GutTracker reports co-occurrence โ€” "you logged rice in 3 meals; 2 of them came with discomfort" โ€” never cause.

GutTracker is a logging and observation tool, not medical advice, and it says so in the UI.

Demo

Live: guttracker.onrender.com

Heads-up on what you're seeing: the deployed instance has no local model next to it, and the challenge rules say no cloud AI โ€” so Render runs the deterministic fallback path, and the status chip honestly reads offline. That's the same fallback that runs whenever my sister's laptop can't reach the model, so the demo is the degraded path working, not a broken feature.

Locally it's the full thing:

./start.sh   # backend :8081, UI :5173, model on :8080
Enter fullscreen mode Exit fullscreen mode

The status chip polls /api/v1/health/ai every 30s: green means the open-weight model is answering, grey means Java's deterministic tokenizer is.

Code

GutTracker

CI License: MIT

โ–ถ Live demo: https://guttracker.onrender.com ยท Code ยท CI runs

A privacy-first food and symptom journal that runs entirely on your own machine with an open-weight model doing the writing and Java doing the maths.

You type a meal the way you'd say it out loud โ€” "two idlis and a bowl of sambar with filter coffee" โ€” and get back structured ingredients. You type how you feel later โ€” "mild bloating about an hour later" โ€” and get back correlations over your own records, phrased in a way that never pretends to be a diagnosis.

Nothing leaves the laptop. No account, no sync, no telemetry, no cloud AI.


Why open

This challenge asks why open matters for what you built. For GutTracker it is not a footnote, it is the whole design:

  • The data is the point. A food-and-symptom log is medical-adjacent and deeply personal. Putting it on aโ€ฆ

How I Built It

Stack: Spring Boot 3.5 + JPA + SQLite (REST only, single-user, zero auth because it's one person's notebook) ยท React + Vite with three tabs (Log / History / Insights) ยท llama.cpp serving Ternary Bonsai 2 27B, an open-weight model, on 127.0.0.1:8080 behind an OpenAI-compatible API ยท Docker โ†’ GitHub Actions โ†’ Render.

flowchart LR
  A[React UI :5173] -->|/api/v1| B[Spring Boot :8081]
  B --> S[(SQLite)]
  B -->|parse / restate| C[llama.cpp :8080
Ternary Bonsai 2 27B]
  C -->|out of reach| D[FallbackTokenizer +
deterministic template]
  B --> E[LanguageGuard
accept or reject, never rewrite]

Where the model is allowed to work โ€” three narrow jobs, all of them language, never maths:

Input Model returns Fallback if it can't
A free-text meal JSON: title, ingredients, category, amount fixed vocabulary tokenizer (~100 foods, plurals folded)
A free-text symptom canonical name + severity symptom vocabulary, severity 5 flagged estimated
Java's computed facts 2โ€“4 neutral sentences restating them template built from the same facts

Four bugs worth writing down, because they were the actual work:

1. The reasoning model ate its own answer. Parsing worked but summaries came back empty. finish_reason: "length", content: "", and a full reasoning_content โ€” the model's chain-of-thought was burning the entire max_tokens budget before the real answer started. It looked exactly like "the AI is broken"; it was configuration. Fixed by sending chat_template_kwargs: {"enable_thinking": false} per request and raising the budget 512 โ†’ 1024.

2. The guard caught nothing in the present tense. My pattern was \bcaus(?:ed|ation|causative)s?\b โ€” "caused", "causation", caught. "This causes bloating"? Straight through the safety rail, and that's the most likely sentence a model would ever write. The test that caught it (expected hits to contain 'causation') is the one I trust most in the repo. Closed the siblings too: triggering, diagnose, treats, recommends, increases. The policy is reject and retry once, then fall back to a template โ€” the guard never rewrites text, because a rewriting guard is just a model pretending to be a rule.

3. "No problem" was a symptom. An entry literally saying "no problem, fine" was counted as a symptom report, so no problem ร—1 appeared in the insights. Real reports like that stay in the history โ€” they're true entries โ€” but are excluded from every statistic.

4. My denominator was wrong. The insights ranked "ingredients to watch" by symptom rate across the whole window, so foods appearing once with a symptom (rate 1.0) crowded out rice (2 of 3, and she'd actually notice that). Fixed twice: denominator is now meals containing that ingredient, ranked by number of reported problems first, with the associated symptom chips shown inline.

Verification: 17 unit tests (mvn verify) + a CI job that builds the Docker image, boots it and smoke-tests every endpoint. CI ran green on the first push; Render auto-deploys from main via the render.yaml blueprint.

Why Does Open Innovation Matter?

The honest version: I could not have built this part with a closed API.

  • The privacy story is real only if it's local. A closed API means every meal she types is a request to someone else's servers. This one never leaves the laptop. The health-tech community argues the same point from the other side โ€” open-weight AI in healthcare and privacy-first offline health tools both make on-prem the default rather than the exception โ€” while a good counterpoint piece reminds us that "local" isn't a security guarantee by itself. Fair. Here, the useful property is narrower and honest: the data isn't sent anywhere.

  • Open weights let me change the model's behaviour, not just call it. enable_thinking was a per-request template flag, not a support ticket. The LanguageGuard, the FallbackTokenizer, the deterministic summary template โ€” I can read, audit and patch all of them. With a closed API, the two things that made this safe to hand my sister (no causal claims, no silent failure) would be someone else's policy page.

  • The fallback exists because the source exists. When the model is unreachable, a ~60-line tokenizer in the repo keeps the app useful. No API key, no vendor outage, no notebook.

What open innovation made possible that a closed API wouldn't: an app whose worst case is "the AI is off and it still logs my meals", instead of "the API changed its pricing and my sister's food diary stopped working".

My Agent Session

Building GutTracker: local LLM parsing, a guard that caught nothing, and a reasoning model that ate its own answer
You

Build GutTracker: a privacy-first local food + symptom tracker for the Hacktoberfest Weekend Challenge. Java Spring Boot + SQLite backend, React/Vite UI with 3 tabs, local Ternary Bonsai 2 27B on llama.cpp for parsing, deterministic Java for all stats, and a LanguageGuard that rejects causal or diagnostic wording from the model.

Agent

Backend written but never compiled. Starting with the build and fixing what falls out.

You

Smoke-test the API.

Agent

Not a crash in the parser: MealRequest requires rawText and my curl sent structured fields. The frontend already sends rawText correctly โ€” the bug was in my test. But two genuine bugs fell out of the same pass: unknown routes returned 500 instead of 404, and bean-validation failures returned 500 instead of 400. Added handlers for NoResourceFoundException and MethodArgumentNotValidException.

Agent

AI summary for /insights returns nothing while meal parsing works. Instrumenting the client rather than guessing.

Root cause: Bonsai is a reasoning model and the chain-of-thought consumed the entire max_tokens budget, so content came back empty. This looked like 'the AI is broken' but was pure config. Fix: chat_template_kwargs {"enable_thinking": false} per request, max_tokens 512 โ†’ 1024. Verified the JSON meal-parse path still works with thinking off.

You

In Insights, can you add which potential ingredients cause problems?

Agent

Four issues behind what looks like one feature request:

1) Denominator changed to meals containing that ingredient. 2) Each row now carries the symptoms reported alongside it with counts. 3) Ranking moved from raw rate to number of reported problems first โ€” rate-only ranking let six 1-of-1 ingredients push out rice at 2 of 3, the most informative row in the data. 4) A test entry literally saying 'no problem' was being counted as a symptom report, so 'no problem ร—1' appeared in rice's symptom chips. 'Nothing was wrong' reports stay in the log but are excluded from every statistic.

Agent

Writing tests for the guard. One failed, and the failure was the interesting part.

The guard pattern was \bcaus(?:ed|ation|ative)s?\b โ€” it caught 'caused' and 'causation' but not the present tense 'causes'. The one sentence most likely to appear in a model's output slipped straight through the safety rail. Closed that and the sibling gaps: triggering, diagnose, treats, recommends, increases.

You

Push to GitHub, set up CI/CD and give me a Render link.

Agent

Dockerfile used temurin-17 while the pom targets Java 21. CI had the same mistake waiting. Fixed both.

Deployed via render.yaml blueprint. The container has no llama.cpp and the brief forbids cloud AI, so the hosted demo runs the deterministic fallback โ€” which is exactly the path the fallback tests cover.

The debugging beats, end to end: the reasoning-model token bug, the present-tense guard gap, the "no problem" statistics bug, and the Java 21 mismatch that broke the first Docker build.

Top comments (0)