This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend.
What I Built
Gym Log Buddy โ A Local AI Workout Logger Built for a Friend
My friend goes to the gym regularly, but he has one surprisingly annoying problem:
He forgets what weight he lifted last time.
He doesn't want a complicated fitness app with endless menus and forms. During a workout, he just wants to quickly write something like:
bench 60kg 8 reps 3 sets, incline db 22 x 10 x 3, squat 80 5x5
So I built Gym Log Buddy around the way he already takes notes.
Instead of forcing him to structure his workout, the app lets him write naturally. A local Gemma 3 model running through Ollama converts the messy note into structured workout data:
- Exercise
- Weight
- Reps
- Sets
Before anything is saved, the extracted data is shown in a review table so my friend can verify and correct it.
The confirmed workout is then stored locally in SQLite.
On the next session, Gym Log Buddy can show what he did previously and suggest a target for his next workout.
AI where it helps. Deterministic code where it doesn't.
I deliberately didn't use AI for everything.
Gemma 3 handles the part that requires language understanding: turning messy human text into structured data.
The progressive-overload calculation is plain Python:
- If the previous workout reached 10+ reps, increase the weight by 2.5 kg and target 8 reps.
- Otherwise, keep the weight and target one additional rep.
This keeps the predictable part of the application deterministic, testable, and easy to understand.
And the most important feedback?
"Bahut accha hai, keep it up."
My friend said it was very good and encouraged me to keep going.
That's the user requirement that mattered most.
Demo
Gym Log Buddy is intentionally a local application, not a hosted AI service.
It runs on the laptop with Ollama and Gemma 3 installed.
๐ฅ Video Demo
The demo shows the local Streamlit application processing workout notes and turning them into structured workout data.
Code
The complete project is open source on GitHub:
Shubham-cyber-prog
/
Gym-Log-Buddy
Local-first gym log: Gemma 3 (via Ollama) turns messy workout notes into structured data.
Gym Log Buddy
Gym Log Buddy is a local-first workout tracking application that extracts structured exercises, weights, reps, and sets from natural language workout notes It is built for a gym buddy who wants quick, friction-free workout logging without manual forms, accounts, or subscriptions It runs entirely on your local machine using an open-source language model (Gemma 3 via Ollama) for text extraction, deterministic Python rules for progressive overload recommendations, and SQLite for storage.
Features
- Plain English input: Type workout notes naturally (e.g. "bench 60kg 8 reps 3 sets, incline db 22 x 10 x 3, squat 80 5x5").
- Local open-source model: Uses Gemma 3 via Ollama. No cloud APIs, no API keys, and no paid services.
- Deterministic progressive overload
- If last reps >= 10: suggest +2.5 kg and 8 reps
- Otherwise: suggest same weight and +1 rep
- Dual interfaces
- CLI (cli.py): log, last, suggest, history commands.
- Streamlit UI (app.py)โฆ
Tech Stack
- Python
- Gemma 3
- Ollama
- SQLite
- Streamlit
- Requests
The model name is kept as a single configuration constant, making it straightforward to experiment with other local models.
How I Built It
The architecture is intentionally simple:
Messy workout note
โ
Gemma 3 / Ollama
โ
Structured JSON
โ
Validation & safeguards
โ
Review / confirmation
โ
SQLite
โ
Workout history
โ
Progressive overload target
The entire AI inference pipeline runs locally.
No workout note needs to be sent to a cloud AI provider.
Why Gemma 3?
The core problem isn't mathematical.
It's language understanding.
A person can write:
bench 60 8x3
or:
did bench today, 60kg, 8 reps x 3
or:
bench 60kg 8 reps 3 sets
A useful workout logger needs to understand that these are describing the same kind of information.
Gemma 3 provides the natural-language understanding needed to turn those messy inputs into a structured representation.
Python then takes over for everything deterministic.
I Tested the Real Model โ Including Its Failures
One of the most important parts of this project was not hiding model failures.
Failure #1 โ 60 Became 27.2 kg
Input:
bench 60 8x3
Gemma initially interpreted 60 as pounds and converted it to:
{
"exercise": "Bench Press",
"weight_kg": 27.2,
"reps": 8,
"sets": 3
}
The dangerous part was that the output looked completely valid.
There was no exception.
No malformed JSON.
Just the wrong workout.
I fixed this by making kilograms the explicit default and allowing conversion only when the input explicitly contains lb, lbs, or pounds.
Failure #2 โ The Model Invented an Exercise
Input:
did legs today, felt good
The model invented a Squat with null values.
The input contained no exercise.
I added a prompt rule requiring an empty exercise list for non-workout commentary, along with a programmatic validation rule that removes records where weight, reps, and sets are all null.
Failure #3 โ My Schema Couldn't Represent 8,8,6
Input:
bb row 70 8,8,6
The model repeatedly interpreted this incorrectly because my schema only supports one reps value per exercise.
The actual workout has different repetitions across sets:
8 reps
8 reps
6 reps
Instead of manipulating the benchmark to make the result look better, I kept the failure and documented it as a schema limitation.
This also reinforced why the application has an explicit confirmation step.
The AI can suggest. The user confirms.
Then I Found Benchmark Leakage
At one point, the benchmark reached:
18/18 โ 100%
That looked great.
It was also invalid.
I had accidentally included exact test inputs as few-shot examples in the prompt.
After removing those examples, I discovered a second leakage issue: parts of the same test inputs were still present in the prompt's rules section.
I removed those as well and added an automated test that checks for test-input leakage.
The final evaluation was split honestly:
- Held-out inputs: 15/18 = 83.3%
- Original inputs: 18/18, but not counted as evidence because the prompt had been developed around them
- 36 total live model inferences
I am deliberately not presenting 91.7% as a general AI accuracy number.
The benchmark is small and project-specific.
The important result for me was that the evaluation became cleaner and more trustworthy.
Testing Beyond the Model
I also discovered a bug that the original test suite completely missed.
The parser and database tests were passing, but the Streamlit application could crash on a fresh database with:
NameError: name 'selected_chart_ex' is not defined
The problem was that my original tests never actually executed the Streamlit application.
So I added Streamlit's AppTest framework and created tests for:
- Empty database rendering
- Workout history rendering
- Chart rendering
- Exercise-name normalization
The project now has 20 automated pytest tests.
The parser tests use a mocked model, while the actual Gemma 3 evaluation lives separately in:
tests/live_check.py
This distinction matters because mocked tests verify application behavior, while live tests verify the behavior of the actual local model.
Running Gemma 3 Without a Dedicated GPU
I ran the project on a laptop with:
- Intel Iris Xe integrated graphics
- 16 GB RAM
- Windows 11
- No dedicated GPU
The final live evaluation averaged approximately:
17.9 seconds for first-run inference.
Repeated runs averaged around:
9.5 seconds.
I didn't perform controlled hardware benchmarking, so these should be considered practical observations rather than formal performance claims.
It's perfectly acceptable for logging a workout after training.
It would not be ideal for real-time interaction between sets.
And that's an important trade-off to acknowledge.
Why Does Open Innovation Matter?
For this project, open innovation wasn't just about using an open model because it was interesting.
It changed what I could build.
๐ Privacy
Workout history is personal data.
With local inference, the workout notes stay on the user's machine instead of being sent to a third-party AI API.
There is:
- No cloud AI endpoint
- No API key
- No per-request AI cost
๐ Offline Capability
Once Gemma 3 is downloaded, the core parsing workflow can run without an internet connection.
That's particularly useful for gyms where mobile connectivity can be unreliable.
๐ฐ Zero Ongoing AI Cost
A hosted AI API would introduce usage costs and external dependencies.
Running Gemma 3 locally means the model can be used without paying for every workout entry.
๐ง Model Freedom
The application isn't locked to one proprietary AI provider.
The model is configurable, so I can experiment with different open models as they improve.
๐งช Inspectability
Because the inference pipeline runs locally, I can inspect the prompt, model output, validation logic, and evaluation process myself.
I can test failures instead of simply trusting an API response.
The honest trade-off
A larger hosted model would probably be faster and more accurate.
Gemma 3 running locally on my hardware isn't.
I chose:
Privacy + offline capability + zero per-request cost
over:
Maximum speed + maximum model accuracy
The confirmation workflow is how I make that trade-off practical.
One Important Privacy Caveat
I also discovered something important while writing the project documentation.
My project directory is currently inside OneDrive.
That means the SQLite database can potentially be synchronized by OneDrive.
So "local AI" does not automatically mean "private forever."
For the database to remain truly local, the project/database should be stored outside cloud-synchronized folders.
I documented this limitation rather than hiding it.
My Agent Session
I built Gym Log Buddy using Google Antigravity as the coding agent.
The agent helped with implementation and refactoring, while I reviewed the generated code, ran the tests, evaluated the real model, and investigated failures.
A second AI assistant was also used as a reviewer.
That review process helped uncover two particularly important issues:
- Benchmark contamination / prompt leakage
- The Streamlit
NameErrorthat the original tests couldn't detect
The goal wasn't to blindly accept generated code.
It was to use the agent to accelerate development while still treating testing, evaluation, and review as engineering responsibilities.
Prize Categories
๐ Best Use of Gemma
Gemma 3 is the core AI component of Gym Log Buddy.
It runs locally through Ollama and converts unstructured workout notes into structured workout data.
The project demonstrates a practical use of an open-weight model where local inference provides meaningful benefits in privacy, offline capability, cost, and model flexibility.
Final Thoughts
Gym Log Buddy started with a very small problem:
My friend keeps forgetting his gym weights.
It ended up becoming an experiment in local AI, evaluation, prompt leakage, schema design, testing, privacy, and honest benchmarking.
I didn't build a huge fitness platform.
I built a small tool for one real person.
And that's exactly what this challenge was about.
Build for a friend. Build something useful. Then test whether it actually works.
Thanks for reading, and happy Hacktoberfest! ๐

Top comments (0)