DEV Community

BARI ANKIT VINOD
BARI ANKIT VINOD

Posted on

Spotter: a private workout form coach I built for a friend who trains at home

Hacktoberfest Weekend Challenge: Build for a Friend Submission ๐Ÿค

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

Spotter is a workout form coach for my friend, who trains at home. There's no gym nearby and no budget for a private trainer, so a phone propped on a shelf is the only feedback they get. Every "AI coach" app they tried either wanted a subscription, sent their workout videos to a company's servers, or repeated generic YouTube advice.

So I built something different. They upload a short clip, add a little training context, and get back:

  • the detected exercise and a confidence score
  • a rep-by-rep analysis
  • valid variations separated from real form issues
  • an annotated video plus short clips of each issue
  • a grounded coach summary with fixes and a next-session plan

It currently supports squats, push-ups and shoulder presses. It also says "I can't tell" on purpose: if a clip isn't one of those movements, Spotter rejects it instead of forcing a guess. It is a practice companion, not a medical device. It doesn't diagnose injuries or claim to prevent them.

My friend said, "I have good shoulder strength but it helps me in squats".

Demo

Video link - https://youtu.be/DImVwhrrkFA

Code

title Spotter
emoji ๐Ÿ‹๏ธ
colorFrom green
colorTo blue
sdk gradio
sdk_version 6.17.3
python_version 3.10
app_file app.py
fullWidth true
short_description Workout form coach for a friend, built on open-weight small models.
tags
gradio
computer-vision
pose-estimation
fitness
video-analysis
llama-cpp
open-source
local-inference
hacktoberfest

Spotter

Spotter is a small-model workout form coach I built for a friend.

My friend trains at home: no gym nearby, no budget for a private trainer, and a phone propped on a shelf is the only feedback they get. Every "AI coach" app they tried wanted a subscription, shipped their workout videos to a company's servers, or paraphrased generic YouTube advice. So I built Spotter for them instead.

A user uploads a short exercise video, adds basic training context, and gets a structured form-review report with rep counts, movement notes, annotated video, and a grounded coach summary โ€” produced by open-weight models that cost nothing to run and neverโ€ฆ

How I Built It

Spotter is a pipeline of small, checkable steps rather than one big chatbot:

video + profile
  -> video quality check
  -> pose extraction (MediaPipe Pose Landmarker Lite)
  -> pose cleaning
  -> exercise router (custom PyTorch BiLSTM, 182,796 parameters)
  -> exercise-specific rep counter
  -> per-rep analysis -> variation detection -> issue markers
  -> annotated video + issue clips
  -> coach summary (Nemotron-3-Nano-4B + LoRA)
  -> verifier
Enter fullscreen mode Exit fullscreen mode

The models. The router is a tiny BiLSTM over 30-frame windows of pose landmarks. The coach summary comes from nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16, fine-tuned with a LoRA on task-specific data. Training and publishing run as Modal jobs. For fully local use, the summary can run through llama.cpp with a GGUF build, so MediaPipe, the router and the coach summary all run on a laptop.

Evidence first, language second. The numbers (reps, depth, issue counts) come from deterministic code. The language model only phrases them. A verifier then checks the summary against that evidence and blocks claims it can't ground or that sound like a diagnosis. I'd rather the coach say less than say something it can't back up.

Optional extras, all opt-in. Spotter's core promise is that it runs locally with no account, so anything cloud-based is off by default and falls back quietly:

  • Voice coach (ElevenLabs). Short spoken cues for the top issues, like "rep 3: knees drifted inward". It only speaks text that already passed the verifier, caches audio on disk so repeat runs cost nothing, and enforces a per-run character budget. Only text is sent, never video.
  • Cross-session memory (Backboard). A "vs last session" comparison, so my friend can see whether a fault is fading. SQLite stays the source of truth. Only derived metrics (date, exercise, rep count, issue counts) are synced, never video, frames or landmarks.
  • Fine-tuning pipeline (Tinker). A LoRA fine-tune of a small Qwen model for the progress-plan step, with a dry-run mode that prints token counts and cost before spending anything.
  • Deployment (Render). A Render Blueprint with a health check and a persistent disk for history. Live require credit card (I don't have).

I built this with Claude Code in a long session of small, verified steps.

Why Does Open Innovation Matter?

Three things my friend cares about would have been impossible, or paid for, with a closed API:

  1. Their workout videos stay on their machine. Videos of someone training at home, often in their living room, are personal. With open weights and local inference, there is no server that has to see them.
  2. It costs nothing to run and needs no account. There is no subscription and no per-request bill. That matters when the whole point is that a trainer is too expensive.
  3. I can change it for them. The coach's voice and focus come from a LoRA I trained, so when a better 4B model appears or my friend wants a different style, I retrain the adapter instead of waiting for a vendor. The small router and the verifier mean every claim can be inspected, not just trusted.

Top comments (0)