DEV Community

Santiago Gonzalez Toral
Santiago Gonzalez Toral

Posted on

A setup coach for my friend's RC car that stays on his laptop

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

My friend races a 1/10 touring car, a Yokomo BD12. On practice days he changes something, drives a few laps, changes something else, and by the afternoon he can't say which change helped. When the championship race comes he has no record of what worked at which track.

That is what he asked me for. In his words: "this app will be very useful for practice days so I can quickly store my settings and get feedback from what works, what might have affected my run."

RC Pit Companion does three things:

  • The coach: he says how the car feels ("the track is very slippery", "once i hit the breaks into the corner the car oversteers"). It picks one change, says what to expect and how to check it over the next few laps, and shows the change on a small 3D car.
  • The practice log: each change is stored with a better, same or worse outcome, next to the track conditions. It tells him whether a lap-time difference is bigger than the usual run-to-run noise.
  • Race day: he enters the conditions and gets his saved setups ranked by how close the track was to the one he's racing.

Here is one real turn. Session on a low-grip carpet track, and I said: "the back steps out when I get on the power out of the hairpin." The model labelled it loose on power (confidence 0.95). A hand-written table then picked the lever: softer rear diff oil, toward 3,000 to 4,000 cSt. The model read it back as: "We are softening the rear diff oil to around 3,000 to 4,000 cSt to let the rear wheels turn independently for more side grip. This might feel looser into the corner, so check power-on exits over three to five laps while confirming turn-in hasn't gotten looser." The whole turn took 4.0 seconds on my Mac with the E4B model.

Here is the app from first launch to race day, on a clean database with the small E2B model that runs on his laptop, one screenshot per step. The track ("Club carpet") and the "better" verdict are made up for the screenshots. Nothing in the images is edited.

1. Start a session. It asks for the track, the surface and the grip, because some advice depends on them: the softer diff oil row only applies on low grip.

The first-run form

2. Say what the car feels like. I typed "the back steps out when I get on the power out of the hairpin". It reads that as loose on power, lists three quick checks to do first (belts, tyres, cross-tweak), then offers one change with the trade-off, how to check it and the notes it came from. It doesn't know his current diff oil, so it tells him to enter it in Setup. It won't guess a number.

The Coach screen after the typed sentence: Loose on power, three quick checks, and the suggestion Softer rear diff oil toward 3,000-4,000 cSt with its trade-off, check and note links, plus an Open Setup button

3. Or tap a chip. Chips skip the model's understanding step, which makes them faster on his laptop. "Push in" gives "Reduce caster, 4 to 3 degrees", the trade-off ("more twitchy on the straight"), and the 3D view of the change.

The Coach screen after tapping Push in: Reduce caster from 4 to 3 degrees with the trade-off and check, a schematic 3D view labelled

4. Apply it, drive 3 to 5 laps, say how it felt.

The Coach screen after pressing Apply, marked Applied, with the question

5. If it was better, save the setup with today's conditions.

The outcome panel after pressing Better, saying

6. Every change is kept with its outcome.

The Session screen listing the change

With runs logged, each run shows the change made before it, and "What affected run 3?" asks the app. This one is the demo data from npm run seed:demo, labelled DEMO in the app.

The Session screen with three demo runs, their ratings and best laps, and the setup change recorded with each

7. On race day, rank the saved setups for the conditions he enters. This frame was taken after loading the labelled DEMO setups, so my one live setup ("Club carpet", score 100) sits next to them. Setups saved on a different surface are listed separately and never rank above a matching one.

The Race day screen: today's conditions, the saved setup Club carpet scored 100 for same track and same grip, a DEMO carpet setup scored 50, and a separate group for other surfaces

The 3D view is how the app explains a setting, because a number like "rear shock position 2 to 3" doesn't mean much until you can see it. Each suggestion plays on a small schematic car: the faint copy is the setting before and the blue one is after. There are twelve of these explainers, ten in 3D and two flat diagrams. Four of them, captured from the scene's own test page:

Front camber going from 1.5 to 1.0 degrees negative, seen from the front: the faint tyres lean in more and the blue ones lean in less. The label reads

Camber, seen from the front. The tops of the tyres lean in less.

Front toe-out going from 1.0 to 1.5 degrees, seen from above, with the front wheel fronts pointing slightly further apart. The label reads

Toe, seen from above. Front toe-out points the wheel fronts out.

Rear ride height going from 5.0 to 5.8 millimetres, seen from the side, with the chassis sitting higher on the rear. The label reads

Ride height, seen from the side.

Rear shock position going from hole 2 to hole 3, seen from the front, with the blue shocks standing more upright than the faint ones. The label reads

Shock position: the outer hole stands the shock more upright. This is the move behind one of the reviewed rows in the table.

Every frame is labelled "Schematic, not to scale" and says how much the movement is exaggerated, because a 0.5 degree change would be invisible at true size.

It is a laptop app. He would have preferred his phone, and I tried: Gemma 4 E2B downloaded in a phone browser, but it lagged too much to use. I wrote that down, dropped the phone, and built for the 2019 MacBook Pro he also has.

I installed it on his laptop and recorded the demo there. He was very happy with it, and he can't wait to use it at the next practice sessions.

Demo

The video was recorded on his laptop. It is a walkthrough of the app with Wi-Fi turned off, using voice in and out. The screenshots above show the same steps one by one, and the sections below say what I did and did not verify.

Code

Hacktoberfest 2026

Monorepo for my Hacktoberfest 2026 challenge entries and the knowledge base behind them.

Hacktoberfest 2026 is run by MLH and DEV, presented by DigitalOcean. The theme is "AI belongs to everyone": build with open-weight models and open-source AI. Pull requests no longer count. Rewards come from virtual stickers (see STICKERS.md), and the DEV Challenges carry cash prizes.

Challenges

Each challenge lives in its own folder under challenges/. I'm starting with the weekend challenge and adding the others as they open.

# Challenge Window (UTC) Status
01 Weekend: Build for a Friend: RC Pit Companion Oct 2 02:00 → Oct 5 06:59 ✅ Published
02 Week 1: Open-Source AI Oct 5 → Oct 12 🔵 Opens Oct 5
03 Week 2: Open-Source AI Oct 12 → Oct 19 🔵 Upcoming
04 Week 3: Open-Source AI Oct 19 → Oct 26 🔵 Upcoming
05 Week
…

The project is in the challenges/01-weekend-build-for-a-friend folder of that repo. MIT licensed. The README has run steps (Node 22.13+, Ollama, two commands) and the full list of limits, and docs/GETTING-STARTED.md walks through install, first launch, a practice day, voice and the common problems.

How I Built It

The model does two small jobs. Everything else is plain code.

  1. Gemma 4 turns what he said into one of 12 symptoms, or into "not something my notes cover."
  2. After the advice is chosen, Gemma phrases it as a short spoken paragraph.

Between those two steps there is no model. A table of 61 adjustments (levers.json) says what to change, in which direction, by how much, and what it costs elsewhere on the car. I reviewed 12 of those rows by hand, and the app marks the other 49 as drafts. The model never makes up a number: a guard checks the wording against the table, and a template answer replaces the model's text if it strays. Every answer cites the notes it came from. A question outside the notes gets a refusal, not a guess.

Here is how a turn moves through the app. Only the two green boxes call the model. Everything else is plain code, and everything runs on his laptop.

Architecture diagram: Chrome with the Preact screens, 3D explainers and voice talks over HTTP to a local Hono server; the server runs the six-step coach workflow (refuse, classify, pre-checks, pick lever, explain, pause) with Mastra and a plain-TypeScript fallback, calls Ollama on localhost for the two Gemma steps, and reads and writes files on disk (levers.json, the notes snapshot, the vault and the paused-turn store)

The stack:

  • Gemma 4 E4B (QAT build) on my Mac, served by Ollama, using Ollama's JSON-schema output so the symptom label is always a valid id. His 2019 laptop has no usable GPU for this, so it runs on the CPU. I asked him for a raw speed test with the small E2B model (gemma4:e2b-it-qat) and got one reading: about 9 tokens per second writing and about 39 tokens per second reading a prompt.
  • Mastra runs the coach as a workflow with two pause points: one waiting for him to accept the suggestion, one waiting for the outcome after the laps. There is a plain TypeScript fallback behind the same interface, in case the workflow misbehaves.
  • A Hono server and Preact UI on localhost, three.js for the 3D scenes, one LibSQL file for the vault, MiniSearch for finding the right passage in the notes.
  • Voice: push-to-talk in, speech out. In the demo video both work with Wi-Fi off. Speech recognition runs on the device in desktop Chrome, after a one-time download of the English speech pack.

The touring-car notes are not new. They are a frozen snapshot of my earlier notes, which I distilled from other people's videos and manuals, and the repo credits that separately and checks it against a hash. All the code is new and written during the window. I did not reuse my earlier 3D guide or setup-sheet app.

On testing: the symptom test has 25 cases, 20 in my wording and 5 in my friend's. I ran it on the bigger E4B model first, and Gemma got 24 right. Then I installed it on my friend's laptop, which runs the smaller E2B model, and hit a bug my test had hidden: whatever I asked, the coach answered "reduce caster". When the small model can't place a sentence it falls back to the same symptom, even for "hello". On the same 25 cases E2B gets 21 right, and it refused only 6 of 12 off-topic sentences I tried.

I added a plain-code check in front of the model. A sentence with no word about driving or setup gets the refusal without calling the model at all, which also saves a few seconds on his laptop. With it E2B gets 23 of 25, and "hello", "what time is it" and "how should I run cli commands on a macbook" are all refused. A sentence that does contain a driving word, like "how many laps is the main race", still goes through to the model. I widened the word list once after it wrongly refused one of my slangy test cases, so that 23 is partly tuned on the test and not a clean number.

On his own five sentences E2B got 5 right and E4B got 4. The E4B miss was "it's understeering out of the corner": I'd labelled it as pushing on corner entry, the model said corner exit, and I think the model's reading is fair. A plain keyword matcher got 17 of 25. On 14 fresh phrasings I'd written down answers for before running them, E4B got 13. Finding the right passage in the notes hit 9 of 10 on questions I wrote before tuning the search. These are small tests, written by me, run on an Apple M4 Max with the model loaded. They say the app works for the sentences I tried; they do not say it understands everyone. His laptop is several times slower than my Mac, as the next list shows.

What fell short:

  • It is slow on his laptop. A typed sentence goes through the model twice, and from his one speed reading I estimate about 25 seconds for a full answer, and about 15 seconds for a symptom chip, which skips one of the two model steps. Text streams in while it is written. That is an estimate from the speed test, not a timing of the app, which I haven't done on his laptop yet. There is a setting that turns the model's wording off and shows the plain template text instead: the advice and numbers are the same, a chip answers straight away, and a typed sentence only waits for the model to understand it.
  • The 3D scenes are schematic. The directions come from the notes and some sizes are exaggerated so you can see them. The label says so.
  • Voice needs desktop Chrome and the on-device English speech pack, which downloads once. The embedded browser I used while building reported on-device recognition as unavailable, so I only saw it work in the demo video, once. Typing and tapping a symptom chip work without it.
  • Setup needs Node and Ollama. There's no installer.

Why Does Open Innovation Matter?

Three things here depend on the model being open.

First, it runs where he races. A pit table is not a place I can count on a good connection. The model, the notes, the vault and the search are all on his laptop, so the app doesn't need a connection to answer. In the demo video, recorded on his laptop, Wi-Fi is off and it still answers. I also checked that the server and Ollama opened only local connections during a full loop on my Mac. Installing it needs the internet once, for the packages, the model and the speech pack.

Second, his setups stay with him. What he runs at a track, and how it goes, is his own record. The vault is one file on his disk.

Third, it costs nothing per question, and I can see why it answers what it answers. A closed API would bill per turn and would send his habits somewhere. With open weights I could also shrink the prompt until it fit a slow CPU (252 tokens, down from 389, with accuracy unchanged on my 25 cases) and find out that Ollama's JSON schema doesn't reach the model: a version that relied on the schema alone scored 3 of 25. The advice itself isn't hidden in model weights either. It's in a table and a set of notes he or I can open and correct in an editor. I did not fine-tune Gemma, because tuned weights can't cite a source or be fixed by editing a file.

My Agent Session

Building RC Pit Companion: an offline setup coach for a friend's RC car with Gemma on his laptop (Hacktoberfest 2026)
You

Help me setup the initial repository structure (README.md, AGENTS.md, agent skills, etc), knowledge base for the Hacktoberfest challenges:

  • Use the devrelay skills to gather information about the hacktoberfest 2026 challenges. The current dir will act as a monorepo and will assign a separate folder per challenges under the challenges subfolder
  • A knowledge base will be created using the okf-wiki skill in order to maintain some context regarding brainstorming ideas, tools, frameworks, etc
  • …
Agent

The repo is set up. The weekend challenge, "Build for a Friend", closes Oct 5 at 06:59 UTC, which is Sun Oct 4, 11:59 PM PDT. When I checked it was about 25 hours away, and nothing has been built for it yet.

What's in place
- Root files: README.md, AGENTS.md (the harness-neutral entry point), and STICKERS.md. CLAUDE.md is a symlink to AGENTS.md. I ran git init but made no commit.
- Skills: everything lives in .agents/skills/, and .claude/skills is a symlink back to it. I copied in the nine devrelay-* skills and okf-wiki, and wrote one new skill, hf26-challenge. It covers scaffolding a challenge folder, a …

You

Previous prompt was incomplete. Go over the following list and complete what's left:

  • Use the devrelay skills to gather information about the hacktoberfest 2026 challenges. The current dir will act as a monorepo and will assign a separate folder per challenges under the challenges subfolder
  • A knowledge base will be created using the okf-wiki skill in order to maintain some context regarding brainstorming ideas, tools, frameworks, etc.
  • This repository will be used by different harnesses …
Agent

The partner report is written and the challenge AGENTS.md is in place. About 25 hours remained when I checked the clock before launching the subagent, so the deadline (Oct 5, 06:59 UTC) is getting close.

What's new
- AGENTS.md in the challenge folder has the brief, hard requirements, judging weights, the 16-partner list, the credit rules, the submission steps, and the official template verbatim. CLAUDE.md there is a symlink to it, and the root AGENTS.md now tells agents to read it first.
- SUBMISSION.md now uses t …

You

do you think the parter report is enough to cover tooling capabilities. It might be worth exploring each parter tool separately and adding them to the knowledge base. Give me your thoughts

Agent

Short answer: no, the report isn't enough for building, but it does its job as a triage tool. Per-partner pages in the knowledge base are worth doing, with some caveats about timing.

Why the report falls short
- It gives each partner one table row. That's enough to decide whether a partner fits an idea, but not to integrate it. It has no setup steps, SDK or API shapes, model and version specifics, quotas, or known failure modes.
- It lists about ten unverified items, and several of them are the ones that decide whether an idea works. Examples are whether Qwen3.5 adapters can be served locally and whether Ollama accepts audio input for Ge …

You

okay. let's explore the following ideas. Assign an Opus 5.5 subagent to dive deep into each, expand the ideas, which parter tools best fit into it, and analyze how it should fit into a rapid prototype to be submitted to the current challenge before the closed deadline. Finally generate a final report.

Community Skunkworks agent

  • Skunkworks within a dev-shop agency like RaidGuild, is a small, autonomous team operating outside normal bureaucratic structures to r …
Agent

Both deep dives and the final report are written, and I recommend building the RC Pit Companion. About 14 hours remain until the 06:59 UTC deadline. The reports come from Opus subagents; I read them in full but didn't re-verify their web and package claims.

ideas/FINAL-REPORT.md is the comparison. The two deep dives are 01-community-skunkworks-agent.md and 02-rc-pit-companion.md.

Why RC edges ahead
- Theme fit: one clubmate at the pit table asking "why does the rear step out on power?" is one person and one moment. Skunkworks reads …

You

let's go with the rc pit companion idea. Here's my friend persona:

  • he has a 1/10th RC Touring car (Yokomo BD12) likes to go to the track on practice sessions. He wants to make these practices worth the effort, know what settings to tweak the car based on feeling, get setup recommendations and store the best settings and track conditions so these can be retrieved during race championships to know which setup settings have work
  • he prefer using the agent companion on this mobile phone, but als …
Agent

The scaffold is committed as 657599b on main, and the working tree is clean. npm run build (typecheck plus production build) passes. The commit has 77 files, and I found no secrets when I scanned them.

What's in the commit
- Monorepo: AGENTS.md (with CLAUDE.md as a symlink), the READMEs, the sticker tracker, the skills in .agents/skills/ (with .claude/skills as a symlink), the knowledge base, and challenges/.
- RC Pit Companion scaffold:
- A Vite + TypeScript PWA shell with a manifest, icons and a service worker, and a mobile-first Coach / Setup vault / About layout.
- Typed stubs only, with no feature logic y …

You

what do you mean by notes snapshot?

Agent

The "notes snapshot" is the touring-car knowledge the coach needs to give advice. I mean a frozen copy of the RC touring-car pages from your OpenGrid wiki.

What it is
- Your OpenGrid wiki at ~/OpenGrid/knowledge-base/opengrid/wiki/ has pages like touring-car-setup-procedure.md, touring-car-suspension-tuning.md, touring-car-steering-geometry.md, touring-car-weight-balance.md and yokomo-bd12.md. They come from your earlier work, including the NotebookLM material.
- The coach has to answer from the RC knowledge those pages hold, such as which adjustments change entry understeer and what each does. So I'd copy only the touring-ca …

You

let's start by testing if we can run the agent with gemma on a phone, otherwise we should stick to laptop and that opens up the opportunity to integrate other parter tools like Mastra and other knowledge-base management tool to store the field knowledge base

Agent

The spike page is ready and served over HTTPS on this Mac. I haven't run the full model test myself. It needs a 2 GB download, and I'd rather you decide that than have me pull it into the in-app browser. The page also isn't committed yet.

What I verified here
- The page loads over HTTPS from this Mac.
- The device check runs. The in-app Chromium has WebGPU with shader-f16 and subgroups, which the model runtime wants.
- The WASM runtime loads from this server, with no request to jsDelivr. The package's default is jsDelivr, which would have broken any offline claim, so the files are self-hosted.

What the page tests on the phone
1. …

You

page is not reachable on my phone using the same network

Agent

Nothing on the Mac side explains it: the firewall is off, and both servers answer on 192.168.x.x when I request them from this Mac. The ARP table also lists five other devices on the same network, several with randomized MAC addresses. Some of those are probably phones, so the router isn't blocking device-to-device traffic outright. I can't tell which one is yours.

I started a second server with no HTTPS, to separate a network problem from a certificate problem. Please try these on the phone in order:

  1. http://192.168.x.x:5174/, which is plain http, port 5174.
    • If it loads, t …
You

model is downloading on my phone. I'll send the report once it ends

Agent

Good. A few things will help the test come out cleanly:

  • Keep the page open and the screen awake until the download finishes. Mobile browsers often suspend or kill background tabs, and a 2 GB download can't resume. If it dies, you restart from zero.
  • Order matters for the offline test: after the download, press 3 (init runtime) while still on Wi-Fi, because the WASM files come from this Mac. Press 4 (create engine from cache), then turn Wi-Fi and mobile data off and press 5.
  • If step 4 fails or the tab reloads, copy the report anyway (step 6). A crash on the phone is a useful result, and the error text says which l …
You

my phone is too slow and lags when running step 4 and 5. Given the time we have, let's stick with a consumer laptop solution. Store this agent session to DEV so we can tell how we tested mobile compatibility when writing the report

Agent

The session is saved to DEV as private session 504, but DEV shows 0 messages parsed for it, so I can't yet say it renders. We're now on the laptop-only plan. The time is about 20:12 UTC, so roughly 10.5 hours remain.

Saved session
- Where: the session on DEV. It is private and unpublished.
- What's in it: your 10 prompts and my replies, with 235 transcript records. It is cut at your save request.
- What I removed: my reasoning blocks, the harness system reminders, personal paths and your usern …

You

go with ollama. Then analyze Mastra and other partner tools for field knowledge base snapshot.

For the notes snapshot, do you think it is better to capture the knowledge base from OpenGrid or better right from the notebooklm. For the latter, we'll need to install the same notebooklm skill used on the other project in order to plugin it into the project. That can be marked as no pre-existing work

Finally, how useful and feasible (given the short deadline) you think will be to tune the gemma mod …

Agent

Ollama is now the runtime; I recorded it in the README, AGENTS.md and the provider stub, and the typecheck passes. On your three questions: use Mastra with one local SQLite file, take the wiki as the snapshot (NotebookLM only as a short extra), and skip fine-tuning.

1. Mastra and partner tools for the field knowledge

The knowledge base is tiny, about 10k words or 150 chunks. It is not a vector-database problem. There are two separate kinds of data:

Data Needs Fit
Field knowledge (frozen, cited) Search, with citations Keyword search (BM25 or SQLite full-text) is enough for v1. Embeddings are optional. …
You

Here's my feedback:

  • yes, download the gemma model you believe fits best for the solution
  • Let's skip model fine tuning on this challenge
  • notebookLM extras after the loop works, sounds perfect. It would double check any useful info is not missing
  • How are you going to store the kb snaphot? Using the same okf-wiki format or other ?
  • In Mastra, does embedding generation adds added value to the kb snaphost?
Agent

The pull is running in the background (6.1 GB; I'll run a baseline check when it finishes). I'll take your two questions in order.

How the snapshot is stored

I'd use the same OKF format for the frozen copy and a different format for what the app reads at runtime.

Layer Format Why
Frozen snapshot (kb/source/) OKF markdown, byte-for-byte copies of the OpenGrid touring-car pages with their frontmatter It preserves provenance exactly: kb/PROVENANCE.md lists each page with its sha256 and original date. It stays readable by your friend and a judge, and a corrected fact is a diff to a text file.
** …
You

yes, import the snapshot while the model downloads

Agent

The snapshot is imported and verified, but nothing is committed yet. 11 pages are frozen in kb/source/ and compile to 106 retrieval chunks (about 7,500 words, median 68 words per chunk, no duplicate ids).

What's in place
- Scope: the 8 touring-car pages, vehicle-dynamics fundamentals, the BD12 page and the lap-time statistics page. I left out the ESC, motor, telemetry and logger pages and the raw/ extractions.
- Frozen: files are byte-for-byte copies. kb/snapshot-manifest.json has a sha256 for each. npm run kb:verify fails if one changes, and I tested that it catches a tampered file.
- Build: npm run kb:build compil …

You

yes, make the two commits

Agent

Both commits are made, and the working tree is clean:

c0c2d08 feat(kb): import frozen, credited snapshot of prior touring-car notes
c4bd117 feat: phone feasibility spike, switch to laptop + Ollama
657599b chore: initial monorepo structure and RC Pit Companion scaffold
  • c4bd117 has the phone spike (page, self-hosted WASM script, docs/PHONE-SPIKE.md), the Ollama decision, and the doc and config updates. It contains nothing from kb/.
  • c0c2d08 has only the 11 frozen pages, the sha256 manifest, kb/PROVENANCE.md, kb/additions/, and the two scripts (kb:verify, kb:build). History now shows the credited prior knowle …
You

maintain a challenge memory that saves relevant ideas/findings that should be included in the final report, so these don't get lost throughout implementation. Set it as a rule in the main AGENTS.md

Agent

gemma4:e4b-it-qat classified 10 of 10 test utterances correctly. That result is much weaker evidence than the number suggests. I logged it in the challenge memory with its caveats.

Baseline on this Mac (M4 Max, 64 GB, Ollama 0.35.1, temperature 0, Ollama JSON-schema output, model warmed up first)

Measure Result
Correct symptom id 10/10
Median time to first token 138 ms (first call 379 ms)
Median decode speed 85.8 tokens/s
Memory in use 3.7 GB, 100% on GPU

Why not to quote 100% yet
- Small and easy: there are only 10 cases, written in plain wording, not the friend's real phrasing.
- …

You

regarding the open items mentioned before:

  • my friend's laptop is a Macbook Pro 2019 - (CPU): 9th-generation Intel Core i7 (6-core 2.6 GHz) and Memory (RAM): 16 GB of 2666 MHz DDR4 memory
  • Friend's quote: "this app will be very useful for practice days so I can quickly store my settings and get feedback from what works, what might have affected my run"
Agent

Both items are logged. The laptop is the bigger of the two, because it changes the model choice and the latency budget.

Quote: it's in the challenge memory as a QUOTE entry. It stresses that the setup vault and "what might have affected my run" matter as much as the symptom coach, and that storing settings has to be quick. Does he agree to be quoted in the public post, and do you want him named? I left that as a TODO-REPORT.

His laptop
- No GPU acceleration: per Ollama's docs, Intel Macs run CPU-only and need macOS 14 (Sonoma) or newer. I don't yet know his macOS version or GPU. If he is on something older, Ollama won't inst …

You

I forgot to mention that my friend's Macbook has a AMD Radeon Pro 5300M (4GB GDDR6)alongside Intel UHD Graphics 630

Right now I cannot reach to him. Let's leave the benchmarking for later

Agent

I've recorded the GPU and the deferred benchmark in the challenge memory and in the challenge AGENTS.md.

What the GPU means for the plan
- No change to the plan: Ollama's docs only promise CPU inference on Intel Macs. Any AMD GPU speed-up would be an unverified bonus, and 4 GB of video memory couldn't hold the E4B model anyway.
- Which model: I'm still planning on Gemma E2B QAT for his machine.
- Which MacBook: the 5300M suggests the 16-inch 2019 model. That is my inference, not confirmed, and that model should be able to run macOS 14 or newer, so Ollama should be installable. Please ask him for his macOS version when you n …

You

let's spin up a Opus 5.5 project planner agent that reviews the current state of the challenge idea and code and build a full implementation plan, decides tooling, libraries to use, UI setup settings, settings interactive visual explanation views, etc. This final plan should be then used for implementation (sequencially or split to various subagents (Sonnet 5.5 or Opus for harder features like visual interactive explainers)

Agent

The plan is written and logged in the challenge memory. It is docs/IMPLEMENTATION-PLAN.md, about 17,000 words including the JSON appendices. The planner read everything I pointed it to and did not open your earlier 3D guide. I read most of it myself, but I haven't re-verified its web-sourced claims (voice APIs, package versions, speech-pack support).

Headline decisions
- Shape: the planner proposes a local Node server (Hono) plus a Preact UI in Chrome on localhost, with Ollama on the same machine. It drops the PWA and service worker, since they add stale-cache bugs and no value here.
- Mastra: it ru …

You

here's my feedback:

  • commit the pending work before starting with implementation
  • Keep entire partner tool out due to workload. We'll use Devrelay agent session tool when writing the submission article

Regarding data questions:

  • Keep the generic droop gauge
  • explain the contradictions found regarding diff oil and rear traction so I can decide which one is right and use it as a lever

After clearing the above, let's review the lever rows you mention

Agent

Everything is committed and your Entire and droop decisions are recorded. The diff oil question is below, for your decision. The data question: I found three statements about diff oil in your notes, not two, and they only clash on low grip.

Done
- The pending work is committed in three commits, plus one that logs the diff-oil finding. The tree is clean. I also replaced one local path in the plan before committing.
- Entire is dropped: the plan, the memory file and the partner list now say so. Categories are Gemma, plus Mastra only if the coach workflow actually ships on it. The session is shared through DevRelay's agent-session tool …

You

here's the answer from notebooklm:


<pasted_content id="1efc">
### Comparison: Thinner vs. Thicker Rear Gear Diff Oil on Asphalt

The handling effects of rear gear differential oil viscosity depend heavily on track traction levels, affecting on-power side grip, exit acceleration, and corner rotation:

| Performance Metric | Thinner / Softer Diff Oil (e.g., 3,000 cSt) | Thicker / Harder Diff Oil (e.g., 7,000–10,000 cSt) |
| :--- | :--- | :--- |
| **On-Power Rear Side Grip** | **Increases si …
Agent

Everything is committed and the tree is clean. Here are the rows to review: the first three for each of the four demo symptoms (13 rows in all). That is the part the demo actually shows, and the remaining rows rarely surface because the engine offers only a primary change and two alternatives.

exit-oversteer (loose on power out of corners)

# Change Effect Trade-off Check
1 More rear toe-in, +0.5° (2.0–3.5°) More rear traction on power Entry response vs straight-line stability The BD12 factory value is already 3.5°, the top of the range. The engine will skip this row on a stock BD …
You

approve all 13 rows, architecture approved, start T0

Agent
You

continue where you left off

Agent

The app is built and committed, about 7.5 hours before the deadline (it's 23:30 UTC). T10 is committed as 9d9f8ea, and I checked it myself, including a clean-clone test from the commit.

Clean-clone test
- What I ran: a fresh clone, npm ci (3 seconds), then npm start.
- What worked: the knowledge index built automatically, the server started on the Mastra engine, and /api/health, /api/meta and the built UI all answered. A chip turn and a typed turn both completed.
- What T10 had to fix first: before its changes, a fresh clone failed to build and then crashed on the first coach turn, because the generated knowledge index wasn't in the repo. The install and start commands now build it, and the server exits with one clear line if it is missing.
- Other checks: whole-repo typecheck is clean, 92 tests pass, and the build succeeds.

Also in T10
- Out-of-range values: a current value outside the notes' range is now flagged for confirmation. It is no longer clamped into a huge jump.
- UI and scene: the caster view frames the whole car, scene buttons are 48 px, and OPEN SETUP lands on the right Setup row.
- Demo data: npm run seed:demo loads clearly labelled DEMO data.
- The runs show one within-noise comparison and one above-noise comparison.
- Race day ranks the demo setups 100 and 38, with the carpet setup under "Other surface".
- I deleted the dev database: it held only an agent test session, so the demo starts empty.

A real accuracy signal, and it's less flattering than 100%.
- Setup: the 100% score is on cases we wrote. I tested 14 fresh phrasings, with the expected answers written before I ran them.
- Result: 13 of 14 were right. The miss was "the car pushes wide on the way into the hairpin", labelled mid-understeer. I'd read it as entry-understeer, and it reproduced.

You

how much time left we have before the deadline. Make sure you use my local timezone

Agent

About 4 hours 55 minutes left. It's Sun Oct 4, 21:04 local time (UTC−5). The deadline is Mon Oct 5, 01:59 AM your time, which is 06:59 UTC.

My earlier "≈7.5 h" figure came from a 23:30 UTC clock reading that was off. The system clock now reads 02:04 UTC, so the real margin is about 2.5 hours shorter than I said.

You

what about the Reword the five eval cases you mentioned before?

Agent

Those five cases are in the 25-case classifier test, data/eval/symptom-utterances.json. I wrote their wording, so they test my phrasing and not your friend's. The result is only worth quoting if the sentences sound like him.

You

here are my friend's own wording:

  • it's understeering out of the corner
  • once i hit the breaks into the corner the car oversteers
  • the track is very slippery
  • the car pushing power goes down in the last few laps
  • steering feels deviated to one side
Agent

Your friend's five phrases scored 4 of 5, and I've committed the result (559af20).

Case His words Result
c11 "it's understeering out of the corner" Missed: I labelled it entry-understeer and the model said exit-understeer
c12 "once i hit the breaks into the corner the car oversteers" ok (entry-oversteer)
c16 "the track is very slippery" ok (low-grip)
c17 "the car pushing power goes down in the last few laps" ok (fade-late-run)
c18 "steering feels deviated to one side" ok (left-right-difference)
  • The miss is partly my label's fault. I wrote c11 for "pushing into the hairpin" (corner entry). "Out of the corner" most naturally means corner exit, so exit-understeer is a defensible answer.
You

continue

Agent

The README and the first full draft of the DEV post are both committed, and the humanizer pass is done. I haven't staged anything on DEV. It's 22:24 your time, so about 3 h 35 min remain (deadline 01:59).

My git add -A slip: it also committed the humanizer skill folder, which you or another process had installed, plus skills-lock.json. I didn't intend that. It's MIT-licensed, nothing is pushed, and I can split it into its own commit if you'd rather.

You

I added the repo remote url to git. Push it so I can run it on my friend's laptop.

Agent

The repo is pushed to main and public: https://github.com/santteegt/HacktoberFest2026. It includes one extra commit, 96f0c84, which adds the clone step to the README.

I checked what went up. There are no keys or tokens, ideas/ and your .env and database files are not tracked.

You

while I'm installing the project on my friend's laptop, stage the draft on DEV as unpublished, work on an HTML artifact that can help users getting started to use the app.

Regarding post submission, are you including screenshots on how the app works?

Agent

All three are done. It's 22:52 your time, so about 3 h 07 min remain until 01:59.

1. Draft staged on DEV (unpublished). It's article 4798987, tagged devchallenge, weekendchallenge and hf26challenge, and marked some_ai.

3. Screenshots: yes, the post now has three. They show the Coach screen, the Session log, and Race day. They come from the app running on its labelled demo data, and the post says so.

You

update the submission post to include a bit more of the visual artifacts that explain the setup settings

Agent

The staged DEV draft (article 4798987) now has a section with four 3D explainer frames, and it's still unpublished. The new section explains that a number like "rear shock position 2 to 3" doesn't mean much until you can see it. The app plays each suggestion on a small schematic car, with the faint copy as the setting before and blue as after.

You

here are the model performance on my friend's laptop:

Prompt: how should I run cli commands on a macbook

Results

total duration: 2m39s, prompt eval rate 38.69 tokens/s, eval count 1462 tokens, eval rate 9.21 tokens/s
Agent

Your friend's laptop gave the first real reading. His laptop reads a prompt at about 39 tokens/s and writes at 9.2 tokens/s. That is within the range I predicted yesterday (5 to 15 for writing, 20 to 50 for reading), and about 6× slower at writing than my Mac on CPU alone.

My estimate, not measured: a typed sentence takes about 25 s from end to end. A symptom chip takes about 15 s, which skips one of the two model steps.

You

Here's my feedback:

  • yes, the reading was from gemma4:e2b-it-qat
  • BTW, you have permissions to quote my friend

Before running a demo on my friend's laptop In the app, I cannot use the STT feature, when I click the hold to talk button I get the following message: Voice input is off in Settings. Type or tap a chip. Is this a bug or should I do something to enable this in a macbook

Agent

It wasn't a crash. It was a bad default, and I've fixed it. Hold to talk printed that message because voice input started out set to Off, and nothing told a new user where to turn it on. That made the hands-free feature a dead end on first run.

What I changed (pushed as b73888f, 92 tests pass): voice input now starts on On this computer; the off message now says where to turn it on.

You

when using the app, I can ask anything, the coach agent always recommends "reduce caster"

Agent

Found and fixed. The coach recommended "reduce caster" for everything because your friend's laptop runs the small model, and I had only ever tested the bigger one.

What was wrong
- The two models: he runs gemma4:e2b-it-qat. Every accuracy test I'd run, including the 24 of 25, used the larger gemma4:e4b-it-qat model.
- The reproduction: I pulled E2B to my Mac and tested it. When the small model can't place a sentence, it falls back to the first symptom on its list, and that symptom's first suggestion is "reduce caster". "hello" and "what time is it" both got it.
- The off-topic test: E4B refused all 12 off-topic sentences I tried. E2B refused only 6 of 12.

What I changed: a plain-code check now runs before the model. A sentence with no word about driving or setup gets the "not in my notes" refusal immediately, with no model call. With it E2B scores 23 of 25 on the test cases.

You

fix worked.

Using the shot list you suggested above, update the app screenshots in the submission post to give an image-based demo walkthrough

Agent

The post now has an image walkthrough built from your shot list, and the DEV draft is synced and still unpublished.

I scripted a pass through the real app: a fresh database, the real buttons, and the gemma4:e2b-it-qat model your friend runs. Nothing in the images is edited. Seven steps, one frame each: start a session, say it, tap a chip, apply and rate, save setup, session log, race day.

You

here's my feedback:

  • Yes, video demo shows wifi is off, and voice input/output working

Check if there are any remaining TODOs in the post draft

Agent

There are two TODO comments left in the DEV draft, and I've fixed the wording they were guarding.

What's left
1. Agent session embed (My Agent Session). The post needs a public DevRelay session link.
2. His reaction (end of "What I Built"). The post says he hasn't tried it himself yet.

You
  • the video was recorded on his laptop.
  • My friend was very happy with the product. He cannot wait to use it on the next practice sessions.
  • Upload the agent session. Include it if needed in the submission post and publish it

I built this with Claude Code: one planner agent, then builder and reviewer subagents working on separate parts of the code. The session shows the phone test that failed, the frozen-notes decision, and how I checked the agents' claims, including a fresh-clone test that found the project didn't install cleanly and led to a fix.

This is a curated copy, not the raw log. It keeps every prompt I wrote and the assistant's final reply to each, shortened, and leaves out the subagents' reports, tool output and reasoning. Paths, my email and local IP addresses are scrubbed. It ends where I asked for the upload, so the last few steps (the 3D explainers, the architecture diagram and the Wi-Fi-off demo check) are in the post and the repo's docs/CHALLENGE-MEMORY.md, not in the session.

Prize Categories

  • Gemma: Gemma 4 (E4B QAT on my machine, E2B QAT on his) is the model behind both language steps, served by Ollama.
  • Mastra: Mastra 1.74 runs the coach workflow, including the two pause points.

I looked at the other partners and left them out. Most would have meant a cloud service in a project that is meant to run on one person's laptop.

Top comments (0)