DEV Community

Cover image for Plumb: I built my friend a tutor for her own codebase, because "the AI writes the code now"
Chestly Ace
Chestly Ace

Posted on

Plumb: I built my friend a tutor for her own codebase, because "the AI writes the code now"

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

A friend of mine, a software engineer, asked me a question I couldn't shake: why should she keep learning when an AI agent can write the code?

I didn't want to answer with a lecture, so I built her an answer instead. Plumb is a tutor for her own codebase. It doesn't write code. It makes sure she stays the person who understands it.

You run tutor in a repo and it opens a chat:

  • Before a change, describe what you're about to build ("let users mark items as favorites"). Plumb reads the real files, explains what the change touches and which patterns to follow, then asks 2–3 questions about the design choices. Each has 2–3 reasonable answers, never trick distractors.
  • "I don't understand" is always an option. Plumb then explains the concept using her own code, and checks with one question.
  • After the change, /review reads the diff and compares it with the decisions she made earlier: followed, changed, or not addressed.
  • /tour walks the repo from an entry point outward, one file per stop.
  • Skip is free. Any question can be skipped. Skipped topics come back later, at a calm moment, never as a nag.

Two things matter more than any feature:

  1. Every claim cites file:lines, and Plumb checks every citation against the real file before she sees it. Wrong files, lines past the end, or lazy ranges like models.py:14-113 are struck through on screen.
  2. Every "why" is labeled. documented means a commit message or doc said so, and Plumb verifies it actually saw that commit. inferred means it's the model's reading. confirmed means she agreed. A guess is never presented as fact.

What she knows, what she skipped and what she decided live in a .tutor/ folder of plain JSON files in her project, which she can open, edit or delete.

Demo

Plumb walking through a change in the FastAPI full-stack template

That's a real, unedited session on the public FastAPI full-stack template. Waits are fast-forwarded. In order:

  • the request is typed;
  • the explanation streams in, with citations highlighted;
  • two citations get rejected as too broad;
  • a design question about the N+1 problem;
  • "I don't understand" on the next one;
  • an explanation built from the repo's own User / UserPublic models;
  • a check question.

This recording used the hosted gemma4:31b-cloud through Ollama. On her laptop it runs on a local Gemma.

Code

GitHub logo chestlyace / Plumb-cli

Plumb-cli is a local-first command-line tool that teaches a developer her own codebase, and checks she understands each architectural decision before and after code is written. It is built for one real person: a software-engineer friend who asked why she should keep learning when an AI agent can write the code.

Plumb

Plumb walking through a change in the FastAPI full-stack template

Plumb is a local-first tutor for your own codebase. Before you build a feature, it walks you through what the change touches in your code and asks you a few questions about the design choices. After you've built it, it reviews what changed against what you decided. It can also give you a guided tour of the repo, one file at a time.

It was built for one real person: a software-engineer friend who asked why she should keep learning when an AI agent can write the code. Plumb's answer let the agent write code if you like, but stay the person who understands it.

  • Runs on your laptop. An open-weight model (Gemma) served by Ollama Your code stays on your machine.
  • Grounded in your files. Every claim cites file:lines, and Plumb checks each citation against the real file before you see it.
  • Honest about the "why".…

She's on Windows, so installing is one line in PowerShell:

irm https://raw.githubusercontent.com/chestlyace/Plumb-cli/main/install.ps1 | iex
Enter fullscreen mode Exit fullscreen mode

Then tutor in any project. The first run is a short guided setup:

  1. It finds Ollama, or offers to install it and starts it.
  2. It reads her RAM and recommends a model: gemma4:e2b (4.6 GB) on an 8 GB laptop, gemma4:e4b (6.6 GB) with more. Then it downloads her pick with a progress bar.
  3. It lets her choose a cloud fallback, or none. It says plainly that a cloud model means her files leave the laptop while it's in use, and it walks her through signing in to Ollama if she picks one.

(macOS and Linux get the same thing with curl … install.sh | sh.)

How I Built It

Gemma 4 via Ollama. Her laptop has no GPU, so Plumb runs a small Gemma 4 on the CPU and picks the size from her memory. The model is one setting. If the local model fails, Plumb switches to the fallback she chose (Gemma 4 31B on Ollama's cloud, by default) for the rest of the session, and tells her so.

Code does the heavy lifting, and each model call gets one small job. A small model gets a narrow context and a single task:

  • a repo map: tree-sitter symbols, an import graph across Python, JS/TS, Dart and a dozen other languages, git history and entry points, cached by file hash;
  • six read-only tools (read_file, grep, list_dir, git_log, repo_map, memory) that refuse anything outside the repo, ignored files and .env;
  • narrow calls: explore and explain, write questions, explain a concept, and a tiny classifier that tells "describe a change" apart from "ask a question".

PydanticAI as the agent loop, behind my own driver interface, with guardrails I wrote around it:

  • a step limit;
  • one retry on invalid output;
  • loop detection, so the model can't re-read the same file forever;
  • a scripted fake model, so the 187 tests run without Ollama.

Textual for the chat UI. The engine never imports it, and --plain gives the same chat as plain text.

Things Gemma taught me

  • Gemma answered in prose instead of calling an output tool. With PydanticAI's default tool-based structured output, the model kept writing the question as nicely formatted prose and only produced valid output on its retry. Switching to prompted JSON output gave 3 out of 3 with no retries. That one change is what makes the questions reliable.
  • It cites lines confidently and imprecisely. In review mode, citations into a diff were off by one or two lines, because a diff only gives a starting line per hunk. Numbering every diff line with its line in the current file fixed that. Now the review points at exactly the right lines.
  • It cites whole files when it's lazy, so citations wider than 60 lines are rejected. A tutor that says "see lines 1-400" isn't teaching.

Why Does Open Innovation Matter?

  • Her code stays on her laptop. A tool that reads your whole codebase should not need to send it to a server you don't control. With an open-weight model running locally, it doesn't.
  • It costs nothing to run, so she can ask "I don't understand" as many times as she wants. A tutor you hesitate to use because of a bill is a bad tutor.
  • The model is swappable. It's one line in a config file. When a better small open model ships, she switches. I could also test exactly how Gemma behaves (prose instead of tool calls, imprecise line numbers) and design around it, instead of hoping an API keeps behaving the same.
  • Everything is inspectable. The model, the agent loop, the tools and the memory files are all open. When it tells her something about her code, she can check where it came from. That's the whole point of a tutor.

My Agent Session

Prize Categories

Best Use of Gemma. Gemma 4 is the only model family Plumb uses: a local gemma4:e2b or gemma4:e4b sized to her laptop, and gemma4:31b-cloud as the default fallback. The design is shaped around what a small Gemma can do reliably: narrow calls, prompted JSON, checked citations.

Top comments (0)