This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
Aman is my classmate. We're both in final-year B.Tech, and like everyone in our batch he's preparing for campus placements: SDE-1 roles, which means technical rounds on DSA, DBMS, Operating Systems, Networks and OOP, plus an HR round.
When I asked him how his prep was going, he wasn't sure. He was putting in the hours, but he couldn't tell which topics were actually weak. Nobody was asking him questions and telling him what he'd missed, so he kept revising the topics he was already comfortable with.
The usual answers didn't fit him either:
- He practises late, around 11 PM, when nobody is free to give him a mock interview.
- The hostel Wi-Fi is unreliable, and paid mock-interview platforms need a stable connection (and money).
- Interview answers are personal. A "tell me about yourself" answer shouldn't have to go to someone else's server.
So I built Prep Buddy: a mock-interview coach that runs entirely on his laptop. It asks him a placement question, grades his typed answer against the points a real interviewer listens for, tells him what he missed, and remembers his weak topics so the next question goes after them.
₹0 to run · no API keys · no account · works with Wi-Fi off · his answers never leave his laptop
A practice session works like this:
- Prep Buddy asks a question from a bank of 56 questions across 7 topics (DSA, CS fundamentals, DBMS, OS, CN, OOP, HR), reworded the way a friendly interviewer would ask it.
- Aman types an answer and gets a score out of 10, what went well, what was missing, an outline of a strong answer and the follow-up question an interviewer would likely ask next.
- If he missed something, the next question is picked to target that gap.
- At the end, he gets a summary with his strongest topic, the one to work on next and 3 concrete next steps.
After a few sessions, the dashboard answers the question he couldn't: where should I spend my next ten minutes?
Demo
The Wi-Fi is off in this recording. Every question, grade and summary comes from Gemma running locally on my Apple M2 laptop with 8 GB of RAM.
There's no hosted link, on purpose. Hosting Prep Buddy on a server would break the one promise it makes: your answers stay on your laptop. The repo below runs on any 8 GB+ machine.
The badge in the top-right corner says AI ready only when the database is up, Ollama is running and both models are downloaded. If something is missing, it shows the exact command to fix it (for example ollama pull gemma3:4b) instead of an error page.
What I Measured
Machine: Apple M2, 8 GB
Chat model: Gemma 3 4B
Question bank: 56 questions across 7 topics
| Step | Cold | Warm average |
|---|---|---|
| Next question | 41.15 s | 20.37 s |
| Grade answer | 25.50 s | 14.81 s |
| Session summary | 23.81 s | 15.08 s |
For the golden-answer test, all 5 questions passed the required
Excellent > Average > Wrong ordering.
The prompt-injection test scored 2/10 and was classified as weak.
Code
Prep Buddy
An offline, private mock-interview coach for campus placements, powered by Google Gemma running on your own laptop.
Built for Aman for the DEV Hacktoberfest Weekend Challenge: Build for a Friend.
₹0 to run · no API keys · works with Wi-Fi off · your answers never leave your laptop
What it does
- Asks you placement interview questions from a hand-written bank of 56 questions across 7 topics (DSA, CS fundamentals, DBMS, OS, CN, OOP, HR), in a friendly interviewer tone.
- Grades your typed answer: a score out of 10, strengths, gaps, an outline of a strong answer and a follow-up question.
- Remembers your weak topics and picks the next question to target them (EmbeddingGemma + pgvector).
- Ends each session with a short summary and 3 concrete next steps.
Why local, open-source AI
Every model call goes to Ollama…
The setup is three steps: start Ollama and pull the two models, run docker compose up for the database, then npm run dev. The README has the full commands. The repo also has the PRD, the architecture, the API reference, the AI design notes and the test results in docs/. The code is MIT-licensed.
How I Built It
The open-source AI at the core:
| Piece | What it does in Prep Buddy |
|---|---|
| Gemma 3 4B (open-weight, 3.3 GB) via Ollama | Rewords questions, grades answers, writes the session summary |
| EmbeddingGemma (622 MB, 768-dim vectors) via Ollama | Turns questions and "what you missed" into vectors |
| Mastra (open-source TypeScript agent framework) | Three agents: interviewer, evaluator, coach |
| PostgreSQL + pgvector | Stores progress and does the "find a question about this gap" search |
The rest of the stack is Next.js 16, React 19, TypeScript, Tailwind, Prisma 7 and Zod.
Browser (localhost:3000)
| fetch JSON
v
Next.js API routes --> services: pick next question, grade, summarise
| |
Mastra agents ollama JS client
(interviewer, evaluator, coach) (embeddings)
| |
v v
Ollama on localhost:11434: gemma3:4b + embeddinggemma
PostgreSQL 17 + pgvector (Docker, 127.0.0.1): profiles, attempts, topic stats, question vectors
Every arrow stays on the laptop. There is no cloud service in this diagram.
Here are the decisions that made a 4B model good enough to trust with someone's interview prep.
1. One job and one JSON schema per call, with no tool calling
Small models get unreliable when you ask them to do several things at once, and Gemma 3 isn't listed as tool-capable on Ollama. So each agent does exactly one thing and must return JSON that matches a Zod schema. Mastra passes the schema to Ollama as its format field, so the decoding itself is constrained to valid JSON. Then I validate it again with Zod anyway:
const res = await agent.generate([{ role: "user", content: prompt }], {
structuredOutput: { schema, errorStrategy: "strict" },
providerOptions: { ollama: { options: { num_predict: 800 } } },
abortSignal: AbortSignal.timeout(60_000),
});
return schema.parse(res.object); // never show half-broken JSON to the user
A failed call is retried once, and then the user sees a clear error. If Ollama isn't running, the API returns 503 MODEL_UNAVAILABLE with the command to start it, so the UI never just crashes.
2. Facts come from the question bank, not the model
The 56 questions were written for this app, not copied from LeetCode or GFG, and each one has 2 to 6 key points (for example, normalization → reduces redundancy, avoids insert/update/delete anomalies, 1NF/2NF/3NF, trade-off: more joins). The model never invents a question. It only rewords one, and the evaluator grades against those key points as a rubric. That's how a 4B model can grade like an interviewer without making up what the "right" answer is.
3. Targeting weak spots with EmbeddingGemma and pgvector
When the evaluator returns gaps like "didn't mention update anomalies", I embed those gaps and ask Postgres for the closest unasked question in the same topic:
SELECT id, prompt, embedding <=> $gaps::vector AS distance
FROM "Question"
WHERE topic = $topic AND id <> ALL($alreadyAsked)
ORDER BY distance
LIMIT 1;
On top of that, there's a plain difficulty ladder. A topic starts at EASY, moves to MEDIUM once his average is 7 or higher, and to HARD at 8.5. In "Auto" mode, each session starts with his weakest focus topic. The vectors and the progress data live in the same database, so I didn't need a separate vector store.
4. Not trusting the model with things code can do
- Arithmetic: the overall session score is computed in code. The model's number is overwritten.
- Consistency: Gemma sometimes said "strong" next to a score of 5, so the UI derives the verdict from the score.
-
Double-clicks: every submit sends an
Idempotency-Key. The attempt insert and the topic-average update run in one database transaction, so an answer can't be graded twice.
Every number here was measured on the laptop itself, not on a cloud GPU. To be honest about it: "next question" misses my 15-second target, because rewording costs an extra model call. Setting REPHRASE_QUESTIONS=false skips it. On a laptop with less RAM, CHAT_MODEL=gemma3:1b is a one-line switch.
Things that went wrong (and what I did about them)
-
The model got stuck in a loop. Sometimes Gemma repeated
0} 0} 0}…until the context filled up, and a request hung for minutes. I tried cappingmaxOutputTokens, and it did nothing. Reading the provider's source showed why: that setting is sent as a field Ollama's/api/chatignores. Ollama's ownnum_predictworks, so I used that, plus a 60-second timeout. -
The interviewer answered its own questions. When asked to reword a question, the model would sometimes answer it instead. Now a reworded question is only used if it still ends with
?. Otherwise Prep Buddy shows the original wording. -
Prompt injection. Answers are wrapped in
<answer>tags and treated as data. While tuning, I found that Gemma 3 4B followed injected text more when my rules came after the answer, so the answer now goes last. I also had to tell it to penalise answers that use the right words but state the facts wrong.
Why Does Open Innovation Matter?
For Aman, it's the difference between a tool he can use and one he can't.
| What Aman needs | With Gemma on his laptop | With a typical closed AI API |
|---|---|---|
| Practise at 11 PM on hostel Wi-Fi | Works fully offline once the models are downloaded | Needs a working connection for every answer |
| Practise every night as a student | ₹0 per answer, so 50 questions cost nothing | Paid per token, or a subscription |
| Keep his answers private | Nothing leaves the laptop: no account, no API key | Every answer goes to a third-party server |
| Run on his machine | Swap gemma3:4b ↔ gemma3:1b with one env var |
Whatever model the vendor serves |
There are two less obvious ways open mattered.
I could read the code when things broke. Both of my worst bugs (the token loop and the ignored output cap) were solved by reading the source of the Ollama provider and checking what was actually being sent. With a closed API, all I'd have seen was a request that took four minutes.
The privacy promise shaped the product. I planned voice input, then dropped it. The browser's speech recognition in Chrome sends audio to a server, and that would have broken the one promise this app makes.
To be fair to closed models: a large hosted model would give richer feedback, and it would give it faster. But Aman can't use that one in the hostel at 11 PM, for free, without sending his answers to someone else's server. A 4B model he can use every night beats a better model he won't.
Handing it to Aman
I gave it to Aman and sat with him while he did a practice session. His verdict:
"What I liked about Prep Buddy was that after answering the questions, I was provided with instant feedback and it helped me to identify my weak areas. It helped me to know about my strengths and weaknesses prior to the interview."
Finding his weak areas before the interview is the part I built it for. He was already putting in the hours. What he didn't have was someone telling him where to put them, and now his laptop does that, with the Wi-Fi off.
Prize Categories
- Best Use of Gemma: Gemma is the product. Gemma 3 4B runs all three agents (rewording, grading against a rubric, coaching), and EmbeddingGemma powers the weak-spot targeting. Both run locally through Ollama, on an 8 GB laptop.
-
Best Use of Mastra: the interviewer, evaluator and coach are Mastra agents. Mastra's
structuredOutputwith Zod schemas is how every model call returns validated, schema-constrained JSON from a local model.
If you're preparing for placements too, clone it, add your own questions to data/questions.json, and practise without needing Wi-Fi.





Top comments (0)