DEV Community

Isha Vishwakarma
Isha Vishwakarma

Posted on

I Fine-Tuned an Open Model to Be My Friend's Patient Interview Coach

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

InterviewBuddy is a patient mock interviewer for my friend Rose, who is preparing for software engineering interviews.

Rose knows her stuff, but in interviews she freezes up and forgets to mention the result of her stories. Practicing with a real person is awkward and hard to schedule, and generic chatbots either give the answer away or just say "Great answer!" to everything.

InterviewBuddy works like a good practice partner:

  • Asks one question at a time. It covers behavioral, technical, and system design.
  • Grades every answer in the same 4-part format:
    • ✅ What worked: specific things they did well
    • 🔧 To improve: 1–2 concrete fixes
    • 💡 Stronger version: their own story, retold better (not a made-up one)
    • ➡️ Next question
  • Handles the human moments. If they say "I'm blanking," it gives a hint instead of the answer. If they apologize for being bad at this, it reassures them and keeps going. If they ramble, it teaches them to trim.
  • Remembers them between sessions. When they click End session, it writes a short summary of their strengths and weak spots. Next time, it opens with those notes and picks questions that target the weak spots ("Last time you skipped the result. Nice job including it today.").

Demo

Screenshots from app
Landing page showing summary from previous session using backboard memory

landing page

Say Hi to start. The first question is picked using notes from the previous session.

start prompt

Get Curated feedbacks on answers using fine tuned Qwen3-8B model

answer

Get complete session summary once you end current session

end session

Code

GitHub logo ishavishwakarma29 / Interview-Buddy

Hacktober fest week 1 challenge

🎤 InterviewBuddy

A patient mock interviewer, fine-tuned on an open-weight model, that remembers what you struggle with.

Built for the Hacktoberfest Weekend Challenge: Build for a Friend. I made it for a friend preparing for software engineering interviews.

InterviewBuddy asks one question at a time and grades every answer in the same format:

What worked:      specific things you did well
To improve:       1–2 concrete fixes
Stronger version: your own story, retold better
Next question

When you're stuck, it gives a hint instead of the answer. When you're nervous, it slows down. When you end a session, it writes a summary of your strengths and weak spots. Next time, it uses that summary to pick questions that target your weak spots.


How it works

Browser (chat page)
  ├─ Session start → Backboard: load notes from past sessions
  ├─ Every answer  → Qwen3-8B + my LoRA adapter (Tinker): feedback +
…

How I Built It

The interviewer is an open-weight model I fine-tuned myself. Here's how the pieces fit together:

Friend opens the web page
  ├─ Session start  → Backboard: load their past-session notes
  ├─ Every answer   → my fine-tuned open model (Tinker): feedback + next question
  └─ End session    → the same model writes a summary → Backboard: save it
Enter fullscreen mode Exit fullscreen mode

1. The dataset (where most of the work went). I created 105 multi-turn mock-interview conversations (drafted with an AI agent, then reviewed and validated: 95 for training, 10 held out for testing) across four categories:

Category Conversations What it teaches
Behavioral 30 Coaching the STAR format and putting numbers on results
Technical concepts 25 Gently catching factual mistakes (e.g. hash map complexity)
System design 20 Asking clarifying questions before jumping to a database
Hard moments 30 "I don't know", nervousness, rambling, "just tell me the answer"

That adds up to 153 graded feedback turns. 27 conversations include "Notes from past sessions," which teaches the model to use memory instead of ignoring it. A validation script checks every line (correct format, alternating roles, exact system prompt) before training.

2. Fine-tuning on Tinker.

I trained a LoRA adapter (rank 32) on top of Qwen3-8B for 5 epochs. The loss is calculated only on the interviewer's turns, so the model learns to be the interviewer, not to imitate the candidate. I watched the held-out test loss after every epoch to catch memorization. The whole run was about 150K training tokens, which cost a small fraction of my credits.

3. Memory with Backboard.

Backboard stores what the model learns about my friend across sessions. To keep it cheap, the app calls Backboard only twice per practice session: one read at the start and one write at the end. The fine-tuned model does all the interviewing and writes the summary that gets saved, so Backboard is purely the memory layer.

4. The app.

A small Flask backend and a single-page chat UI.

Bugs I hit along the way (keeping it honest):

  • The model name I picked wasn't available on the server. I wrote a small script that asks Tinker which models are actually supported.
  • Training crashed at the final save because a log folder didn't exist yet. Luckily the weights had already been uploaded, so I didn't lose the run.
  • I got "Event loop is closed" errors on Windows. The async Backboard client was being reused across separate asyncio.run() calls. I fixed it by running every call on one long-lived event loop.

Why Does Open Innovation Matter?

I could change the model itself, not just the prompt.
With a closed API, all I could do is write a longer prompt and hope. With open weights I trained the behavior in: the 4-part feedback format, refusing to give away answers, and the patience when someone is nervous.

I own the result.
The LoRA adapter is mine. As Rose improves, I can add her real sticking points to the dataset and retrain. If I want to, I can download the weights and run the interviewer locally, so practice sessions never leave our machines. No vendor can deprecate the model or change its behavior under us.

It's tailored to one person.
A general-purpose closed model is built to be good for everyone. This one is built to be good for Rose: her interview, her weak spots, and the tone that works for them. Fine-tuning made that cheap enough for a weekend project.

It was cheap.
The whole fine-tune was about 150K tokens on a small open model. That's affordable enough to retrain whenever the dataset improves.

My Agent Session

I built this with an AI coding agent (Claude Code). It helped design the architecture, draft and validate the training data, write the training and app code, and debug the Windows event-loop crash.

Prize Categories

Best Use of Tinker

Tinker is the core of the project: the interviewer is the model I fine-tuned on it.

  • What I trained: a LoRA adapter (rank 32) on Qwen3-8B model, using 95 hand-checked mock-interview conversations. Another 10 were held back to test the model on conversations it had never seen.
  • How I trained it: I wrote my own training loop with Tinker's Python SDK (forward_backward + optim_step) instead of a one-click recipe. That let me control:
    • What the model learns from. Loss counts only on the interviewer's turns, so the model learns to coach, not to imitate a nervous candidate.
    • When to stop. I measured held-out test loss after each of the 3 epochs to catch the model memorizing examples instead of learning the style.
    • Learning rate. It starts at Tinker's recommended rate for this model size and decreases linearly over training.
  • What fine-tuning bought me that prompting couldn't: a consistent 4-part feedback format (What worked / To improve / Stronger version / Next question), refusing to give away answers before the candidate tries, hints instead of answers when they're stuck, and staying patient when they're nervous. These behaviors are in the weights, not in a fragile 2-page prompt.
  • Serving: the same fine-tuned checkpoint answers every turn in the live app through Tinker's sampling client. It also writes the end-of-session summary that becomes my friend's memory.
  • Cost: the full training run was about 150K tokens, a small slice of my 10 credits. That's cheap enough to retrain as my friend's real weak spots show up.

Best Use of Backboard

Backboard is what makes InterviewBuddy a practice partner instead of a one-off quiz: it remembers my friend between sessions.

  • What it stores: at the end of each session, my fine-tuned model writes a short coaching summary (2 strengths and 1–2 things to practice next). The app saves it with Backboard's add_memory on a dedicated assistant that acts as my friend's memory.
  • How it's used: when a new session starts, the app loads the most recent notes with get_memories and passes them to the model as "Notes from past sessions". The page also shows them to my friend in a "From your last session" box, so they can see their own progress.
  • Why the model actually uses the memory: I didn't just paste notes into a prompt and hope. 27 of the 105 training conversations include past-session notes, with interviewer replies that act on them: picking questions that target known weak spots and calling out improvement ("Last time you skipped the result. Nice job including it today.").
  • Clear split of jobs: the open model does all the thinking and Backboard does all the remembering. Swapping or retraining the model doesn't wipe my friend's history.

Top comments (0)