This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
My friend is a software engineer applying for backend roles. She's switching after a long time, so interviews feel new again. She has the skills. What makes her hesitant is the conversation itself: saying what she knows out loud, to a stranger, while they listen. She's been practicing on her own, and practicing interviews alone doesn't really work.
You read a question, think of an answer in your head, and move on. You never say it out loud, nobody stops you halfway, and nobody asks "okay, but what did you do there?" Asking a friend to play interviewer works once or twice, then it gets awkward for both of you.
So I built her Interview Room: a mock interviewer that actually talks.
- You pick a track (Frontend, Backend or Machine learning), a round (behavioural, technical, HR, or a full loop of five questions), your level, how long your answers should be, and an interviewer: Ananya or Aarav, friendly, neutral or tough.
- It asks out loud. You answer out loud, and you can pause to think.
- It listens to what you said and asks one follow-up on it. If you said "we" the whole time, it asks what you did yourself. If you never said how it ended, it asks.
- At the end you review your own answers, one at a time, next to what a strong answer usually covers. Then you get a report built from your own marks and from what was measured: how long you spoke, your pace, your filler phrases, how often you said "I" vs "we". There are no scores. The app never decides if your answer was good. You do.
You can do it as a video interview with a 3D interviewer who lip syncs, or as a phone screen with just a voice, like a recruiter call. If you want it to feel like a real video call, you can turn your own camera on: a small mirror in the corner, off by default, never recorded or sent anywhere.
Everything runs in the browser. Hearing when you talk, turning your speech into text, the interviewer, its voice and the avatar all run on your own laptop. What you say is never uploaded. The network is only used to download the models the first time (they stay on your laptop after that), plus a few runtime files and a font.
Demo
Live: https://interview-room-iooj.onrender.com/
Before you click:
- It needs a laptop or desktop with WebGPU, so a recent Chrome or Edge.
- The first visit downloads about 1.6 GB. On my connection that took 4 to 8 minutes. After that it's ready in 2 to 3 seconds.
- The home page checks your device first and shows what will download and how big it is, before anything starts.
Code
927tanmay
/
interview-room
A spoken mock interviewer that runs entirely in your browser: Whisper, Gemma, Kokoro and a lip-synced avatar, all on-device.
How I Built It
The pieces
| Job | Model | Size |
|---|---|---|
| Hears when you start and stop talking | Silero VAD | 1.8 MB |
| Turns your speech into text | Whisper base | 295 MB |
| The interviewer | Gemma 3 1B (q4) | 880 MB |
| The interviewer's voice | Kokoro 82M | 326 MB |
| Runs it all | ONNX Runtime Web, on WebGPU | 67 MB |
All open-weight, all in the browser through transformers.js. The microphone, the voice, lip sync and interrupting the interviewer come from react-ai-voice-avatar, an npm package I maintain. Interview Room itself is new. I started it on Saturday.
Gemma is the interviewer, but it's not in charge
My first idea was to give Gemma the conversation and let it run the interview. With a 1B model that didn't work. In my first tests it said "That's a good start" before nearly every question, wrapped its lines in quote marks, sometimes answered as the candidate, and once invented a "15% increase in session duration" that nobody had mentioned.
So I split the job three ways:
- Plain code decides what to ask about. Rules read the answer and pick an angle: what you did yourself if you only said "we", the outcome if you never gave one, specifics if it was vague, a missing key point, an edge case.
- Gemma only words it. It gets the conversation as real turns, plus a short note like "ask what they did personally".
- A guard checks every line before it's spoken. It strips quotes and markdown, drops praise, drops anything copied from the question, and keeps exactly one question. If nothing usable is left, or Gemma takes more than 6 seconds, the interviewer says a written line for that angle instead.
With that, 8 of 12 test follow-ups were usable straight from Gemma and the other 4 fell back cleanly, so you never hear a broken line. Each takes about a second. One real follow-up, after an answer that only said "we":
What specific action did you take to improve the dashboard speed?
Two Gemmas, side by side
Before building anything, I wrote a small eval page that runs the same prompts on any model in the browser, and ran it on Gemma 3 1B and Gemma 4 E2B on my M4 MacBook:
| Gemma 3 1B | Gemma 4 E2B | |
|---|---|---|
| Download | 859 MB | 3.11 GB |
| First load, download included | 4 min | 14 min |
| Follow-ups on the right angle | 8 of 12 | 12 of 12 |
| Time per follow-up | 1.0 to 1.5 s | 1.0 to 1.6 s |
| Code review (2 cases) | 0 right | 2 right |
| Judging a STAR answer (2 cases) | 0 right | 1 right, by calling everything present |
Two things surprised me. Gemma 4 E2B was as fast as the 1B, because its replies are shorter. And it loads with the plain text-generation pipeline in transformers.js, which only pulls the text parts of the model. The vision and audio encoders are never downloaded.
The table also shaped the whole app. Neither model could tell me whether a STAR answer was complete. If a model can't judge that reliably, I shouldn't hand my friend a score from it. That's where the "no scores" rule came from.
The report: you judge, the app measures
Instead of a model grading answers, the report has two halves.
Your own marks. When the interview ends you go through your answers one at a time. First, a quick "how did that one feel?": good, okay or rough. Then the points a strong answer usually covers appear next to your own words, and you mark each one covered, partly or missed. Every question in the bank has 3 to 5 of these points, 206 in total, each tagged as a result, your own role, an example, a trade-off and so on. It takes a minute or two, and works with the keyboard.
What was measured. Plain code, no model, so it's instant:
| Number | How |
|---|---|
| Answer length | Your first word to your last, from the voice detector, across thinking pauses |
| Pace | Words heard divided by the time you were actually talking, next to a rough 120 to 160 a minute |
| Time to first word | From the end of the question to when you started |
| Long pauses | Gaps over 3 seconds |
| Filler phrases | "you know", "I mean", "basically", "kind of", "sort of", "like" next to a comma |
| "I" vs "we" | Counted |
| Numbers | Figures and spoken numbers. Vague ones like "one or two" don't count |
Then simple rules put the two together. At the top you get up to three things to work on, and each says where it came from: "you left out the result in 3 of 4 stories", or "you felt good about this one but marked most points missed". Where your marks and the numbers disagree, it says so gently. If you marked your own role as covered but said "we" twelve times and "I" twice, it asks you to take another look. Every answer ends with a short Next time list made from the points you missed, and a Practice this one again button.
One thing I had to be honest about: Whisper is trained to leave out "um" and "uh", so the app can't count them. It counts the filler phrases Whisper does keep, highlights every one in your transcript so you can see if it got one wrong, and says plainly that "um" and "uh" aren't counted.
Letting people think
Real answers have long pauses. A voice assistant tuned for chat would cut you off mid-thought, which is the fastest way to make this feel fake.
- A pause doesn't end your answer. It waits for about 5 seconds of quiet, or for you to press I'm done.
- If you say nothing at all, it nudges you after 12 seconds and moves on after 25.
- You can talk to it: "can you repeat that?", "what do you mean?", "I don't know", "can we skip this one?", "give me a minute". It handles those out loud instead of counting them as your answer.
Keeping 1.6 GB in a browser
transformers.js caches models in the browser's Cache API, but Chrome won't store a single file of 256 MB or more there, so the 859 MB Gemma file was downloaded again on every visit. My package already kept its own models in the Origin Private File System, so I exported that cache and used it for Gemma as well. Now a second visit is ready in 2 to 3 seconds.
Fixing my own package along the way
Building a real app on my own package found two bugs in it. Whisper was called without chunking, so if you spoke for more than 30 seconds without a pause, the rest was cut off. And if the app's reply was empty, the hook got stuck on "thinking" and never listened again. For an interview app both are deal breakers. I fixed them in the package, added the speaking time to onSubmit for the pace, exported the model cache, and released it as 0.7.0 during the weekend.
Hosting on Render
It's a free static site on Render. There's no server because there's nothing to run on one. What matters is the cross-origin isolation headers, which ONNX Runtime needs for multithreaded WebAssembly. They live in render.yaml next to the code, and every push to main deploys.
What I got wrong
- Letting Gemma run the interview. Covered above. A small model is a good writer and a bad judge, so I gave it the writing and kept the judging away from it.
- Trusting Whisper's transcript for fillers. My first filler count was confidently low, because "um" and "uh" were never in the text.
- The first real run. Speaking to it for real showed things no test had: a spoken "skip" didn't skip, "sorry, I don't know" wasn't understood, some follow-ups drifted off the point they were meant to ask about, "like" was counted as a filler when it wasn't one, and "one or two" was counted as a number. All fixed before submitting.
- Downloads fail sometimes. At 1 am, testing the live site, Whisper and Kokoro lost their connection halfway through. The page sat at "Loading models 21%" with no message, and the voice quietly switched to a backup. Now it says which model failed and offers Try again, which only fetches what's missing.
What I didn't do
- No scores, on purpose. Not out of 10, not a percentage, not a tick grid.
- No STAR judging. Neither model got it right reliably.
- Heavy mode isn't there yet. Gemma 4 E2B gave better follow-ups in my tests and reviewed code correctly, but running it alongside Whisper, Kokoro and the avatar needs more memory and more testing than I had time for. It shows as "coming soon".
- System design questions are written but parked. A system design round is a conversation, not a question and an answer, and it deserves its own design.
What my friend said
I sent her the link while I was still building. She ran it on her own laptop, start to finish: everything local, no hiccups, and nothing to pay for.
Her reaction was really positive. What she liked:
- It's good for practicing, and it helps with the anxiety. Saying your answers out loud to an interviewer, before the real one, takes some of the fear out of it.
- She saw the gaps in her answers straight away. Her favourite part of the report was the side-by-side view of what she covered and what she missed.
- The level of the questions felt right to her.
- A proper mock interview, free.
She also gave me a list, which is the most useful thing a friend can do. Two of the four were small enough to add before submitting:
| She asked for | What I think |
|---|---|
| More than one follow-up per question | Fair. Real interviewers dig twice when the first answer is thin. The engine already picks the angle, so a second round is mostly a rule change. |
| The report should say what to improve, not only what was missed | She was right, and this one was quick. I added it before submitting: every answer card now ends with a short "Next time" list, built from the points she marked missed or partly, each turned into something to do. |
| Adding her own questions | Her idea: upload her own list of questions and answers, and practice them in random order. That turns it from a demo into something she'd use every day. |
| Her camera on | Seeing yourself makes it feel like a real call. Added before submitting: a "show my camera" switch on the interview screen, off by default. It's a mirror in the corner, it only asks for the camera when you turn it on, and like everything else it stays on your laptop: never recorded, never sent. |
Why Does Open Innovation Matter?
A mock interview is private. You stumble, you say the wrong thing, you talk about your current job and why you want to leave. I didn't want any of that going to a server, and with open-weight models it doesn't have to. Whisper, Gemma and Kokoro run on your own laptop, and the code is open, so you can check that nothing else leaves it.
Open models also meant I could test them myself. I ran the same eval on two Gemma models, saw exactly where each one failed, and built the app around that. With a hosted API I'd be testing a model that could change under me next week.
And it keeps working. There's no API key, no account and no bill that stops the demo when the challenge ends. It's a static page and some open models, so it's just a link my friend can keep using. It stands on open pieces (transformers.js, ONNX Runtime, Silero, Whisper, Kokoro, Gemma, three.js), and the app and my package are open too, so anyone can take it further.
My Agent Session
I did the planning and the first model tests in a private chat with Claude, because it has my personal notes in it. What came out of it is all in the repo: docs/PLAN.md, docs/MODEL-TESTS.md and the first two commits.
The build itself was one Claude Code session, recorded with Entire: more than 30 checkpoints, each tied to a commit, with the prompts and reasoning behind it. I set the direction and the rules, Claude Code planned and wrote the code, and every step stopped for me to check.
The full session on Entire (the repo on Entire)
Three checkpoints worth opening:
- Testing Gemma 4 E2B against Gemma 3 1B, and choosing the interviewer
- Building the report: things to work on, second looks, charts, answer cards
- Her feedback, shipped: the "Next time" list on every answer
One tip for other Entire users: Claude Code includes your account email in the session context, and Entire's email redaction is off by default. Turn it on in .entire/settings.json before you push to a public repo.
Prize Categories
- Best Use of Gemma: Gemma 3 1B is the live interviewer, running in the browser on WebGPU. It words every follow-up, with rules choosing the angle and a guard checking the line. I tested Gemma 4 E2B against it on the same eval before choosing.
-
Best Use of Render: Render hosts the front end as a static site, with the cross-origin isolation headers the in-browser models need set in
render.yaml. Every push deploys. - Best Use of Entire: the whole build is one recorded session, more than 30 checkpoints, each tied to its commit and public on Entire and in the repo. Linked above.
What's next
Her list comes first:
- Your own questions and answers, uploaded and practiced in random order.
- A second follow-up when the first answer is still thin.
Then mine:
- Heavy mode with Gemma 4 E2B, and a deeper review on the device: code review and possible gaps, as suggestions, never scores.
- A small embedding model that, after you mark a point, highlights the sentence in your answer that comes closest, so you can check your own marks.
- Progress across sessions: pace, fillers and answer length over time, kept on your laptop.
- Hosting the voice detector and ONNX runtime files on the site, so it doesn't depend on a CDN.




Top comments (5)
"Practicing interviews alone doesn't really work — you never say it out loud, nobody stops you halfway" — you identified the actual problem, and it isn't knowledge, it's performance under observation. The interruption angle ("okay, but what did you do there?") is what separates a mock interviewer from a flashcard app. Technical question since it's all in the browser: how's the latency on Gemma between her answer and the follow-up? Anything clever you had to do to keep the conversation feeling alive?
Thanks, that's exactly it. Most people I know who struggle in interviews know the material. What's hard is saying it out loud while someone listens and pushes back.
On latency: on my M4, Gemma 3 1B takes about 1 to 1.5 seconds to write a follow-up. With the voice starting up, it's roughly 2 seconds from when she finishes to when the interviewer starts talking. That's about how long a real interviewer pauses anyway, so it doesn't feel laggy.
A few things keep it that fast:
Gemma does very little. Plain code reads the answer and picks the angle (ownership, outcome, specifics...). Gemma only words one short question. A small prompt and a short reply are much faster than asking it to reason about the whole answer.
Whisper works while you're still talking. Each stretch of speech is transcribed when you pause, so when you finish, most of the answer is already text. There's no big transcription wait at the end.
Everything stays loaded. All the models live in Web Workers that stay up from setup to the end of the interview, and Gemma gets a one-token warm-up at load so the first reply isn't cold.
It never stalls. If Gemma takes more than 6 seconds or says something off, a guard throws it away and the interviewer says a written line for that angle. You never sit waiting.
The longest wait is actually on purpose: it waits about 5 seconds of quiet before treating your answer as finished, so it doesn't cut you off while you think. If you're done sooner, there's an "I'm done" button that skips the wait.
That's a really clean division of labor — plain code owns the reasoning (angle selection), the model only does the phrasing. Every token you don't ask the LLM to produce is latency you never pay, and 2s end-to-end with voice is genuinely 'natural pause' territory.
Curious about the angle picker: is it rule-based on the answer's structure (no numbers → ask for specifics, no result → ask for outcome), or something more learned? And did you ever try letting Gemma pick the angle too, then rolled it back for speed?
Some comments may only be visible to logged-in visitors. Sign in to view all comments.