DEV Community

Jacob Zhang
Jacob Zhang

Posted on

3 Hard Lessons from Building a Voice-Based AI Mock Interviewer

I spent the last few months building a voice-based mock interview tool for ML engineers. The idea sounded simple: an AI interviewer that talks to you, asks follow-ups, and scores your answers.

The reality was harder than expected. Here are the three lessons that cost me the most time.

1. Latency is the whole product

In text chat, users forgive a 3-second reply. In voice, a 2-second pause kills the conversation. The interview feels broken.

I used the OpenAI Realtime API, which streams audio with low latency. But the API alone was not enough. I had to:

  • Start generating the reply while the user is still finishing their sentence (careful turn detection).
  • Keep the spoken replies short. Long answers sound robotic and make the user wait.
  • Handle interruptions. Real interviewers get interrupted. The system needs to stop talking and listen.

If you are building anything voice-first, budget half your engineering time for latency. It is not a detail. It is the product.

2. Scoring AI answers needs guardrails, or it lies to you

My first scoring version gave a 7/10 to a 5-second "I don't know." That is worse than no score at all — it teaches the user nothing and destroys trust.

What worked:

  • Hard gate: no score unless the user spoke substantively for at least 60 seconds.
  • A strict rubric: 5/10 is a solid average, 8+ is genuinely excellent. The prompt explicitly forbids inflating scores for vague answers.
  • Five separate dimensions (correctness, depth, clarity, structure, communication) instead of one number. A single score hides everything.

Honest scoring is a feature. Users can tell when they are being flattered.

3. Follow-up questions must come from the conversation, not a script

A fixed list of follow-ups feels like a quiz. A real interviewer listens and probes deeper into what you just said.

I generate follow-ups from the full conversation history. The interviewer tracks the candidate's reasoning and asks about the weakest point. This was the hardest part to get right — it needs enough context to be smart, but not so much that it rambles.

The test I use: if you removed the AI label, would this feel like a human interviewer? If yes, ship it.


These lessons come from building DeepOffer (https://getdeepoffer.com), a voice mock interview platform for MLE interviews. If you are preparing for interviews, the free question bank is a good place to practice. And if you are building voice AI yourself, I would love to compare notes in the comments.

Top comments (0)