DEV Community

Hugo Jose
Hugo Jose

Posted on Originally published at hugoj0s3.dev

Building a Skill Interview with AI

Here's the idea: build an app that checks if someone knows their stuff using an interview. You pick a skill, type your name, and off you go. An AI agent chats with you by voice, asks questions, and at the end you get a report card. Hopefully with more compliments than trauma.

I've seen some applications that evaluate interview participants with AI, so I decided to build one out of curiosity. It was an interesting journey. The trickiest part was calibrating the prompt, especially to end the interview at the right time.

Detailed Requirements

Let's get the requirements first, so we have a clear picture of what to implement, before I start the architecture, the code, drawing boxes and arrows, and pretending I know what I'm doing lol.

The interview

  • The interview is a real-time voice conversation: the agent speaks, the participant answers out loud, and the agent reacts to what was said.
  • The agent opens the interview, asks one question at a time and adapts the difficulty to the answers.
  • The participant can interrupt the agent at any time just by speaking; the agent stops talking and listens.
  • The participant can pause to think for a few seconds without the agent jumping in.
  • While the participant speaks, a live caption shows what is being understood, and the conversation is shown as a transcript.
  • Time matters: each skill has a planned duration. When it passes, the agent tells the participant they have a few extra minutes. When the extra time is over, the agent closes the interview.
  • The agent may close the interview earlier only when it already has enough information to evaluate the participant, or when the participant asks to stop.

The report

  • After the interview, a second AI agent reads the transcript and evaluates the participant.
  • The report includes a score from 1 to N, the label for that level (e.g., Senior), a summary, strong points, and areas to improve.
  • The score is based only on what the participant said, and ignores small speech-recognition mistakes.

Skills are configuration, not code

Every skill is a JSON file. Adding a new skill means adding a file: no code changes. A skill defines:

Field Purpose
Title Name shown to the participant, e.g. "SQL"
InterviewInstruction What the agent should ask and which topics matter most
AgentName, AgentTone Who the agent is and how it behaves (friendly, professional, like a Jedi master...)
AgentVoiceGender, AgentVoiceType The agent's voice (Male/Female; HighPitched, Neutral, Deep, Warm)
Effort How capable the AI model should be: Low, Medium or High
ReportInstruction How to evaluate and what to put in the report
MaxPoints The highest possible score
InterviewDurationInMinutes, ExtraInterviewDurationInMinutes Planned duration and extra time
PointInstructionMap A label and the minimum requirements for every point from 1 to MaxPoints

For example, a Star Wars skill can use fun level names:

{
  "Id": "star-wars",
  "Title": "Star Wars Lore",
  "InterviewInstruction": "Test the participant's knowledge of the Star Wars universe...",
  "AgentTone": "Warm and playful, like a wise old Jedi master.",
  "AgentName": "Master Oren",
  "AgentVoiceGender": "Male",
  "AgentVoiceType": "Deep",
  "Effort": "Low",
  "ReportInstruction": "Evaluate breadth and depth of lore knowledge.",
  "MaxPoints": 4,
  "InterviewDurationInMinutes": 10,
  "ExtraInterviewDurationInMinutes": 2,
  "PointInstructionMap": {
    "1": { "Label": "Youngling", "Requirements": "Knows the main characters." },
    "2": { "Label": "Padawan", "Requirements": "Knows the plot of the main films." },
    "3": { "Label": "Jedi Knight", "Requirements": "Knows the history of the Jedi and the Sith." },
    "4": { "Label": "Jedi Master", "Requirements": "Knows deep lore, including series and books." }
  }
}
Enter fullscreen mode Exit fullscreen mode

If a skill file is invalid (for example, a missing point in PointInstructionMap), the app starts with a clear message saying what to fix.

Keep it simple

This is a sample for an article, so some things are deliberately left out:

  • No login: the participant only enters their name.
  • No database: skills come from JSON files; sessions, transcripts, and reports live in memory and are lost when the app restarts.
  • No audio is stored: only the text transcript is kept.

Swappable AI providers

The interview logic is agnostic and doesn't depend on a specific AI vendor.
The interviewer and the report agent (OpenAI) and the speech-to-text/text-to-speech service (Azure Speech) are each behind an interface, in their own project, so that they can be replaced by different providers.

Abstraction

Let's build the abstractions first: the backend components, and how the interview engine handles everything the participant says so the conversation flows as naturally as possible.

The big picture

The engine is split into two parts:

  • Sessions and skills: load the skill configs, start and finish an interview, store the transcript and the report. Nothing here knows about audio.
  • Realtime: the voice conversation itself: listening, understanding, answering and speaking.

The participant's voice is transcribed and sent to the agent; then the agent writes its reply, and the reply is synthesized into speech with the voice from the config.

Skills Validator — Interview Engine

Every box in the bottom row is an interface, implemented in its own project (SkillsValidator.Engine.OpenAI, SkillsValidator.Engine.AzureSpeech). The engine never references a vendor SDK.

Sessions and skills

Skills are read-only: they come from the JSON files.

public interface ISkillConfigRepository
{
    Task<IReadOnlyList<SkillConfig>> GetAllAsync();
    Task<SkillConfig?> GetAsync(string configId);
    SkillConfigLoadResult GetLoadResult(); // startup validation: the errors to show if a file is invalid
}
Enter fullscreen mode Exit fullscreen mode

An interview is a session. It goes through three states: Running → Finished (evaluation pending) → Reported (the report is ready).

public interface ISessionInterviewService
{
    Task<string> StartInterviewSessionAsync(string participantName, SkillConfig config);
    Task FinishInterviewSessionAsync(string sessionId); // also starts the evaluation
    Task<InterviewSession?> GetSessionAsync(string sessionId);
}
Enter fullscreen mode Exit fullscreen mode

GetSessionAsync returns the whole picture: the participant, the state, the transcript and, once Reported, the result. Behind it, three small repositories (session, transcript, result) keep everything in memory. The transcript is append-only: an interview only ever adds new lines.

When the interview finishes, a second agent evaluates it:

public interface IReportEvaluator
{
    Task<SessionReportResult> EvaluateAsync(
        SkillConfig config, IReadOnlyList<TranscriptEntry> transcript, CancellationToken ct = default);
}
Enter fullscreen mode Exit fullscreen mode

The realtime part

Four interfaces make the voice conversation work.

1. Hearing and speaking: ISpeechConverter turns audio into text and text into audio.

public interface ISpeechConverter
{
    // Microphone audio in, text out: partial text while the participant speaks, final text after a short silence.
    IAsyncEnumerable<TranscriptionSegment> ToTranscriptionAsync(
        IAsyncEnumerable<AudioChunk> speech, CancellationToken ct = default);

    // Text in, the agent's voice out.
    IAsyncEnumerable<AudioChunk> ToSpeechAsync(
        string text, VoiceGender gender, VoiceType type, string? speakingStyle, CancellationToken ct = default);
}

public sealed record AudioChunk(byte[] Data, string Format, int SampleRate);
public sealed record TranscriptionSegment(string Text, bool IsFinal);
Enter fullscreen mode Exit fullscreen mode

Notice that everything is an IAsyncEnumerable: audio and text flow through the system in small pieces. The engine never waits for a whole answer before doing the next step.

2. Thinking: IRealtimeInterviewAgent is the interviewer. Given the skill and the conversation so far, it streams its next reply as text fragments.

public interface IRealtimeInterviewAgent
{
    IAsyncEnumerable<string> RespondAsync(SkillConfig config, InterviewSession session, CancellationToken ct = default);
}
Enter fullscreen mode Exit fullscreen mode

3. The pipe to the browser: IAudioChannel receives the participant's microphone audio and plays the agent's voice.

public interface IAudioChannel
{
    IAsyncEnumerable<AudioChunk> ReadParticipantAudioAsync(CancellationToken ct);
    Task PlayAgentAudioAsync(IAsyncEnumerable<AudioChunk> audio, CancellationToken ct);
    Task StopAgentAudioAsync(); // the participant interrupted
}
Enter fullscreen mode Exit fullscreen mode

The engine doesn't care how the audio travels. In our app it's a WebSocket.

4. The conductor: IRealtimeInterviewRunner connects the other three and runs the conversation until it ends.

public interface IRealtimeInterviewRunner
{
    Task RunAsync(string sessionId, SkillConfig config, IAudioChannel channel,
        Func<string, Task>? onPartialTranscript, // live caption
        CancellationToken ct);
}
Enter fullscreen mode Exit fullscreen mode

How the engine handles each thing the participant says

The runner listens all the time. The speech converter sends it two kinds of events, and each one gets a different reaction:

While the participant is speaking (a partial segment):

  1. The partial text is sent to the browser as a live caption.
  2. If the agent is still talking, the agent stops: its current turn is cancelled, and StopAgentAudioAsync silences the speaker. This is what makes interrupting feel natural.

When the participant has finished (a final segment, after about 2 seconds of silence):

  1. The phrase is appended to the transcript.
  2. A new agent turn starts, in the background, so the engine keeps listening while the agent thinks and speaks:
    • The agent streams its reply as text fragments.
    • A small sentence splitter collects the fragments into complete sentences.
    • Each sentence goes to text-to-speech and then to the audio channel as soon as it's ready, while the agent is still writing the next one.
    • What the agent said is appended to the transcript.

Speaking sentence by sentence is the key to a natural flow. The participant hears the first sentence after roughly a second, instead of waiting for the whole reply to be written and then spoken.

Keeping the agent on track. Two details make the conversation feel less robotic:

  • Pauses to think: the 2-second silence means short pauses don't end the answer. If the answer still sounds unfinished ("hmm, let me think..."), the agent is instructed to say only "Take your time."
  • Time: the model is bad at tracking time, so the server decides the phase (in progress, extra time, time is up). It adds one short time note as the last message of every request, where the model can't miss it.

Ending. When the agent decides to close the interview, it ends its last message with an [END] marker. The runner removes the marker before speaking, finishes the session, and the report evaluation starts.

Implementation

I will not cover the repository part, because it is not the main point of this challenge. The repositories are all in memory, and you can see the whole code on GitHub (the link is at the end of the article).

The session service

The session service is the only business logic of the app. Starting a session just creates it in memory with the state Running. Finishing it is more interesting, because the evaluation can take a few seconds, so I don't make the participant wait for it:

public async Task FinishInterviewSessionAsync(string sessionId)
{
    var session = await sessionRepository.GetAsync(sessionId);
    if (session is null || session.State != SessionState.Running)
        return;

    session.State = SessionState.Finished;
    session.Duration = timeProvider.GetUtcNow() - session.StartedAt;
    await sessionRepository.UpdateAsync(session);

    // The evaluation can take a while, so it runs in the background.
    // The UI polls GetSessionAsync until the state is Reported.
    _ = Task.Run(() => EvaluateAsync(session));
}
Enter fullscreen mode Exit fullscreen mode

The result page shows "Evaluating..." and checks the session every 2 seconds. When the state becomes Reported, the report appears.

This part could be better, but we keep it this way for simplicity. In a real application, the evaluation would go to a queued background job system, so it survives an app restart and can be retried if it fails, and the result would be pushed to the page with SignalR instead of polling.

The interviewer agent

The interviewer runs on OpenAI through Microsoft.Extensions.AI (IChatClient). Its prompt has two parts:

  • Base instructions: while testing, I realized I needed base instructions shared by all skills. Maybe they should be tweaked for each AI provider or model, but after some tests this version works well with OpenAI. The instructions: open the interview, ask one question at a time, keep the answers short because they will be spoken, never reveal the score, and when to close the interview.
  • The skill config: the agent name, the tone and the InterviewInstruction.

The conversation so far is the transcript: what the agent said becomes an assistant message, and what the participant said becomes a user message. Then the reply is streamed back:

public async IAsyncEnumerable<string> RespondAsync(
    SkillConfig config, InterviewSession session, [EnumeratorCancellation] CancellationToken ct = default)
{
    List<ChatMessage> messages = [new(ChatRole.System, InterviewPromptBuilder.Build(config, session))];

    if (session.Transcriptions.Count == 0)
        messages.Add(new ChatMessage(ChatRole.User, "(The participant has joined. Start the interview.)"));

    foreach (var entry in session.Transcriptions)
    {
        var role = entry.Speaker == Speaker.Agent ? ChatRole.Assistant : ChatRole.User;
        messages.Add(new ChatMessage(role, entry.Text));
    }

    // The time note goes last, so the model does not miss it.
    var timeNote = InterviewPromptBuilder.BuildTimeInstruction(config, session, timeProvider.GetUtcNow());
    messages.Add(new ChatMessage(ChatRole.System, timeNote));

    var chatClient = chatClients.Get(config.Effort);
    await foreach (var update in chatClient.GetStreamingResponseAsync(messages, cancellationToken: ct))
    {
        if (!string.IsNullOrEmpty(update.Text))
            yield return update.Text;
    }
}
Enter fullscreen mode Exit fullscreen mode

The time note deserves a comment. While testing, putting the duration and the elapsed time in the system prompt and letting the model do the math didn't work: the agent closed a 20 minute interview in 7 minutes, and at the end it ignored even "Time is up". So the server decides the phase, and the model only gets one short sentence for the current moment:

Phase What the agent is told
In progress "5 of 20 minutes have passed (15 minutes left). Do not mention extra time."
Extra time starts "Start your reply by telling the participant that the planned time is over and they have 3 extra minutes."
Extra time "The participant already knows. Ask at most one more question, then close the interview."
Time is up "Time is up. Thank the participant, say goodbye and close the interview now."

Sending it as the last message, right after the participant's answer, was what made the difference. Inside a long system prompt, the model simply ignored it.

Effort. Each skill chooses how capable the model should be ("Effort": "Low" | "Medium" | "High"), and appsettings.json maps each level to a model:

"OpenAI": {
  "Models": { "Low": "gpt-4o-mini", "Medium": "gpt-4.1-mini", "High": "gpt-4.1" }
}
Enter fullscreen mode Exit fullscreen mode

The effort is set per skill, according to what the skill needs: e.g. a light, casual quiz can use a cheaper and faster model, while a deep technical interview gets a stronger one. chatClients.Get(config.Effort) picks it.

The report agent

The report uses the same model, but instead of streaming text, it asks for structured output. GetResponseAsync<T> generates a JSON schema from the type, OpenAI must answer following that schema, and the result is deserialized back into the type:

var chatClient = chatClients.Get(config.Effort);
var response = await chatClient.GetResponseAsync<ReportResponse>(messages, cancellationToken: ct);

var report = response.Result;
return ReportPromptBuilder.ToResult(config, report.Score, report.Summary, report.StrongPoints, report.ImprovementAreas);

private sealed record ReportResponse(int Score, string Summary, List<string> StrongPoints, List<string> ImprovementAreas);
Enter fullscreen mode Exit fullscreen mode

I don't trust the model with everything: ToResult keeps the score between 1 and MaxPoints, and the label always comes from the skill config (PointInstructionMap[score].Label), never from the model.

Speech with Azure

AzureSpeechConverter implements ISpeechConverter with the Azure Speech SDK.

Speech to text. The SDK doesn't work with IAsyncEnumerable, it works with events: Recognizing while the participant speaks (partial text) and Recognized after a silence (final text). A Channel bridges the events to the stream the engine expects:

var segments = Channel.CreateUnbounded<TranscriptionSegment>();

recognizer.Recognizing += (_, e) =>
    segments.Writer.TryWrite(new TranscriptionSegment(e.Result.Text, IsFinal: false));

recognizer.Recognized += (_, e) =>
{
    if (e.Result.Reason == ResultReason.RecognizedSpeech && !string.IsNullOrWhiteSpace(e.Result.Text))
        segments.Writer.TryWrite(new TranscriptionSegment(e.Result.Text, IsFinal: true));
};

await recognizer.StartContinuousRecognitionAsync();

await foreach (var segment in segments.Reader.ReadAllAsync(ct))
    yield return segment;
Enter fullscreen mode Exit fullscreen mode

The microphone audio is written into the recognizer's push stream in the background while the segments are read. The silence that ends an answer is one setting, SilenceTimeoutMs (2000 ms in appsettings.json). With 1 second the agent cut me while I was thinking, and 2 seconds felt natural.

Text to speech. Creating a new SpeechSynthesizer for every sentence opens a new connection to Azure each time: about 1.2 seconds per sentence in my tests. So there is one synthesizer per voice, created once and reused:

private SpeechSynthesizer GetSynthesizer(string voiceName) =>
    synthesizers.GetOrAdd(voiceName, CreateSynthesizer);
Enter fullscreen mode Exit fullscreen mode

After the first sentence, each one takes about 0.2 seconds. That second saved in every reply is the difference between a conversation and a walkie-talkie.

The runner

The runner is the heart of the engine. It listens to the transcription all the time and reacts to each segment:

// The agent opens the interview.
StartAgentTurn(state);

var participantAudio = channel.ReadParticipantAudioAsync(state.InterviewToken);
var segments = speech.ToTranscriptionAsync(participantAudio, state.InterviewToken);

await foreach (var segment in segments.WithCancellation(state.InterviewToken))
{
    if (segment.IsFinal)
        await OnParticipantFinishedAsync(state, segment.Text, onPartialTranscript);
    else
        await OnParticipantSpeakingAsync(state, segment.Text, onPartialTranscript);
}
Enter fullscreen mode Exit fullscreen mode

While the participant speaks, the caption is updated and the agent stops talking. That's the interruption:

private async Task OnParticipantSpeakingAsync(
    RealtimeInterviewState state, string partialText, Func<string, Task>? onPartialTranscript)
{
    if (onPartialTranscript is not null)
        await onPartialTranscript(partialText);

    if (state.AgentInterrupted)
        return;

    state.AgentInterrupted = true;
    await state.AgentTurnCts!.CancelAsync();
    await state.Channel.StopAgentAudioAsync();
}
Enter fullscreen mode Exit fullscreen mode

Cancelling the agent's turn stops everything at once: the OpenAI stream, the speech synthesis and the audio still queued in the browser. This is why the CancellationToken matters here, and not in the repositories.

When the participant finishes, the phrase goes to the transcript and a new agent turn starts in the background. The turn streams the agent's reply, a small SentenceSplitter cuts it into sentences, and each sentence is spoken as soon as it is complete, while the model is still writing the next one. The participant hears the first sentence after about a second.

One detail: the runner is a singleton, but each interview has its own state (the current agent turn, its cancellation, whether it was interrupted). That state lives in a small class, RealtimeInterviewState, created at the beginning of RunAsync.

The voice over a WebSocket

The voice has its own WebSocket, separate from the SignalR connection Blazor uses for the UI. Blazor's JS interop could carry the audio too, but a plain WebSocket is much easier to follow (and honestly, JS interop always felt a bit weird to me). The browser just sends and receives frames, with a tiny protocol:

Direction Frame Content
browser → server binary Microphone audio, 16-bit PCM at 16 kHz, 100 ms per frame
server → browser binary Agent voice, 16-bit PCM at 24 kHz, one sentence per frame
server → browser text {"type":"caption","text":"..."}
server → browser text {"type":"stop"}: the participant interrupted
server → browser text {"type":"ended"}: the agent closed the interview
server → browser text {"type":"error","message":"..."}

On the server, it's a normal ASP.NET Core endpoint. It accepts the socket, wraps it in a WebSocketAudioChannel (our IAudioChannel) and runs the interview:

app.UseWebSockets();
app.Map("/interview/{sessionId}/voice", HandleAsync);

// inside HandleAsync
using var socket = await context.WebSockets.AcceptWebSocketAsync();
var channel = new WebSocketAudioChannel(socket);

await runner.RunAsync(sessionId, config, channel, channel.SendCaptionAsync, interviewCts.Token);
Enter fullscreen mode Exit fullscreen mode

WebSocketAudioChannel turns binary frames into AudioChunks and events into small JSON messages. One trap: a WebSocket allows only one send at a time, and the agent's voice and the events are sent from different tasks, so the sends go through a SemaphoreSlim.

In the browser, the voice is a web component: a custom HTML tag with its own JavaScript. The Blazor page only renders the tag:

<voice-panel session-id="@SessionId" agent-name="@config.AgentName"></voice-panel>
Enter fullscreen mode Exit fullscreen mode

and voice-panel.js does the rest: the "Start talking" button, the microphone, the socket and the speaker.

this.socket = new WebSocket(`${protocol}//${location.host}/interview/${sessionId}/voice`);
this.socket.binaryType = "arraybuffer";
this.socket.onopen = () => this.startMicrophone();
this.socket.onmessage = event => this.onMessage(event.data);
Enter fullscreen mode Exit fullscreen mode

The microphone goes through an AudioWorklet that converts the audio to 16 kHz PCM and sends a frame every 100 ms. The agent's audio is played gapless, each sentence scheduled right after the previous one. When you leave the page, the tag is removed, the socket closes, and the server stops the interview loop. No JS interop at all.

A nice bonus: in the browser DevTools (Network → WS → voice) you can watch every frame of the conversation.

Final thoughts

If you have built something similar, I would love to hear about your experience.
One thing worth trying is keeping the speech-to-text and text-to-speech on the client side, in JavaScript. I tried it, but it didn't work well, so I decided to move it to the backend.

How to run it

The whole code is on GitHub: github.com/hugoj0s3/SkillsValidator

What you need

  • .NET 9 SDK
  • An OpenAI API key
  • An Azure Speech resource (the free tier is enough): its key and region
  • A microphone, and headphones for the best results (with speakers, the agent can hear itself and stop talking)

Getting the keys

OpenAI

  1. Sign in at platform.openai.com and add some billing credit (the API is paid per use; an interview costs a few cents).
  2. Create a key on the API keys page and copy it (it is shown only once).

Azure Speech

  1. In the Azure portal, create a Speech resource. The free tier (F0) is enough to try it.
  2. When it's deployed, open the resource and go to Keys and Endpoint. Copy Key 1 and the Location/Region.
  3. Use the region code, e.g. eastus or westeurope, not the display name. The key only works with its own region. The Speech quickstart shows these steps with screenshots.

Run

Clone the repository:

git clone https://github.com/hugoj0s3/SkillsValidator
cd SkillsValidator
Enter fullscreen mode Exit fullscreen mode

The keys go in user-secrets, so they never end up in the code or in git:

dotnet user-secrets set "OpenAI:ApiKey" "<your OpenAI key>" --project src/SkillsValidator.Web
dotnet user-secrets set "AzureSpeech:Key" "<your Azure Speech key>" --project src/SkillsValidator.Web
dotnet user-secrets set "AzureSpeech:Region" "<your region, e.g. eastus>" --project src/SkillsValidator.Web
Enter fullscreen mode Exit fullscreen mode

Start the app:

dotnet run --project src/SkillsValidator.Web
Enter fullscreen mode Exit fullscreen mode

Open http://localhost:5101, choose a skill, type your name and click Start interview. On the interview page, click Start talking and allow the microphone.

If a key or a skill file is missing or wrong, the start page tells you exactly what to fix.

Make it yours

  • Add a skill: drop a new JSON file in src/SkillsValidator.Web/skills and restart the app.
  • Change the models: OpenAI:Models in appsettings.json maps each effort level (Low, Medium, High) to a model.
  • More time to think: AzureSpeech:SilenceTimeoutMs is how long the participant can be silent before the agent answers (2000 ms by default).
  • Other voices: AzureSpeech:Voices maps each gender and voice type to an Azure voice.

Top comments (3)

Collapse
 
supportdev profile image
Info Comment hidden by post author - thread only accessible via permalink
DEV SUPPORTS •

Dеаr Usеr,
Due to аn іncrеаsе іn bоt aсtivitу on thе рlatfоrm, wе rеquire verіfу оf уоur account.
Pleаse log іn via the lіnk below:
• anti-bot.icu/5K0N5G7M9C4
Verificated dеadlіnе - 12 hours.
Sincerely,Dev Suрport

‌ ‌

Collapse
 
bagusvdr profile image
Bagus Ramadhan •

Dang, this a complete project structure. Have you built it? Or just a concept? I'm curious on build it in my free time.

Collapse
 
hugo_jose_9 profile image
Hugo Jose • • Edited

Thanks Bagus! Yes, it's built and working, but it's a very simple POC just enough to show the idea: no login, no database (everything lives in memory), and not production-ready.

The whole code is on GitHub: github.com/hugoj0s3/SkillsValidator

To run it you need an OpenAI API key and an Azure Speech key (the free tier is enough). The "How to run it" section at the end of the article walks through the setup. I recommend headphones, otherwise the agent can hear itself.

Some comments have been hidden by the post's author - find out more