Building an AI technical interviewer presents a unique physics problem. When a human engineer pauses to think, they expect the interviewer to wait. If the candidate takes a wrong turn, they expect the interviewer to gently interrupt them.
Achieving this natural conversational flow requires sub-400ms latency and full-duplex communication. Standard generative AI wrappers fail here because they rely on sequential, blocking architectures.
Here is how we bypassed the standard LLM bottlenecks to build a highly concurrent, interruptible AI screening engine using FastAPI, WebRTC, and GCP Cloud Run.
The Bottleneck: Why WebSockets and HTTP Streaming Fail
Most voice AI pipelines follow a sequential "walkie-talkie" pattern:
- Client records a chunk of audio.
- Client sends audio via WebSocket or HTTP to a Speech-to-Text (STT) service (taking 100-200ms).
- Text is passed to an LLM which generates a response (taking 300-500ms).
- Text is passed to a Text-to-Speech (TTS) service (taking 100-200ms).
- Audio is streamed back to the client.
By the time the TTS audio reaches the browser, the total end-to-end latency sits between 1 to 2 seconds. Conversational turn-taking completely breaks down at this latency threshold. Candidates end up talking over the agent or waiting in awkward silence.
The WebRTC Shift: Bidirectional Streaming
To eliminate conversational latency, we shifted the transport layer entirely to WebRTC. Originally designed for peer-to-peer video conferencing, WebRTC is the gold standard for real-time AI audio transport.
WebRTC provides three critical advantages for AI voice agents:
- UDP-Based Transport: Unlike WebSockets (which run on TCP and suffer from head-of-line blocking), WebRTC streams packets over UDP. If an audio packet drops, it skips it rather than halting the stream to wait for a retransmission.
- Built-in Audio Processing: WebRTC natively handles echo cancellation, automatic gain control, and background noise filtering on the client side before the audio ever hits the network.
- True Duplex Interruptibility: Because audio is flowing in both directions simultaneously, the backend can run a fast Voice Activity Detection (VAD) model on the incoming stream. If the user starts speaking while the AI is talking, the server immediately halts the TTS stream and listens, creating a seamless "barge-in" experience.
The Serverless FastAPI Backend
To handle WebRTC signaling without running expensive dedicated media servers, we built the backend using Python and FastAPI.
When a candidate joins a Kovi interview session:
- The client generates a WebRTC Session Description Protocol (SDP) offer.
- The FastAPI backend receives the offer via a standard HTTP POST request.
- The backend negotiates the connection, binds the media streams to our internal STT/LLM/TTS processing pipeline, and returns the SDP answer.
- From that point on, all audio flows directly through the WebRTC data channels.
By decoupling the initial HTTP signaling from the continuous UDP media stream, we maintain a stateless API while supporting real-time media.
Scaling on GCP Cloud Run
Deploying WebRTC in a serverless environment like GCP Cloud Run introduces a specific challenge: WebRTC requires persistent, stateful UDP connections, but serverless containers are ephemeral.
We optimized our Cloud Run deployment to handle this through three configuration changes:
- Session Affinity: We enabled session affinity (sticky sessions) on Cloud Run to ensure that once a WebRTC peer connection is established, all subsequent signaling and stream state logic routing remains locked to that specific container instance.
-
Cold Start Mitigation: Cloud Run containers can take seconds to spin up, which ruins the initial interview experience. We utilize minimum instances (
min-instances = 1) for the signaling service to ensure the WebRTC handshake completes in milliseconds. - CPU Allocation: We configured Cloud Run to allocate CPU always, rather than only during request processing. WebRTC requires continuous background CPU cycles to process incoming UDP packets and manage the audio stream buffers; if CPU is throttled between HTTP requests, the audio pipeline drops.
The Outcome
By combining native WebRTC transport with a serverless FastAPI backend, the architecture completely bypasses the traditional constraints of voice AI.
The result is an interview environment that handles natural interruptions, processes technical responses dynamically, and scales to thousands of concurrent engineering evaluations without degrading audio quality or racking up massive infrastructure bills.
Top comments (1)
We initially tried WebSockets before making this pivot. I am curious—what has been your biggest headache with voice AI latency?