When building an autonomous agent to conduct technical engineering interviews, the margin for error is zero.
If an AI customer service bot hallucinates a return policy, it is an inconvenience. If an AI interviewer hallucinates a candidate's technical score or gets tricked by prompt injection into discussing philosophical hypotheticals instead of system design, it destroys the integrity of the hiring pipeline.
Standard conversational AI wrappers—where a massive system prompt is fed into a single LLM call—are fundamentally unsuited for objective technical evaluations. They suffer from conversational drift, they are easily manipulated, and they struggle to enforce strict grading rubrics.
To solve this at TechEval.ai, we abandoned the single-prompt approach. Instead, we architected Kovi as a deterministic state machine using LangGraph, deployed entirely on serverless infrastructure.
Here is a look under the hood at how we built a hallucination-free AI interviewer capable of running 1,000+ concurrent WebRTC screens with sub-400ms latency.
1. The Supervisor-Worker State Machine
Rather than relying on one omnipotent AI model to handle conversation, proctoring, and evaluation simultaneously, Kovi utilizes a strict LangGraph supervisor-worker architecture.
- The Supervisor Node: This node controls the flow of the interview. It acts as a rigid state machine that tracks the 60-minute timer, monitors the required tech stack parameters, and ensures all "Must-Ask" questions are covered before the session terminates.
- The Context Engine (RAG): Before the interview begins, Kovi parses the candidate's resume and the specific Job Description. Using Pinecone and Neon DB, this worker retrieves exact candidate experiences to tailor the opening questions.
- The Deep-Dive Worker: If a candidate provides a superficial answer about a technology (e.g., FastAPI or microservices), the supervisor routes the state to this worker, which dynamically generates a highly specific architectural follow-up question to test true depth.
- The Isolated Evaluator: Grading happens asynchronously on an entirely separate model. The conversational voice model never sees the 10-point scoring rubric, ensuring the candidate cannot manipulate the AI into revealing the expected answer.
2. Sub-400ms Conversational Latency (WebRTC + Serverless)
The biggest giveaway of a poorly built AI voice agent is the "walkie-talkie" delay—waiting 2 to 3 seconds for the AI to respond.
To achieve natural, real-time interruptions and conversational flow, we bypassed standard WebSocket limitations and built a native WebRTC streaming layer. The backend is orchestrated via a high-throughput FastAPI gateway.
To keep economics lean while supporting massive concurrency, the entire compute layer is hosted on serverless GCP Cloud Run and Vercel. By strictly localizing our infrastructure (compute, Neon DB, and Pinecone) in the Mumbai (Asia-South1) region, we bypassed the massive latency penalty of routing through US servers.
The result? Kovi delivers sub-400ms audio response times across major Tier 1 Indian tech hubs (Bengaluru, Hyderabad, Mumbai, Delhi, Chennai).
3. Defeating AI Copilots with Conversational Telemetry
Remote technical screens face an epidemic of LLM cheating. Traditional proctoring relies on screen sharing or browser lockdowns, which are easily bypassed with a secondary device.
We built Kovi’s proctoring engine around conversational telemetry:
- Rhythm & Cadence Analysis: When candidates read answers generated by ChatGPT, their speech cadence flattens, and pauses align abnormally with AI token generation speeds. Kovi monitors this audio buffer in real-time.
- Dynamic Curveballs: When high-probability scripting is detected, the LangGraph supervisor injects a hyper-specific, undocumented follow-up question. A candidate reading a script will freeze; a genuine engineer will seamlessly whiteboard the solution verbally.
- Non-Blocking Evidence: Instead of aggressively terminating the interview and risking false-positive disputes, Kovi silently logs hardware telemetry (tab switches, gaze estimation via webcam) and Copilot flags, delivering concrete evidence directly to the HR scorecard.
4. DPDP-Ready from Day One
Enterprise hiring data is highly sensitive. Because our architecture is completely localized in India, no candidate voice data, video telemetry, or resume context ever crosses international borders. Kovi aligns natively with India's Digital Personal Data Protection (DPDP) Act, providing automated data expiration and strict isolation between tenant environments.
The Future of Automated Screening
By moving away from stochastic text generators to deterministic, state-driven agent architectures, we can finally evaluate engineering talent objectively, at an unlimited scale, and for a flat rate of just ₹150 per interview.
Ready to explore the architecture or integrate it into your ATS?
- Read the full system design documentation at docs.techeval.ai
- Install the Python SDK to schedule your first automated interview:
pip install kovi-sdk
What are your thoughts on using state machines for AI agents instead of single-prompt wrappers? Have you run into conversational drift with standard LLM tools? Let's discuss in the comments.
Top comments (0)