DEV Community

Cover image for I built an AI mock interviewer with LangGraph.js and DeepSeek — here's what I learned
ShivWad
ShivWad

Posted on

I built an AI mock interviewer with LangGraph.js and DeepSeek — here's what I learned

DevGrill

Static interview prep tools have a fundamental problem: they show you the question, you think about it, you read the model answer, and you feel like you're ready.

You're not.

Real interviews are conversations. An interviewer probes your reasoning, challenges your trade-offs, asks "why not X instead?", and doesn't let you off the hook for a vague answer. No question bank replicates that — and that's the gap I built DevGrill to fill.

This is a technical breakdown of how it works.


What DevGrill does

You upload your resume and paste a job description. DevGrill generates targeted interview questions — not generic ones, but questions that probe the gap between where you are and what the role requires. Then it runs a live, multi-phase interview against you, scores your responses, and gives you a detailed feedback report.

Two interview types right now: System Design and Technical (coding + CS fundamentals).

The key design constraint I set from day one: the AI interviewer cannot validate weak answers. Most AI tools are sycophantic by default — they find something positive to say about everything. That's actively harmful for interview prep. If your system design has a single point of failure and you didn't mention it, Mr. Grill (the interviewer persona) will ask about it.


The architecture: why LangGraph.js

The core challenge is that an interview is a stateful, multi-turn conversation with conditional branching. You can't model that with a simple chat loop. Phases need to transition at the right time, human input needs to interrupt execution and resume cleanly, and the final evaluation needs access to the entire conversation history.

LangGraph.js solves this with a typed state graph where each node reads from and writes to a shared state object. The graph persists across turns using a checkpointer, so the entire interview state survives between HTTP requests.

Here's the graph I ended up with — eight nodes, each with a single responsibility:

question_generator → setup → interviewer ⇄ human_input → phase_evaluator → judge → report_generator → persist_result
Enter fullscreen mode Exit fullscreen mode

Let me walk through the interesting ones.


question_generator: two-stage prompting

Question generation is a two-stage DeepSeek call using the Pro model (better reasoning for this task):

  1. Stage 1 — Extract the candidate's experience level, tech stack, and role signals from the resume + JD
  2. Stage 2 — Generate interview questions that specifically target the delta between the candidate's background and the role requirements This is why the questions feel personalised. A senior backend engineer applying for a principal role gets probed on system-level thinking and org impact. A mid-level engineer targeting a role with heavy Kafka usage gets Kafka questions even if it's not on their resume — because the JD signals it's expected.

The interviewer/human_input split

This was the trickiest architectural decision. Early versions had a single node handling both AI response generation and waiting for user input — which caused a double-invocation bug where the model would generate a question and then answer it itself.

The fix was splitting into two nodes:

// interviewer node — AI generates its response and yields
async function interviewer(state: InterviewState) {
  const response = await llm.invoke(buildInterviewerMessages(state));
  return {
    messages: [...state.messages, { role: "assistant", content: response.content }],
    awaitingInput: true,
  };
}

// human_input node — interrupts execution and waits for real input
async function humanInput(state: InterviewState) {
  // LangGraph interrupt() pauses execution here
  // The graph is resumed externally when the user submits their answer
  const userAnswer = interrupt("Awaiting candidate response");
  return {
    messages: [...state.messages, { role: "user", content: userAnswer }],
    awaitingInput: false,
  };
}
Enter fullscreen mode Exit fullscreen mode

The interrupt() call is LangGraph's mechanism for pausing graph execution mid-run and resuming it later with external input. The graph state (including full conversation history) is persisted to Postgres via PostgresSaver between turns, so there's no in-memory state to manage on the server.


Anti-sycophancy enforcement

This lives in the system prompt for the interviewer node, and it's non-negotiable:

  • No unsolicited compliments ("Great answer!", "That's a solid approach!")
  • If an answer is incomplete or wrong, follow up with a probe — don't move on
  • Only advance the phase when the candidate has demonstrated understanding, not just attempted an answer
  • Challenge specific decisions: "You chose a relational DB here — walk me through why NoSQL wouldn't work" The model is DeepSeek V4 Flash for interviewer turns (fast, cheap, good instruction-following). I found that explicit negative examples in the prompt ("do NOT say things like...") significantly reduced sycophantic outputs compared to just positive instructions alone.

phase_evaluator and the judge

After each phase (e.g., Requirements Gathering → High-Level Design → Deep Dive), phase_evaluator decides whether to advance or loop back.

judge runs at the very end with temperature set to 0.1 — much lower than the conversational turns. Scoring needs to be consistent and deterministic. It outputs structured scores across rubric dimensions (communication, technical depth, trade-off reasoning, etc.) plus specific quotes from the transcript to support each score.

Keeping the judge as a separate node with its own temperature setting was important. Early versions used the same model config for evaluation as for conversation, which produced score variance across identical interviews.


Checkpointing with PostgresSaver

Every graph turn is checkpointed to Postgres:

import { PostgresSaver } from "@langchain/langgraph-checkpoint-postgres";

const checkpointer = PostgresSaver.fromConnString(process.env.DATABASE_URL!);
const graph = interviewGraph.compile({ checkpointer });

// Resume an existing interview
const result = await graph.invoke(
  { userInput: answer },
  { configurable: { thread_id: interviewId } }
);
Enter fullscreen mode Exit fullscreen mode

thread_id maps to an interview session. This means:

  • The full interview survives server restarts
  • Multiple concurrent interviews don't share state

- The report generator has access to every message in the conversation

Stack summary

Layer Tech
Frontend Next.js 15 App Router, shadcn/ui, Tailwind, Monaco editor
Agent service Express.js, LangGraph.js
LLM DeepSeek V4 Pro (question gen) + V4 Flash (interviewer/judge)
DB PostgreSQL on Neon, Drizzle ORM
Auth Clerk
Monorepo Turborepo + pnpm workspaces
Infra Vercel (web), Railway (agent service)

What a full interview costs

A complete system design interview (question generation + full multi-phase conversation + scoring + report) costs roughly $0.008–0.01 using DeepSeek's pricing. That's the entire session, not per turn. This made it viable to offer free access without burning through API budget.


What I'd do differently

Streaming. The interviewer response currently lands all at once after a network call. Adding streaming would make the experience feel significantly more like a real conversation.

Voice input. Typing answers to interview questions is unnatural. The plan is to add Whisper-based transcription so you can actually speak your answers.

More interview types. Behavioural (STAR-format) and HR rounds are next.


Try it

If you're prepping for a technical role and want something that actually pushes back — grill.shivwad.in. Free to use, no credit card.

I'm actively iterating on it — feedback (especially critical feedback) is genuinely useful. Drop a comment or reach out directly.

Top comments (0)