<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: ShivWad</title>
    <description>The latest articles on DEV Community by ShivWad (@shivwad).</description>
    <link>https://dev.to/shivwad</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4005036%2Fdb49faad-0ffc-4958-8392-62ee9a177b89.jpg</url>
      <title>DEV Community: ShivWad</title>
      <link>https://dev.to/shivwad</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/shivwad"/>
    <language>en</language>
    <item>
      <title>I built an AI mock interviewer with LangGraph.js and DeepSeek — here's what I learned</title>
      <dc:creator>ShivWad</dc:creator>
      <pubDate>Tue, 30 Jun 2026 14:32:18 +0000</pubDate>
      <link>https://dev.to/shivwad/i-built-an-ai-mock-interviewer-with-langgraphjs-and-deepseek-heres-what-i-learned-202b</link>
      <guid>https://dev.to/shivwad/i-built-an-ai-mock-interviewer-with-langgraphjs-and-deepseek-heres-what-i-learned-202b</guid>
      <description>&lt;p&gt;&lt;a href="https://grill.shivwad.in" rel="noopener noreferrer"&gt;DevGrill&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Static interview prep tools have a fundamental problem: they show you the question, you think about it, you read the model answer, and you feel like you're ready.&lt;/p&gt;

&lt;p&gt;You're not.&lt;/p&gt;

&lt;p&gt;Real interviews are conversations. An interviewer probes your reasoning, challenges your trade-offs, asks "why not X instead?", and doesn't let you off the hook for a vague answer. No question bank replicates that — and that's the gap I built &lt;a href="https://grill.shivwad.in" rel="noopener noreferrer"&gt;DevGrill&lt;/a&gt; to fill.&lt;/p&gt;

&lt;p&gt;This is a technical breakdown of how it works.&lt;/p&gt;




&lt;h2&gt;
  
  
  What DevGrill does
&lt;/h2&gt;

&lt;p&gt;You upload your resume and paste a job description. DevGrill generates targeted interview questions — not generic ones, but questions that probe the &lt;em&gt;gap&lt;/em&gt; between where you are and what the role requires. Then it runs a live, multi-phase interview against you, scores your responses, and gives you a detailed feedback report.&lt;/p&gt;

&lt;p&gt;Two interview types right now: &lt;strong&gt;System Design&lt;/strong&gt; and &lt;strong&gt;Technical&lt;/strong&gt; (coding + CS fundamentals).&lt;/p&gt;

&lt;p&gt;The key design constraint I set from day one: &lt;strong&gt;the AI interviewer cannot validate weak answers&lt;/strong&gt;. Most AI tools are sycophantic by default — they find something positive to say about everything. That's actively harmful for interview prep. If your system design has a single point of failure and you didn't mention it, Mr. Grill (the interviewer persona) will ask about it.&lt;/p&gt;




&lt;h2&gt;
  
  
  The architecture: why LangGraph.js
&lt;/h2&gt;

&lt;p&gt;The core challenge is that an interview is a &lt;em&gt;stateful, multi-turn conversation with conditional branching&lt;/em&gt;. You can't model that with a simple chat loop. Phases need to transition at the right time, human input needs to interrupt execution and resume cleanly, and the final evaluation needs access to the entire conversation history.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://langchain-ai.github.io/langgraphjs/" rel="noopener noreferrer"&gt;LangGraph.js&lt;/a&gt; solves this with a typed state graph where each node reads from and writes to a shared state object. The graph persists across turns using a checkpointer, so the entire interview state survives between HTTP requests.&lt;/p&gt;

&lt;p&gt;Here's the graph I ended up with — eight nodes, each with a single responsibility:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;question_generator → setup → interviewer ⇄ human_input → phase_evaluator → judge → report_generator → persist_result
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Let me walk through the interesting ones.&lt;/p&gt;




&lt;h2&gt;
  
  
  question_generator: two-stage prompting
&lt;/h2&gt;

&lt;p&gt;Question generation is a two-stage DeepSeek call using the Pro model (better reasoning for this task):&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Stage 1&lt;/strong&gt; — Extract the candidate's experience level, tech stack, and role signals from the resume + JD&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stage 2&lt;/strong&gt; — Generate interview questions that specifically target the delta between the candidate's background and the role requirements
This is why the questions feel personalised. A senior backend engineer applying for a principal role gets probed on system-level thinking and org impact. A mid-level engineer targeting a role with heavy Kafka usage gets Kafka questions even if it's not on their resume — because the JD signals it's expected.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  The interviewer/human_input split
&lt;/h2&gt;

&lt;p&gt;This was the trickiest architectural decision. Early versions had a single node handling both AI response generation and waiting for user input — which caused a double-invocation bug where the model would generate a question &lt;em&gt;and then answer it itself&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;The fix was splitting into two nodes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// interviewer node — AI generates its response and yields&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;interviewer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;InterviewState&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;llm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;invoke&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;buildInterviewerMessages&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[...&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;assistant&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;content&lt;/span&gt; &lt;span class="p"&gt;}],&lt;/span&gt;
    &lt;span class="na"&gt;awaitingInput&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;};&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;// human_input node — interrupts execution and waits for real input&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;humanInput&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;InterviewState&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="c1"&gt;// LangGraph interrupt() pauses execution here&lt;/span&gt;
  &lt;span class="c1"&gt;// The graph is resumed externally when the user submits their answer&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;userAnswer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;interrupt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Awaiting candidate response&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[...&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;user&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;userAnswer&lt;/span&gt; &lt;span class="p"&gt;}],&lt;/span&gt;
    &lt;span class="na"&gt;awaitingInput&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;};&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;interrupt()&lt;/code&gt; call is LangGraph's mechanism for pausing graph execution mid-run and resuming it later with external input. The graph state (including full conversation history) is persisted to Postgres via &lt;code&gt;PostgresSaver&lt;/code&gt; between turns, so there's no in-memory state to manage on the server.&lt;/p&gt;




&lt;h2&gt;
  
  
  Anti-sycophancy enforcement
&lt;/h2&gt;

&lt;p&gt;This lives in the system prompt for the interviewer node, and it's non-negotiable:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;No unsolicited compliments ("Great answer!", "That's a solid approach!")&lt;/li&gt;
&lt;li&gt;If an answer is incomplete or wrong, follow up with a probe — don't move on&lt;/li&gt;
&lt;li&gt;Only advance the phase when the candidate has &lt;em&gt;demonstrated&lt;/em&gt; understanding, not just attempted an answer&lt;/li&gt;
&lt;li&gt;Challenge specific decisions: "You chose a relational DB here — walk me through why NoSQL wouldn't work"
The model is DeepSeek V4 Flash for interviewer turns (fast, cheap, good instruction-following). I found that explicit negative examples in the prompt ("do NOT say things like...") significantly reduced sycophantic outputs compared to just positive instructions alone.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  phase_evaluator and the judge
&lt;/h2&gt;

&lt;p&gt;After each phase (e.g., Requirements Gathering → High-Level Design → Deep Dive), &lt;code&gt;phase_evaluator&lt;/code&gt; decides whether to advance or loop back.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;judge&lt;/code&gt; runs at the very end with temperature set to 0.1 — much lower than the conversational turns. Scoring needs to be consistent and deterministic. It outputs structured scores across rubric dimensions (communication, technical depth, trade-off reasoning, etc.) plus specific quotes from the transcript to support each score.&lt;/p&gt;

&lt;p&gt;Keeping the judge as a separate node with its own temperature setting was important. Early versions used the same model config for evaluation as for conversation, which produced score variance across identical interviews.&lt;/p&gt;




&lt;h2&gt;
  
  
  Checkpointing with PostgresSaver
&lt;/h2&gt;

&lt;p&gt;Every graph turn is checkpointed to Postgres:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;PostgresSaver&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;@langchain/langgraph-checkpoint-postgres&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;checkpointer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;PostgresSaver&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fromConnString&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;DATABASE_URL&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;graph&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;interviewGraph&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;compile&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;checkpointer&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="c1"&gt;// Resume an existing interview&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;graph&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;invoke&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;userInput&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;answer&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;configurable&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;thread_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;interviewId&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;thread_id&lt;/code&gt; maps to an interview session. This means:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The full interview survives server restarts&lt;/li&gt;
&lt;li&gt;Multiple concurrent interviews don't share state&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  - The report generator has access to every message in the conversation
&lt;/h2&gt;

&lt;h2&gt;
  
  
  Stack summary
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Tech&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Frontend&lt;/td&gt;
&lt;td&gt;Next.js 15 App Router, shadcn/ui, Tailwind, Monaco editor&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agent service&lt;/td&gt;
&lt;td&gt;Express.js, LangGraph.js&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LLM&lt;/td&gt;
&lt;td&gt;DeepSeek V4 Pro (question gen) + V4 Flash (interviewer/judge)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DB&lt;/td&gt;
&lt;td&gt;PostgreSQL on Neon, Drizzle ORM&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Auth&lt;/td&gt;
&lt;td&gt;Clerk&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Monorepo&lt;/td&gt;
&lt;td&gt;Turborepo + pnpm workspaces&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Infra&lt;/td&gt;
&lt;td&gt;Vercel (web), Railway (agent service)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  What a full interview costs
&lt;/h2&gt;

&lt;p&gt;A complete system design interview (question generation + full multi-phase conversation + scoring + report) costs roughly &lt;strong&gt;$0.008–0.01&lt;/strong&gt; using DeepSeek's pricing. That's the entire session, not per turn. This made it viable to offer free access without burning through API budget.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I'd do differently
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Streaming.&lt;/strong&gt; The interviewer response currently lands all at once after a network call. Adding streaming would make the experience feel significantly more like a real conversation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Voice input.&lt;/strong&gt; Typing answers to interview questions is unnatural. The plan is to add Whisper-based transcription so you can actually speak your answers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;More interview types.&lt;/strong&gt; Behavioural (STAR-format) and HR rounds are next.&lt;/p&gt;




&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;If you're prepping for a technical role and want something that actually pushes back — &lt;a href="https://grill.shivwad.in" rel="noopener noreferrer"&gt;grill.shivwad.in&lt;/a&gt;. Free to use, no credit card.&lt;/p&gt;

&lt;p&gt;I'm actively iterating on it — feedback (especially critical feedback) is genuinely useful. Drop a comment or reach out directly.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>typescript</category>
    </item>
  </channel>
</rss>
