<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ansuj Kumar Meher</title>
    <description>The latest articles on DEV Community by Ansuj Kumar Meher (@ansujkm).</description>
    <link>https://dev.to/ansujkm</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3794184%2F0439d60f-c2fa-4a48-a6bd-aa4bcba500e6.jpg</url>
      <title>DEV Community: Ansuj Kumar Meher</title>
      <link>https://dev.to/ansujkm</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ansujkm"/>
    <language>en</language>
    <item>
      <title>My Git History Called Me a Rage Coder (And Then Explained Why)</title>
      <dc:creator>Ansuj Kumar Meher</dc:creator>
      <pubDate>Mon, 13 Jul 2026 10:37:57 +0000</pubDate>
      <link>https://dev.to/ansujkm/my-git-history-called-me-a-rage-coder-and-then-explained-why-4o3o</link>
      <guid>https://dev.to/ansujkm/my-git-history-called-me-a-rage-coder-and-then-explained-why-4o3o</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for &lt;a href="https://dev.to/challenges/weekend-2026-07-09"&gt;Weekend Challenge: Passion Edition&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;I ran it on my own repo before anyone else's, and it did not go easy on me. StreamSync — 116 commits — got filed as a "Rage Coder," 35% of my commit messages apparently reading like the transcript of a bad night. The AI wasn't being generous about it, either.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkj40y5j5v0hm1sfpvumc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkj40y5j5v0hm1sfpvumc.png" alt="Github Repo" width="800" height="712"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Commit Confessions&lt;/strong&gt; takes a public GitHub repo and reads its commit history back to you like a confession — not a dashboard, not a wrapped-style stat card, an actual short piece of writing that cites your real commit messages, your real busiest hour, your real longest streak, and tells you what it thinks that says about you.&lt;/p&gt;

&lt;p&gt;Paste in a repo, and it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Pulls up to 300 commits and works out when you actually code — hour by hour, weekends vs. weekdays, your longest streak, your longest disappearing act, and how often your commit messages sound like you were mid-crisis&lt;/li&gt;
&lt;li&gt;Assigns you a developer persona and a D&amp;amp;D-style alignment based on that behavior, not a random quiz&lt;/li&gt;
&lt;li&gt;Has Gemini write a genuinely specific narrative about it — one that has to cite real detail from your data or it doesn't count&lt;/li&gt;
&lt;li&gt;Narrates that out loud with a documentary-grade ElevenLabs voice over a low atmospheric loop&lt;/li&gt;
&lt;li&gt;Generates a shareable card with AI cover art that's actually read your README, plus a constellation pattern built from your real commit punchcard&lt;/li&gt;
&lt;li&gt;If you want proof, mints the whole thing as a permanent, free memo transaction on Solana Devnet — a receipt that this repo, this obsession, existed&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It's built to be run on your own repos first. It's more fun, and slightly more brutal, than you'd expect.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;Live: &lt;a href="https://commit-confession.vercel.app" rel="noopener noreferrer"&gt;commit-confession.vercel.app&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdpge5ptucqqwfyz0s29t.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdpge5ptucqqwfyz0s29t.gif" alt="SHORT DEMO" width="600" height="302"&gt;&lt;/a&gt;&lt;/p&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/ANSUJKMEHER" rel="noopener noreferrer"&gt;
        ANSUJKMEHER
      &lt;/a&gt; / &lt;a href="https://github.com/ANSUJKMEHER/Commit-Confession" rel="noopener noreferrer"&gt;
        Commit-Confession
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;Commit Confessions&lt;/h1&gt;
&lt;/div&gt;
&lt;blockquote&gt;
&lt;p&gt;Your git history, read back to you like a confession.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Feed it a public GitHub repo. It pulls the commit history — timestamps, streaks, late-night patterns, message samples — and sends the data through Gemini to produce a short, honest narrative about the coding passion hiding in those commits. Optionally, ElevenLabs narrates it back like a movie trailer.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Built for DEV Weekend Challenge: Passion Edition (July 2026)&lt;/strong&gt;&lt;/p&gt;

&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;Live Demo&lt;/h2&gt;
&lt;/div&gt;
&lt;p&gt;→ &lt;a href="https://commit-confessions.vercel.app" rel="nofollow noopener noreferrer"&gt;commit-confessions.vercel.app&lt;/a&gt;&lt;/p&gt;

&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;How it works&lt;/h2&gt;
&lt;/div&gt;
&lt;div class="snippet-clipboard-content notranslate position-relative overflow-auto"&gt;&lt;pre class="notranslate"&gt;&lt;code&gt;GitHub REST API  →  pattern analysis  →  Gemini 2.0 Flash  →  narrative
                                                     ↘  ElevenLabs (optional voiceover)
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;GitHub API&lt;/strong&gt;: fetches up to 300 commits (no auth needed for public repos, but adding a token raises the rate limit from 60 to 5,000 requests/hr)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;api/analyze.js&lt;/code&gt;&lt;/strong&gt; (Vercel serverless): computes stats — hourly distribution, late-night %, weekend %, streak/gap lengths, sample messages — then calls Gemini&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gemini 2.0 Flash&lt;/strong&gt;: one well-crafted prompt → &lt;code&gt;{ title, narrative,&lt;/code&gt;…&lt;/li&gt;
&lt;/ol&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/ANSUJKMEHER/Commit-Confession" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;p&gt;Good repos to try it on: your own (obviously), &lt;code&gt;torvalds/linux&lt;/code&gt; if you want to see what 50,000+ commits of late-night density looks like, or &lt;code&gt;antirez/redis&lt;/code&gt; for a solo project with genuinely vivid commit messages.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/ANSUJKMEHER" rel="noopener noreferrer"&gt;
        ANSUJKMEHER
      &lt;/a&gt; / &lt;a href="https://github.com/ANSUJKMEHER/Commit-Confession" rel="noopener noreferrer"&gt;
        Commit-Confession
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;Commit Confessions&lt;/h1&gt;
&lt;/div&gt;
&lt;blockquote&gt;
&lt;p&gt;Your git history, read back to you like a confession.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Feed it a public GitHub repo. It pulls the commit history — timestamps, streaks, late-night patterns, message samples — and sends the data through Gemini to produce a short, honest narrative about the coding passion hiding in those commits. Optionally, ElevenLabs narrates it back like a movie trailer.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Built for DEV Weekend Challenge: Passion Edition (July 2026)&lt;/strong&gt;&lt;/p&gt;

&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;Live Demo&lt;/h2&gt;
&lt;/div&gt;
&lt;p&gt;→ &lt;a href="https://commit-confessions.vercel.app" rel="nofollow noopener noreferrer"&gt;commit-confessions.vercel.app&lt;/a&gt;&lt;/p&gt;

&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;How it works&lt;/h2&gt;
&lt;/div&gt;
&lt;div class="snippet-clipboard-content notranslate position-relative overflow-auto"&gt;&lt;pre class="notranslate"&gt;&lt;code&gt;GitHub REST API  →  pattern analysis  →  Gemini 2.0 Flash  →  narrative
                                                     ↘  ElevenLabs (optional voiceover)
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;GitHub API&lt;/strong&gt;: fetches up to 300 commits (no auth needed for public repos, but adding a token raises the rate limit from 60 to 5,000 requests/hr)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;api/analyze.js&lt;/code&gt;&lt;/strong&gt; (Vercel serverless): computes stats — hourly distribution, late-night %, weekend %, streak/gap lengths, sample messages — then calls Gemini&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gemini 2.0 Flash&lt;/strong&gt;: one well-crafted prompt → &lt;code&gt;{ title, narrative,&lt;/code&gt;…&lt;/li&gt;
&lt;/ol&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/ANSUJKMEHER/Commit-Confession" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The shape of it
&lt;/h3&gt;

&lt;p&gt;React + Vite on the frontend, three serverless functions doing the actual work, GitHub's REST API as the only real data dependency. No database — there's nothing to persist, every confession is generated fresh.&lt;/p&gt;

&lt;h3&gt;
  
  
  Reading a repo like a diary
&lt;/h3&gt;

&lt;p&gt;The stats engine pulls up to 300 commits and turns raw timestamps and messages into something with a shape: an hourly distribution, a full week × hour punchcard, late-night and weekend percentages, longest streak, longest gap, and a "rage score" — a regex pass over every commit message hunting for &lt;code&gt;fix&lt;/code&gt;, &lt;code&gt;revert&lt;/code&gt;, &lt;code&gt;wtf&lt;/code&gt;, &lt;code&gt;stupid&lt;/code&gt;, &lt;code&gt;please&lt;/code&gt;, and their friends. Whichever persona (Night Owl, Rage Coder, Weekend Warrior, and four others) scores highest on the resulting weighted formula wins, and the same numbers feed a D&amp;amp;D-style alignment — Chaotic Evil, for the record, requires late nights over 25%, weekends over 20%, and a rage score over 15%. You know who you are.&lt;/p&gt;

&lt;p&gt;The part I actually care about is what happens next: that data, plus the first 1,500 characters of the repo's README and a sample of real commit messages, gets handed to Gemini with a prompt that's less "write something nice" and more "write something true." It has to cite at least three specific details — an exact commit message, the exact busiest hour, the exact streak length — and it's explicitly told to avoid every hackathon-blurb cliché ("shows dedication," "journey of growth"). The output comes back as structured JSON — title, narrative, verdict, and an image prompt — with a three-model fallback chain (&lt;code&gt;gemini-2.5-flash&lt;/code&gt; → &lt;code&gt;gemini-2.0-flash&lt;/code&gt; → &lt;code&gt;gemini-2.0-flash-lite&lt;/code&gt;) so a quota limit on one model doesn't take the whole thing down.&lt;/p&gt;

&lt;p&gt;The README-awareness is what I'm most pleased with: because Gemini actually reads what the project is &lt;em&gt;for&lt;/em&gt;, the generated cover-art prompt describes something specific to it — glowing cursors on a shared document for a collaborative editor, not generic circuit-board stock art.&lt;/p&gt;

&lt;h3&gt;
  
  
  Giving it a voice
&lt;/h3&gt;

&lt;p&gt;The narrative gets narrated back through ElevenLabs' &lt;code&gt;eleven_flash_v2_5&lt;/code&gt; model with a deep documentary-narrator voice, an atmospheric urban loop fading in underneath at low volume and pausing itself when the narration ends. On mobile, the audio and the share card get bundled together as actual &lt;code&gt;File&lt;/code&gt; objects through the Web Share API, so a confession can go straight to WhatsApp or LinkedIn with both attached — not just a link.&lt;/p&gt;

&lt;h3&gt;
  
  
  The part where I deleted a dependency
&lt;/h3&gt;

&lt;p&gt;The most interesting engineering decision here wasn't a feature — it was a removal. I wanted an on-chain "proof of passion": mint a small, free, permanent record of your stats on Solana Devnet. The standard way to do that is &lt;code&gt;@solana/web3.js&lt;/code&gt;, and &lt;code&gt;@solana/web3.js&lt;/code&gt; does not want to run inside a Vercel serverless function — its WebSocket bindings and circular dependencies crashed the bundler with a 500 error four different ways before I stopped trying to fix the import and just took the SDK out entirely.&lt;/p&gt;

&lt;p&gt;What's left is 279 lines that only use what Node already has. Ed25519 signing through the native &lt;code&gt;crypto&lt;/code&gt; module, with a hand-wrapped PKCS#8 DER header around the raw key. A 50-line Base58 encoder I wrote because I didn't want to pull in a package for it. The Solana transaction format — header, account keys, blockhash, instructions, compact-u16 length prefixes — serialized byte by byte, by hand. Everything talks to Devnet through plain &lt;code&gt;fetch&lt;/code&gt; calls to the JSON-RPC endpoint. Zero dependencies, instant cold starts, a transaction confirmed in under three seconds.&lt;/p&gt;

&lt;p&gt;The user never sees any of this. They click one button and get back a Solscan link to a memo transaction that just says, permanently: this repo, these stats, this moment, happened.&lt;/p&gt;

&lt;h3&gt;
  
  
  Making it look like it means it
&lt;/h3&gt;

&lt;p&gt;Dark, editorial, a little moody — Fraunces serif for anything the app is "saying" to you, JetBrains Mono for anything that's raw data, amber for the late-night hours everywhere they show up: the commit clock's ticks, the verdict text, the glow behind the persona. The commit clock itself is a D3 radial chart, one tick per hour, length proportional to commit count, animating in over about a second the first time you see your result. The share card layers persona, stats, verdict, AI cover art, and a constellation pattern generated from your actual punchcard data — real commit hours turned into stars, connected if they're close enough together. Four theme variants, if the default doesn't match your vibe.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prize Categories
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Best Use of Google AI&lt;/strong&gt; — Gemini 2.5 Flash is the narrative engine, the persona/alignment writer, and the image-prompt generator, all off one README-aware, structured-JSON call with a three-model fallback for quota safety.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Best Use of ElevenLabs&lt;/strong&gt; — &lt;code&gt;eleven_flash_v2_5&lt;/code&gt; turns the narrative into a genuinely cinematic narration, layered with ambient audio and bundled straight into native share sheets alongside the visual card.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Best Use of Solana&lt;/strong&gt; — A real Devnet memo transaction, built and signed without &lt;code&gt;@solana/web3.js&lt;/code&gt; after it broke the serverless bundler — native crypto, hand-written Base58, manual transaction serialization, confirmed in under three seconds, with zero wallet friction for the end user.&lt;/p&gt;

&lt;p&gt;Built solo, over a weekend that — appropriately enough — involved several of the exact behaviors this tool now measures.&lt;/p&gt;

&lt;p&gt;If you run it on your own repo, I want to know what it calls you. Drop your persona and your worst commit message in the comments — I'll go first: Rage Coder, 35%, no regrets.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
    </item>
    <item>
      <title>I Built an AI Interview Coach That Turns Any Resume Into a Personalized Prep Package — No API Keys Needed</title>
      <dc:creator>Ansuj Kumar Meher</dc:creator>
      <pubDate>Sun, 31 May 2026 06:10:19 +0000</pubDate>
      <link>https://dev.to/ansujkm/i-built-an-ai-interview-coach-that-turns-any-resume-into-a-personalized-prep-package-no-api-keys-400m</link>
      <guid>https://dev.to/ansujkm/i-built-an-ai-interview-coach-that-turns-any-resume-into-a-personalized-prep-package-no-api-keys-400m</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/hermes-agent-2026-05-15"&gt;Hermes Agent Challenge&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;Interview preparation is broken. You Google "top 50 interview questions," get a wall of generic prompts that have nothing to do with your background, and somehow you're supposed to feel prepared. If you've built real-time chat systems with Socket.io and trained CNNs with TensorFlow, why are you practicing questions about linked lists and nothing else?&lt;/p&gt;

&lt;p&gt;I built &lt;strong&gt;AI Interview Coach&lt;/strong&gt; — a full-stack web application that takes your resume PDF and produces a complete, personalized interview preparation package in under five seconds.&lt;/p&gt;

&lt;p&gt;Here's what happens when you upload a resume:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Resume Score (out of 10)&lt;/strong&gt; — Your resume is scored across five dimensions: Formatting &amp;amp; Readability, Content Completeness, Skills Relevance, Experience Impact, and Education &amp;amp; Certs. Each dimension gets a score, a color-coded progress bar, and specific notes explaining the rating.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Skills Summary&lt;/strong&gt; — The engine auto-detects technologies from your resume and groups them into categories: Languages, Frameworks, ML &amp;amp; AI, Cloud &amp;amp; DevOps, Databases, and Tools &amp;amp; Practices. It scans against 90+ technology keywords.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;20 Technical Interview Questions&lt;/strong&gt; — These are generated based on what's &lt;em&gt;actually on your resume&lt;/em&gt;. If you list TensorFlow, you'll get questions about batch normalization and vanishing gradients. If you list React, you'll get questions about &lt;code&gt;useEffect&lt;/code&gt; dependency arrays and virtualization. Each question includes a "why" explanation connecting it to something specific on your resume.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;10 HR &amp;amp; Behavioral Questions&lt;/strong&gt; — Adapted based on whether you're a student, have internship experience, show leadership signals, or have multiple projects. Each comes with a STAR-method framing tip.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Improvement Suggestions&lt;/strong&gt; — Prioritized into High, Medium, and Low buckets. Things like "Quantify your achievements" or "Add certifications" — not vague advice, but specific actions mapped to scoring gaps.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The entire thing runs locally. No OpenAI key. No GPT calls. No cost per request. No data leaves your machine.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why does this matter?&lt;/strong&gt; Because most interview prep tools either cost money, require API keys, or give you the same generic output regardless of your background. This one is free, private, and genuinely personalized.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;📹 Video Walkthrough:&lt;/strong&gt;&lt;br&gt;
  &lt;iframe src="https://www.youtube.com/embed/6DIINkZniIc"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Upload Zone&lt;/strong&gt; — Drag &amp;amp; drop your resume PDF:&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F4e9rm8mm6fo1udp27cyv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F4e9rm8mm6fo1udp27cyv.png" alt="Drag &amp;amp; drop your resume PDF" width="800" height="448"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Resume Score&lt;/strong&gt; — Circular indicator + five-dimension breakdown:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fl2t2ypsudqfowquvs38b.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fl2t2ypsudqfowquvs38b.png" alt="Resume Score" width="800" height="539"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Skills Summary&lt;/strong&gt; — Auto-detected technologies grouped by category:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fxj8bzyerf8ykhnqrfz3x.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fxj8bzyerf8ykhnqrfz3x.png" alt="Skills Summary" width="800" height="412"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Technical Questions&lt;/strong&gt; — Searchable, filterable, with difficulty badges:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fwy1kb517i096ct8g5ge6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fwy1kb517i096ct8g5ge6.png" alt="Technical Questions" width="799" height="654"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;HR Questions&lt;/strong&gt; — Category-tagged with STAR-method tips:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fhm3s6pesihuntroskgkd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fhm3s6pesihuntroskgkd.png" alt="HR Questions" width="799" height="599"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Improvement Suggestions&lt;/strong&gt; — Prioritized by impact:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fpntvyd8b2k8v0th10gv0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fpntvyd8b2k8v0th10gv0.png" alt="Improvement Suggestions" width="800" height="580"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Analysis engine output&lt;/strong&gt; — scores, skills, and tailored questions in the terminal:&lt;/p&gt;


&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/ANSUJKMEHER" rel="noopener noreferrer"&gt;
        ANSUJKMEHER
      &lt;/a&gt; / &lt;a href="https://github.com/ANSUJKMEHER/Interview-Coach-App" rel="noopener noreferrer"&gt;
        Interview-Coach-App
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;AI Interview Coach&lt;/h1&gt;
&lt;/div&gt;
&lt;p&gt;A full-stack web application that analyzes resumes and generates personalized interview preparation packages. Built with React + Express for the Hermes Agent Challenge.&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;Overview&lt;/h2&gt;
&lt;/div&gt;
&lt;p&gt;Upload a resume PDF and instantly receive:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Resume Score&lt;/strong&gt; — rated across 5 dimensions (formatting, completeness, skills relevance, experience impact, education)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Skills Summary&lt;/strong&gt; — auto-detected and grouped by category&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;20 Technical Interview Questions&lt;/strong&gt; — tailored to your actual skills and projects&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;10 HR / Behavioral Questions&lt;/strong&gt; — with STAR-method framing tips&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Improvement Suggestions&lt;/strong&gt; — prioritized by impact&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;No API keys or external services required. Everything runs locally.&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;Architecture&lt;/h2&gt;
&lt;/div&gt;
&lt;div class="snippet-clipboard-content notranslate position-relative overflow-auto"&gt;
&lt;pre class="notranslate"&gt;&lt;code&gt;interview-coach-app/
├── backend/
│   ├── server.js              # Express server — upload handling, analysis engine
│   ├── package.json           # Dependencies: express, multer, cors
│   └── tmp/                   # Temp directory for PDF processing
├── frontend/
│   ├── src/
│   │   ├── App.jsx            # Main app — theme toggle, upload, results orchestration
│   │   ├── main.jsx           # React&lt;/code&gt;&lt;/pre&gt;…&lt;/div&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/ANSUJKMEHER/Interview-Coach-App" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;h3&gt;
  
  
  My Tech Stack
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Frontend:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;React 18 — Functional components with hooks (&lt;code&gt;useState&lt;/code&gt;, &lt;code&gt;useRef&lt;/code&gt;, &lt;code&gt;useCallback&lt;/code&gt;). Five purpose-built components: &lt;code&gt;FileUpload&lt;/code&gt;, &lt;code&gt;ResumeScore&lt;/code&gt;, &lt;code&gt;SkillsSummary&lt;/code&gt;, &lt;code&gt;QuestionsList&lt;/code&gt;, and &lt;code&gt;Suggestions&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Vite 5 — Dev server with API proxy to the backend, plus production builds served directly by Express.&lt;/li&gt;
&lt;li&gt;CSS Custom Properties — A complete dark and light theme system with 25+ CSS variables. No Tailwind, no CSS framework — every style is hand-written with transitions for smooth theme switching.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Backend:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Node.js + Express 4 — Handles file uploads via Multer (disk storage, PDF-only filter, 10MB limit), spawns the Python extraction process, runs the analysis engine, and serves the built frontend in production.&lt;/li&gt;
&lt;li&gt;Custom request logger — Every request gets a unique ID and timing for debugging.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;PDF Extraction:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Python 3 + pymupdf — Extracts raw text from PDF resumes with proper Unicode handling. The script outputs structured JSON including page-level text, page count, and document metadata.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Analysis Engine:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A custom rule-based engine written entirely in JavaScript — no external API calls. It includes:

&lt;ul&gt;
&lt;li&gt;A multi-dimensional scoring system with regex-based pattern matching&lt;/li&gt;
&lt;li&gt;Skill extraction across 90+ technology keywords organized into 6 categories&lt;/li&gt;
&lt;li&gt;Conditional question generation with 30+ templates spanning Deep Learning, NLP, Computer Vision, React, Node.js, Databases, Cloud, Security, and more&lt;/li&gt;
&lt;li&gt;HR question adaptation based on inferred career stage&lt;/li&gt;
&lt;li&gt;Priority-ranked improvement suggestions mapped to scoring gaps&lt;/li&gt;
&lt;/ul&gt;


&lt;/li&gt;

&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Build &amp;amp; Dev Tooling:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Vite for frontend bundling&lt;/li&gt;
&lt;li&gt;Express serves the production build from &lt;code&gt;frontend/dist/&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;uv&lt;/code&gt; for Python virtual environment and dependency management&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How I Used Hermes Agent
&lt;/h2&gt;

&lt;p&gt;Hermes Agent was involved at every stage of building this project — not as a magic "generate my app" button, but as a collaborative development partner that I could iterate with in real time.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F0lznys7mld0kt6yc71oa.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F0lznys7mld0kt6yc71oa.png" alt="Hermes Agent CLI" width="800" height="415"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Architecture Decisions
&lt;/h3&gt;

&lt;p&gt;The first thing I worked through with Hermes was the fundamental architecture question: should the analysis engine call an external LLM, or should it be rule-based?&lt;/p&gt;

&lt;p&gt;Hermes helped me reason through the tradeoffs. An LLM-based approach would generate more varied questions, but it would require API keys (friction for users), add latency (seconds per request), introduce cost (real money per upload), and make the output non-deterministic. Hermes helped me design a rule-based engine that's fast, free, deterministic, and transparent. Every score has an explanation. Every question has a "why."&lt;/p&gt;

&lt;p&gt;Hermes also helped me decide on the three-layer architecture — React frontend, Express backend, Python extraction subprocess. The alternative was doing PDF extraction in Node.js, but Hermes pointed out that pymupdf has significantly better handling of ligatures, embedded fonts, and Unicode edge cases compared to the Node.js PDF libraries available.&lt;/p&gt;

&lt;p&gt;Here's Hermes planning the approach and setting up the Python environment:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fn0yoz3a93tq9gij46bws.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fn0yoz3a93tq9gij46bws.png" alt="Architecture Decisions" width="800" height="398"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Implementation: The Analysis Engine
&lt;/h3&gt;

&lt;p&gt;Hermes scaffolded all 16 source files, installed dependencies, and verified the frontend build — all in a single session:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fbm49nsp2h5ryqglxa3vj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fbm49nsp2h5ryqglxa3vj.png" alt="Implementation: The Analysis Engine" width="800" height="353"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The heart of the application is the &lt;code&gt;analyzeResume()&lt;/code&gt; function in &lt;code&gt;server.js&lt;/code&gt; — around 300 lines of scoring, extraction, and question generation logic. Hermes helped me build this incrementally.&lt;/p&gt;

&lt;p&gt;For the scoring system, I started with a basic keyword count approach, but Hermes helped me evolve it into a five-dimension model. Take the Experience Impact dimension:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;impact&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="se"&gt;\d&lt;/span&gt;&lt;span class="sr"&gt;+%/&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="nx"&gt;impact&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="se"&gt;\d&lt;/span&gt;&lt;span class="sr"&gt;+&lt;/span&gt;&lt;span class="se"&gt;\+&lt;/span&gt;&lt;span class="sr"&gt; &lt;/span&gt;&lt;span class="se"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;users|requests|events|bugs|projects|modules&lt;/span&gt;&lt;span class="se"&gt;)&lt;/span&gt;&lt;span class="sr"&gt;/i&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="nx"&gt;impact&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="mf"&gt;1.5&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/reduced|improved|increased|optimized|accelerated/i&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="nx"&gt;impact&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="mf"&gt;1.5&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/led|owned|driven|spearheaded|initiated/i&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="nx"&gt;impact&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="se"&gt;\d&lt;/span&gt;&lt;span class="sr"&gt;+K&lt;/span&gt;&lt;span class="se"&gt;\+&lt;/span&gt;&lt;span class="sr"&gt;|&lt;/span&gt;&lt;span class="se"&gt;\d&lt;/span&gt;&lt;span class="sr"&gt;+M&lt;/span&gt;&lt;span class="se"&gt;\+&lt;/span&gt;&lt;span class="sr"&gt;|&lt;/span&gt;&lt;span class="se"&gt;\d&lt;/span&gt;&lt;span class="sr"&gt;+TB|&lt;/span&gt;&lt;span class="se"&gt;\$[\d&lt;/span&gt;&lt;span class="sr"&gt;,&lt;/span&gt;&lt;span class="se"&gt;]&lt;/span&gt;&lt;span class="sr"&gt;+/i&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="nx"&gt;impact&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="nx"&gt;impact&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;round&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;impact&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Hermes helped me identify which regex patterns actually correlate with strong resumes — looking for quantified achievements, scale indicators, and action verbs that signal ownership. We iterated on the weights until the scores felt calibrated against real resumes.&lt;/p&gt;

&lt;p&gt;For question generation, Hermes helped me design the conditional template system. Instead of random questions, each template is gated on what's actually in the resume:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;has&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;tensorflow&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;keras&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Deep Learning&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Hard&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;You used transfer learning with MobileNetV2. Explain why you'd freeze early layers vs fine-tune them all.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Your EarthSense-AI project used transfer learning — interviewers will probe the reasoning.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This produces questions that reference the user's &lt;em&gt;actual projects&lt;/em&gt; and explain &lt;em&gt;why an interviewer would ask this&lt;/em&gt;. Hermes helped me write templates for 15+ technology areas.&lt;/p&gt;

&lt;h3&gt;
  
  
  Debugging: The Unicode Incident
&lt;/h3&gt;

&lt;p&gt;The first real-world resume I uploaded crashed the app. The error:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;UnicodeEncodeError: 'charmap' codec can't encode character '\ufb01'
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's the "fi" ligature — common in PDFs generated by LaTeX. On Windows, Python's default stdout uses cp1252 encoding, which doesn't support it.&lt;/p&gt;

&lt;p&gt;Hermes immediately identified the root cause and proposed a two-layer fix. On the Node.js side, force the spawned Python process to use UTF-8:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;proc&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;spawn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;pythonExe&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;scriptPath&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;pdfPath&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;--json&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;env&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="p"&gt;...&lt;/span&gt;&lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;PYTHONIOENCODING&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;utf-8&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On the Python side, wrap stdout if the console encoding isn't UTF-8:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stdout&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;encoding&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;casefold&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stdout&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;io&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;TextIOWrapper&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;stdout&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;buffer&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;encoding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;errors&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;replace&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This fixed the crash without losing any text content. Hermes knew exactly where to intervene because it understood both the Node.js child process API and Python's encoding stack.&lt;/p&gt;

&lt;h3&gt;
  
  
  Feature Refinement
&lt;/h3&gt;

&lt;p&gt;Hermes drove several UX improvements that I wouldn't have prioritized on my own. Here's Hermes refactoring the &lt;code&gt;FileUpload&lt;/code&gt; component to add progress tracking and cancel support — you can see the diff of what changed:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fld71cexycl6j4465uev1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fld71cexycl6j4465uev1.png" alt="Feature Refinement" width="799" height="232"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Duplicate detection&lt;/strong&gt; — When multiple skill categories triggered similar questions, Hermes suggested the deduplication step with normalized comparison, then padding with fallback questions to guarantee exactly 20 technical and 10 HR.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AbortController integration&lt;/strong&gt; — Hermes wired up the upload cancel button to actually abort the in-flight fetch request and clean up state, rather than just hiding the progress bar.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Export formatting&lt;/strong&gt; — The "Download Report" feature generates a structured text file with headers, scores, and all questions. Hermes helped format the output so it reads well in any text editor.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CSS custom properties for theming&lt;/strong&gt; — Instead of class-based theme switching, Hermes designed a CSS variable system with 25+ tokens that made adding the light theme a one-block-of-CSS addition rather than rewriting every component.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Development Velocity
&lt;/h3&gt;

&lt;p&gt;Working with Hermes, I went from an empty directory to a functioning full-stack app in a single development session. The agent handled the scaffolding — Vite config, Express boilerplate, Multer setup — while I focused on the analysis logic and UI design decisions.&lt;/p&gt;

&lt;p&gt;After writing the files, Hermes automatically verified the build and ran syntax checks before moving on:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Flyhh9lien60vcg9qnyf0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Flyhh9lien60vcg9qnyf0.png" alt="Hermes build verification" width="800" height="315"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When I hit the Unicode bug, Hermes fixed it in minutes rather than the hour I'd have spent reading Python encoding documentation. When I wanted to add search and filtering to the questions list, Hermes built the entire &lt;code&gt;QuestionsList&lt;/code&gt; component with search, difficulty filter, topic filter, copy-per-question, and copy-all in one pass.&lt;/p&gt;

&lt;p&gt;That's the real value: Hermes doesn't replace thinking, but it removes the friction between having an idea and seeing it work.&lt;/p&gt;

&lt;h2&gt;
  
  
  Technical Challenges
&lt;/h2&gt;

&lt;h3&gt;
  
  
  PDF Extraction Is Messier Than You Think
&lt;/h3&gt;

&lt;p&gt;PDF is a &lt;em&gt;page layout&lt;/em&gt; format, not a &lt;em&gt;document&lt;/em&gt; format. Text inside a PDF can be fragmented across drawing commands, use arbitrary fonts with custom ligature tables, and contain Unicode characters that look normal on screen but break when piped through a process boundary. The "fi" ligature is just the most common offender — there are dozens more in PDFs generated by professional typesetting tools.&lt;/p&gt;

&lt;p&gt;The fix required interventions at two layers: the Node.js process spawner (setting &lt;code&gt;PYTHONIOENCODING&lt;/code&gt;) and the Python script itself (wrapping &lt;code&gt;sys.stdout&lt;/code&gt;). A single-layer fix wouldn't have been robust.&lt;/p&gt;

&lt;h3&gt;
  
  
  Deduplication Without Losing Personality
&lt;/h3&gt;

&lt;p&gt;When a resume lists React, Node.js, and TypeScript, the question generator might fire templates from the "React" bucket, the "Backend" bucket, and the "TypeScript" bucket — some of which produce overlapping questions. Simple string deduplication would miss near-duplicates, so I normalize the text (lowercase, collapse whitespace) before comparison:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;deduplicateQuestions&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;questions&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;seen&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Set&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[];&lt;/span&gt;
  &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;q&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;questions&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;q&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;question&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;toLowerCase&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="se"&gt;\s&lt;/span&gt;&lt;span class="sr"&gt;+/g&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt; &lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;seen&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;has&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;key&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;seen&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;key&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
      &lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;push&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;q&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After deduplication, I pad with fallback questions to guarantee the target count. The fallbacks are intentionally generic ("Describe a challenging bug you encountered") so they never feel redundant with the skill-specific ones.&lt;/p&gt;

&lt;h3&gt;
  
  
  Making a Rule-Based Engine Feel Intelligent
&lt;/h3&gt;

&lt;p&gt;The hardest design challenge wasn't technical — it was making a rule-based system produce output that feels genuinely tailored rather than randomly selected. Three things made the difference:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Every question explains why it's being asked&lt;/strong&gt; — "You have TensorFlow and Keras on your resume" or "Your EarthSense-AI project used transfer learning — interviewers will probe the reasoning." This transparency makes the output feel personalized even though it's template-driven.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;HR questions adapt to career stage&lt;/strong&gt; — The engine checks whether the text matches student patterns, internship signals, leadership keywords, or project count — and selects different question variants accordingly.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Suggestions map to scoring gaps&lt;/strong&gt; — If the Experience Impact score is low, the suggestion to "Quantify your achievements" appears in the High Priority bucket. The suggestions aren't random — they're tied to the dimensions that scored weakest.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Frontend UX: Small Touches, Big Difference
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Collapsible sections&lt;/strong&gt; — Let you focus on what matters. When you're practicing technical questions, you don't need the suggestions panel open.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Search and filter on questions&lt;/strong&gt; — Type a keyword, filter by difficulty (Easy, Medium, Hard), filter by topic (React, Databases, Security). The filter count updates in real time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Copy to clipboard&lt;/strong&gt; — Per question or all at once, useful for pasting into a study doc or sharing with a friend.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Progress simulation&lt;/strong&gt; — The progress bar advances in random increments during the server round-trip, creating the feel of active processing rather than a dead spinner.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Lessons Learned
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Rule-based doesn't mean dumb&lt;/strong&gt; — A well-designed rule engine with good templates produces output that's consistent, explainable, and free. For structured tasks like scoring and question generation, you don't always need an LLM.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Cross-process encoding is a minefield&lt;/strong&gt; — Node.js spawning Python on Windows has encoding assumptions at every layer: the child process environment, Python's stdout wrapper, and Node's buffer-to-string decoding. You need to control all of them.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Frontend polish is what separates a demo from a tool&lt;/strong&gt; — Search, filter, copy, download, dark mode, collapsible sections — none of these are technically hard, but together they make the difference between "interesting prototype" and "I'd actually use this."&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Error messages are user experience&lt;/strong&gt; — Every error in this app tells you what went wrong and what to do: "Only PDF files are allowed," "File too large (max 10 MB)," "Could not extract meaningful text — it may be image-based." No stack traces in the browser.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Future Improvements
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;LLM-powered question generation&lt;/strong&gt; — Integrate an optional local LLM (via Ollama or similar) to generate more varied, context-aware questions while keeping the API-free default.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Answer scaffolding&lt;/strong&gt; — For each question, generate a starter answer outline based on the resume content.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Job description matching&lt;/strong&gt; — Upload a job description alongside your resume and get a gap analysis: skills you have vs. skills they want, plus targeted questions for the gaps.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;PDF report export&lt;/strong&gt; — Generate a formatted PDF report instead of plain text, with the score visualization and color-coded question cards.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Session history&lt;/strong&gt; — Save past analyses to localStorage so you can track how your resume improves over time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ATS keyword checker&lt;/strong&gt; — Compare your resume against common ATS keyword databases for specific job titles.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;AI Interview Coach is a focused tool that solves a real problem: turning a generic resume PDF into a personalized interview preparation package — with scoring, skill extraction, tailored questions, and prioritized improvement suggestions. It runs entirely locally, requires no API keys, and produces results in seconds.&lt;/p&gt;

&lt;p&gt;Hermes Agent was instrumental throughout development. It helped me make the right architecture decisions (rule-based over LLM, Python for PDF extraction), built the analysis engine incrementally with proper scoring weights and conditional templates, solved the cross-platform Unicode encoding bug in minutes, and drove UX refinements like deduplication, abort handling, and the CSS theming system. The collaboration wasn't about generating boilerplate — it was about making better design decisions faster and shipping a more polished product than I would have built alone.&lt;/p&gt;

&lt;p&gt;The project is open source under MIT. Upload your resume and see what comes back — I'd love to hear whether the questions actually match your background.&lt;/p&gt;

</description>
      <category>hermesagentchallenge</category>
      <category>devchallenge</category>
      <category>agents</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Vectorized Conversations: Building a Quick RAG Chat Assistant Using Elasticsearch as a Vector Database</title>
      <dc:creator>Ansuj Kumar Meher</dc:creator>
      <pubDate>Sat, 28 Feb 2026 06:03:47 +0000</pubDate>
      <link>https://dev.to/ansujkm/vectorized-conversations-building-a-quick-rag-chat-assistant-using-elasticsearch-as-a-vector-1fg3</link>
      <guid>https://dev.to/ansujkm/vectorized-conversations-building-a-quick-rag-chat-assistant-using-elasticsearch-as-a-vector-1fg3</guid>
      <description>&lt;p&gt;This blog post was submitted to the Elastic Blogathon Contest and is eligible to win a prize.&lt;/p&gt;

&lt;h2&gt;
  
  
  Author Introduction
&lt;/h2&gt;

&lt;p&gt;Hi, I’m Ansuj Kumar Meher, a developer deeply interested in search systems, distributed architecture, and AI-driven applications. Over the past few months, I’ve been exploring how vector search can transform traditional retrieval systems. For the Elastic Blogathon 2026, I wanted to move beyond theory and build something practical — a fully working Hybrid RAG assistant powered by Elasticsearch on Elastic Cloud.&lt;/p&gt;

&lt;p&gt;This blog documents that journey — including architecture decisions, implementation details, and real observations from building the system.&lt;/p&gt;




&lt;h2&gt;
  
  
  Abstract
&lt;/h2&gt;

&lt;p&gt;Large language models are impressive, but without grounding, they hallucinate. In this blog, I demonstrate how to build a Hybrid Retrieval-Augmented Generation (RAG) system using Elasticsearch as a vector database on Elastic Cloud. By combining BM25 keyword search, HNSW-based vector similarity, and Gemini 2.5 Flash for generation, we create a scalable and production-ready semantic search assistant.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. The Problem: LLMs Without Grounding
&lt;/h2&gt;

&lt;p&gt;When I first built a chatbot using an LLM, it felt magical. It could answer questions fluently and summarize information beautifully.&lt;/p&gt;

&lt;p&gt;But then I asked it something outside its knowledge scope.&lt;/p&gt;

&lt;p&gt;It still answered — confidently — and incorrectly.&lt;/p&gt;

&lt;p&gt;That’s when I realized:&lt;/p&gt;

&lt;p&gt;LLMs are great language generators.&lt;br&gt;&lt;br&gt;
They are not reliable retrieval systems.&lt;/p&gt;

&lt;p&gt;If we want trustworthy AI systems, they must retrieve real information before generating answers.&lt;/p&gt;

&lt;p&gt;This is where Retrieval-Augmented Generation (RAG) becomes essential.&lt;/p&gt;


&lt;h2&gt;
  
  
  2. Why Hybrid Search Instead of Just Vector Search?
&lt;/h2&gt;

&lt;p&gt;While building this project, I tested three approaches:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Keyword-only search (BM25)
&lt;/li&gt;
&lt;li&gt;Vector-only search
&lt;/li&gt;
&lt;li&gt;Hybrid search (BM25 + vector similarity)
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Each had strengths and weaknesses.&lt;/p&gt;
&lt;h3&gt;
  
  
  Keyword Search (BM25)
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Excellent at lexical precision
&lt;/li&gt;
&lt;li&gt;Struggles with semantic intent
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;
  
  
  Vector Search
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Understands meaning
&lt;/li&gt;
&lt;li&gt;Sometimes ignores important keywords
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;
  
  
  Hybrid Search
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Balances precision and semantic understanding
&lt;/li&gt;
&lt;li&gt;Produces more stable ranking
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For RAG systems, retrieval quality determines generation quality. Hybrid search consistently produced better grounding.&lt;/p&gt;

&lt;p&gt;That’s why I built the system using Elasticsearch’s hybrid search capabilities.&lt;/p&gt;


&lt;h2&gt;
  
  
  3. System Architecture
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F0iv9s5jd0melf3hofm32.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F0iv9s5jd0melf3hofm32.png" alt="Hybrid RAG architecture showing user query, embedding generation, Elasticsearch hybrid search, and LLM output" width="800" height="1474"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The flow is simple but powerful:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;User submits a query.
&lt;/li&gt;
&lt;li&gt;Query is converted into a 384-dimensional embedding using MiniLM.
&lt;/li&gt;
&lt;li&gt;Elasticsearch performs hybrid retrieval:

&lt;ul&gt;
&lt;li&gt;BM25 keyword match
&lt;/li&gt;
&lt;li&gt;kNN vector search using HNSW + cosine similarity
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Top documents are retrieved.
&lt;/li&gt;
&lt;li&gt;Retrieved context is injected into the LLM prompt.
&lt;/li&gt;
&lt;li&gt;Gemini generates a grounded response.
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The key idea is that the LLM never answers without context.&lt;/p&gt;


&lt;h2&gt;
  
  
  4. Deploying Elasticsearch on Elastic Cloud
&lt;/h2&gt;

&lt;p&gt;Instead of running Elasticsearch locally, I deployed it on Elastic Cloud to simulate a production-ready environment.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fqlu03g9fgkpnaw0n3zy1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fqlu03g9fgkpnaw0n3zy1.png" alt="Elastic Cloud deployment dashboard showing healthy Elasticsearch cluster" width="800" height="194"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  Why Elastic Cloud?
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Managed cluster
&lt;/li&gt;
&lt;li&gt;Built-in security
&lt;/li&gt;
&lt;li&gt;Automatic scaling
&lt;/li&gt;
&lt;li&gt;Production-grade infrastructure
&lt;/li&gt;
&lt;li&gt;Native vector search support
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This ensures the architecture reflects real-world deployment patterns, not just a demo setup.&lt;/p&gt;


&lt;h2&gt;
  
  
  5. Configuring Elasticsearch as a Vector Database
&lt;/h2&gt;

&lt;p&gt;To enable vector search, I created an index with a &lt;code&gt;dense_vector&lt;/code&gt; field:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="err"&gt;PUT&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;chat_index&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"mappings"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"properties"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"text"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"embedding"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"dense_vector"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"dims"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;384&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"index"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"similarity"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"cosine"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F1tobfa7pddsc9ogagtft.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F1tobfa7pddsc9ogagtft.png" alt="Elasticsearch index mapping displaying dense_vector field with 384 dimensions and cosine similarity" width="474" height="655"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Key points:
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;dims: 384&lt;/code&gt; matches MiniLM embedding output.
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;similarity: cosine&lt;/code&gt; aligns with semantic embedding comparison.
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;index: true&lt;/code&gt; enables HNSW approximate nearest neighbor search.
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This transforms Elasticsearch into a scalable vector database.&lt;/p&gt;




&lt;h2&gt;
  
  
  6. Understanding HNSW and Why It Matters
&lt;/h2&gt;

&lt;p&gt;One challenge in vector search is performance.&lt;/p&gt;

&lt;p&gt;Naively comparing a query vector against every stored vector is computationally expensive.&lt;/p&gt;

&lt;p&gt;Elasticsearch solves this using HNSW (Hierarchical Navigable Small World graphs).&lt;/p&gt;

&lt;p&gt;In simple terms:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Vectors are organized into graph layers.
&lt;/li&gt;
&lt;li&gt;Search navigates these layers efficiently.
&lt;/li&gt;
&lt;li&gt;Retrieval becomes fast even at scale.
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This makes hybrid search practical for large datasets and production systems.&lt;/p&gt;




&lt;h2&gt;
  
  
  7. Implementing Hybrid Search
&lt;/h2&gt;

&lt;p&gt;Here is the hybrid query I used:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="err"&gt;GET&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;chat_index/_search&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"query"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"bool"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"must"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"match"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Elasticsearch search"&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"knn"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"field"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"embedding"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"query_vector"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;...&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;384&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;values&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;...&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"k"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"num_candidates"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F2cgj83ieepmh8xrwp6sp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F2cgj83ieepmh8xrwp6sp.png" alt="Elasticsearch Dev Tools console showing hybrid search query combining match and knn" width="800" height="682"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;_score&lt;/code&gt; reflects combined lexical and semantic relevance.&lt;/p&gt;

&lt;p&gt;In my testing, hybrid retrieval produced more balanced and reliable results compared to vector-only search.&lt;/p&gt;

&lt;p&gt;During experimentation, I also observed how tuning k and num_candidates impacted performance. Increasing num_candidates improved recall by exploring more potential nearest neighbors, but slightly increased latency. Similarly, raising k provided broader context for generation, but too many retrieved documents sometimes diluted answer precision. For small datasets this tradeoff is minimal, but at production scale, ANN parameter tuning becomes critical for balancing speed and retrieval quality.&lt;/p&gt;




&lt;h2&gt;
  
  
  8. Building the RAG Pipeline in Python
&lt;/h2&gt;

&lt;p&gt;The full implementation is available here:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;GitHub Repository:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
&lt;a href="https://github.com/ANSUJKMEHER/RagChat" rel="noopener noreferrer"&gt;https://github.com/ANSUJKMEHER/RagChat&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 1: Generate Embeddings
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;sentence_transformers&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;SentenceTransformer&lt;/span&gt;

&lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;SentenceTransformer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;all-MiniLM-L6-v2&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;query_embedding&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;tolist&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h3&gt;
  
  
  Step 2: Hybrid Retrieval
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;es&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;search&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;index&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;chat_index&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;bool&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;must&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;match&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
            &lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="n"&gt;knn&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;field&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;embedding&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;query_vector&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;query_embedding&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;k&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;num_candidates&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fpsj1lbku0hba55eq6hpf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fpsj1lbku0hba55eq6hpf.png" alt="Terminal output showing ranked hybrid search results with relevance scores" width="799" height="236"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h3&gt;
  
  
  Step 3: Construct Context
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;retrieved_docs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;hit&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;_source&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;hit&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;hits&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;hits&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]]&lt;/span&gt;
&lt;span class="n"&gt;context&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;retrieved_docs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This context is passed into the LLM prompt, ensuring grounded generation.&lt;/p&gt;




&lt;h2&gt;
  
  
  9. Integrating Gemini 2.5 Flash
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;

&lt;span class="n"&gt;GEMINI_API_KEY&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getenv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;GEMINI_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;url&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://generativelanguage.googleapis.com/v1/models/gemini-2.5-flash:generateContent?key=&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;GEMINI_API_KEY&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="n"&gt;payload&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;contents&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;parts&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;final_prompt&lt;/span&gt;&lt;span class="p"&gt;}]}&lt;/span&gt;
    &lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F8n4vkx5fsxxgxjyt9nqi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F8n4vkx5fsxxgxjyt9nqi.png" alt="Terminal output displaying final LLM-generated answer grounded in Elasticsearch retrieval" width="800" height="386"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The final answer now reflects retrieved Elastic documents rather than hallucinated content.&lt;/p&gt;




&lt;h2&gt;
  
  
  10. Sample Output
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Example Query:&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;How does Elasticsearch improve search?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;Example Output:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Supports hybrid search combining BM25 and vector similarity
&lt;/li&gt;
&lt;li&gt;Enables semantic similarity using dense vector embeddings
&lt;/li&gt;
&lt;li&gt;Scales efficiently using distributed architecture
&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  11. Practical Applications
&lt;/h2&gt;

&lt;p&gt;This architecture applies directly to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Enterprise knowledge assistants
&lt;/li&gt;
&lt;li&gt;AI customer support bots
&lt;/li&gt;
&lt;li&gt;E-commerce semantic search
&lt;/li&gt;
&lt;li&gt;Log analysis in observability platforms
&lt;/li&gt;
&lt;li&gt;AI copilots grounded in proprietary data
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For example, in an enterprise knowledge base, hybrid search prevents LLMs from fabricating internal policy details by forcing responses to rely strictly on indexed documentation. This significantly reduces hallucination risk while maintaining conversational fluency.&lt;/p&gt;

&lt;p&gt;Hybrid search ensures meaning and precision coexist.&lt;/p&gt;




&lt;h2&gt;
  
  
  12. Key Observations from Building This
&lt;/h2&gt;

&lt;p&gt;While building and testing this system, I observed:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Hybrid retrieval significantly improves answer grounding.
&lt;/li&gt;
&lt;li&gt;Retrieval quality impacts LLM output more than generation parameters.
&lt;/li&gt;
&lt;li&gt;Elastic Cloud simplifies scaling concerns.
&lt;/li&gt;
&lt;li&gt;Even small improvements in retrieval ranking dramatically improve answer quality.
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One important realization:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;In RAG systems, retrieval matters more than generation.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Conclusion + Takeaways
&lt;/h2&gt;

&lt;p&gt;Vectorized thinking is not about replacing keyword search.&lt;/p&gt;

&lt;p&gt;It is about enhancing it.&lt;/p&gt;

&lt;p&gt;By combining:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Dense vector indexing
&lt;/li&gt;
&lt;li&gt;Hybrid search
&lt;/li&gt;
&lt;li&gt;HNSW-based ANN
&lt;/li&gt;
&lt;li&gt;Elastic Cloud deployment
&lt;/li&gt;
&lt;li&gt;RAG architecture
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We create AI systems that are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Reliable
&lt;/li&gt;
&lt;li&gt;Scalable
&lt;/li&gt;
&lt;li&gt;Context-aware
&lt;/li&gt;
&lt;li&gt;Production-ready
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What surprised me most during this implementation was how much retrieval quality shapes generation quality. Adjusting embedding strategy and hybrid scoring had a larger impact on answer correctness than tweaking LLM temperature or prompt structure. This reinforced an important lesson: in production RAG systems, search engineering is not optional — it is foundational.&lt;/p&gt;

&lt;p&gt;Elasticsearch demonstrates that search and vectors do not compete — they complement each other.&lt;/p&gt;

&lt;p&gt;This project is a small step toward building grounded, trustworthy AI systems powered by Elastic.&lt;/p&gt;




&lt;h2&gt;
  
  
  GitHub Repository
&lt;/h2&gt;

&lt;p&gt;Full source code available at:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/ANSUJKMEHER/RagChat" rel="noopener noreferrer"&gt;https://github.com/ANSUJKMEHER/RagChat&lt;/a&gt;&lt;/p&gt;




</description>
      <category>vectorswithelastic</category>
      <category>searchwithvectors</category>
      <category>writewithelastic</category>
      <category>storiesinsearch</category>
    </item>
  </channel>
</rss>
