<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ambuj Tripathi</title>
    <description>The latest articles on DEV Community by Ambuj Tripathi (@ambuj_tripathi).</description>
    <link>https://dev.to/ambuj_tripathi</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4012395%2Fcc638ce4-1156-4fb3-ba3b-db3d1f0701f9.jpeg</url>
      <title>DEV Community: Ambuj Tripathi</title>
      <link>https://dev.to/ambuj_tripathi</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ambuj_tripathi"/>
    <language>en</language>
    <item>
      <title>AI tools kept inventing skills on my resume, so I built an anti-hallucination MCP pipeline</title>
      <dc:creator>Ambuj Tripathi</dc:creator>
      <pubDate>Mon, 28 Sep 2026 17:48:06 +0000</pubDate>
      <link>https://dev.to/ambuj_tripathi/ai-tools-kept-inventing-skills-on-my-resume-so-i-built-an-anti-hallucination-mcp-pipeline-2ko9</link>
      <guid>https://dev.to/ambuj_tripathi/ai-tools-kept-inventing-skills-on-my-resume-so-i-built-an-anti-hallucination-mcp-pipeline-2ko9</guid>
      <description>&lt;p&gt;I was applying for jobs last month and tried 3 different "AI cover letter generators." Every single one invented skills I don't have.&lt;/p&gt;

&lt;p&gt;One told a recruiter I had "5+ years of Kubernetes orchestration experience." I've never touched Kubernetes in production. Another claimed I "led a team of 40 engineers." I've never managed anyone. These tools are liability machines.&lt;/p&gt;

&lt;p&gt;So I built CoverCraft — an open-source cover letter generator where the AI literally cannot make a claim that isn't traceable to your actual resume text. If your resume doesn't say it, the letter doesn't either. Period.&lt;/p&gt;

&lt;p&gt;But here's the thing that made this genuinely hard to build: the hallucination problem in cover letters is way worse than in chatbots. In a chatbot, a wrong answer is annoying. In a cover letter, a fabricated claim gets you into an interview you can't survive. The interviewer asks about your "Kubernetes experience" and you're sitting there like 🫠&lt;/p&gt;

&lt;p&gt;Let me walk through the architecture because the anti-hallucination pipeline is where it gets interesting.&lt;/p&gt;

&lt;p&gt;The 6-Stage Deterministic Pipeline (with Live GitHub Code Proof &amp;amp; MCP)&lt;br&gt;
Most AI cover letter tools work like this:&lt;/p&gt;

&lt;p&gt;Resume + JD → single GPT prompt → "here's your letter lol"&lt;br&gt;
CoverCraft works like this:&lt;/p&gt;

&lt;p&gt;Resume → Parse (3-tier: unpdf → Gemini Vision OCR → raw binary fallback)&lt;br&gt;
       + Auto-Extract candidate Name, Email, LinkedIn &amp;amp; GitHub (Zero Phone PII!)&lt;br&gt;
       ↓&lt;br&gt;
JD → Competency Extraction + Auto-Detect Target Role &amp;amp; Company from pasted text&lt;br&gt;
       ↓&lt;br&gt;
Company → Real-time Tavily search → Jina Reader deep scrape →&lt;br&gt;
          Source credibility scoring → Stale content filter (&amp;lt;9m/&amp;lt;18m)&lt;br&gt;
       ↓&lt;br&gt;
GitHub MCP Proofer (Tool #7) → Live repo inspection → Commit tree &amp;amp; SHA hash grounding (4,999 req/hr $0 PAT)&lt;br&gt;
       ↓&lt;br&gt;
Human Approval Gate (you review company intel &amp;amp; repo proofs BEFORE it enters the prompt)&lt;br&gt;
       ↓&lt;br&gt;
Generation → Adversarial Red-Teamer audit → Langfuse Observability → Clean export&lt;br&gt;
Every stage has its own API route, its own system prompt, and its own guardrails. The LLM never sees raw unprocessed input. Let me break down each one.&lt;/p&gt;

&lt;p&gt;Stage 1: Resume Parsing — Why I Need 3 Fallback Tiers&lt;br&gt;
I assumed PDF parsing was a solved problem. It absolutely is not.&lt;/p&gt;

&lt;p&gt;People upload resumes in every possible state: Canva-exported PDFs with custom font encodings that produce mojibake. Scanned printed copies that are literally images. Password-protected files from HR portals. .docx files disguised as .pdf because they renamed the extension.&lt;/p&gt;

&lt;p&gt;Here's the actual parser cascade in production:&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp4xqtud6fm5ta3qk0vwi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp4xqtud6fm5ta3qk0vwi.png" alt=" " width="800" height="1100"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;// parse-resume/route.js — 3-Tier Extraction Cascade&lt;br&gt;
async function parsePdfBuffer(buffer) {&lt;br&gt;
  // Tier 1: unpdf (serverless-safe, pure JS, instant)&lt;br&gt;
  try {&lt;br&gt;
    const { getDocumentProxy, extractText } = await import("unpdf");&lt;br&gt;
    const pdf = await getDocumentProxy(new Uint8Array(buffer));&lt;br&gt;
    const { text } = await extractText(pdf, { mergePages: true });&lt;br&gt;
    if (text?.trim().length &amp;gt; 20) return text.trim();&lt;br&gt;
  } catch (err) { /* fall through */ }&lt;/p&gt;

&lt;p&gt;// Tier 2: Gemini Vision OCR (handles scanned/image-only PDFs)&lt;br&gt;
  try {&lt;br&gt;
    const base64Data = buffer.toString("base64");&lt;br&gt;
    const result = await model.generateContent([&lt;br&gt;
      { inlineData: { data: base64Data, mimeType: "application/pdf" } },&lt;br&gt;
      "Extract all readable text from this resume..."&lt;br&gt;
    ]);&lt;br&gt;
    const geminiText = result?.response?.text();&lt;br&gt;
    if (geminiText?.trim().length &amp;gt; 20) return geminiText.trim();&lt;br&gt;
  } catch (err) { /* fall through */ }&lt;/p&gt;

&lt;p&gt;// Tier 3: Raw binary token extraction (last resort)&lt;br&gt;
  const str = buffer.toString("binary");&lt;br&gt;
  const regex = /(([^)]+))\s*Tj/g;&lt;br&gt;
  // Extract raw PDF text operators...&lt;br&gt;
}&lt;br&gt;
Tier 3 is my favorite. It literally reads PDF binary operators (Tj is the "show text" operator in the PDF spec). It's ugly, but it works when absolutely everything else fails. The number of "My resume won't upload" support tickets dropped to zero after adding this.&lt;/p&gt;

&lt;p&gt;Also: We permanently deleted the Phone Number field.&lt;/p&gt;

&lt;p&gt;Why do cover letter tools ask candidates for their phone numbers? There is zero architectural reason for an AI generator to hold personal contact digits. We deleted the field completely to protect candidate PII. Instead, when you upload your resume PDF, our regex + entity parser instantly auto-extracts your Name, Email, LinkedIn URL, and GitHub Profile URL. And when you paste a Job Description, a heuristic detector auto-extracts the Target Role and Target Company automatically. No manual re-typing, zero unnecessary data retention.&lt;/p&gt;

&lt;p&gt;Stage 2: Skill Matching — Application-Layer Scoring, Not LLM Scoring&lt;br&gt;
This is a deliberate design choice that most AI tools get wrong. They let the LLM evaluate AND score the candidate fit. That's asking the same model that hallucinates to also be the judge of truth.&lt;/p&gt;

&lt;p&gt;I split it: LLM classifies, application code scores.&lt;/p&gt;

&lt;p&gt;// analyze/route.js — Deterministic Application-Layer Scoring&lt;br&gt;
const weights = {&lt;br&gt;
  STRONG_MATCH: 1.0,   // Direct resume evidence&lt;br&gt;
  PARTIAL_MATCH: 0.6,  // Related but not exact&lt;br&gt;
  TRANSFERABLE: 0.4,   // Adjacent skill&lt;br&gt;
  MISSING: 0.0,        // Honest zero&lt;br&gt;
};&lt;/p&gt;

&lt;p&gt;const rawScore = skills.reduce((sum, item) =&amp;gt;&lt;br&gt;
  sum + (weights[item.status] ?? 0), 0&lt;br&gt;
) / totalSkills;&lt;/p&gt;

&lt;p&gt;data.overall_match = Math.round(rawScore * 100);&lt;br&gt;
The LLM's job is classification: "Does this resume mention Docker? STRONG_MATCH / PARTIAL / MISSING." The score formula is pure math. No temperature. No stochastic variance. Same resume + same JD = same score every single time. If you run it 100 times, you get 100 identical numbers.&lt;/p&gt;

&lt;p&gt;Why does this matter? Because I've seen other tools give the same candidate a "92% match" on Monday and "78% match" on Tuesday for the same job. That's not analysis — that's a random number generator with extra steps.&lt;/p&gt;

&lt;p&gt;Stage 3: Company Research — The Part That Scared Me Most&lt;br&gt;
Here's the pipeline that keeps me up at night from a hallucination perspective. We scrape real-time company intelligence to make the letter relevant. But scraped web data is the single largest hallucination vector.&lt;/p&gt;

&lt;p&gt;The research pipeline:&lt;/p&gt;

&lt;p&gt;Tavily Advanced Search → 8 results for "{company} tech stack engineering hiring careers"&lt;/p&gt;

&lt;p&gt;Jina AI Reader → Deep-scrapes top 2 URLs for full markdown extraction (not just snippets)&lt;/p&gt;

&lt;p&gt;Source Credibility Scoring → Every URL gets a trust tier:&lt;/p&gt;

&lt;p&gt;// research/route.js — Deterministic Source Credibility Filter&lt;br&gt;
// Tier 1 (95 trust): Official company domains, SEC filings, Greenhouse/Lever&lt;br&gt;
// Tier 2 (75 trust): Naukri, AmbitionBox, Indeed, Glassdoor, LinkedIn&lt;br&gt;
// Tier 3 (55 trust): Supporting media&lt;br&gt;
// BLOCKED: Pinterest, Quora, Facebook, Instagram, TikTok → discarded&lt;/p&gt;

&lt;p&gt;// Stale content filter:&lt;br&gt;
// &amp;lt;9 months: CURRENT → full weight&lt;br&gt;
// &amp;lt;18 months: AGING → -15 trust penalty&lt;br&gt;
// &amp;gt;18 months: STALE → flagged, deprioritized&lt;br&gt;
Human Approval Gate — This is the non-negotiable part. The researched company intel gets shown to the user in structured cards BEFORE it enters the generation prompt. You can approve, reject, or exclude any piece of intelligence. Zero unverified company claims enter the LLM.&lt;/p&gt;

&lt;p&gt;I've seen tools write "Google's recent pivot to quantum computing" in cover letters for Google — sourced from a 2019 blog post. Our stale content filter + human gate makes this physically impossible.&lt;/p&gt;

&lt;p&gt;Stage 4: Live GitHub Code Proof via MCP Tool #7 — Don't Just Check Resume Text, Verify the Commits&lt;br&gt;
This is the feature that made recruiters and senior engineers stop and say "wait, that's actually genius."&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F01mz6pztkgcv2ua266di.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F01mz6pztkgcv2ua266di.png" alt=" " width="800" height="1100"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Anyone can write words on a resume: "Architected 11-node cyclic LangGraph agent with checkpointing." In a normal AI tool, the LLM just echoes that phrase with fancier adjectives.&lt;/p&gt;

&lt;p&gt;CoverCraft doesn't do that. It uses our 7th registered Model Context Protocol (MCP) tool (github_proofer in /mcp-server/tools/github_proofer.py) to connect directly to the candidate's public GitHub account:&lt;/p&gt;

&lt;p&gt;Live Repository &amp;amp; Commit Tree Inspection: It pulls your public repositories, descriptions, primary languages, and recent commit history via the GitHub REST API.&lt;/p&gt;

&lt;p&gt;Verifiable Commit SHA Anchoring: When the candidate claims they built a cyclic agent or an MCP server, CoverCraft grabs the actual latest commit hash (sha: 7f2a1b9) and commit message: "feat: implement cyclic state machine checkpointing &amp;amp; live audit logging".&lt;/p&gt;

&lt;p&gt;Grounding in the Letter: The generated cover letter anchors technical assertions directly to real commit evidence: "Directly demonstrated production agentic architecture in agentic-rag-financial-parser (commit 7f2a1b9), implementing cyclic state persistence."&lt;/p&gt;

&lt;p&gt;$0 Cost / Bulletproof Free Tier: We authenticated the GitHub MCP tool using a fine-grained GitHub Personal Access Token (PAT) with read-only public repository permissions. This gives 5,000 requests per hour for $0 free, completely eliminating the risk of shared unauthenticated IP rate limits (60 req/hr) without costing a single penny.&lt;/p&gt;

&lt;p&gt;Adversarial Overclaim Sentinel: If you claim you have deep production expertise in Kubernetes or Rust, but your GitHub has zero repos and zero commits touching it, the pipeline flags the discrepancy and prompts you for specific clarification before generation. No fake claims slip through.&lt;/p&gt;

&lt;p&gt;Stage 5: Generation — The Anti-Hallucination Prompt Engineering&lt;br&gt;
The system prompt is 250+ lines. Not because I like writing essays, but because every line exists to close a hallucination vector I found in testing.&lt;/p&gt;

&lt;p&gt;The one I'm most proud of:&lt;/p&gt;

&lt;p&gt;MUST-HAVE GAP TRANSPARENCY:&lt;br&gt;
When the JD explicitly specifies a "MUST HAVE" technology&lt;br&gt;
that has NO evidence in the candidate's resume:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;NEVER fabricate direct production experience.&lt;/li&gt;
&lt;li&gt;DO NOT silently omit critical requirements.&lt;/li&gt;
&lt;li&gt;INSTEAD: Transparently acknowledge using this pattern:&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;"While I may not have direct production experience in&lt;br&gt;
[Target Technology], my deep background in [Verified Skill]&lt;br&gt;
gives me the exact technical foundation to rapidly adapt&lt;br&gt;
and execute in [Target Ecosystem] with zero friction."&lt;br&gt;
Most tools either lie about the skill or silently skip it. Both are bad. Lying gets you caught in interviews. Silently skipping leaves an obvious hole the recruiter notices. The transparent acknowledgment is actually more impressive — it shows self-awareness and honesty that 99% of applicants don't have.&lt;/p&gt;

&lt;p&gt;The Overclaim Red-Teamer runs post-generation:&lt;/p&gt;

&lt;p&gt;// generate/route.js — Adversarial Seniority Overclaim Detector&lt;br&gt;
const OVERCLAIM_PATTERNS = [&lt;br&gt;
  { trigger: /\bspearheaded\b/gi, safe: "led implementation of" },&lt;br&gt;
  { trigger: /\bpioneered\b/gi, safe: "developed" },&lt;br&gt;
  { trigger: /\bsolely architected\b/gi, safe: "architected" },&lt;br&gt;
  { trigger: /\bcommanded the\b/gi, safe: "coordinated the" },&lt;br&gt;
];&lt;/p&gt;

&lt;p&gt;// For each pattern: if the verb appears in the letter&lt;br&gt;
// but NOT in the original resume → flag and replace.&lt;br&gt;
// You said "developed" in your resume?&lt;br&gt;
// The letter says "developed." Not "spearheaded." Not "pioneered."&lt;br&gt;
This catches the classic LLM seniority inflation pattern. GPT loves to turn "worked on" into "spearheaded." It's a deterministic regex pass — no LLM involved.&lt;/p&gt;

&lt;p&gt;Stage 6: Evidence Tracing — Proof Badges&lt;br&gt;
Every claim in the generated letter has an inline citation marker:&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1rg42105cumjwj1fs262.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1rg42105cumjwj1fs262.png" alt=" " width="800" height="1100"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;"Engineered official MCP servers using standard stdio transport [Resume: official MCP servers with stdio transport]"&lt;/p&gt;

&lt;p&gt;The frontend parses these into interactive clickable badges. Click one → it opens a modal showing the exact resume line that grounds the claim. It's basically Perplexity's citation system, but for cover letters.&lt;/p&gt;

&lt;p&gt;If the AI couldn't find resume evidence for a sentence, the citation is missing, and the overclaim detector flags it. No orphan claims survive the pipeline.&lt;/p&gt;

&lt;p&gt;Production Observability: Langfuse Full-Trace Telemetry &amp;amp; The Circuit Breaker&lt;br&gt;
To understand what the multi-agent pipeline is doing in real-time, CoverCraft instruments end-to-end tracing with Langfuse (@langfuse/core / langfuse Node SDK):&lt;/p&gt;

&lt;p&gt;Child Spans for Every MCP Tool: - Discrete Traces Across Every Pipeline Stage: When a generation request flows through CoverCraft, Langfuse logs dedicated execution traces across all key stages — Resume Parsing, JD Competency Extraction, Company Web Recon, Anti-Hallucination Cover Letter Synthesis, and Interview Defense generation.&lt;/p&gt;

&lt;p&gt;Extraction, Tavily Web Recon, GitHub Commit Proofer, Gemini Generation, and ATS Scoring.&lt;/p&gt;

&lt;p&gt;Latency &amp;amp; Token Telemetry: Captures exact prompt tokens, completion tokens, latency per tool, and model parameters for full visibility.&lt;/p&gt;

&lt;p&gt;Graceful Non-Blocking Execution: If Langfuse servers take longer or reach timeouts, the tracing wrapper catches errors silently so user letter generation is never blocked or delayed.&lt;/p&gt;

&lt;p&gt;And for LLM API reliability, we have our Circuit Breaker state machine:&lt;/p&gt;

&lt;p&gt;The entire tool runs on free-tier APIs. Gemini Flash Lite for generation, Tavily for search, Jina for extraction. Free tiers go down. A lot.&lt;/p&gt;

&lt;p&gt;// lib/gemini.js — Circuit Breaker State Machine&lt;br&gt;
const circuit = {&lt;br&gt;
  state: "CLOSED",          // Normal operation&lt;br&gt;
  failureCount: 0,&lt;br&gt;
  failureThreshold: 5,      // Trip after 5 consecutive failures&lt;br&gt;
  cooldownPeriodMs: 30000,  // 30s cooldown&lt;br&gt;
  successThreshold: 2,      // 2 canary successes to reset&lt;br&gt;
};&lt;/p&gt;

&lt;p&gt;// CLOSED → request flows normally&lt;br&gt;
// 5 failures → OPEN → instant fail-fast (no wasted API calls)&lt;br&gt;
// 30s cooldown → HALF_OPEN → canary request&lt;br&gt;
// 2 canary successes → CLOSED again&lt;br&gt;
Primary model fails? Automatic 3-tier multi-provider fallback: Gemini 3.5 Flash Lite → Gemini 3.1 Flash Lite Preview → OpenRouter (NVIDIA Nemotron). Each model gets 2 retries with exponential backoff + jitter before falling over.&lt;/p&gt;

&lt;p&gt;Why jitter? Because without it, if the API recovers, every instance retries at the exact same moment and kills it again. The Math.random() * 250 spread prevents thundering herd.&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpe4i6vmtcvqxbbecto3m.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpe4i6vmtcvqxbbecto3m.png" alt=" " width="800" height="1100"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The Stack&lt;br&gt;
Frontend: Next.js 16 (App Router) + Recharts + Tailwind CSS&lt;/p&gt;

&lt;p&gt;Backend: Next.js Serverless Routes + Python 3.11 Model Context Protocol (MCP) Server (7 Tools)&lt;/p&gt;

&lt;p&gt;Code Grounding: GitHub REST API &amp;amp; MCP Proofer (5,000 req/hr free PAT, live commit SHA hashing)&lt;/p&gt;

&lt;p&gt;LLM: - LLM: 3-Tier Multi-Provider Pipeline — Gemini 3.5 Flash Lite (primary) + Gemini 3.1 Flash Lite Preview (secondary) + OpenRouter NVIDIA Nemotron (tertiary backup)&lt;/p&gt;

&lt;p&gt;Company Research: Tavily Search API (8-page crawl) + Jina AI Reader (markdown extractor)&lt;/p&gt;

&lt;p&gt;Observability: Langfuse (full-trace latency, token consumption &amp;amp; child spans) + MongoDB Atlas&lt;/p&gt;

&lt;p&gt;Auth &amp;amp; Privacy: NextAuth.js (Google OAuth) + Zero Phone Number Collection (Dual Auto-Extraction)&lt;/p&gt;

&lt;p&gt;Rate Limiting: IP-based sliding window (in-memory)&lt;/p&gt;

&lt;p&gt;Deployment: - Deployment: Vercel Serverless Edge + Docker Multi-Stage Standalone&lt;/p&gt;

&lt;p&gt;What I'd Do Differently&lt;br&gt;
Streaming. Right now it waits for the full JSON response before rendering. SSE streaming would feel way snappier. But structured JSON output and streaming are awkward together with Gemini.&lt;/p&gt;

&lt;p&gt;Resume Profile Persistence. Currently you re-upload your resume every session. Should store a parsed profile and only re-parse on changes.&lt;/p&gt;

&lt;p&gt;A/B Letter Variants. Generate 2-3 variants with different emphasis (technical depth vs. leadership vs. culture fit) and let the user pick.&lt;/p&gt;

&lt;p&gt;Links&lt;br&gt;
GitHub: &lt;a href="https://github.com/Ambuj123-lab/career-workspace-ambujsystems" rel="noopener noreferrer"&gt;https://github.com/Ambuj123-lab/career-workspace-ambujsystems&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Built this as a side project while job hunting. The irony of building a tool to help you get hired while you need to get hired is not lost on me.&lt;/p&gt;

&lt;p&gt;Happy to answer any questions about the anti-hallucination pipeline, the circuit breaker, or why PDF parsing is a war crime. AMA.&lt;/p&gt;

&lt;p&gt;Honest question: Do you actually read cover letters when hiring? Or is this entire document class dead and I wasted my time?&lt;/p&gt;

&lt;p&gt;The hallucination tradeoff: I chose "transparent gap acknowledgment" over "skip the missing skill." Some people think showing weakness is suicide. Others think it shows maturity. What's your take?&lt;/p&gt;

&lt;p&gt;Circuit breakers for LLM APIs: Anyone else running production on free-tier LLM APIs? What's your failure pattern look like? I'm seeing ~2-3 503s per day from Gemini.&lt;/p&gt;

&lt;p&gt;Source credibility scoring: My stale content filter uses 9-month / 18-month thresholds. Too aggressive? Too lenient? What would you use?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>opensource</category>
      <category>python</category>
    </item>
    <item>
      <title>PyPDFLoader, LlamaParse, Custom Regex — I Tried Everything on Indian Government PDFs. Here's What Actually Worked.</title>
      <dc:creator>Ambuj Tripathi</dc:creator>
      <pubDate>Thu, 02 Jul 2026 15:09:42 +0000</pubDate>
      <link>https://dev.to/ambuj_tripathi/pypdfloader-llamaparse-custom-regex-i-tried-everything-on-indian-government-pdfs-heres-what-58ej</link>
      <guid>https://dev.to/ambuj_tripathi/pypdfloader-llamaparse-custom-regex-i-tried-everything-on-indian-government-pdfs-heres-what-58ej</guid>
      <description>&lt;p&gt;Six months ago I asked the same questions you're asking. "How do I handle merged cells?" "Why does my table extraction break?" "Which parser should I use?"&lt;/p&gt;

&lt;p&gt;I tried &lt;strong&gt;every popular approach&lt;/strong&gt; — PyPDFLoader, Unstructured, LlamaParse, custom regex — on some of the most painful PDFs you can imagine: Indian Government Budget documents, Finance Bills, and the &lt;strong&gt;Constitution of India&lt;/strong&gt; (400+ pages of dense legal text with footnotes on every page).&lt;/p&gt;

&lt;p&gt;This article is an honest post-mortem of what went wrong, why, and the &lt;strong&gt;only architecture that actually survived production.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;🤯 &lt;strong&gt;The Document From Hell&lt;/strong&gt;&lt;br&gt;
Most RAG tutorials use clean, simple PDFs. The Constitution of India is not that.&lt;/p&gt;

&lt;p&gt;Here's what you're dealing with on every single page:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;19. Protection of certain rights regarding freedom of speech, etc.—
(1) All citizens shall have the right—
    (a) to freedom of speech and expression;
    (b) to assemble peaceably and without arms;
...
______________________________________________
1. Subs. by the Constitution (First Amendment) Act, 1951, s. 3
2. Ins. by the Constitution (Forty-fourth Amendment) Act, 1978.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every page has three zones:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Article content&lt;/strong&gt; (what users actually want)&lt;br&gt;
&lt;strong&gt;A separator line&lt;/strong&gt; (______)&lt;br&gt;
&lt;strong&gt;Footnotes&lt;/strong&gt; (amendment citations that ALSO begin with numbers like&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1., 19., 34.)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Those footnotes start with the same numbers as real Articles. Embedding models encode them with equal weight. This is where hallucinations are born.&lt;/p&gt;

&lt;p&gt;❌ &lt;strong&gt;Attempt 1:&lt;/strong&gt; &lt;strong&gt;LlamaParse (Agentic Tier) — The Expensive Failure&lt;/strong&gt;&lt;br&gt;
My initial setup: LlamaParse at Agentic tier (10 credits/page) + LangChain's MarkdownHeaderTextSplitter.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What I expected:&lt;/strong&gt; Clean, hierarchically separated chunks per Article.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What I got:&lt;/strong&gt; 624 giant chunks from a 402-page document.&lt;/p&gt;

&lt;p&gt;LlamaParse is excellent for tables, invoices, and structured forms. But for dense continuous legal text with hundreds of numbered items, &lt;strong&gt;it merged multiple pages into single Markdown blocks.&lt;/strong&gt; Article 19 wasn't a standalone chunk — it was buried inside a 5,000-character blob alongside Articles 17, 18, 20, and a dozen footnotes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Hallucination Test:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Query: "What is Article 19?"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Vector similarity matched a footnote&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;(19. Ins. by Constitution (Forty-fourth Amendment)...)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;higher than actual Article 19 text. The LLM received garbage context and returned garbage output.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cost damage:&lt;/strong&gt; 402 pages × 10 credits = 4,020 credits per sync. Multiple debugging iterations = 30K+ credits burned.&lt;/p&gt;

&lt;p&gt;🛡️ &lt;strong&gt;The Idempotency Layer: Never Waste an API Call Twice&lt;/strong&gt;&lt;br&gt;
Before fixing retrieval, I built a safety net. After burning 30K+ credits on debugging, I swore: never again.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SHA-256 File Hashing&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;python&lt;/span&gt;

&lt;span class="c1"&gt;# sync.py — Hash every PDF before processing
&lt;/span&gt;&lt;span class="n"&gt;current_hash&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;hashlib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sha256&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;file_path&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;()).&lt;/span&gt;&lt;span class="nf"&gt;hexdigest&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;registry_hash&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;supabase&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_registry_entry&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;filename&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;file_hash&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;current_hash&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;registry_hash&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="c1"&gt;# File unchanged — skip entirely. Zero API calls.
&lt;/span&gt;    &lt;span class="k"&gt;pass&lt;/span&gt;
&lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="c1"&gt;# File changed — delete old vectors, re-process
&lt;/span&gt;    &lt;span class="n"&gt;pinecone&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;delete_vectors&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;filter&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;source_file&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;filename&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;
    &lt;span class="nf"&gt;reprocess&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;file_path&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;supabase&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;update_hash&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;filename&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;current_hash&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every PDF is hashed with &lt;strong&gt;SHA-256&lt;/strong&gt; before processing. Hash stored in Supabase. On re-sync, if hash matches → entire file skipped. Zero parsing, zero embedding, zero Pinecone calls.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Deterministic Chunk IDs&lt;/strong&gt;&lt;br&gt;
python&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# chunker.py — Same input = Same IDs, always
&lt;/span&gt;&lt;span class="n"&gt;parent_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;source_file&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;_&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;page_number&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;_&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;parent_index&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="n"&gt;child_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;hashlib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;md5&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;parent_id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;_&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;child_index&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;()).&lt;/span&gt;&lt;span class="nf"&gt;hexdigest&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No random UUIDs. Chunk IDs derived from file name + page + position. Re-syncing same file = identical IDs. Pinecone upsert overwrites instead of duplicating.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;This is the difference between a script that works once and a system you can safely run in production every day&lt;br&gt;
.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;✅ &lt;strong&gt;Attempt 2:&lt;/strong&gt; The Deterministic Pipeline (What Actually Worked)&lt;br&gt;
I asked a fundamental question: "For this specific document, do I actually need an LLM to parse it?"&lt;/p&gt;

&lt;p&gt;No. The Constitution has a completely predictable structure:&lt;/p&gt;

&lt;p&gt;Articles always start with&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;\n[number]. [Title]—
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Footnotes are always after underscores&lt;br&gt;
Page headers always say "THE CONSTITUTION OF INDIA"&lt;br&gt;
&lt;strong&gt;This is regex territory, not LLM territory.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 1:&lt;/strong&gt; &lt;strong&gt;Aggressive Footnote Removal&lt;/strong&gt; &lt;br&gt;
python&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# parser.py
&lt;/span&gt;&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;page_num&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;doc&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;page_count&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;doc&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;page_num&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;get_text&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="c1"&gt;# Remove page headers
&lt;/span&gt;    &lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sub&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;THE CONSTITUTION OF\s*INDIA\n\(Part.*?\)&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;''&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="c1"&gt;# Split at footnote separator — discard everything below
&lt;/span&gt;    &lt;span class="n"&gt;parts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;_{10,}&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;clean_text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;parts&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;  &lt;span class="c1"&gt;# Only main text survives
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Result:&lt;/strong&gt; Zero footnotes in the vector index.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 2:&lt;/strong&gt; &lt;strong&gt;Article-Boundary Chunking&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;python&lt;/span&gt;

&lt;span class="c1"&gt;# chunker.py — Split at Article boundaries, not character counts
&lt;/span&gt;&lt;span class="n"&gt;raw_splits&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;\n(?=\d{1,3}[A-Z]*\.\s+[A-Z])&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;page_text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;split&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;raw_splits&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="c1"&gt;# Each split = exactly one Article
&lt;/span&gt;    &lt;span class="n"&gt;article_match&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;match&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;^(\d{1,3}[A-Z]*)\.&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;split&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;article_num&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;article_match&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;group&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;article_match&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
    &lt;span class="c1"&gt;# e.g., "19", "21A", "370"
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Result:&lt;/strong&gt; 624 messy blobs → 3,248 precise chunks, each one Article.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 3:&lt;/strong&gt; &lt;strong&gt;Metadata Injection into Pinecone&lt;/strong&gt;&lt;br&gt;
python&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;chunk_metadata&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;source_file&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;constitution of india.pdf&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;chunk_type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;parent_child&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;is_omitted&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;is_omitted&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;article_number&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;article_num&lt;/span&gt;  &lt;span class="c1"&gt;# Hard-tagged at ingestion
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every chunk carries its Article identity in Pinecone. Not inferred. Not guessed. Deterministically tagged.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 4:&lt;/strong&gt; &lt;strong&gt;Smart LangGraph Routing&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;python&lt;/span&gt;

&lt;span class="c1"&gt;# graph.py — LangGraph Retriever Node
&lt;/span&gt;&lt;span class="n"&gt;target_article&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;intent&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;article_number&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;target_article&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;target_article&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;null&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;none&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="c1"&gt;# Bypass vector similarity — database-level equality filter
&lt;/span&gt;    &lt;span class="n"&gt;pinecone_filter&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;$and&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;article_number&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;$eq&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;target_article&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;})&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is WHERE&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;article_number&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;19&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;in SQL. The vector index &lt;strong&gt;cannot&lt;/strong&gt; return chunks from any other Article.&lt;/p&gt;

&lt;p&gt;🎯 &lt;strong&gt;Validation:&lt;/strong&gt; &lt;strong&gt;The Hallucination Test Suite&lt;/strong&gt;&lt;br&gt;
Results independently scored by a third-party LLM evaluator:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Query&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
What is Article 20? &lt;br&gt;
&lt;strong&gt;Key Behavior&lt;/strong&gt;&lt;br&gt;
Returned all 3 safeguards (Ex Post Facto, Double Jeopardy, Self-Incrimination) precisely&lt;br&gt;
&lt;strong&gt;Score&lt;/strong&gt;   9/10&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffzdbg3tchlm6kvul06it.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffzdbg3tchlm6kvul06it.png" alt=" " width="799" height="427"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is Article 34?&lt;/strong&gt; &lt;br&gt;
&lt;strong&gt;Key Behavior&lt;/strong&gt;&lt;br&gt;
Correctly retrieved martial law provisions with no Schedule noise   *&lt;em&gt;Score *&lt;/em&gt;           9/10&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fskqh0obj5p80njndxyci.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fskqh0obj5p80njndxyci.png" alt=" " width="800" height="374"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frhjv01piiisq5xk6wqhc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frhjv01piiisq5xk6wqhc.png" alt=" " width="799" height="384"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Query&lt;/strong&gt;&lt;br&gt;
Article 31C + Kesavananda Bharati?&lt;br&gt;&lt;br&gt;
&lt;strong&gt;Key Behavior&lt;/strong&gt;&lt;br&gt;
Retrieved 31C accurately; correctly refused to hallucinate case law *&lt;em&gt;Score *&lt;/em&gt;            92/100&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F49ricznfvkeliw3nqiy1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F49ricznfvkeliw3nqiy1.png" alt=" " width="800" height="408"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffzo9q10h8klrvi9ac70s.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffzo9q10h8klrvi9ac70s.png" alt=" " width="800" height="455"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Query&lt;/strong&gt;&lt;br&gt;
Basic Structure Doctrine?&lt;br&gt;&lt;br&gt;
&lt;strong&gt;Key Behavior&lt;/strong&gt;&lt;br&gt;
Identified as judicial principle; stated it appears in no constitutional article    Pass&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4dtsllqf27t2h05p55v2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4dtsllqf27t2h05p55v2.png" alt=" " width="800" height="376"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Query&lt;/strong&gt;&lt;br&gt;
Article 31B + Ninth Schedule?&lt;br&gt;&lt;br&gt;
&lt;strong&gt;Key Behavior&lt;/strong&gt;&lt;br&gt;
Correctly framed the Basic Structure vs Ninth Schedule tension  8.8/10&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Filp38i3rb90vz2glp9i7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Filp38i3rb90vz2glp9i7.png" alt=" " width="799" height="378"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqj18e7ddn54cskra35c1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqj18e7ddn54cskra35c1.png" alt=" " width="800" height="387"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The most significant result is from Query 3. The system responded:&lt;br&gt;
_&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"The provided documents do not contain specific details regarding the Kesavananda Bharati case."_&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;em&gt;&lt;strong&gt;That's not a failure. That's correct, production-grade RAG behavior. A null response is a success. A hallucinated response is a disaster.&lt;/strong&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;🏗️ &lt;strong&gt;The Full Architecture&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Query: "What is Article 19?"
         ↓
   [LLM Classifier Node]
   → Extracts: article_number = "19"
         ↓
   [Retriever Node]
   → pinecone_filter = {
       "$and": [
         {"source_file": {"$eq": "constitution of india.pdf"}},
         {"article_number": {"$eq": "19"}}
       ]
     }
         ↓
   [Pinecone — Database lookup, NOT vector similarity]
         ↓
   [LLM Generator — clean, precise context]
         ↓
   Accurate response. Hallucination-resistant.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;⚠️ &lt;strong&gt;Known Limitations (Being Honest)&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The Seventh Schedule Overlap The Schedule uses numbered entries
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;(19. Price control, 34. Betting and gambling)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;. The regex tags these as&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;article_number&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;19"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;. Current impact: Low — &lt;em&gt;LLM differentiates them in generation.&lt;/em&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;*&lt;em&gt;General Conceptual Queries *&lt;/em&gt;&lt;em&gt;"What are all Fundamental Rights?"&lt;/em&gt; doesn't trigger metadata filter. Falls back to semantic search.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;No Cross-Article Relationships&lt;/strong&gt; The system doesn't model that Article 32 enforces Article 19. Each Article indexed independently.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;🔧 &lt;strong&gt;Tech Stack&lt;/strong&gt;&lt;br&gt;
&lt;strong&gt;Parser&lt;/strong&gt;:** PyMuPDF (free, local)&lt;br&gt;
&lt;strong&gt;Chunker&lt;/strong&gt;:** Custom regex-based hierarchical chunker&lt;br&gt;
&lt;strong&gt;Embeddings:&lt;/strong&gt; Jina AI v3 (MRL: 1024→256 dims, 75% storage savings)&lt;br&gt;
&lt;strong&gt;Vector DB:&lt;/strong&gt; Pinecone Serverless (with metadata filtering)&lt;br&gt;
&lt;strong&gt;Orchestration:&lt;/strong&gt; LangGraph (8-node agentic pipeline)&lt;br&gt;
&lt;strong&gt;LLM:&lt;/strong&gt; Google Gemini&lt;br&gt;
&lt;strong&gt;Registry:&lt;/strong&gt; Supabase (file hashing + sync tracking)&lt;br&gt;
&lt;strong&gt;Monitoring:&lt;/strong&gt; Langfuse (LLM observability)&lt;br&gt;
💡 &lt;strong&gt;Three Takeaways&lt;/strong&gt;&lt;br&gt;
&lt;strong&gt;Assess document structure before choosing a parser.&lt;/strong&gt; LlamaParse is excellent for semi-structured documents. For continuous legal text with predictable patterns, a custom regex parser gives you more control at zero cost.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Design for metadata from day one.&lt;/strong&gt; Vector similarity is a fallback, not a first choice.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Test the hallucination boundary, not just the happy path.&lt;/strong&gt; Asking your RAG system about things that aren't in the documents is as important as asking about things that are.&lt;/p&gt;

&lt;p&gt;📊 &lt;strong&gt;Community Response&lt;/strong&gt;&lt;br&gt;
This approach got significant traction in the AI community:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reddit (r/LangChain):&lt;/strong&gt; 50,000+ views, 500+ shares across two posts&lt;br&gt;
&lt;strong&gt;GitHub:&lt;/strong&gt; 64 stars, 22 forks&lt;br&gt;
&lt;strong&gt;HuggingFace:&lt;/strong&gt; 3 published fine-tuned models (1B, 3B, 8B) with 5,500+ downloads&lt;br&gt;
🔗 Links&lt;br&gt;
&lt;strong&gt;GitHub (Full Source Code):&lt;/strong&gt; github.com/Ambuj123-lab/agentic-rag-financial-parser&lt;br&gt;
&lt;strong&gt;Live Demo:&lt;/strong&gt; ambuj-portfolio-v2.netlify.app&lt;br&gt;
&lt;strong&gt;LinkedIn:&lt;/strong&gt; linkedin.com/in/ambuj-tripathi-042b4a118&lt;br&gt;
_&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;&lt;em&gt;Has anyone else dealt with footnote-heavy PDFs or failed LlamaParse attempts? How did you handle them? Drop your approach in the comments — I'd love to compare notes.&lt;/em&gt;&lt;/strong&gt;&lt;br&gt;
_&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If you found this useful, drop a ❤️ and follow for more production RAG content!&lt;/p&gt;

</description>
      <category>ai</category>
      <category>rag</category>
      <category>langchain</category>
      <category>python</category>
    </item>
  </channel>
</rss>
