<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: DHARANEE</title>
    <description>The latest articles on DEV Community by DHARANEE (@dharanee_dd).</description>
    <link>https://dev.to/dharanee_dd</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4110762%2Fcd550a8d-fd20-4e44-bf88-9a7ed84d149f.png</url>
      <title>DEV Community: DHARANEE</title>
      <link>https://dev.to/dharanee_dd</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/dharanee_dd"/>
    <language>en</language>
    <item>
      <title>AI is Evolving</title>
      <dc:creator>DHARANEE</dc:creator>
      <pubDate>Sat, 05 Sep 2026 06:58:16 +0000</pubDate>
      <link>https://dev.to/dharanee_dd/-8jg</link>
      <guid>https://dev.to/dharanee_dd/-8jg</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/dharanee_dd/building-an-ai-powered-invoice-processing-pipeline-ocr-meets-llms-4pe5" class="crayons-story__hidden-navigation-link"&gt;Building an AI-Powered Invoice Processing Pipeline: OCR Meets LLMs&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/dharanee_dd" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4110762%2Fcd550a8d-fd20-4e44-bf88-9a7ed84d149f.png" alt="dharanee_dd profile" class="crayons-avatar__image" width="96" height="96"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/dharanee_dd" class="crayons-story__secondary fw-medium m:hidden"&gt;
              DHARANEE
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                DHARANEE
                
                
              
              &lt;div id="story-author-preview-content-4580352" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/dharanee_dd" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4110762%2Fcd550a8d-fd20-4e44-bf88-9a7ed84d149f.png" class="crayons-avatar__image" alt="" width="96" height="96"&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;DHARANEE&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/dharanee_dd/building-an-ai-powered-invoice-processing-pipeline-ocr-meets-llms-4pe5" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Sep 5&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/dharanee_dd/building-an-ai-powered-invoice-processing-pipeline-ocr-meets-llms-4pe5" id="article-link-4580352"&gt;
          Building an AI-Powered Invoice Processing Pipeline: OCR Meets LLMs
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/productivity"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;productivity&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/architecture"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;architecture&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/machinelearning"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;machinelearning&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
            &lt;a href="https://dev.to/dharanee_dd/building-an-ai-powered-invoice-processing-pipeline-ocr-meets-llms-4pe5#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              &lt;span class="hidden s:inline"&gt;Add&amp;nbsp;Comment&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            7 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
    </item>
    <item>
      <title>Building an AI-Powered Invoice Processing Pipeline: OCR Meets LLMs</title>
      <dc:creator>DHARANEE</dc:creator>
      <pubDate>Sat, 05 Sep 2026 06:57:48 +0000</pubDate>
      <link>https://dev.to/dharanee_dd/building-an-ai-powered-invoice-processing-pipeline-ocr-meets-llms-4pe5</link>
      <guid>https://dev.to/dharanee_dd/building-an-ai-powered-invoice-processing-pipeline-ocr-meets-llms-4pe5</guid>
      <description>&lt;h1&gt;
  
  
  Building an AI-Powered Invoice Processing Pipeline: OCR Meets LLMs
&lt;/h1&gt;

&lt;p&gt;Invoice processing sounds simple until you actually deal with invoices coming from different vendors.&lt;/p&gt;

&lt;p&gt;Different layouts, different field names, scanned documents, inconsistent formatting, missing information, duplicate invoices, and even suspicious transactions can make manual processing slow and error-prone.&lt;/p&gt;

&lt;p&gt;As part of our project, we worked on an &lt;strong&gt;AI-powered invoice processing pipeline&lt;/strong&gt; designed to automate much of this process — from extracting information from invoices to validating the data and identifying potential duplicates or fraud.&lt;/p&gt;

&lt;p&gt;I worked as the &lt;strong&gt;Lead Frontend Developer&lt;/strong&gt;, primarily responsible for the frontend experience and its integration with the backend services. At the same time, I worked closely with the team to understand the core processing pipeline and how the different services connected together.&lt;/p&gt;

&lt;p&gt;This article walks through how the system works and some of the engineering decisions behind it.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. The Problem
&lt;/h2&gt;

&lt;p&gt;In a typical supply-chain environment, invoices can arrive in many different formats. Some may be digitally generated PDFs, while others may be scanned documents or images.&lt;/p&gt;

&lt;p&gt;Processing them manually means employees have to read the documents, identify important fields, enter the information into a system, verify the calculations, and check whether the invoice has already been processed.&lt;/p&gt;

&lt;p&gt;This becomes time-consuming and introduces opportunities for human error.&lt;/p&gt;

&lt;p&gt;Our goal was to build a system that could automate as much of this workflow as possible while still keeping validation and human review in the loop when the system was uncertain.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. The High-Level Idea
&lt;/h2&gt;

&lt;p&gt;The basic idea behind our system was to create a pipeline where each component has a specific responsibility.&lt;/p&gt;

&lt;p&gt;An invoice first enters the system as a PDF or image. Instead of sending the document directly to a language model, we first use &lt;strong&gt;OCR (Optical Character Recognition)&lt;/strong&gt; to extract the text from the document.&lt;/p&gt;

&lt;p&gt;That extracted text is then passed to an &lt;strong&gt;LLM&lt;/strong&gt;, specifically Llama 3.3 70B, which understands the context of the invoice and converts the unstructured OCR output into structured information such as the vendor name, invoice number, dates, tax values, and total amount.&lt;/p&gt;

&lt;p&gt;The structured result is then passed through validation and business checks. The system also performs duplicate detection and fraud-oriented checks before the final information is stored and made available through the application dashboard.&lt;/p&gt;

&lt;p&gt;In simplified form:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Invoice
   ↓
OCR
   ↓
Raw Text
   ↓
Llama 3.3 70B
   ↓
Structured JSON
   ↓
Validation
   ↓
Duplicate &amp;amp; Fraud Checks
   ↓
Clean Structured Data
   ↓
Database + Dashboard
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important idea here is that &lt;strong&gt;the LLM isn't responsible for everything&lt;/strong&gt;. It is one component inside a larger processing pipeline.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Tech Stack &amp;amp; Why
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Tesseract 5 for OCR
&lt;/h3&gt;

&lt;p&gt;The first major component of the pipeline is OCR.&lt;/p&gt;

&lt;p&gt;We used &lt;strong&gt;Tesseract 5&lt;/strong&gt; to extract text from invoice documents.&lt;/p&gt;

&lt;p&gt;You might wonder:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Why not simply send the invoice directly to an LLM?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's a reasonable question, especially because modern vision-language models can process images directly.&lt;/p&gt;

&lt;p&gt;For our architecture, separating OCR from language understanding gave us more control over the pipeline.&lt;/p&gt;

&lt;p&gt;OCR converts the visual document into machine-readable text:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Invoice Image/PDF
       ↓
     Tesseract
       ↓
Extracted Text
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once we have the text, the language model can focus on understanding the content rather than performing the initial character recognition.&lt;/p&gt;

&lt;p&gt;This separation also makes debugging easier. If something goes wrong, we can determine whether the problem came from the OCR stage or from the LLM's interpretation of the extracted text.&lt;/p&gt;




&lt;h3&gt;
  
  
  Llama 3.3 70B via Groq Cloud
&lt;/h3&gt;

&lt;p&gt;Once OCR produces raw text, we need to understand what that text actually means.&lt;/p&gt;

&lt;p&gt;This is where &lt;strong&gt;Llama 3.3 70B&lt;/strong&gt; comes in.&lt;/p&gt;

&lt;p&gt;Invoices are not standardized documents. One vendor might use:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Invoice Number
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;while another might use:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Bill No.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and another could use:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Reference ID
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A purely rule-based parser would require us to continuously add rules for these variations.&lt;/p&gt;

&lt;p&gt;An LLM provides a more flexible approach because it can understand the context and semantic meaning of the extracted text.&lt;/p&gt;

&lt;p&gt;We used Llama 3.3 70B through &lt;strong&gt;Groq Cloud&lt;/strong&gt; and designed prompts that instructed the model to return structured information in a predefined JSON format.&lt;/p&gt;

&lt;p&gt;For example, instead of simply asking:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Extract the invoice details.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;we could provide explicit instructions about the expected fields, formatting requirements, missing values, and output structure.&lt;/p&gt;

&lt;p&gt;The result is something like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"invoice_number"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"INV-1024"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"vendor"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ABC Technologies"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"subtotal"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;20000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"tax"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;3600&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"total"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;23600&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The key point is that the LLM isn't treated as an unquestionable source of truth. Its output goes through additional validation before being accepted by the system.&lt;/p&gt;




&lt;h3&gt;
  
  
  FastAPI
&lt;/h3&gt;

&lt;p&gt;For the backend API layer, we used &lt;strong&gt;FastAPI&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The frontend and backend were developed as separate layers, so the frontend communicates with the backend through APIs.&lt;/p&gt;

&lt;p&gt;FastAPI was a good fit because the backend was primarily API-driven and needed to interact with several processing components, including OCR, the LLM service, database operations, and file storage.&lt;/p&gt;

&lt;p&gt;It also provides automatic API documentation and request validation through Pydantic, which makes development and testing more convenient.&lt;/p&gt;

&lt;p&gt;At a high level:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Next.js Frontend
       ↓
   REST API
       ↓
    FastAPI
       ↓
Processing Services
       ↓
Database / Storage
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This separation also allowed the frontend and backend to evolve independently.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Architecture
&lt;/h2&gt;

&lt;p&gt;One thing we wanted to avoid was building the entire application as one large script.&lt;/p&gt;

&lt;p&gt;Instead, we designed the application as a &lt;strong&gt;modular, multi-tenant system with nine service layers&lt;/strong&gt;, where different responsibilities were separated into appropriate components.&lt;/p&gt;

&lt;p&gt;At a high level, the architecture can be viewed as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                Frontend
                   ↓
              API Layer
                   ↓
          Authentication / RBAC
                   ↓
          Invoice Processing
                   ↓
          OCR + LLM Pipeline
                   ↓
       Validation / Fraud Checks
                   ↓
       Duplicate Detection Layer
                   ↓
          Database / Storage
                   ↓
        Notifications / Reporting
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The exact responsibilities are separated across the application's service layers rather than putting everything into a single processing function.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;multi-tenant design&lt;/strong&gt; is particularly important because invoices and users belonging to one organization should remain logically separated from another organization's data.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Organization A
 ├── Users
 ├── Invoices
 └── Reports

Organization B
 ├── Users
 ├── Invoices
 └── Reports
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This makes the system more suitable for a SaaS-style environment where multiple organizations can use the same application.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. The Hard Part #1: Duplicate Detection
&lt;/h2&gt;

&lt;p&gt;One of the more interesting parts of the project was duplicate invoice detection.&lt;/p&gt;

&lt;p&gt;At first glance, it might seem easy:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Just compare the invoice number.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;But invoice numbers aren't globally unique.&lt;/p&gt;

&lt;p&gt;Two different vendors can legitimately generate:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;INV-1001
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So invoice number alone isn't sufficient.&lt;/p&gt;

&lt;p&gt;We therefore used a &lt;strong&gt;two-tier approach&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Tier 1: SHA-256 Hashing
&lt;/h3&gt;

&lt;p&gt;The first level handles exact duplicate files.&lt;/p&gt;

&lt;p&gt;When an invoice file is uploaded, we generate a &lt;strong&gt;SHA-256 hash&lt;/strong&gt; for the file.&lt;/p&gt;

&lt;p&gt;The same file will produce the same hash.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Invoice.pdf
     ↓
 SHA-256
     ↓
ABC123...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If exactly the same file is uploaded again, its SHA-256 value will match the existing record.&lt;/p&gt;

&lt;p&gt;This allows us to detect exact file duplicates efficiently without having to compare the complete contents of every document.&lt;/p&gt;

&lt;h3&gt;
  
  
  Tier 2: Logical Triplet Matching
&lt;/h3&gt;

&lt;p&gt;But there is a problem with relying only on file hashes.&lt;/p&gt;

&lt;p&gt;Imagine the same invoice is downloaded and re-saved as another PDF.&lt;/p&gt;

&lt;p&gt;Maybe:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the metadata changed,&lt;/li&gt;
&lt;li&gt;the PDF was compressed,&lt;/li&gt;
&lt;li&gt;the file was scanned again,&lt;/li&gt;
&lt;li&gt;or the document was generated slightly differently.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The content may represent the same invoice, but the binary file itself is different.&lt;/p&gt;

&lt;p&gt;In that case:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;File A → SHA-256 → Hash A

File B → SHA-256 → Hash B
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The hashes won't match.&lt;/p&gt;

&lt;p&gt;That's where &lt;strong&gt;logical matching&lt;/strong&gt; comes in.&lt;/p&gt;

&lt;p&gt;We compare a combination of important invoice attributes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Vendor Name
+
Invoice Number
+
Total Amount
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is our logical triplet.&lt;/p&gt;

&lt;p&gt;If these values match an existing invoice, the system can flag it as a potential duplicate even when the physical files are different.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why two methods?
&lt;/h3&gt;

&lt;p&gt;Because each method solves a different problem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SHA-256:&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Is this exactly the same file?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;Logical triplet matching:&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Does this appear to represent the same invoice?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Using both gives us a stronger duplicate-detection mechanism than relying on either one alone.&lt;/p&gt;

&lt;p&gt;It also illustrates an important engineering principle: &lt;strong&gt;sometimes the best solution isn't one sophisticated algorithm, but multiple simpler checks working together.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  6. The Hard Part #2: Fraud Heuristics
&lt;/h2&gt;

&lt;p&gt;Duplicate detection is only one part of invoice verification.&lt;/p&gt;

&lt;p&gt;We also implemented rule-based checks to identify potentially suspicious invoices.&lt;/p&gt;

&lt;p&gt;The purpose wasn't to claim that the system could definitively prove an invoice was fraudulent.&lt;/p&gt;

&lt;p&gt;Instead, the system identifies &lt;strong&gt;anomalies and risk indicators&lt;/strong&gt; that deserve additional attention.&lt;/p&gt;

&lt;p&gt;Examples of checks include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Duplicate invoice detection&lt;/li&gt;
&lt;li&gt;Unusual invoice amounts&lt;/li&gt;
&lt;li&gt;Suspicious patterns in invoice information&lt;/li&gt;
&lt;li&gt;Missing or inconsistent fields&lt;/li&gt;
&lt;li&gt;Mathematical inconsistencies&lt;/li&gt;
&lt;li&gt;Policy violations&lt;/li&gt;
&lt;li&gt;Unexpected invoice patterns&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For example, if:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Subtotal = ₹10,000
Tax      = ₹1,800
Total    = ₹25,000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;the numbers don't add up.&lt;/p&gt;

&lt;p&gt;That doesn't automatically mean the invoice is fraudulent, but it is a strong reason to flag it for review.&lt;/p&gt;

&lt;p&gt;This is why we treated fraud detection as a &lt;strong&gt;risk and decision-support mechanism&lt;/strong&gt;, rather than an absolute "fraud/not fraud" verdict.&lt;/p&gt;

&lt;p&gt;The system can assign risk indicators and allow suspicious invoices to move toward manual review.&lt;/p&gt;




&lt;h2&gt;
  
  
  7. What I'd Do Differently / Lessons Learned
&lt;/h2&gt;

&lt;p&gt;Building the project taught me that integrating AI into an application is very different from simply calling an AI API.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Don't blindly trust model output
&lt;/h3&gt;

&lt;p&gt;An LLM can produce convincing but incorrect information.&lt;/p&gt;

&lt;p&gt;This is particularly dangerous when processing financial documents.&lt;/p&gt;

&lt;p&gt;Our solution was to place validation after the LLM rather than treating its output as ground truth.&lt;/p&gt;

&lt;p&gt;If something doesn't make sense, the system should be able to catch it.&lt;/p&gt;




&lt;h3&gt;
  
  
  2. Separate responsibilities between components
&lt;/h3&gt;

&lt;p&gt;It is tempting to make one AI model responsible for everything.&lt;/p&gt;

&lt;p&gt;But a better architecture is often:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;OCR → Extraction
LLM → Understanding
Rules → Validation
Hashing → Exact Duplicate Detection
Database → Persistence
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each component has a clear responsibility.&lt;/p&gt;

&lt;p&gt;This makes the system easier to debug and maintain.&lt;/p&gt;




&lt;h3&gt;
  
  
  3. The frontend is more than just UI
&lt;/h3&gt;

&lt;p&gt;As the Lead Frontend Developer, one of my biggest takeaways was that frontend development in a system like this isn't only about creating screens.&lt;/p&gt;

&lt;p&gt;The frontend has to understand:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;API states&lt;/li&gt;
&lt;li&gt;Authentication&lt;/li&gt;
&lt;li&gt;Error handling&lt;/li&gt;
&lt;li&gt;Loading states&lt;/li&gt;
&lt;li&gt;Processing states&lt;/li&gt;
&lt;li&gt;Role-based access&lt;/li&gt;
&lt;li&gt;Backend responses&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For an invoice-processing application, a user needs to know whether an invoice is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Uploaded
   ↓
Processing
   ↓
Approved
   ↓
Rejected
   ↓
Or Requires Review
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A good frontend needs to communicate that entire lifecycle clearly.&lt;/p&gt;




&lt;h2&gt;
  
  
  8. Wrap-up
&lt;/h2&gt;

&lt;p&gt;Building this project gave us a practical look at how &lt;strong&gt;OCR, LLMs, APIs, cloud storage, databases, validation, and frontend systems can work together as one application&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The most important lesson for me was that AI doesn't replace traditional software engineering — it becomes one component within it.&lt;/p&gt;

&lt;p&gt;In our case:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Computer Vision
      +
Generative AI
      +
Backend Engineering
      +
Database
      +
Cloud Storage
      +
Rule-Based Validation
      +
Frontend
      ↓
Complete Invoice Processing System
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is still plenty that could be improved, especially around more advanced fraud detection, near-duplicate matching, model evaluation, and asynchronous processing at larger scale.&lt;/p&gt;

&lt;p&gt;But that's also what makes engineering projects interesting: the first working system is not the final system.&lt;/p&gt;

&lt;p&gt;I'd love to hear how others would approach invoice automation, duplicate detection, or AI-assisted document processing. Feel free to share your thoughts or questions in the comments.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>architecture</category>
      <category>machinelearning</category>
    </item>
  </channel>
</rss>
