<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: V.karthik sai ram</title>
    <description>The latest articles on DEV Community by V.karthik sai ram (@vkarthik_sairam_65258a1).</description>
    <link>https://dev.to/vkarthik_sairam_65258a1</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4148473%2F86cfe9d8-a524-433e-9ea5-ac291d8fe9a9.webp</url>
      <title>DEV Community: V.karthik sai ram</title>
      <link>https://dev.to/vkarthik_sairam_65258a1</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/vkarthik_sairam_65258a1"/>
    <language>en</language>
    <item>
      <title>An OCR'd wrong date becomes a confidently wrong Hindsight memory</title>
      <dc:creator>V.karthik sai ram</dc:creator>
      <pubDate>Tue, 29 Sep 2026 04:07:31 +0000</pubDate>
      <link>https://dev.to/vkarthik_sairam_65258a1/an-ocrd-wrong-date-becomes-a-confidently-wrong-hindsight-memory-31o5</link>
      <guid>https://dev.to/vkarthik_sairam_65258a1/an-ocrd-wrong-date-becomes-a-confidently-wrong-hindsight-memory-31o5</guid>
      <description>&lt;p&gt;A vision model reads a smudged diary page and gets the 8 in "18 June" wrong. It reads it as a 3. On its own, that's a typo. Once it's stored in an agent's long-term memory, it's something worse: a fact the agent will cite, with a source chip, in a calm voice, to a lawyer standing in front of a judge.&lt;br&gt;
I spent more time on the path between "the model read a photo" and "this is now a memory" than on any other part of Tareekh. This article is about that path, and why I put a person in the middle of it.&lt;br&gt;
What goes in&lt;br&gt;
Tareekh is a memory for a litigator's practice. The inputs are whatever the lawyer already produces, which in an Indian district court means a lot of paper:&lt;br&gt;
phone photos of pocket-diary pages, handwritten, often shot at an angle under a tube light&lt;br&gt;
certified copies of order sheets, with several dated rows per page&lt;br&gt;
deeds, 1-B land records, encumbrance certificates, legal notices, police petition receipts, a WhatsApp export&lt;br&gt;
typed notes and pasted text&lt;br&gt;
Each upload is read, split into one entry per hearing, filed under the right case and retained in Hindsight. An agent answers questions from that memory with citations.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9hmtsui5prueg79e3rfy.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9hmtsui5prueg79e3rfy.jpg" alt=" " width="725" height="315"&gt;&lt;/a&gt;&lt;br&gt;
Tareekh's architecture. Everything on the left passes through extract, segment, repair and review before anything is retained.&lt;br&gt;
The test practice I use has 79 uploads across five cases. The diary photos are deliberately rough: barrel distortion, chromatic aberration, vignetting, soft corners, sensor noise in the shadows, and the occasional hand-shake blur. If the pipeline only works on clean scans, it doesn't work.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fswv0874ob9l7igv2s4rd.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fswv0874ob9l7igv2s4rd.jpg" alt=" " width="800" height="1120"&gt;&lt;/a&gt;&lt;br&gt;
A diary page from the test practice, as a phone camera would take it.&lt;br&gt;
Why memory makes extraction errors worse&lt;br&gt;
In a search system, an OCR error costs you a missed hit. In an agent with memory, it's much more expensive, for three reasons.&lt;br&gt;
First, retain is not a dumb insert. Hindsight runs extraction on what you store, pulls out facts and entities, and later consolidates them into observations and mental models. A wrong date doesn't sit in one row. It gets built into a summary of how the case has gone.&lt;br&gt;
Second, the agent is good at sounding sure. Tareekh's bank has a "Cite sources" directive, and every factual line in an answer carries a [n] marker that resolves to a real upload. That's the feature. It also means a wrong fact arrives with a citation attached.&lt;br&gt;
Third, nobody re-reads their own diary. The whole point is that the lawyer stops keeping it in their head. If the memory drifts from the page, nobody finds out until it matters.&lt;br&gt;
So I treated the retain call as a commit to a system of record, not a cache write.&lt;br&gt;
The pipeline, stage by stage&lt;br&gt;
Extract&lt;br&gt;
Text files and .docx are read directly. PDFs use their text layer if there is one; scanned PDFs are rendered page by page and sent to OCR. Images go to a vision model with a prompt that's mostly about what not to do:&lt;br&gt;
OCR_PROMPT = (&lt;br&gt;
    "Transcribe all text in this image exactly as written. It is either a lawyer's handwritten court diary page, "&lt;br&gt;
    "a notebook page, or a scanned court document. Rules: keep original line order and wording, keep abbreviations "&lt;br&gt;
    "and numbers exactly (case numbers, dates, amounts, diary numbers), SKIP words that are struck through, "&lt;br&gt;
    "keep any printed date header. Output only the transcription, no commentary."&lt;br&gt;
)&lt;br&gt;
"Skip struck-through words" is there because diary pages are full of corrections. If a lawyer crosses out "Tuesday" and writes "Thursday", a model that faithfully transcribes both has recorded two days for one hearing.&lt;br&gt;
Segment and resolve&lt;br&gt;
One LLM call gets the raw text and a compact listing of the case registry. It returns entries shaped like {case_id, hearing_date, author, doc_type, text, confidence, reason}.&lt;br&gt;
A diary page covering three matters becomes three entries, and an order sheet with three dated rows becomes three entries. Lawyers write "beach land" or "Gorle" or "OS 214/24", and the model has to map those to a case id without inventing one.&lt;br&gt;
Repair&lt;br&gt;
Then plain code checks the model's work against things that are known to be true:&lt;br&gt;
def repair(entries, source_file, hints, known_ids):&lt;br&gt;
    """Validate the LLM's output against the registry and fill gaps from hints/filename. Pure function."""&lt;br&gt;
    for e in entries:&lt;br&gt;
        conf = float(e.get("confidence") or 0.5)&lt;br&gt;
        cid = e.get("case_id")&lt;br&gt;
        if cid not in known_ids:&lt;br&gt;
            cid = None # the model named a case that doesn't exist&lt;br&gt;
        if not cid:&lt;br&gt;
            found = registry.find_cases(text[:300], limit=1)&lt;br&gt;
            if found:&lt;br&gt;
                cid = found[0]["case_id"]&lt;br&gt;
                conf = min(conf, found[0]["score"] / 100)&lt;br&gt;
        date = &lt;em&gt;valid_date(e.get("hearing_date"))&lt;br&gt;
        if not date and fallback_date: # IMG_20250618&lt;/em&gt;... gives us a date&lt;br&gt;
            date = fallback_date&lt;br&gt;
            conf = min(conf, 0.75)&lt;br&gt;
        if not cid or not date:&lt;br&gt;
            conf = min(conf, 0.3)&lt;br&gt;
(Trimmed from ingest/segment.py .) Notice that repair only ever lowers confidence. A fuzzy alias match can't claim more certainty than its match score. A date taken from the filename is capped at 0.75. An entry missing either a case or a date drops to 0.3.&lt;br&gt;
The file type also overrides the model on one point. If the upload is an image, the entry is a handwritten note. The model is only trusted to tell order sheets and documents apart from notes, because that's a judgment about content, and "was this handwritten" is a fact about the file.&lt;br&gt;
Review&lt;br&gt;
This is the part I'd defend hardest.&lt;br&gt;
The review gate&lt;br&gt;
Every entry lands in a review state first. Auto-confirm only happens when the upload asked for it and every entry in it clears the bar:&lt;/p&gt;

&lt;h1&gt;
  
  
  ingest/pipeline.py
&lt;/h1&gt;

&lt;p&gt;confident = entries and all(e["case_id"] and e["hearing_date"] and e["confidence"] &amp;gt;=&lt;br&gt;
    settings.auto_confirm_threshold&lt;br&gt;
    for e in entries)&lt;br&gt;
if hints.get("auto_confirm") and confident:&lt;br&gt;
    confirm(upload_id, [])&lt;br&gt;
else:&lt;br&gt;
    _status(upload_id, "review")&lt;br&gt;
The threshold is 0.8 by default. It's all or nothing per upload: one doubtful entry on a diary page holds back the other two as well, because they came from the same photo and the same OCR pass, and a page that confused the model once may have confused it twice.&lt;br&gt;
Confirm accepts edits (case, date, text, author, or reject) and only then builds Hindsight items. Two details there came from things breaking. A multi-page document yields one entry per page with the same case and date, and Hindsight rejects two items with the same document_id in one batch, so entries are merged by (case_id, hearing_date) before retaining. And retain always runs with retain_async=True, so Hindsight queues the work and retries on provider rate limits instead of failing the confirm.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbmit0j5dm2jfzdkxe3tw.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbmit0j5dm2jfzdkxe3tw.jpg" alt=" " width="739" height="1600"&gt;&lt;/a&gt;&lt;br&gt;
The Add notes screen. Tareekh reads the upload, files each hearing under its case, and asks you to check anything it isn't sure about.&lt;br&gt;
Before and after&lt;br&gt;
Here's the failure the gate exists for. A diary page has a faint header, and the model takes its best guess at the hearing date. The most tempting guess is the next date written at the bottom, because it's often the most prominent date on the page ("next date 5 Oct", in capitals). Without a gate, that becomes a memory saying a hearing happened on a day it was only scheduled for. Ask "what happened last time" and you get a cited, fluent account of a hearing on the wrong day. Nothing in the system would ever flag it.&lt;br&gt;
After the gate, the same page shows up in review, flagged. The reason field says "date from filename/hint", confidence is 0.75, and the lawyer taps once to accept or fix. The segmentation prompt also now says in so many words: Dates like "next date 5 Oct" are NOT the hearing date.&lt;br&gt;
The uploads that should end up in review are exactly the ones I'd want a human to look at anyway: blurred headers, pages mentioning two cases by the same party name, documents that concern three suits at once.&lt;br&gt;
Source-linked memories&lt;br&gt;
Every memory keeps a link to its source. The note reader shows the original photo above the text that was read from it, so a wrong read is easy to spot.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6wm5lctnulhrm7g9c21k.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6wm5lctnulhrm7g9c21k.jpg" alt=" " width="800" height="500"&gt;&lt;/a&gt;&lt;br&gt;
What I learned&lt;br&gt;
Treat retain like a database commit. The memory layer is going to extract, link and summarize whatever you give it. Garbage in turns into confident garbage out. Validate before you retain, not after.&lt;br&gt;
Let code lower the model's confidence, never raise it. The LLM's self-reported confidence isn't calibrated. I don't trust it as a number. What I do trust is a rule like "if the case id isn't in the registry, it's wrong." Deterministic checks should only ever pull the number down.&lt;br&gt;
Keep the source attached forever. metadata.source_file and upload_id go into every Hindsight item and come back with every recalled fact. When something looks off, the lawyer is one tap from the original photo. That's the real safety net, more than any threshold.&lt;br&gt;
Review is a cost, so spend it carefully. A tool that asks you to confirm every entry won't get used. All-or-nothing per upload, with a 0.8 bar, is my current compromise, and I'm not sure it's the right number. It's an environment variable for that reason.&lt;br&gt;
Make re-uploads safe. document_id = upload:case:date means fixing a mistake and re-confirming replaces the memory rather than adding a second, contradictory one next to it.&lt;br&gt;
If you're feeding documents into agent memory, the extraction step deserves as much design as retrieval. The Hindsight docs explain what retain does with an item once it's stored, which is the best argument for being careful about what you store. For the bigger picture of why this layer is different from a vector index, Vectorize's piece on what agent memory is is worth ten minutes.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>llm</category>
    </item>
  </channel>
</rss>
