DEV Community

Aditya D
Aditya D

Posted on

Samajh — a reader for the letters my mother couldn't read

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

Samajh — a reader for the letters my mother couldn't read

The letter

Every few weeks, an official letter arrives at home — a bank notice, an insurance renewal, a hospital bill, or another document written in formal English.

My mother can understand the underlying idea, but the language of the document makes the important part difficult to follow.

The problem is not always complexity. Sometimes one sentence determines whether something needs to be paid, when it is due, or what action is expected.

I kept seeing someone else read and explain these documents.

Understanding depended on another person.

I wanted to build something that could help with that first step.

That became Samajh.


What I built

Samajh is a document reader and explainer designed for people who may understand their own language better than the formal language used in official documents.

You provide a document, and Samajh explains it in simple language and in the reader's preferred language.

Instead of producing one large AI-generated summary, it focuses on five things:

  1. What is this?
  2. What is the gist?
  3. What does it want from me, and by when?
  4. What numbers and dates matter?
  5. What is easy to miss, and what is uncertain?

Every factual explanation is connected back to the source document.

Input formats

Samajh can work with:

  • Plain text
  • .txt
  • .md
  • .csv
  • PDF documents
  • Photos/images when the configured AI model supports image input

It supports multiple languages, including English and Telugu, so the explanation can be presented in the reader's own language.

There is also deliberately no account, dashboard, or document history.

The goal is simple:

Understand the document first.


Why quotes matter

Large language models can produce fluent answers while still being wrong.

That becomes a serious problem when the document contains:

  • payment amounts
  • deadlines
  • dates
  • policy conditions
  • instructions
  • warnings

So Samajh does not treat the model's explanation as the final truth.

For each factual claim, the system tries to preserve evidence from the original document.

The flow is roughly:

Claim → Source quote → Verification

The quote is independently checked against the extracted document text.

If the system cannot verify the claim, it should not quietly present it as a fact.

Instead, it can mark information as uncertain or leave it out.

This was one of the most important design decisions in the project.

The goal is not to make AI sound confident. The goal is to make the result checkable.


How it works

flowchart TD
    A[Upload document] --> B[Extract text]
    B --> C[Structured AI analysis]
    C --> D[Validate model output]
    D --> E[Normalize and recover output]
    E --> F[Ground claims to source]
    F --> G[Reader-friendly explanation]

Top comments (2)

Collapse
 
junyoung_arche profile image
Junyoung Park •

"The goal is not to make AI sound confident. The goal is to make the result checkable." This is the line I'd frame.

We built something close to this (an offline helper that explains English emails in Korean) and hit one trap you might already be guarding against: the quote existed in the source, so verification passed, but the quote didn't actually support the claim. Our model once said "interview required" and cited a sentence about the application deadline. What fixed most of it for us was requiring the quote to also contain a word tied to the field it supports (a date for "by when", a currency/number for "how much"), not just to appear in the document.

The other thing small local models kept getting wrong in our tests was frequency, not dates: "every two weeks" came back as "twice a week". Does Samajh's "what numbers and dates matter" section catch that kind of rewording, or does it only check that the number itself appears?

Really nice that there's no account or history. For letters like these that's the right default.

Collapse
 
koda2026 profile image
Harun - solo dev •

@aditya_d This hit home. Hard. 💯

I’m Harun (12yo founder of HYNAWEB), and I’m building KODA—an AI coding mentor designed for the exact same demographic: users who understand their native language (Tamil/Hindi/Telugu) better than formal English, and who often lack expensive hardware.

Your core insight—"The goal is not to make AI sound confident. The goal is to make the result checkable"—is literally Article 4 (Truthfulness) and Article 5 (Code Honesty) of my KODA Constitution.

We face identical traps:

  1. Hallucinated Dates/Amounts: Like Junyoung mentioned, models love to rephrase frequencies ("every two weeks" → "twice a week"). In KODA, we enforce strict grounding: if the model extracts a date/number, it must cite the exact substring from the source code/doc. No citation = Refusal.
  2. Language Barrier: Most dev tools assume English fluency. KODA replies in the user’s native script but keeps code/syntax in English. Just like Samajh explains bills in Telugu but keeps the numbers intact.

I’d love to hear how you handled the OCR extraction quality for handwritten vs. printed letters? That’s our biggest bottleneck for mobile users taking photos of error logs.

Keep building for your mom. Tools like this are why we code. 🐯❤️