DEV Community

Cover image for I Built Letterbox to Explain Difficult Letters. It Caught an AI Adding a Deadline.
Tombri Bowei
Tombri Bowei Subscriber

Posted on

I Built Letterbox to Explain Difficult Letters. It Caught an AI Adding a Deadline.

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

The hardest part of a letter is sometimes what happens after you finish reading it.

You understood the individual words. You read it twice. But you are still wondering whether that date is a deadline, whether that instruction applies to you, and whether you have missed something important.

So you ask someone you trust:

“Can you look at this for me?”

That is the moment I wanted Letterbox to help with.

When I shared it with my cousin, her mom used it herself. My cousin helped her a little and afterwards told me it really helped her mum.

Her mom used it herself. That detail matters to me.

The project reached someone beyond my own test cases. And the little bit of help from my cousin matters too: family support was still there. A tool like this should make it easier to help each other while giving the person holding the letter a way to explore it themselves.

I want Letterbox to make room for both: understanding something on your own and asking someone you trust to check it with you.

But adding AI creates another problem. A confident explanation can be easier to read—and still be wrong.

During testing, mine added a deadline to an instruction that did not contain one.

That became the most important engineering story in this build.

Letterbox is a letter reader where the explanation comes with evidence you can inspect.

Upload a one-page photo or paste text. It identifies dates, amounts, and requested actions, then uses Gemma to simplify selected sentences.

Each explanation sits beside its supporting quote. Click it, and that passage is highlighted in the original.

There is read-aloud support, an optional lookup for official contact pages, and a reminder flow that requires you to check the date before proceeding.

The goal is simple: make the letter easier to understand while keeping the person reading it in control.

Demo

Try Letterbox.

Watch on YouTube →

Try the fictional appointment walkthrough and select a source quote to follow it back to the letter.

Then paste fictional text and choose Explain my letter to exercise, Gemma. The built-in examples use curated explanations and are labelled as walkthroughs.

The film combines fictional walkthroughs with recorded tests and sponsor evidence. It is a product demo, rather than a recording of my cousin or her mum. Narration was generated with ElevenLabs.

The public app runs on Render independently of my laptop. Its free instance may take about a minute to wake after inactivity.

Use fictional or redacted input in the public test. Photos and text go to Render; selected source sentences go to Google’s hosted Gemma API. Google’s free API may use submitted content to improve its products. The installed local edition has a separate Ollama-based processing path.

Code

Explore the source, setup, and automated checks. →

The repository contains the deployed cloud edition: a Python server, browser interface, English Tesseract OCR, hosted Gemma adapter, and operational integrations.

The app code is MIT licensed; model and provider terms apply separately.

How I Built It

The quote was correct. The explanation added a deadline.

In a real local Gemma 3 1B test, a fictional payment letter contained:

If you have already paid, please contact our team so we can check your account.

The proposed explanation added:

before October 20, 2026

That date appeared elsewhere in the letter. It did not appear in the sentence supporting this instruction.

The quote matched perfectly. The explanation sounded plausible. Yet it had attached a deadline to an instruction that did not state one.

Checking that a quote exists would not have caught this.

Letterbox’s numeric check flagged the added date. The app showed the source sentence instead of the proposed paraphrase and retained the warning.

That failure explains the design better than a flawless demo ever could. The checks need to sit outside the model, where I can inspect and change them.

A narrow job for the model

I built the pipeline around five steps:

  1. Read: Tesseract transcribes a photo, or the user pastes text.
  2. Select: Python identifies candidate sentences containing dates, amounts, or actions.
  3. Simplify: Gemma paraphrases those selected sentences in a structured response.
  4. Check: code checks source matches, numbers, and selected condition wording.
  5. Review: the interface connects explanations to their sources. Reminders require date confirmation.

The model helps with wording. The code performs specific checks. The reader can inspect the original.

These checks support review; they do not prove correctness. OCR can misread a photo, and matching quotes and numbers can still accompany a mistaken interpretation.

What I verified

Test Observed result
Three fictional letters, local Gemma 3 1B 10/10 supporting quotes matched; one proposed explanation flagged for added numbers
One fictional request through public hosted Gemma 4/4 quotes matched in 4.65 seconds; no numeric or condition flags in that run
One fictional English photo through public OCR 11 text lines in 20.76 seconds, including the test date and arrival instruction
Public sponsor endpoints ElevenLabs returned speech; SerpApi returned five allowed official-domain results
Sentry cloud trace Four real spans covering analysis, source selection, Gemma inference, and quote checking
GitHub Actions 20 tests, JavaScript syntax checks, Docker build, and non-root startup verification with Tesseract available

I also tested two separate local workflow implementations:

  • Mastra: a three-step workflow loaded a fictional source, ran a checked Gemma analysis, and prepared the result for human review.
  • Temporal: after a worker stopped during a simulated transient failure, a replacement worker resumed the same workflow and completed its activity on attempt two.

Neither performed a calendar action. These are local developer experiments, separate from the public Render container.

The test set is small. My cousin’s feedback is a useful first handover, not a measured usability result.

Next, I want to observe where people hesitate, which explanations help, and what still sends them back to a family member for clarification. Those moments will tell me what to improve.

Why Does Open Innovation Matter?

For a letter reader, where the letter goes is part of the product.

The installed edition uses Gemma’s open weights through Ollama. After setup and model download, the core reader can operate without sending the letter to a hosted model.

English photo transcription uses Tesseract, an open-source OCR engine.

A closed API could power the public experience. It would not provide the same downloadable model for that local processing path.

Open weights let me test smaller models on available hardware. Keeping the evidence checks outside the inference adapter lets me change models while preserving the review workflow.

The public demo makes a different tradeoff: easier access, with cloud processing disclosed. Optional cloud speech and search require separate consent. Sentry and Atlas receive selected operational metrics rather than submitted letter text.

That freedom is practical:

  • Choose where processing happens.
  • Inspect the surrounding code.
  • Change the model to fit the hardware.
  • Improve the checks when testing exposes a failure.

The added-deadline example made that last point concrete. I could identify a failure and change the surrounding behavior instead of hoping the next answer would be better.

My Agent Session

I built Letterbox with an AI coding assistant, which helped implement the interface, processing pipeline, tests, and deployment.

The public repository and GitHub Actions history show the implementation and automated checks.

These are built evidence rather than a published DevRelay session recording.

Prize Categories

Verified in the public deployment:

  • Best Use of Gemma: source-sentence simplification with checks and visible supporting quotations.
  • Best Use of Render: the public Docker service provides the interface, OCR, and hosted-model gateway independently of my laptop.
  • Best Use of ElevenLabs: optional read-aloud speech verified through the public endpoint, plus demo narration.
  • Best Use of Sentry Agent Tracing: real analysis traces, timings, and token counts, without submitted letter text.
  • Best Use of SerpApi: official-contact search with HTTPS results filtered to allowed NHS and government domains. It locates pages; it does not authenticate letters.
  • Best Use of GitHub Copilot: entry through the category’s GitHub Actions route, covering tests, syntax checks, Docker build, and runtime verification.

Demonstrated separately in local testing:

  • Best Use of MongoDB Atlas: anonymous metrics insertion and readback passed locally. Cloud writes remain unverified while hosting network access is pending.
  • Best Use of Mastra: a real SDK workflow orchestrated source loading, checked Gemma analysis, and prepared for human review.
  • Best Use of Temporal: a real SDK recovery experiment resumed after a worker restart and completed the retried activity.

Top comments (1)

Collapse
 
kevinpruett023_kevinpruet profile image
Lee •

Great!
I wanna have meaningful conversation about collaboration with you.
How about you?