This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
My mom gets a lot of paperwork: electricity bills from CFE, letters from the bank, notices from the SAT (Mexico's tax authority), and every so often something scary-looking like a court summons. They're all written in dense bureaucratic Spanish, and they all end the same way: a phone call to me. "Mijo, ¿qué es esto? ¿Tengo que pagar algo?"
So I built her Traductor de Papeles (Paperwork Translator): a mobile web app where she takes a photo of a document, or asks a question out loud, and gets an answer in plain, warm Mexican Spanish:
- What it is and who sent it.
- When and how much: the deadline and the amount, but only if they're actually printed on the page.
- What to do, in a few short steps.
- A clear fraud warning if it asks for personal data, pushes fake urgency, or looks off.
If there's a deadline, one tap on "Agregar a mi calendario" downloads a reminder that goes off 3 days and 1 day before. "Mis papeles" shows her past documents and what's coming due.
Everything is designed for her, not for me: two huge buttons on the home screen ("Tomar foto" and "Preguntar con voz"), big text, high contrast, a loading message that tells her how long to wait, and no passwords, just a family code entered once. It installs to her home screen like any other app.
The model has strict rules: it explains, it doesn't advise. It never makes up an amount or a date, never asks for passwords, PINs or card numbers, and for anything serious it tells her to call her son. If the photo is blurry, it says so and asks for another one instead of guessing.
Demo
🔗 https://paperwork-web-lyry.onrender.com: live Oct 4–7, 2026 only. Gemma 4 runs on a rented GPU, so I'm keeping the demo up for three days to keep costs bounded. Family code for judges: brisa-girasol-colibri-cactus
| Home | Result | History | Voice |
|---|---|---|---|
![]() |
![]() |
![]() |
![]() |
Code
RZEROSTERN
/
hacktoberfest-week1
The first development for the MLH Hacktoberfest 2026. A solution for a friend.
Paperwork Translator · Traductor de Papeles
Live demo: Sun Oct 4 – Wed Oct 7, 2026 only. Gemma 4 runs on a rented GPU, so the demo is up for 3 days to keep costs bounded. After that, everything here still runs on your own machine (see Run it locally).
My mom gets a lot of paperwork: electricity bills from CFE, letters from the bank, notices from the SAT (the tax authority), and now and then something scary-looking like a court summons. Every one of them is written in a dense bureaucratic Spanish, and every one of them ends up as a phone call to me: "Mijo, ¿qué es esto? ¿Tengo que pagar algo?"
Paperwork Translator is a mobile web app built for her. She takes a photo of the document (or asks a question out loud) and gets an answer in plain, warm Mexican Spanish:
- What it is…
How I Built It
The open-source AI at the core:
-
Gemma 4 (
gemma4:e4b, Apache 2.0) via Ollama reads the photo and fills a strict JSON schema: document type, issuer, deadline, amount, steps, fraud flag and reason, confidence, and a plain-Spanish explanation. I pass the Pydantic schema to Ollama'sformatfield, validate the output, retry once, and if it's still invalid return an honest "No pude leer bien su papel" instead of guessing. -
faster-whisper (Whisper
small, MIT) transcribes her spoken questions in Spanish. Gemma 4 then answers using only the structured data of the document she's asking about.
Around it: a FastAPI backend (Python, SQLModel/SQLite) and a Nuxt 4 PWA. The Nuxt server proxies every API call, so the API address and the family code never reach the browser.
Things I learned by testing instead of trusting:
- Before building anything, I ran a model spike on anonymized samples. Gemma 4 read amounts and issuers correctly, but thinking mode doubled latency (~1,100 hidden tokens per photo). Turning it off cut a photo from ~60 s to ~20 s on my M2 Pro, but then the model mistook a document's issue date for a deadline. One explicit prompt rule ("issue dates and 'inmediato' are not deadlines") fixed it. That's the kind of mistake that would make my mom pay something she doesn't owe, so it now has its own test.
- Gemma 4 lists an audio capability, but Ollama doesn't document audio input yet, so I went with Whisper → text → Gemma 4.
- Privacy is enforced in code, not just promised: photos and recordings are processed in memory (even Whisper decodes from an in-memory buffer) and discarded. Only the extracted fields are saved, and logs never contain document text.
Deployment: the web app and the private API (including Whisper) run on Render from a Blueprint. Render has no GPUs, and the smaller Gemma 4 on CPU took 1–6 minutes per photo in my tests, so Gemma 4 runs on a DigitalOcean GPU Droplet. Ollama there listens only on localhost, behind Caddy with HTTPS and a bearer token.
Every technical decision, with the measurements behind it, is in docs/DECISIONS.md.
Why Does Open Innovation Matter?
Because these are my mom's bills, bank letters and tax notices.
- Privacy. With open weights, her documents go to a model I run, on hardware I rent or own. They're not sent to a third party's API to be logged or kept under someone else's retention policy. For this app, that's the whole point.
- No meter running. There are no per-image or per-token fees. She can photograph every envelope that arrives and I never think about a bill. On my own laptop, the cost per request is zero.
- I control the model. I pinned the exact version, measured it on real documents, found its failure mode, and fixed it. I tuned it (thinking off) for speed. Nobody can deprecate it, change its behavior overnight, or raise its price.
- It runs anywhere. The same code runs on my MacBook, a rented GPU, or (slowly) a plain CPU server. If the demo GPU goes away on Wednesday, the app doesn't. It goes back to running at home.
A closed API would have meant sending a family's financial paperwork to someone else's servers, paying for every photo, and hoping the model didn't change underneath us.
My Agent Session
I built this with Claude Code as a pair programmer. It helped with the scaffolding, tests, docs, and a strict GitFlow with one PR per step (12 PRs total). That was during development only: at runtime the product uses no closed-model APIs.
Prize Categories
-
Best Use of Render: Render hosts the user-facing front end and a private API service that runs the open Whisper speech-to-text model, all deployed from a
render.yamlBlueprint.




Top comments (0)