This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend.
What I Built
My grandmother gets letters from her bank that she cannot make sense of. It is not one letter, it keeps happening. Each time she does one of three things: she ignores it, she pays without understanding what she is paying for, or she asks someone to read it to her.
A regular chatbot was never going to help. She does not know how to use one, she does not trust them, and she is afraid of exposing her data on the internet.
So I built Papel Claro ("clear paper") for her. You paste the text of a difficult document, or take a photo of it, and it answers the questions a person actually has:
- What is this? The type of document and a summary in two to four short sentences.
- What do I have to do? A numbered list of actions.
- By when, and how much? Dates and amounts, copied as they are written.
- What should I watch out for? Fines, interest, a name sent to a credit bureau.
- What do these words mean? A small glossary of the hard words in the document.
After that you can ask follow-up questions ("can I pay this in instalments?"), and the app answers only from the document. If the document does not say, it tells you so. There is also a button that reads the explanation aloud.
Brazilian contracts and bank letters are written in juridiquês, a dialect of Portuguese that seems designed to be read by lawyers only. A sentence like "o não pagamento poderá ensejar o vencimento antecipado das parcelas vincendas" means "if you do not pay, the bank can demand all the remaining instalments at once". Nobody should need a lawyer to learn that.
Her fear about her data is reasonable. These letters carry a full name, an address, a CPF (the Brazilian tax ID) and the exact size of a debt. That is the last text I want my grandmother to paste into a website. So the model runs on a computer at home, and she opens the page on her phone over the home Wi-Fi. The document never leaves the house.
She has already used it that way, and she thought it was wonderful.
Demo
The app runs locally, so there is no hosted link. The video shows it running on my own computer.
Code
MathFe
/
papel-claro
Explains difficult documents in plain Brazilian Portuguese, running Gemma locally through Ollama.
Papel Claro
Paste a confusing document, or take a photo of it, and get it explained in plain Brazilian Portuguese. Everything runs on your own computer with Gemma through Ollama. The document never leaves the machine.
Built for the DEV Hacktoberfest Weekend Challenge: Build for a Friend.
What it does
Bank letters, rental contracts, collection notices and official notifications in Brazil are written in dense legal Portuguese ("juridiquês"). Papel Claro reads the document and answers the questions a person actually has:
- What is this? The type of document and a summary in two to four short sentences.
- What do I have to do? A numbered list of actions.
- By when, and how much? Dates and amounts, copied exactly as they appear.
- What should I watch out for? Fines, interest, automatic renewal, and common signs of a scam.
- What do these words mean? A small glossary of the hard…
MIT licensed. To run it you need Java 17 or newer and Ollama:
ollama pull gemma3:4b
./mvnw spring-boot:run
Then open http://localhost:8080, or http://<computer-ip>:8080 from a phone on the same Wi-Fi.
How I Built It
Browser -> Spring Boot -> Spring AI ChatClient -> Ollama -> Gemma 3 (4B)
The stack is deliberately boring: Java, Spring Boot 4.1, Spring AI 2.0, and a page in plain HTML, CSS and JavaScript. There is no database, because there is nothing to keep.
One model does everything. gemma3:4b reads images as well as text, so a photo of a letter goes straight to the model. There is no separate OCR step, no second model, and no extra service to install. The same model writes the summary, extracts dates and amounts, and answers the follow-up questions. It is a 3.3 GB download, and on my machine (an AMD RX 9060 XT and 16 GB of RAM) an explanation takes 6 to 10 seconds and a follow-up answer 1 to 3 seconds.
The answer is a Java record, not a blob of text. Spring AI asks the model for JSON and maps it onto records, so the screen can show actions, dates and amounts in separate blocks:
return this.chatClient.prompt()
.user(user -> {
user.text(EXPLAIN_TEMPLATE).param("document", documentText(input));
attachImage(user, input);
})
.call()
.entity(DocumentExplanation.class);
The document goes in as a template parameter, so a contract full of {braces} is never read as template syntax.
What a real run taught me. The first version was written against the documentation and had never been compiled. The first time it ran against a real model, four things came up that no documentation would have told me:
- A closed Ollama froze the page for minutes. Spring AI retries failed calls up to ten times with a growing wait. For a developer that is resilience. For my grandmother, tapping a button on her phone while Ollama happens to be closed on the computer, it is a screen that never answers. With a single retry attempt, the page now says "open Ollama and try again" in about two seconds.
- A log line broke the privacy promise. When the model returned JSON in the wrong shape, the exception message quoted the offending value, and that value was a sentence from the document. The code was logging that message. Now the log records only the kind of failure.
- An empty answer looked like success. A small model sometimes returns JSON with none of the expected fields. The records filled in defaults and the screen showed an empty explanation. That now counts as a failure and triggers a second try.
- The prompt was putting ideas in the model's head. The system prompt listed signs of a scam: requests for a password, PIX to a private person, strange links. The 4B model started warning about PIX in documents that never mentioned it. A list of examples is a list of suggestions to a small model. The rule now says "only mention a scam if the document itself shows one of these signs", and each field is told what to copy and from where. The invented warnings stopped.
That last one is the main thing I learned about prompting a 4B model: be literal, say where each value comes from, and never offer an example you would not want to see in the output.
Honest limits. A 4B model is not a lawyer, and the app says so on every result.
- A photo of a full page in small print is not reliable. In testing, a photo of a whole A4 page produced an amount that was not in the document. A close-up of one paragraph was read correctly. So the advice is: photograph the part you want to understand, from close.
- The glossary can be wrong. It explained "pro rata die" incorrectly every time. Dates and amounts come from the document; definitions come from the model's memory, and that is the weaker part.
- Text is limited to 16,000 characters, which fits the 8,192-token context with room for the answer.
Why Does Open Innovation Matter?
Because of who this is for.
With a hosted model, helping my grandmother would mean sending a document with her name, address, CPF and debt to a company, under terms neither of us would read. I would be solving one unreadable contract by agreeing to another. It would also be the exact thing she is afraid of. With an open-weight model on a computer at home, the document travels from her phone to that computer over the home Wi-Fi and nowhere else. The app has no database, keeps nothing once the request is answered, and does not log document content.
Open weights also make this something I can give away. There is no API key to pay for, no account to create, no usage limit, and it keeps working without internet once the model is downloaded. My grandmother does not depend on me keeping a subscription alive, or on a company deciding to keep a product running.
And it can be checked. The prompts are in the repository, in Portuguese, next to the code. If the explanation is wrong, anyone can see exactly what the model was asked. For a tool that tells people what a legal document means, I think that matters more than it does for most software.
None of the pieces here are mine: Gemma, Ollama, Spring AI and Spring Boot are all open. My part was choosing the problem and spending a weekend getting them to work together for one person. That is the point of open innovation for me. The distance between "someone I love has a problem" and "here is a tool for it" has become short enough to cover in a weekend.
My Agent Session
I did not build this alone, and I would rather say so plainly.
- First version. I chose the stack (Java, Spring Boot, a local Gemma) and the problem to solve. Claude wrote the first version of the code in the Claude desktop app while I followed along and it explained each decision. That session could not reach Maven Central, so the code was written against the Spring AI 2.0 documentation and had never been compiled.
- Review. Before submitting, I asked Claude Code to review the project on my machine. It ran the first real build, downloaded the model, called the app with text and photo documents, and found the four problems described above, plus the photo limit.
- This article was drafted with Claude too. Everything in it about my grandmother is mine.
Prize Categories
Best Use of Gemma. Gemma is the whole product, not a feature of it. A single gemma3:4b, running locally through Ollama, does three jobs: it reads the photo of the document, it turns legal Portuguese into structured plain-language output, and it answers follow-up questions grounded in that document. Its size is what makes the privacy argument real: it is small enough to run on an ordinary home computer, so the document never has to leave the house.
Top comments (2)
The three reactions you listed — ignore it, pay without understanding, or ask someone to read it — is the most accurate description of this problem I've seen. We have the exact same dynamic in my family (Portuguese grandparents in Luxembourg, official letters in French). Love that you went with paste-or-photo instead of asking her to learn a chatbot — nobody's grandmother should have to. How are you handling the photo side: Gemma's vision directly, or a separate OCR step? Curious how it copes with crumpled letters and bank letterheads.
Thank you, that means a lot. And it is good (and a bit sad) to hear it is the same in Luxembourg with letters in French. I suspect every family has a version of this.
On the photo side: Gemma's vision directly, no OCR step. The browser shrinks the photo to 1600px on the long side and sends it straight to gemma3:4b through Ollama, together with the same prompt used for pasted text. I liked that it keeps the setup to one model and nothing else to install.
How it copes, honestly: it depends a lot on how much is in the frame. A close-up of one paragraph was read correctly in my tests. A photo of a whole A4 page in small print was not reliable, and once gave me an amount that was not in the document. So the advice I give my grandmother is to photograph only the part she wants to understand, from close.
I have not tested crumpled paper or heavy letterheads properly yet, so I can't promise anything there. If I had to guess, a dedicated OCR pass before the model would help with full pages, but it would cost the "one model, nothing else" simplicity. It is on my list to try.
If you ever want to adapt it, the prompts are plain text in the repo. I have not tried it, but I would be curious to hear how it goes.