This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
My friend writes down everything in class. Every lecture, every assignment, page after page. The problem shows up the night before an exam: he photographs his notes, uploads them to an AI chatbot, asks for a summary, and gets nothing useful back. It can't read his handwriting.
When I ran his pages through a vision model myself, I found something worse than "can't read it". The model read most of it, and the mistakes looked right:
- The Perennial Student (the book he was writing about) became The **Peripheral* Student*
- "satirical" became "spatialical"
- "his reputation" became "her reputation"
- a date in a letter disappeared, and "5:00 to 7:00 pm" became "8:00 to 8:00"
And then there's his shorthand. He writes "acc" for according and "gitq" for given in the question, which no model is ever going to guess.
A tired student skimming a summary at 2 AM would never catch any of that. So I built HandNotes: a handwriting reader that runs entirely on his laptop, lets you correct what it got wrong, and remembers those corrections for the next page.
What it does:
- Reads a photo of his notes live. Text streams in word by word from an open vision model running locally.
- Highlights words it isn't sure about, with one-click jumps to each one so checking a page is fast.
- Learns from corrections. Every fix is compared word by word with what the model wrote. His vocabulary and past misreads go into the prompt for the next page, and a misread you've corrected twice gets fixed automatically.
- Knows his short forms. Teach it "gitq = given in the question" once, either by adding it or just by expanding it while correcting a page, and it expands it on every page after that.
- Shows its work. An accuracy chart per page, a list of his usual misreads, and a "compare with plain model" view that shows, word by word, what the memory changed.
- Turns notes into revision material: a summary with likely exam questions, flip-to-reveal flashcards, and export to PDF, Word or text.
What he said: when I showed it to him, he really liked it, and he told me he's going to use it on a regular basis. Coming from the person who photographs his notes because typing them up is too much effort, that's the review I was hoping for.
Demo
It runs locally (that's the point), so the GIF at the top is a real run recorded on my laptop, sped up 4x: adding his short forms, reading one of his pages live, comparing against the plain model, and making flashcards.
Plain model vs. with his memory, on the same page. Struck through is what the plain model read; green is what it read with his memory:
The revision summary, generated locally from the transcript:
Code
jemankalita
/
Hacktoberfest-Build-for-a-Friend
A handwriting reader that learns one friend's handwriting from corrections. Gemma 3 via Ollama, runs fully offline on a laptop.
HandNotes
A handwriting reader that learns one person's handwriting, running entirely on your own laptop.
My friend writes down everything in class, then photographs his notes the night before an exam and asks an AI chatbot to summarize them. It can't read his handwriting. And when a model does read most of it, the mistakes look right: in our first test a book title turned into a different word, "his" became "her", and a date quietly disappeared.
HandNotes reads his notes with an open vision model (Gemma 3 via Ollama) on his own machine, lets you fix what it got wrong, and remembers those fixes so the next page comes out better. Nothing is uploaded anywhere.
Built for the DEV Hacktoberfest Weekend Challenge: Build for a Friend.
Real run at 4x speed: adding his short forms, reading a page live, comparing with the plain model, and making flashcards.
Features
…Setup is three commands: ollama pull gemma3:4b, pip install -r requirements.txt, python server.py.
How I Built It
The model: Gemma 3 4B, an open-weight vision model, served locally by Ollama. On my laptop (GTX 1650 Ti, 4 GB) Ollama splits it about half GPU, half CPU, and a page takes 30 to 40 seconds.
The app: a small FastAPI server bound to 127.0.0.1, with a hand-written HTML/CSS/JS front end. Fonts and icons are bundled, so even the UI makes no network requests.
The prompt matters more than I expected. My first prompt let the model "tidy up" his notes: it dropped a whole line with a date in it. Telling it to copy literally, skip nothing, and mark unclear words with [?] brought the date back.
The learning is a memory, not retraining. Fine-tuning a vision model in a weekend on a 4 GB GPU wasn't realistic, so HandNotes keeps a small per-writer memory file:
- The model transcribes the page.
- You fix the transcript and save.
- A word-level diff (
difflib) between the model's output and your fix records each misread ("Peripheral" → "Perennial"), the words he uses, and the page's accuracy. - On the next page, those words and misreads are added to the prompt, and misreads confirmed at least twice are replaced automatically. Short words like his/her are never auto-replaced, because they depend on context.
- Short forms are detected separately. If one short word gets replaced by a longer phrase that starts with the same letter and contains its letters in order (
gitq→ given in the question), it's stored as his shorthand. The model is told to copy those exactly as written, and HandNotes expands them after every read.
How well does it work? I ran four of his pages in order, letting the memory build up, and scored each against a transcript checked word by word against the photo. For pages 2 to 4 I also read the page with the plain model for comparison:
| Page | With memory | Plain model |
|---|---|---|
| 1. Literature assessment | 94.1% | (no memory yet) |
| 2. Research overview | 95.5% | 95.5% |
| 3. Formal letter | 96.0% | 95.2% |
| 4. Literature assessment | 90.1% | 91.9% |
That table isn't the neat upward line I hoped for, and it taught me more than one would have:
- The memory does fix names and terms. Re-reading the first page with memory, "Peripheral" became "Perennial" and "spatialical" became "satirical".
- It can also learn the wrong lesson. My friend spells "behavior" on one page and "behaviour" on another. The memory learned one spelling and "fixed" the other into it, which is why page 4 came out worse with memory. On another re-read, it turned a correct "Coles" into "Cole".
- Pronouns are the stubborn error. Even when told to copy literally, the model kept turning "his" into "her".
- Errors travel. On one run the model invented a word at the end of a page ("distressed"), and it flowed straight into the generated summary. That's why the review step comes before the study tools.
I also found a bug in my own scoring along the way: curly and straight apostrophes (student’s vs student's) were counted as misreads. I fixed it with a test and re-ran everything, so the numbers above are from the corrected version.
The next step is clear from that list: let the user mark a change as "his slip, don't learn this" versus "the model misread this", and flag words he genuinely uses instead of auto-replacing them. After that, I want to cut each corrected word out of the page and use that library to fine-tune a small open handwriting model on his writing specifically.
Why Does Open Innovation Matter?
For this project, open isn't a nice-to-have. It's the reason the project can exist at all.
- It can be taught one person. A hosted chatbot can't be taught my friend's handwriting or his shorthand. An open model I run myself, with a memory I control, can. And because every correction is stored locally, his own handwriting dataset builds up as he uses it. With an open model I can actually fine-tune on that later.
-
His notes stay his. These are photos of his assignments and personal letters. Nothing is uploaded anywhere: the server only listens on
127.0.0.1. - It works where he studies. No internet needed once the model is downloaded, which matters in a hostel room the night before an exam.
- It costs nothing to run. No subscription and no API bill. For a student who's already "too lazy" to type up his notes, anything that costs money or effort won't get used.
- I could see and change everything. When the model dropped a line, I could rewrite the prompt. When the memory learned the wrong lesson, I could see exactly why in a JSON file. I could swap models with one environment variable.
My Agent Session
I built this pairing with Claude Code over one long night. It helped with the code, the tests (37 of them, covering the memory, short forms, the API with the model mocked, and the exporters), the demo recording, and the measurements above.
Prize Categories
- Best Use of Gemma: Gemma 3 4B does all the reading, summarizing and flashcard generation, running locally through Ollama.




Top comments (0)