This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
My father prints for the neighbours.
That's the whole setup. Two houses, one inkjet model between them, and every so often someone knocks with a USB stick or an email. A school project. A stack of holiday photos. Someone's wedding invitations that absolutely must not come out banded.
For the jobs he knows, the workflow is fine. He made the file, he knows what's on it. But for a neighbour's file, the first thing he used to do was open everything. Every image, every page, one at a time, guessing at paper and orientation by eye. Only then would he commit a sheet.
That preview pass is the thing I took away from him.
KantoPrint drops that pass. You hand it the customer's file and it tells you, per page, which paper to load, how much ink is on it, whether anything is touching the cut line, and whether to turn it 90°. Before the printer wakes up.
Every page of a PDF, not just the first. Results stream in page by page so a twenty-page file doesn't go quiet for five minutes. Cards fold away when there's more than one page. Filter a batch by tray, copy a load sheet, stop it if it's taking too long.
What he said when I handed it over. Nothing quotable, and I'm not going to invent a line. What I got was an observation: he no longer opens all those files. That was the relief, and it was specifically about the opening — not about any one verdict being right.
That reordered my priorities. I had been treating verdict accuracy as the product. For him, a recommendation he can act on beat a correct one he has to double-check. The tool is only useful if the answer is good enough to trust without opening the file yourself.
Demo
Code
Repository: https://github.com/znarfm/KantoPrint
MIT licensed. Next.js 16, sharp, pdf-lib, Ollama, Zod, vitest, oxlint and oxfmt. Sixty unit tests over the logic that doesn't need a model, plus an opt-in suite that drives the real one against deterministic fixtures. Wired into CI.
git clone https://github.com/znarfm/KantoPrint
pnpm install
ollama pull gemma4:e2b
pnpm fixtures && pnpm dev
How I Built It
gemma4:e2b, running locally through Ollama. It's the only model in the system.
The interesting part wasn't getting it to classify a page. It was finding out how much of it to let near.
gemma4:e2b is a 2B model, and on my own test fixtures it was confidently wrong about pixels. It called a wide-margin service contract bleeding. It called a full-bleed colour invitation safe. It rated a 0.9%-ink form "medium" risk sitting right next to a gauge reading 0.9%.
A coin flip would have been about as reliable. So I stopped asking it things pixels can answer. Ink at the trim edge comes from measuring the edge band. The risk band comes from measuring ink load. Rotation comes from the page dimensions. The model keeps the two questions that are genuinely judgement: what is this job, and what paper goes in the tray.
That split cost some accuracy on paper and bought a lot of it in practice. More interestingly, three fields I originally had the model fill in got deleted outright, because they disagreed with the numbers printed beside them and they were eating about forty of the fifty-five tokens it generated per page.
One more thing I'd keep if I rebuilt it: the operator notes. I asked the model to write them, and it wrote "Set up the material for printing," which is not an instruction to anybody. They're composed from the flags now. Duller, and actually useful.
Why Does Open Innovation Matter?
It runs on a laptop with no internet. Ollama serves the model on localhost. There's no outbound call at all — our router could be unplugged and this would behave the same.
Our neighbours' files never leave the house. This is the one that settled it. When you print other people's work — school projects, ID photos, someone's thesis — a hosted API means every one of those files, plus a render of it, sits on hardware you don't control. For a two-house print shop that's not a compliance conversation. It's the deal we made with the people who trusted us.
Nothing costs money per job. A print job costs pennies. Per-token billing makes no sense at that scale; a one-time 4.6 GB download does.
Being honest about what that buys, on a Core i5 laptop with no GPU: 11 to 17 seconds per page, the first time the model sees it. Same page again is 1.6 to 3.2s, because Ollama caches the vision encoding — which is also why my first benchmark of this project was wrong by about 5x, and why I trust the honest number more than the flattering one.
And it swaps. KANTOPRINT_MODEL is one environment variable. Anything Ollama can serve, this can be pointed at.
What open doesn't fix is worth saying too. At that resolution an 11pt line is about 4.6 pixels tall, so the model can't read body copy, and the card says so. It also doesn't measure banded output or colour accuracy — it measures coverage, not how your printer reproduces it.
My Agent Session
Prize Categories
-
Best Use of Gemma —
gemma4:e2bis the only model here, served locally via Ollama. The build leans on Gemma-specific behaviour: vision cost that scales with input resolution, constrained decoding through theformatgrammar so the reply cannot be prose, and disabled thinking mode for latency.



Top comments (0)