DEV Community

Meinard Francisco
Meinard Francisco

Posted on

My father prints for the neighbours. I built him something better than guessing.

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

My father prints for the neighbours.

That's the whole setup. Two houses, one inkjet model between them, and every so often someone knocks with a USB stick or an email. A school project. A stack of holiday photos. Someone's wedding invitations that absolutely must not come out banded.

For the jobs he knows, the workflow is fine. He made the file, he knows what's on it. But for a neighbour's file, the first thing he used to do was open everything. Every image, every page, one at a time, guessing at paper and orientation by eye. Only then would he commit a sheet.

That preview pass is the thing I took away from him.

KantoPrint drops that pass. You hand it the customer's file and it tells you, per page, which paper to load, how much ink is on it, whether anything is touching the cut line, and whether to turn it 90°. Before the printer wakes up.

Every page of a PDF, not just the first. Results stream in page by page so a twenty-page file doesn't go quiet for five minutes. Cards fold away when there's more than one page. Filter a batch by tray, copy a load sheet, stop it if it's taking too long.

What he said when I handed it over. Nothing quotable, and I'm not going to invent a line. What I got was an observation: he no longer opens all those files. That was the relief, and it was specifically about the opening — not about any one verdict being right.

That reordered my priorities. I had been treating verdict accuracy as the product. For him, a recommendation he can act on beat a correct one he has to double-check. The tool is only useful if the answer is good enough to trust without opening the file yourself.

Demo

Initial state of the homepage

Inputting 3 files at once

Results

Code

Repository: https://github.com/znarfm/KantoPrint

MIT licensed. Next.js 16, sharp, pdf-lib, Ollama, Zod, vitest, oxlint and oxfmt. Sixty unit tests over the logic that doesn't need a model, plus an opt-in suite that drives the real one against deterministic fixtures. Wired into CI.

git clone https://github.com/znarfm/KantoPrint
pnpm install
ollama pull gemma4:e2b
pnpm fixtures && pnpm dev
Enter fullscreen mode Exit fullscreen mode

How I Built It

gemma4:e2b, running locally through Ollama. It's the only model in the system.

The interesting part wasn't getting it to classify a page. It was finding out how much of it to let near.

gemma4:e2b is a 2B model, and on my own test fixtures it was confidently wrong about pixels. It called a wide-margin service contract bleeding. It called a full-bleed colour invitation safe. It rated a 0.9%-ink form "medium" risk sitting right next to a gauge reading 0.9%.

A coin flip would have been about as reliable. So I stopped asking it things pixels can answer. Ink at the trim edge comes from measuring the edge band. The risk band comes from measuring ink load. Rotation comes from the page dimensions. The model keeps the two questions that are genuinely judgement: what is this job, and what paper goes in the tray.

That split cost some accuracy on paper and bought a lot of it in practice. More interestingly, three fields I originally had the model fill in got deleted outright, because they disagreed with the numbers printed beside them and they were eating about forty of the fifty-five tokens it generated per page.

One more thing I'd keep if I rebuilt it: the operator notes. I asked the model to write them, and it wrote "Set up the material for printing," which is not an instruction to anybody. They're composed from the flags now. Duller, and actually useful.

Why Does Open Innovation Matter?

It runs on a laptop with no internet. Ollama serves the model on localhost. There's no outbound call at all — our router could be unplugged and this would behave the same.

Our neighbours' files never leave the house. This is the one that settled it. When you print other people's work — school projects, ID photos, someone's thesis — a hosted API means every one of those files, plus a render of it, sits on hardware you don't control. For a two-house print shop that's not a compliance conversation. It's the deal we made with the people who trusted us.

Nothing costs money per job. A print job costs pennies. Per-token billing makes no sense at that scale; a one-time 4.6 GB download does.

Being honest about what that buys, on a Core i5 laptop with no GPU: 11 to 17 seconds per page, the first time the model sees it. Same page again is 1.6 to 3.2s, because Ollama caches the vision encoding — which is also why my first benchmark of this project was wrong by about 5x, and why I trust the honest number more than the flattering one.

And it swaps. KANTOPRINT_MODEL is one environment variable. Anything Ollama can serve, this can be pointed at.

What open doesn't fix is worth saying too. At that resolution an 11pt line is about 4.6 pixels tall, so the model can't read body copy, and the card says so. It also doesn't measure banded output or colour accuracy — it measures coverage, not how your printer reproduces it.

My Agent Session

Building KantoPrint: teaching a 2B vision model what it is not allowed to guess
You

Build KantoPrint: air-gapped prepress triage for desktop inkjet and cutting workflows, using local Gemma 4 multimodal inference. Every page of a PDF should be analysed, and the card collapsible for multi-page files.

Agent

Starting with the schema foundation, since everything else depends on the contract. Moving from a single page field to pages: PageResult[] so a PDF produces one analysed result per page.

Also dropping inkRiskLevel from the model contract. It was producing a band that disagreed with the gauge printed beside it — a 0.9% ink form rated 'medium'. If pixels measure it, pixels own it. That takes the model schema from 5 fields to 3.

You

Why does the model keep getting bleed wrong? Is the vision call even working?

Agent

Good question to check first. Probing the model directly rather than trusting the pipeline.

Vision works. So the misclassification is a capability limit, not a plumbing bug. The model sees a blank-margin contract and calls it bleeding.

You

The ink load numbers say 11-17 seconds per page now, but earlier you said 3.7-4.4s. Which is it?

Agent

This needs measuring, not arguing. Separating 'first time an image is seen' from 'same image again'.

That's a 5x gap and it isn't noise: Ollama caches the vision encoder output per image. Every earlier benchmark in this project re-ran the same fixture files, so it measured the cache, not the encoder.

The honest number is 11-17s for a page the model has never seen. The 3.7-4.4s figure I'd been quoting was wrong and had to come out of the README. A 20-page PDF is 4-6 minutes, which is what makes streaming, the page cap, and Stop necessary rather than decorative.

You

OK. Then why does the model report wrong bleed?

Agent

Measuring instead of asking. Edge-band ink coverage at ~3% of the short side, which is about 3mm at 150 dpi.

The model was wrong in both directions on the same run. Edge ink above 5% is bleeding; at or below 0.5% the band is provably blank. Between those the model still decides.

Three fields now get overridden by measurement rather than trusted: bleed, ink risk, rotation. The model keeps only what is genuinely judgement — what the job is, and which paper.

You

That's a reasonable split, but operatorNotes was supposed to come from the model.

Agent

It did, for a while. Then I read what it produced:

That's not an instruction to anybody. Composing them from the flags instead: duller, and actually useful. Bonus: those notes were about 25 of the ~55 generated tokens per page, so deleting them was also a latency win.

There's a deeper version of this. The chips on the card and the sentences in the brief were the same predicates written twice, so a flag could appear in one and go missing from the other. One list now feeds both.

You

Run the tests. Did they actually catch anything?

Agent

Two real bugs in the first hour, both of which I'd have shipped:

First: ingest fell back to the browser's declared MIME when magic bytes didn't match. A text file named .pdf was accepted and died later inside pdf-lib with a worse message. PDF, PNG and JPEG all have signatures, so the fallback was only ever weakening validation. Removed it.

Second: the empty-file check ran after the type sniff, so an empty file got the wrong reason. Reordering.

Worth noting the four other failures were my test arithmetic being wrong, not the code. The edge-band one I'd asserted was less than the page average; it's exactly equal, because a black half spanning the full height fills the left edge band completely and the white half fills the right one.

You

You said sharp can't read PDFs. What did you do about it?

Agent

sharp has no poppler, so PDF input is out. Two native rasterisers were already on the box, so: pdftoppm first, ghostscript second, with a clear error naming both if neither exists.

Multi-page needed care — poppler zero-pads its filenames to the digit count of the last page rendered, so pg-1.png and pg-01.png are both possible depending on the page cap. Collecting the directory and sorting numerically is the only version that survives a config change.

Both get a hard timeout. A malformed PDF could otherwise hang the request forever and strand its temp directory.

You

Turn on the shadcn lint rules.

Agent

Two of these were worth having. no-raw-colors caught that the ink gauge had been borrowing raw Tailwind amber and emerald because the theme had no such tokens — ink risk has three bands, so the theme now names all three:

require-static-classes caught something subtler: the gauge colour was built from an indexed lookup, which the linter can't follow — meaning no rule could check those classes at all. Now a conditional between complete class strings, so everything about it is readable.

Used contracts rather than blanket allows where the customisation was intentional, so ProgressLabel can take typography without Progress itself being unrestyleable.

Prize Categories

  • Best Use of Gemma — gemma4:e2b is the only model here, served locally via Ollama. The build leans on Gemma-specific behaviour: vision cost that scales with input resolution, constrained decoding through the format grammar so the reply cannot be prose, and disabled thinking mode for latency.

Top comments (0)