DEV Community

Mininglamp
Mininglamp

Posted on

No API Keys, No Cloud Bills: Running a Document Pipeline Entirely On-Device

No API Keys, No Cloud Bills: Running a Document Pipeline Entirely On-Device

We run a document processing pipeline on a Mac mini. No API keys. No cloud inference. No network calls at all during execution. This post is about what that actually looks like in practice, what went wrong along the way, and where we landed on performance.

Why we went local

Our team builds tooling for due diligence workflows. Contracts, financial statements, scanned images, password-protected archives. The kind of stuff clients will not let you upload anywhere, period. We tried wrapping cloud vision APIs behind a VPN with audit logging, and the compliance team still said no. Fair enough.

So we started looking at on-device GUI agents. The idea: a model that can see the screen, click through file managers, open PDFs, read images, and do it all without any data leaving the machine.

We settled on Mano-P, an Apache 2.0 open-source agent built for edge devices. It uses a 4B parameter model with W8A8 quantization, runs through a vision-only pipeline, and needs no API endpoints or MCP servers. The whole thing installs via brew.

brew tap Mininglamp-AI/tap && brew install mano-cua
Enter fullscreen mode Exit fullscreen mode

The actual pipeline

Here is what a typical session looks like for us. Someone drops a task: process a batch of due diligence files. The agent handles it step by step, all on the local machine.

First, it searches the local filesystem for the target archive. We tested this on nested folder structures with 2000+ files and it consistently found the right one. Then it needed a password to unzip. The password was stored in a separate local note. The agent found that too, entirely on its own, just by navigating the local file system visually.

After unzipping, it reads the contents, including scanned images. The vision model handles layout-heavy documents reasonably well. Not perfect on every table cell, but good enough for extraction tasks where you need 90% accuracy and a human does a final check.

The last step is batch processing. We threw 200 contracts at it. Each one needed the same extraction routine. On our M4 Pro Mac mini with 32GB RAM, the 4B model runs at roughly 80 tokens per second decode speed. That is fast enough that the bottleneck is actually the GUI interaction latency, not the model inference.

Numbers that matter

We ran a head-to-head comparison. 100 real GUI tasks on macOS, local 4B model versus cloud-based Qwen3-VL-Plus.

Local 4B thinking model: 56% pass rate, average 7.9 seconds per step.
Cloud Qwen3-VL-Plus: 39% pass rate, average 10.2 seconds per step.

That was the result that made us commit to local. A model one hundredth the size of typical cloud VL services, running on consumer hardware, beating the cloud option by 17 percentage points on real GUI tasks. The latency was better too because there is zero network round-trip.

On OSWorld specialized benchmarks, the underlying model scores 58.2%, holding the top spot among specialized models. The second place entry scores 45.0%. On web navigation tasks via the WebRetriever protocol, it hits 41.7 NavEval, which edges past Gemini 2.5 Pro at 40.9 and Claude 4.5 at 31.3.

What went wrong

Plenty. Some honest notes.

The first week was rough. Screen recording permissions on macOS are finicky and the model occasionally misclicked on small UI elements. Dense dropdown menus were a consistent failure mode early on. We ended up increasing the resolution of screenshots fed to the model, which helped but also slowed down prefill slightly.

Memory was tight. The 4B quantized model peaks around 4.3GB, which sounds small, but when you also have the target application open plus a file manager plus Preview showing scanned images, you are pushing the 32GB envelope. We had to be deliberate about closing unused windows.

Batch processing at 200 files revealed another issue: the agent occasionally lost track of which file it had already processed. We worked around this by having it maintain a simple checklist in a text file as it went. Not elegant, but effective.

The privacy argument in practice

People talk about data sovereignty in the abstract a lot. For us it is extremely concrete. A client hands over 200 contracts containing names, financial figures, deal terms. In the cloud model, every screenshot of every page goes to a remote server. Even if the provider promises not to retain data, the compliance paperwork alone takes weeks.

With local execution, the conversation is different. We show the client the network monitor. Zero outbound connections during the entire processing run. The audit log is a local file they can inspect. Compliance review took two days instead of two months.

Hardware setup

Mac mini, M4 chip, 32GB RAM. Total hardware cost under $1500. The alternative was a cloud vision API that would have billed us per image per page per document. For 200 contracts averaging 40 pages each, that adds up fast. We estimated roughly $800 per batch at current cloud pricing. The Mac mini paid for itself in two batches.

There is also a compute stick option that connects via USB 4.0 if you want to add local inference capability to an existing machine. We have not tested that path yet.

What I would tell someone considering this

Start with the 4B quantized model. Do not try to run larger models locally unless you have 64GB or more. The 4B model is surprisingly capable for GUI tasks specifically because the training pipeline was built around screen understanding, not general chat.

Test your specific UI. Performance varies a lot depending on the application. Simple file managers and text editors work great. Complex web apps with lots of dynamic elements are harder. The model handles them, but expect more retries.

Keep your expectations calibrated. This is not a cloud-scale model. It will not write your code or summarize a 100-page report in one shot. But for structured, repetitive GUI tasks with sensitive data, it is genuinely better than the cloud alternatives we tried.

Project link: https://github.com/Mininglamp-AI/Mano-P

Top comments (0)