DEV Community

Mininglamp
Mininglamp

Posted on

Run a Local GUI Agent on Your Mac in 10 Minutes with Mano-P and Cider

What This Guide Covers

Mano-P is an open-source GUI agent that sees your screen, reasons about what to do next, and executes real macOS actions: clicking buttons, filling forms, navigating apps. It works through an observation-reasoning-action loop. The agent takes a screenshot, feeds it through a vision-language model, decides what to click or type, executes that action, then repeats until the task is complete. The 4B model runs entirely on-device with MLX 8-bit quantization through Cider SDK. No cloud calls, no API keys, no screenshots leaving your machine.

This tutorial walks through installing Mano-P on an Apple Silicon Mac, activating Cider quantization, and running your first local GUI task. By the end you will have a working setup that can automate real macOS workflows from the terminal. Everything here was tested on an M4 Mac mini with 32GB RAM.

Prerequisites

  • Apple M4 chip or newer
  • 32GB RAM minimum
  • macOS with Screen Recording and Accessibility permissions enabled
  • Homebrew installed

The quantized 4B model peaks at roughly 4.3GB memory. A 32GB machine handles it comfortably with room for your normal workload.

Step 1: Install Mano-CUA via Homebrew

Mano-P ships a CLI tool called mano-cua. Install it through the Mininglamp tap:

brew tap Mininglamp-AI/tap && brew install mano-cua
Enter fullscreen mode Exit fullscreen mode

Verify the installation:

mano-cua --version
Enter fullscreen mode Exit fullscreen mode

Step 2: Run the Environment Check

Before pulling any models, let the checker confirm your system meets requirements:

mano-cua check
Enter fullscreen mode Exit fullscreen mode

This validates your chip generation, available RAM, macOS version, and whether Screen Recording plus Accessibility permissions are granted. If permissions are missing, macOS will prompt you to enable them in System Settings > Privacy & Security.

Grant both permissions before continuing. The agent needs Screen Recording to capture what is on screen and Accessibility to simulate mouse clicks and keyboard input. Without Screen Recording, the agent is blind. Without Accessibility, it can see but cannot interact with anything. Both must be enabled for the specific terminal application you use to launch mano-cua, whether that is Terminal.app, iTerm2, or another shell.

Step 3: Install Cider SDK

Cider is the quantization runtime that makes the 4B model fast on Apple Silicon. It provides online activation quantization for MLX with W8A8 and W4A8 support. The key idea: most quantization tools only quantize model weights, leaving activations in FP16 or BF16. Cider quantizes both, which means the matrix multiplications that dominate inference run in lower precision end to end.

mano-cua install-sdk
Enter fullscreen mode Exit fullscreen mode

This pulls the Cider SDK and configures the MLX backend. The SDK handles weight quantization and activation quantization together, which is why it outperforms naive weight-only quantization approaches.

Some numbers on what Cider delivers on M5 Pro hardware:

  • W8A8 prefill runs roughly 12.7% faster than the W8A16 baseline
  • Compared to standard MLX W4A16, Cider achieves 1.4x to 2.2x prefill speedup
  • These gains come from quantizing both weights and activations rather than weights alone

For a GUI agent, prefill speed matters a lot. Every step involves processing a full-resolution screenshot as input tokens. Faster prefill means faster per-step latency, and the agent typically runs 7 to 8 steps per task. The cumulative savings add up.

You can find the Cider source and benchmarks at github.com/Mininglamp-AI/Cider. The project has 323 stars on GitHub at the time of writing.

Step 4: Download the Model

Pull the quantized Mano-CUA-4B-Thinking model:

mano-cua install-model
Enter fullscreen mode Exit fullscreen mode

This downloads Mano-CUA-4B-Thinking-1.1, an MLX 8-bit quantized model from HuggingFace. The model weights are hosted at huggingface.co/Mininglamp-2718/Mano-CUA-4B-Thinking-1.1.

Expect a download around a few GB. Once cached locally, subsequent launches skip the download entirely.

Step 5: Run Your First Task

With everything installed, point the agent at a real task:

mano-cua run "Open Safari and search for the weather in San Francisco" --local
Enter fullscreen mode Exit fullscreen mode

The --local flag ensures all inference happens on your Mac. No data leaves the device. The agent captures a screenshot, reasons about the current state, decides what action to take, executes it, then repeats until the task is complete or it determines the task is done.

Watch the terminal output. You will see the agent's reasoning steps, the actions it chooses, and its assessment of whether it succeeded. The reasoning trace shows the model thinking through what it sees on screen, which element it should target, and why. This transparency is useful for debugging when the agent makes a wrong move.

On M5 Pro, the 4B model decodes at roughly 80 tokens per second. Each step takes about 7.9 seconds on average, and most tasks complete within 7 to 8 steps. That means a typical task runs from start to finish in under a minute.

Step 6: Try More Complex Tasks

The agent handles multi-step workflows across different applications:

mano-cua run "Open System Settings, go to Wi-Fi, and tell me the network name" --local
Enter fullscreen mode Exit fullscreen mode
mano-cua run "Open Notes, create a new note titled Meeting Agenda, and type three bullet points about project planning" --local
Enter fullscreen mode Exit fullscreen mode
mano-cua run "Open Finder, navigate to the Downloads folder, and sort files by date modified" --local
Enter fullscreen mode Exit fullscreen mode

Each task involves the agent reasoning through multiple screens, identifying UI elements, and executing the right sequence of clicks and keystrokes. The agent navigates between apps using the Dock, menu bars, and keyboard shortcuts just like a human user would. It handles dialog boxes, dropdown menus, and scroll views.

Tasks that require reading text from the screen and reporting back also work. The model can extract information from what it sees and include it in its final response.

Performance and Benchmarks

On a 100-task real macOS GUI test suite, the numbers look like this:

Metric Value
Model Mano-CUA-Thinking-4B
Pass rate 56.0%
Average steps per task 7.5
Average time per step 7.9s
Peak memory 4.3GB

For context, the 4B local model at 56% significantly outperforms cloud-based Qwen3-VL-Plus, which scored 39% on the same test suite. A local 4B model beating a cloud VL model on real GUI tasks is a meaningful result.

On the OSWorld benchmark, Mano-CUA-Thinking-4B achieves 58.2% success rate, the top score among specialized GUI agent models.

GSPruning: Visual Token Efficiency

Mano-P also includes GSPruning, a visual token pruning technique that reduces the number of visual tokens the model processes from each screenshot. Screenshots contain a lot of redundant information. Large uniform regions like backgrounds, menu bars, and empty space do not need as many tokens as the areas where clickable UI elements live. GSPruning identifies and removes these low-information tokens, yielding 2x to 3x throughput improvement without significant accuracy loss.

For GUI agent workloads where every single step involves a full-resolution screenshot, token pruning directly translates to faster step execution. Combined with Cider quantization, the entire inference pipeline is optimized at both the model level and the input level.

Privacy Model

Running with --local means:

  • All inference runs on your Apple Silicon GPU
  • Screenshots are processed in memory and never written to disk or transmitted
  • No telemetry, no cloud fallback
  • Your screen content stays on your machine

This matters for anyone working with sensitive documents, internal tools, or environments where screen data simply cannot leave the device.

Understanding the Action Loop

It helps to understand the cycle the agent runs through on each step:

  1. Screenshot capture: The agent grabs the current screen state via the Screen Recording API
  2. Visual encoding: The screenshot is tokenized and fed through the vision encoder with GSPruning applied
  3. Reasoning: The language model generates a chain-of-thought about what it sees and what action to take next
  4. Action execution: The chosen action (click at coordinates, type text, press a key combination) is executed via the Accessibility API
  5. Repeat: Back to step 1 with the updated screen state

The model sees the raw pixels of your screen. It does not rely on accessibility trees, DOM structures, or any application-specific hooks. This means it works with any macOS application out of the box, including ones that do not expose accessibility metadata.

Troubleshooting

Permission denied errors: Re-check System Settings > Privacy & Security. Both Screen Recording and Accessibility must be enabled for the terminal app or iTerm you are using to launch mano-cua.

Slow first inference: The first run after install takes longer as MLX compiles the compute graph. Subsequent runs are faster.

Out of memory: Close memory-heavy applications. The model needs roughly 4.3GB at peak. On a 32GB machine this is fine, but if you have dozens of browser tabs open, free up some headroom.

Model download stalls: If mano-cua install-model hangs, check your network connection. The model downloads from HuggingFace, so a proxy or VPN might be needed depending on your region.

Wrapping Up

The full pipeline from Homebrew install to running your first GUI task takes under ten minutes. A 4B model that scores 56% on real macOS GUI tasks, runs at 80 tok/s, and fits in 4.3GB of memory is a practical tool, not a demo.

Star the repo if this is useful: github.com/Mininglamp-AI/Mano-P


Top comments (0)