A local model is useful when you can finish the work while the answer is still relevant. Waiting for an oversized model to write a long response is not always the best use of your computer.
OGAD (Off Grid AI Desktop) lets you run and compare local models in one app. The most useful first change is often a smaller model that can still do the job, followed by a shorter context and a more focused request. You keep local processing without assuming every model will be fast on every machine.
Download OGAD for Mac or Windows
What would you like to do with Off Grid AI?
Have a feature or use case you would like us to support? Tell us what you want to do and which device you use.
Write to support@offgridmobileai.co, join our Slack community, or talk to us on Reddit.
Start with the answer you actually need
Compare models on a task from your normal work: rewrite one email, explain one function or extract three decisions from a short note. Keep the input and requested output the same for each trial.
Ask for a bounded result:
Rewrite this update in two short paragraphs. Keep the dates and unresolved questions. Return only the revised text.
That is easier to compare than an open-ended request for an exhaustive analysis. A shorter answer also requires fewer generated tokens, although the initial processing time can still depend on the model and input.
Choose a model that fits comfortably
- Open Models and inspect the memory guidance for your current choice.
- If it is tight on RAM, download a smaller compatible text model.
- Run your sample task in a fresh chat.
- Compare the time to a useful result and the quality of the answer.
- Keep the smallest option that meets the task's needs.
Do not choose on speed alone. A fast answer that needs extensive repair can cost more of your time than a slower accurate draft.
Reduce work the model does not need
Use a practical Context window in the model settings. More context is useful when the task needs long material; it can cost memory and prompt-processing work when it does not.
Start a new conversation with a short checked handoff when the old discussion contains many abandoned versions. Include the current facts, the latest draft and the next task. Do not depend on a summary to preserve every earlier detail.
Unload a model you no longer need from the model picker. Close other heavy work while making your comparison. The first request after loading may include setup time, so compare that separately from repeated requests with the model already available.
Change advanced settings carefully
OGAD exposes CPU threads and Batch size controls. More threads or a larger batch is not a universal speed setting. The result depends on the processor, memory and workload.
Keep a note of the starting values. Change one setting, repeat the same task, and return to the earlier value if it does not help. Use the app's normal settings as a starting point before attempting machine-specific tuning.
This guide does not promise a particular tokens-per-second figure. Your hardware, model, context and other running work all affect the result.
The controls described here are available in OGAD 0.0.51.
Use a comparison that tells you where the wait occurs
Choose one short task and run it several times with the same settings. Note the initial delay before the answer starts and how long the complete usable answer takes. These are different costs: a model can start slowly but then produce text quickly.
For example, use a 150-word meeting note and ask for three decisions plus any missing owner. Keep that same note and output request when trying a smaller model. Do not compare a three-bullet answer from one model with a long essay from another.
In the Text settings, Generation details can show context use, speed, token count and time beneath each answer. Read the result alongside the answer quality. A higher token rate is not a useful improvement if the output omits the decision you needed.
Try changes in a useful order
Start with the request: remove irrelevant history and ask for the output you will actually use. Next compare a smaller model that fits comfortably. Then adjust context if it is larger than the task needs. Advanced tuning comes after those checks.
Keep a short record of the model, context and setting changed. Restore the old value when a change does not help. This prevents several simultaneous changes from leaving you with no idea what caused the difference.
For a second test, use your normal work environment with the apps you usually keep open. A configuration that feels quick only after closing everything may not suit your day. Choose the setup that helps you finish the task with acceptable accuracy and enough room for the rest of your work.
Find your practical local model
Download OGAD and compare two sensible model choices on one real task. Keep the choice that helps you finish the work, with your data processed locally after setup.

Top comments (0)