You have an NVIDIA GPU in your Windows PC. You want to use it for local AI chat, images or transcription, without spending the evening assembling inference engines and their libraries.
OGAD (Off Grid AI Desktop) can download its NVIDIA components from inside the app. Install the optional pack, restart, then run a local model. The app also shows which backend is actually in use, so you can check whether a request used CUDA or a fallback.
This guide uses v0.0.54-beta.108, a Windows x64 prerelease. It covers local models. The app, GPU components and model files need an initial download; local inference can work without internet once setup is complete. Release.
Download the Windows beta | Off Grid AI
What would you like to do with Off Grid AI?
Have a feature or use case you would like us to support? Tell us what you want to do and which device you use.
Write to support@offgridmobileai.co, join our Slack community, or talk to us on Reddit.
What does the NVIDIA download give you?
The optional pack supplies NVIDIA processing engines and runtime libraries for chat and vision, image generation, transcription and the relevant Computer Use engines. It does not supply all the models you might want to run. Download your chosen model separately.
The practical benefit is one place to install these components. You do not need to build the native engines or install the full CUDA developer toolkit to use the released app. You do need a working NVIDIA driver on the host computer.
The Windows pack is about 5.14 GB to download. Allow more disk space for extraction and the installed files. Its size is not a VRAM requirement, and it does not make a model fit into a GPU that lacks enough memory. Pack manifest.
| Component | What the pack helps run |
|---|---|
| Chat and vision | Local answers and supported image understanding |
| Image generation | Local image models |
| Transcription | Audio-to-text models |
| Computer Use engines | Grounding and decision processing used by that feature |
Voice replies and search embeddings use separate processing settings. Installing this pack does not prove that every AI task in the app now uses CUDA. It also does not change the license requirements of a feature that uses one of these engines.
How do you enable NVIDIA processing?
1. Start with the released Windows app
Install the Windows x64 beta linked above. If you already use a stable version, note that these instructions refer to the newer beta controls.
Complete local model setup first, or download one local text model from Models. Choose a model that fits the memory available on your PC. Keep it for both your first CPU or Vulkan request and your later CUDA request; changing models at the same time makes a comparison less useful.
2. Open the GPU download panel
Go to Settings → GPU performance. The Use your NVIDIA GPU panel shows the optional components and their download state.
Choose Download. You can continue to use the app while it downloads. The panel supports pausing and resuming; if the download fails, use the retry action rather than fetching an unrelated archive from elsewhere.
The app checks the expected file size and SHA-256 hash before it installs the pack. Wait for installation to complete, then select Restart app when prompted. A completed download is not the same as the running app having loaded the new engines.
3. Choose the processing preference for the task
In Chat, select the Settings button to open Model settings. For chat and vision, use the backend controls in Text. Image processing has its own controls in Image; audio-to-text has them in Transcription.
Start with Auto, or select CUDA when you want to prefer the NVIDIA engine. A preference takes effect when the model next loads. If the model is already running, unload and load it again before checking the result.
CUDA is unavailable for the pack-controlled tasks until the pack is ready. Selecting it after installation still does not guarantee that the engine can use your particular driver and hardware. The app can try a fallback if the preferred route fails.
How do you check that a real task uses the GPU?
Use a short piece of work you can review. For example, paste this into a local chat:
Turn these notes into a short status update: the import works; the export still needs testing; the next check is Friday. Do not invent a release date.
After the model loads, look at the backend control's Now value. That is the current runtime information. The selected preference describes what you asked the app to try.
Check that the answer is useful as well. It should preserve the incomplete export test and avoid adding a date you did not supply. The goal is a working local task, not merely a GPU label.
To compare responsiveness, use the same model and prompt with the CPU preference and then CUDA. Reload between changes. The first run can include model-loading work, so do not treat it as a reliable performance benchmark. GPU memory, model size and other programs on the machine all affect the result.
What should you do if CUDA is not used?
| Symptom | Check | Next action |
|---|---|---|
| The NVIDIA pack is not offered | Is a working NVIDIA driver installed? | Check the driver on the PC; the app checks for its presence |
| CUDA still says download required | Did the pack finish and the app restart? | Complete installation, then restart |
| You changed the preference but Now is unchanged | Was the model already loaded? | Reload the model |
| CUDA falls back | Can that engine load with the current driver and available memory? | Try a smaller model, then compare with Auto or CPU |
| Speech replies use another backend | Speech has separate settings | Check Voice, not just the GPU pack card |
Do not assume that a card is compatible because it carries the NVIDIA name. The release does not establish one minimum driver version or a measured speed improvement for every GPU. If CPU or Vulkan works, you can keep using local AI while you investigate the CUDA route.
For a failed or slow request, Settings → AI activity can show the model, backend, request status and error. Inspect that one request rather than changing several settings at once. Logs can include your prompt text, so review copied details before sharing them.
Can you use the result without internet?
Yes, for downloaded local models and installed runtime components. Finish the model and GPU downloads, make one successful request, then disconnect and repeat a local task. Web tools and remote providers still need their own connections.
Start with one model and one useful request. Install the Windows beta, add the NVIDIA pack if your PC supports it, and check the actual backend while the model does your work.

Top comments (0)