Quick follow-up to Zero GPU, zero dollars.
My AI crew talks to one address on my laptop: a LiteLLM "brain switch". Every job (chat, think hard, code, look at pictures, read long documents) has a list of free AI services to try in order. If one is out of free uses, the next one answers. If the internet dies, the laptop's own Ollama models take over.
Until today that list had 5 services: Groq, Gemini, Mistral, OpenRouter and Cloudflare Workers AI. Now it has 7.
NVIDIA NIM
build.nvidia.com gives you a free API key for a big menu of open models. The limit is about 40 requests a minute for the whole account, which is plenty for one person.
The star for me is Kimi K3. It now leads my "think hard" job:
- answers in about 3 seconds
- handles tool calls properly (a lot of free models fumble these)
- saves Gemini Pro's small free daily allowance for when I really need it
GLM-5.3 from the same key backs up chat, thinking and coding.
Ollama Cloud
If you already use Ollama locally, the free cloud plan is the same app with bigger models running on their servers. The catch: the free plan only offers a few models (gpt-oss, gemma4:31b, nemotron-3-super), runs one request at a time, and has 5-hour and weekly limits they don't publish.
So it's a backup, not a main brain. It sits at the bottom of the chat list, and Gemma 4 31B is the second choice for looking at pictures.
The order now
| Job | Tries first | Then |
|---|---|---|
| Think hard | Kimi K3 (NVIDIA) | Gemini Pro, GLM-5.3, Nemotron, gpt-oss |
| Chat | gpt-oss on Groq | Gemini Flash, Llama 3.3 on Cloudflare, Gemma 4, Mistral, GLM-5.3, Ollama Cloud |
| Code | Codestral (Mistral) | GLM-5.3, Kimi K3, Groq, Gemini, OpenRouter, Cloudflare |
| Pictures | Gemini Flash | Gemma 4 on Ollama Cloud, OpenRouter |
| Long docs | Gemini Flash | Kimi K3, Nemotron |
Still no graphics card, still no credit card. The whole config is in my crew repo (bring your own keys, none are included).
Zero GPU. Zero dollars. Zero excuses. 🦝
Top comments (0)