DEV Community

Cover image for Kimi K3, DeepSeek and GLM free from NVIDIA: I tested the claim and here is the catch
Kirill Lukyanov
Kirill Lukyanov

Posted on Originally published at klukyanov.ru

Kimi K3, DeepSeek and GLM free from NVIDIA: I tested the claim and here is the catch

A claim is going around: NVIDIA hands out a free API key for Kimi K3, DeepSeek V4.1 Flash and two flavours of GLM 5.3 — no credit card, fifteen minutes. I went through it myself. The models are real, the free tier is more generous than the reposts say, and there is one catch the headline leaves out: a phone number check that does not cover every country.

What NVIDIA gives away

NVIDIA Build is NVIDIA's catalog of hosted open models. The API is OpenAI-compatible: same /v1/chat/completions request, you only swap the base URL and the key. As of October 1, 2026, integrate.api.nvidia.com/v1/models lists 81 models, including the four most interesting open-weight releases of this autumn.

The part most reposts get wrong is the money. New accounts used to get roughly 1,000 credits, and people still quote that number. Credits are gone. NVIDIA's own account verification dialog now says "Unlimited API requests without daily limits". What is limited is the rate: threads on the NVIDIA developer forum converge on about 40 requests per minute per key, and a forum moderator stated in July that the limit depends on the model, use case and overall traffic and cannot be officially raised on the free tier.

40 RPM is 57,600 requests a day if you hammer it nonstop. For prototyping, a personal agent or evaluating a model on your own tasks, that is effectively unlimited. For production, NVIDIA points you to paid options: partner endpoints or self-hosted deployment.

The four models

All four are mixture-of-experts models, and all four ship with a 1,048,576-token context window.

Model API ID Total params Active per token Highlights License
Kimi K3 moonshotai/kimi-k3 ~2.8T not stated Long-horizon agentic coding, tool use, image input; thinking always on Modified MIT
GLM 5.3 z-ai/glm-5.3 753B ~40B Text, reasoning, tool calling MIT with a clause for >$10B-revenue resellers
GLM 5.3 Flash z-ai/glm-5.3-flash 320B 18B Text + images, tools, structured output, thinking budget MIT
DeepSeek V4.1 Flash deepseek-ai/deepseek-v4.1-flash 552B 8B prefill / 16B decode Multimodal, reasoning effort adjustable 1–100 MIT

Source: model cards on build.nvidia.com.

For coding I would start with Kimi K3 — it is the reason for the hype, and running it yourself takes a rack you do not have. Free access on someone else's GPUs is the only sensible way for most of us to try it. GLM 5.3 Flash and DeepSeek V4.1 Flash are for fast, cheap calls where you do not need the smartest answer.

The catch: phone verification

Sign-up itself is smooth: email, password, NVIDIA Developer Program account. No card, as promised. But on the API keys page you get a modal — "We'll need to verify your phone number" — and no key until you enter a one-time SMS code. NVIDIA frames it as fraud and abuse protection.

The country list is not global. Russia is not on it at all; under the list NVIDIA says it is "rapidly expanding worldwide availability". Kazakhstan is listed, but in my attempt with a Kazakh number the "Send Code to Phone" button never became active, and the page showed no error. I could not confirm whether Kazakh numbers go through.

I did not try to get around the check and would not recommend it: virtual numbers and borrowed SIMs violate the terms, and a key obtained that way can be revoked along with the account. So "fifteen minutes, no card" is true only if your phone number is from a supported country.

First request

If your number is supported, the first call takes a minute. Any client that lets you change the base URL works, including coding agents.

export NVIDIA_API_KEY="nvapi-..."

curl https://integrate.api.nvidia.com/v1/chat/completions \
  -H "Authorization: Bearer $NVIDIA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "moonshotai/kimi-k3",
    "messages": [{"role": "user", "content": "Explain Swift async/await in three sentences"}],
    "max_tokens": 1024
  }'
Enter fullscreen mode Exit fullscreen mode

Swap model for any ID from the table. One Kimi K3 detail: thinking is always on, and in multi-turn chats or tool calls you must send back the full previous assistant message, including reasoning_content and tool_calls, or it loses the thread.

My take

The offer itself is great: four of the strongest open models of the season, no payment, no daily cap, standard API. If your phone number qualifies, it is the easiest way to try Kimi K3 without renting hardware. Just know that "no card, fifteen minutes" quietly skips the phone step — and that step is where some of us stop.

Originally published at klukyanov.ru.
Shorter weekly write-ups (in Russian) — on Telegram.

Top comments (0)