Grok 4.6 API Pricing: The 200K Cliff and When 4.5 Wins
Grok 4.6 costs $2.00 per million input tokens and $6.00 output, which is exactly what Grok 4.5 costs. The two model cards are otherwise identical down to the rate limits. One number differs, cached input, and it is the newer model that is more expensive there.
The bigger number is the one that is not a difference between models at all: at 200,000 prompt tokens, every rate doubles, and it doubles for the whole request.
Model IDs: grok-4.6 / grok-4.5 (x-ai/grok-4.6 on gateways)
Context: 500K tokens, both
Under 200K: $2.00 in / $6.00 out both models
$0.50 cached on 4.6, $0.30 cached on 4.5
200K and over: $4.00 in / $12.00 out both models
$1.00 cached on 4.6, $0.60 cached on 4.5
Threshold: applies to ALL tokens in the request, not the excess
Measured: fixed per-request overhead 206 tokens on 4.6, 494 on 4.5
Break-even: ~575 cached tokens; below it 4.6 is cheaper per call
Snapshot: 2026-08-19
How Much Does the Grok 4.6 API Cost?
$2.00 in, $0.50 cached in, $6.00 out per million tokens, until the prompt reaches 200K. Straight from xAI's model list:
| Model | Context | Input | Cached input | Output |
|---|---|---|---|---|
| grok-4.6, prompt under 200K | 500K | $2.00 | $0.50 | $6.00 |
| grok-4.6, prompt 200K and over | 500K | $4.00 | $1.00 | $12.00 |
| grok-4.5, prompt under 200K | 500K | $2.00 | $0.30 | $6.00 |
| grok-4.5, prompt 200K and over | 500K | $4.00 | $0.60 | $12.00 |
| grok-4.3 | 1M | $1.25 | $0.20 | $2.50 |
Grok 4.3 is in that table for a reason. It is 38% cheaper on input and 58% cheaper on output than either 4.x flagship, and it carries a 1M context instead of 500K.
Gateways pass the headline rate through unchanged. OpenRouter and ofox both list x-ai/grok-4.6 at $2.00 and $6.00 with a 500,000-token context. What neither catalog exposes is the second row, so the doubling above 200K is invisible until it lands on your invoice.
What Happens When Your Prompt Crosses 200K Tokens?
The whole request reprices, not the overflow. xAI's wording is unambiguous: requests whose prompt reaches the threshold are billed at the higher rate for all tokens in the request.
Two requests, 2,000 tokens apart, both with 2,000 tokens of output:
| Prompt size | Input cost | Output cost | Total |
|---|---|---|---|
| 199,000 tokens | $0.398 | $0.012 | $0.410 |
| 201,000 tokens | $0.804 | $0.024 | $0.828 |
A 1% larger prompt for a 102% larger bill. There is no gradual slope here, and no partial credit for the tokens below the line.
Two things make that line easier to cross than it looks. First, the threshold counts the prompt, so a long conversation walks toward it one turn at a time. Second, your prompt is not the only thing in the prompt, which is the subject of the next section.
If you are batching, split before the line rather than after it. Two 150K requests at the low tier cost $0.60 in input; one 300K request costs $1.20 for the same tokens.
What Is the Difference Between Grok 4.6 and Grok 4.5?
One number on the published cards.
| Field | Grok 4.6 | Grok 4.5 |
|---|---|---|
| Context window | 500,000 | 500,000 |
| Modalities | text, image to text | text, image to text |
| Function calling | Yes | Yes |
| Structured outputs | Yes | Yes |
| Reasoning | Yes | Yes |
| Requests per second | 150 | 150 |
| Tokens per minute | 50,000,000 | 50,000,000 |
| Regions | us-east-1, us-west-2 | us-east-1, us-west-2 |
| Input / output | $2.00 / $6.00 | $2.00 / $6.00 |
| Cached input | $0.50 | $0.30 |
| Aliases | none listed |
grok-4.5-latest, grok-build-latest
|
Cached input is 67% more expensive on the newer model. That is the opposite of the usual direction and it is the entire published basis for keeping 4.5 around.
The alias row matters if you use Grok Build: grok-build-latest resolves to Grok 4.5, not 4.6.
Why Does the Same Prompt Cost More on Grok 4.5?
Because a fixed block of input tokens rides along with every request, and it is more than twice as large on 4.5.
We sent identical bodies of 1, 50 and 200 repeated words to four models through an OpenAI-compatible gateway on 2026-08-19 and fit the line. Every model tokenized the body at exactly 1.00 token per word. The intercepts did not match:
| Model | Fixed input overhead | Of which reported cached |
|---|---|---|
x-ai/grok-4.6 |
206 tokens | 128 |
x-ai/grok-4.5 |
494 tokens | 384 |
x-ai/grok-4.3 |
4 tokens | none |
z-ai/glm-5.3 |
12 tokens | none |
Four and twelve tokens are an ordinary chat template. Two hundred and four hundred and ninety-four are a preamble. Adding your own system message raised both figures by exactly the size of that message, so this sits underneath anything you send.
We could not test a second route to attribute it, so treat the source as unproven. Either way it is on your invoice, so price it.
| Grok 4.6 | Grok 4.5 | |
|---|---|---|
| Cached portion | 128 × $0.50/M = $0.000064 | 384 × $0.30/M = $0.000115 |
| Uncached portion | 78 × $2.00/M = $0.000156 | 110 × $2.00/M = $0.000220 |
| Per request | $0.000220 | $0.000335 |
| Per 1M requests | $220 | $335 |
It also eats your headroom. On 4.5 you reach the 200K cliff 494 tokens earlier than your own token count suggests. If you are budgeting a prompt at exactly 199,800 tokens, you are already over.
Which One Is Cheaper for Your Workload?
Around 575 cached tokens, the answer flips.
Grok 4.5 saves $0.20 per million cached tokens, which is $0.0000002 per cached token. It loses $0.000115 per request on the larger preamble. Divide one by the other and you get 575 tokens.
- Cached prefix under ~575 tokens: Grok 4.6 costs less per call
- Cached prefix over ~575 tokens: Grok 4.5 costs less, and the gap widens linearly
That covers input. Output is where the two models genuinely diverge. We ran the same code-generation prompt through both, 8 runs each:
| Grok 4.6 | Grok 4.5 | |
|---|---|---|
completion_tokens, median |
216 | 225 |
reasoning_tokens, median |
723 | 30 |
| Billed output, median | 948 | 263 |
| Latency, median | 15.8 s | 5.0 s |
| Cost per 1,000 tasks, all-in | $6.18 | $2.64 |
That last row is the only one in the table that is not an output-side number, so here is its basis. It is output at $6.00 per million plus the full input at $2.00 per million, and the input side includes the fixed per-request overhead: 244 input tokens on 4.6, 532 on 4.5. Output alone would be $5.69 and $1.58. We did not record cache hits on these runs, so input is priced entirely at the uncached rate, making the all-in figure an upper bound.
Watch the field names. completion_tokens excludes reasoning tokens on these models: a response reporting 216 completion tokens had 723 reasoning tokens alongside it, and total_tokens was the sum of prompt, completion and reasoning. Any cost estimator built on prompt_tokens × input + completion_tokens × output undercounts the billed output by 77%.
So the short version: 4.5 for long cached prefixes and short deterministic work, 4.6 when the extra reasoning is the point.
How Do I Call Grok 4.6?
from openai import OpenAI
client = OpenAI(api_key="YOUR_KEY", base_url="https://api.ofox.ai/v1")
r = client.chat.completions.create(
model="x-ai/grok-4.6", # x-ai/grok-4.5 for the cheaper cache tier
messages=[{"role": "user", "content": "Refactor this module and run the tests."}],
)
u = r.usage
billed_output = u.completion_tokens + u.completion_tokens_details.reasoning_tokens
print(u.prompt_tokens, billed_output)
Three things that bite people on the way in:
-
The ID is namespaced on gateways.
grok-4.6works against xAI directly; on OpenRouter and ofox it isx-ai/grok-4.6. -
logprobsfails silently. xAI documents it as unsupported on grok-4.20 and newer. We sentlogprobs: truewithtop_logprobs: 3to both models: HTTP 200, no error, and nologprobsobject in the response. - Both models are in us-east-1 and us-west-2 only. If you have data-residency constraints outside the US, this pair is not the answer regardless of price.
For rate limits across providers, our LLM API rate limits comparison has the numbers side by side.
References
- xAI developer docs: models and pricing
- OpenRouter: x-ai/grok-4.6
- ofox model page: Grok 4.6
- ofox model page: Grok 4.5
Originally published on ofox.ai/blog.
Top comments (0)