DEV Community

Owen
Owen

Posted on Originally published at ofox.ai

Grok 4.6 API Pricing: The 200K Cliff and When 4.5 Wins

Grok 4.6 API Pricing: The 200K Cliff and When 4.5 Wins

Grok 4.6 costs $2.00 per million input tokens and $6.00 output, which is exactly what Grok 4.5 costs. The two model cards are otherwise identical down to the rate limits. One number differs, cached input, and it is the newer model that is more expensive there.

The bigger number is the one that is not a difference between models at all: at 200,000 prompt tokens, every rate doubles, and it doubles for the whole request.

Model IDs:     grok-4.6 / grok-4.5  (x-ai/grok-4.6 on gateways)
Context:       500K tokens, both
Under 200K:    $2.00 in / $6.00 out    both models
                $0.50 cached on 4.6, $0.30 cached on 4.5
200K and over: $4.00 in / $12.00 out   both models
                $1.00 cached on 4.6, $0.60 cached on 4.5
Threshold:     applies to ALL tokens in the request, not the excess
Measured:      fixed per-request overhead 206 tokens on 4.6, 494 on 4.5
Break-even:    ~575 cached tokens; below it 4.6 is cheaper per call
Snapshot:      2026-08-19
Enter fullscreen mode Exit fullscreen mode

How Much Does the Grok 4.6 API Cost?

$2.00 in, $0.50 cached in, $6.00 out per million tokens, until the prompt reaches 200K. Straight from xAI's model list:

Model Context Input Cached input Output
grok-4.6, prompt under 200K 500K $2.00 $0.50 $6.00
grok-4.6, prompt 200K and over 500K $4.00 $1.00 $12.00
grok-4.5, prompt under 200K 500K $2.00 $0.30 $6.00
grok-4.5, prompt 200K and over 500K $4.00 $0.60 $12.00
grok-4.3 1M $1.25 $0.20 $2.50

Grok 4.3 is in that table for a reason. It is 38% cheaper on input and 58% cheaper on output than either 4.x flagship, and it carries a 1M context instead of 500K.

Gateways pass the headline rate through unchanged. OpenRouter and ofox both list x-ai/grok-4.6 at $2.00 and $6.00 with a 500,000-token context. What neither catalog exposes is the second row, so the doubling above 200K is invisible until it lands on your invoice.

What Happens When Your Prompt Crosses 200K Tokens?

The whole request reprices, not the overflow. xAI's wording is unambiguous: requests whose prompt reaches the threshold are billed at the higher rate for all tokens in the request.

Two requests, 2,000 tokens apart, both with 2,000 tokens of output:

Prompt size Input cost Output cost Total
199,000 tokens $0.398 $0.012 $0.410
201,000 tokens $0.804 $0.024 $0.828

A 1% larger prompt for a 102% larger bill. There is no gradual slope here, and no partial credit for the tokens below the line.

Two things make that line easier to cross than it looks. First, the threshold counts the prompt, so a long conversation walks toward it one turn at a time. Second, your prompt is not the only thing in the prompt, which is the subject of the next section.

If you are batching, split before the line rather than after it. Two 150K requests at the low tier cost $0.60 in input; one 300K request costs $1.20 for the same tokens.

What Is the Difference Between Grok 4.6 and Grok 4.5?

One number on the published cards.

Field Grok 4.6 Grok 4.5
Context window 500,000 500,000
Modalities text, image to text text, image to text
Function calling Yes Yes
Structured outputs Yes Yes
Reasoning Yes Yes
Requests per second 150 150
Tokens per minute 50,000,000 50,000,000
Regions us-east-1, us-west-2 us-east-1, us-west-2
Input / output $2.00 / $6.00 $2.00 / $6.00
Cached input $0.50 $0.30
Aliases none listed grok-4.5-latest, grok-build-latest

Cached input is 67% more expensive on the newer model. That is the opposite of the usual direction and it is the entire published basis for keeping 4.5 around.

The alias row matters if you use Grok Build: grok-build-latest resolves to Grok 4.5, not 4.6.

Why Does the Same Prompt Cost More on Grok 4.5?

Because a fixed block of input tokens rides along with every request, and it is more than twice as large on 4.5.

We sent identical bodies of 1, 50 and 200 repeated words to four models through an OpenAI-compatible gateway on 2026-08-19 and fit the line. Every model tokenized the body at exactly 1.00 token per word. The intercepts did not match:

Model Fixed input overhead Of which reported cached
x-ai/grok-4.6 206 tokens 128
x-ai/grok-4.5 494 tokens 384
x-ai/grok-4.3 4 tokens none
z-ai/glm-5.3 12 tokens none

Four and twelve tokens are an ordinary chat template. Two hundred and four hundred and ninety-four are a preamble. Adding your own system message raised both figures by exactly the size of that message, so this sits underneath anything you send.

We could not test a second route to attribute it, so treat the source as unproven. Either way it is on your invoice, so price it.

Grok 4.6 Grok 4.5
Cached portion 128 × $0.50/M = $0.000064 384 × $0.30/M = $0.000115
Uncached portion 78 × $2.00/M = $0.000156 110 × $2.00/M = $0.000220
Per request $0.000220 $0.000335
Per 1M requests $220 $335

It also eats your headroom. On 4.5 you reach the 200K cliff 494 tokens earlier than your own token count suggests. If you are budgeting a prompt at exactly 199,800 tokens, you are already over.

Which One Is Cheaper for Your Workload?

Around 575 cached tokens, the answer flips.

Grok 4.5 saves $0.20 per million cached tokens, which is $0.0000002 per cached token. It loses $0.000115 per request on the larger preamble. Divide one by the other and you get 575 tokens.

  • Cached prefix under ~575 tokens: Grok 4.6 costs less per call
  • Cached prefix over ~575 tokens: Grok 4.5 costs less, and the gap widens linearly

That covers input. Output is where the two models genuinely diverge. We ran the same code-generation prompt through both, 8 runs each:

Grok 4.6 Grok 4.5
completion_tokens, median 216 225
reasoning_tokens, median 723 30
Billed output, median 948 263
Latency, median 15.8 s 5.0 s
Cost per 1,000 tasks, all-in $6.18 $2.64

That last row is the only one in the table that is not an output-side number, so here is its basis. It is output at $6.00 per million plus the full input at $2.00 per million, and the input side includes the fixed per-request overhead: 244 input tokens on 4.6, 532 on 4.5. Output alone would be $5.69 and $1.58. We did not record cache hits on these runs, so input is priced entirely at the uncached rate, making the all-in figure an upper bound.

Watch the field names. completion_tokens excludes reasoning tokens on these models: a response reporting 216 completion tokens had 723 reasoning tokens alongside it, and total_tokens was the sum of prompt, completion and reasoning. Any cost estimator built on prompt_tokens × input + completion_tokens × output undercounts the billed output by 77%.

So the short version: 4.5 for long cached prefixes and short deterministic work, 4.6 when the extra reasoning is the point.

How Do I Call Grok 4.6?

from openai import OpenAI

client = OpenAI(api_key="YOUR_KEY", base_url="https://api.ofox.ai/v1")

r = client.chat.completions.create(
    model="x-ai/grok-4.6",          # x-ai/grok-4.5 for the cheaper cache tier
    messages=[{"role": "user", "content": "Refactor this module and run the tests."}],
)
u = r.usage
billed_output = u.completion_tokens + u.completion_tokens_details.reasoning_tokens
print(u.prompt_tokens, billed_output)
Enter fullscreen mode Exit fullscreen mode

Three things that bite people on the way in:

  • The ID is namespaced on gateways. grok-4.6 works against xAI directly; on OpenRouter and ofox it is x-ai/grok-4.6.
  • logprobs fails silently. xAI documents it as unsupported on grok-4.20 and newer. We sent logprobs: true with top_logprobs: 3 to both models: HTTP 200, no error, and no logprobs object in the response.
  • Both models are in us-east-1 and us-west-2 only. If you have data-residency constraints outside the US, this pair is not the answer regardless of price.

For rate limits across providers, our LLM API rate limits comparison has the numbers side by side.

References


Originally published on ofox.ai/blog.

Top comments (0)