DEV Community

Yong Yu
Yong Yu

Posted on Originally published at yongboyu.hashnode.dev

GPT-6 Sol vs Grok 4.7: Which One Is Actually Cheaper? (Cache Pricing, Long-Context Cliffs, and Token Burn)

Attributed compile (not original research)

Primary source: GPT-6 Sol vs Grok 4.7: Same Input Price, Which Model Costs Less to Finish the Job? by Curtis Pyke (Kingy AI), published 2026-09-23

Second source: GPT-6 Sol vs Grok 4.7: The Cheaper Token That Costs Three Times as Much by Magnus Corvin (OrcaRouter), published 2026-09-22

This post restates their analysis in my own words and adds a small cost calculator. I did not run either model. Every price and benchmark figure below is quoted from a named source and labeled vendor-reported or third-party. Check the rate cards again before you make a decision.


Grok 4.7 shipped on September 21. GPT-6 Sol shipped the next day. Both list $2 per million input tokens, and most quick comparisons stop there: Grok charges $6 per million output tokens and Sol charges $10, so Grok looks 40% cheaper.

Both source posts argue that this is the wrong place to stop, and I agree. What you actually pay depends on four things a headline rate card hides:

  1. Cached input, which covers repeated system prompts, repo summaries, and long agent histories
  2. Long-context cliffs, where a whole request gets repriced once the prompt passes a threshold
  3. How many tokens the model spends on the same task
  4. How often the answer is usable, since a retry costs as much as the first attempt

Below I go through each one, then give you a script so you can do the math on your own traffic instead of trusting someone else's example.

The rate card, including the lines people skip

List prices per 1M tokens, as reported in both source posts, which cite OpenAI's Sol model reference and xAI's pricing page:

GPT-6 Sol Grok 4.7
API model ID gpt-6-sol grok-4.7
Input $2.00 $2.00
Cached input (read) $0.20 $0.50
Output $10.00 $6.00
Long-context trigger prompt over 272K tokens prompt at or above 200K tokens
Long-context rates $4 / $0.40 / $15 $4 / $1 / $12
Context window 1.05M 500K
Reasoning effort none … max (default medium) low … xhigh

Two details matter most:

  • Cached reads cost 2.5× more on Grok than on Sol. Grok wins on output, but Sol wins on the line that dominates long-running agents and chat sessions with a stable prefix.
  • Both long-context surcharges apply to the whole request, not just the tokens past the threshold. Between 200K and 272K tokens, Grok has already switched to its long-context tier while Sol is still on standard pricing.

Curtis Pyke's post also points out that the effort names don't mean the same thing across vendors. Sol medium and Grok medium can spend very different numbers of tokens, so compare what the bill says, not the setting's label.

Why "cheaper per token" can still mean "more expensive per task"

Magnus Corvin's post puts the gap in a single figure. When Artificial Analysis (a third party) ran its Intelligence Index, Grok 4.7 at xhigh produced roughly 240M output tokens, compared with about 77M for GPT-6 Sol. OrcaRouter reports the resulting cost per task as roughly $3.74 for Grok vs $1.06 for Sol, with index scores of 46 vs 48 (OrcaRouter, citing Artificial Analysis). At Grok's high setting, Artificial Analysis's comparison page lists $2.73 per task.

The source flags some caveats, and they belong in any summary. The two models ran at different effort levels (Grok xhigh/high, Sol max). The two listings may also come from different builds of the index. The safe conclusion is "Grok used a lot more tokens on this suite," not "Sol is 2 points smarter." VentureBeat covered the same token-consumption story.

This doesn't make Grok a bad pick. Its vendor-reported launch numbers are strong. xAI reports DeepSWE v1.1 at 71.0% (high effort) and Terminal-Bench 4.0 at 38.0%, and OpenAI reports Sol at 68.8% on DeepSWE v1.1 at max (Kingy AI summary). Pyke is careful to say that two vendor tables with the same benchmark name don't make a controlled head-to-head. We don't know if the harness, effort level, or cost accounting matched.

Four scenarios where the ranking flips

These are the worked examples from Pyke's post, recomputed with the script below. The token counts are hypothetical and chosen to isolate pricing. They aren't measurements.

Scenario Sol Grok Cheaper
100K fresh input, 10K output $0.30 $0.26 Grok (by ~13%, not 40%)
Warm cache: 100K cached + 10K fresh input, 2K output $0.060 $0.082 Sol
250K input, 10K output (Grok past its cliff, Sol not) $0.60 $1.12 Sol
300K input, 20K output (both past their cliffs) $1.50 $1.44 Grok

The same two models change places three times depending on how the request is shaped. That's why a "Grok is cheaper" or "Sol is cheaper" headline tells you nothing about your app.

Pyke also gives a handy break-even rule for cached, short-context work. Assume equal fresh input and equal output length. The difference comes down to $0.30 × cached − $4 × output (in millions of tokens). Solving for zero, Sol comes out cheaper once you read more than about 13 cached tokens per output token, and Grok comes out cheaper below that. A coding agent that re-reads a big repo summary for every short diff falls on Sol's side. A long-form writing job with a small prompt falls on Grok's side. It's a billing rule, not a quality rule, and it assumes both models spend the same number of tokens. As shown above, they often don't.

Do the math on your own traffic

Here's a small calculator that encodes the list prices above, including each vendor's cliff behavior. Keep the price table in one place and update it when the rate cards change.

# cost.py: list prices per 1M tokens as of late Sep 2026. Re-verify before use.
PRICES = {
    "gpt-6-sol": {"cliff": 272_000, "strict": True,    # repriced when prompt > 272K
                  "std": (2.00, 0.20, 10.00), "long": (4.00, 0.40, 15.00)},
    "grok-4.7":  {"cliff": 200_000, "strict": False,   # repriced when prompt >= 200K
                  "std": (2.00, 0.50, 6.00),  "long": (4.00, 1.00, 12.00)},
}

def request_cost(model, fresh_in, cached_in, out):
    """out must include reasoning tokens. Both vendors bill them as output."""
    p = PRICES[model]
    prompt = fresh_in + cached_in
    over = prompt > p["cliff"] if p["strict"] else prompt >= p["cliff"]
    i, c, o = p["long"] if over else p["std"]   # the tier applies to the WHOLE request
    return (fresh_in * i + cached_in * c + out * o) / 1_000_000

def cost_per_accepted(model, runs):
    """runs: [{"tokens": (fresh, cached, out), "passed": bool}, ...]"""
    total = sum(request_cost(model, *r["tokens"]) for r in runs)
    accepted = sum(r["passed"] for r in runs)
    return total / accepted if accepted else float("inf")
Enter fullscreen mode Exit fullscreen mode

How to use it:

  1. Log real usage from both APIs. Each response reports input, cached, and output tokens (reasoning is included in or listed next to output). Save them per request, along with the effort setting.
  2. Run 20–30 of your own tasks on each model with a pass/fail check you can automate: tests pass, the JSON validates, the answer matches.
  3. Compare cost_per_accepted, not request_cost. A model that's 20% cheaper per call but fails twice as often loses.
  4. Watch your prompt sizes. Anything that regularly lands between 200K and 272K tokens gets billed very differently on the two vendors.
  5. Pin the effort level and write it next to every number you report. Most arguments about vendor charts come down to mismatched effort settings.

What this leaves out: cache-write charges (Sol lists $2.50/1M for writes; Grok's rate isn't on the same sheet, per OrcaRouter), tool and web-search fees, Batch/Flex discounts (Pyke notes Sol lists them at half price while Grok's Batch API is unsupported), and reseller pricing that differs from direct API rates.

A practical starting point

Putting both authors' recommendations together:

  • Already on OpenAI, coding agents, heavy prompt caching: start with Sol at high. The cheaper cache reads and lower token use compound.
  • Output-heavy generation with modest prompts: try Grok at high, and check that reasoning tokens aren't eating the output discount.
  • Prompts around 200K–272K tokens, or above 500K: Sol, on price and window size alone. A bigger window doesn't prove better long-document understanding, though.
  • Hard coding tasks where finishing matters more than rates: keep both on the shortlist. There's no published matched-harness result that picks a winner.

Magnus Corvin sums it up well: if a per-token comparison ends with "so Grok is cheaper," it stopped one step too early. The next step is counting tokens, and after that, counting accepted answers.


Compiler (not original author)

Compiled by YongBo Yu, AI engineer in Toronto.

Site: https://yongbo-yu.vercel.app · GitHub: https://github.com/YongBoYu1

Analysis and worked examples credit Curtis Pyke (Kingy AI) and Magnus Corvin (OrcaRouter). Read their originals for full sourcing.

Top comments (0)