DEV Community

Cover image for GPT-6 Sol vs Luna: I’d Route by Cost per Accepted Result
Nathan Brooks
Nathan Brooks

Posted on Originally published at cometapi.com

GPT-6 Sol vs Luna: I’d Route by Cost per Accepted Result

The useful distinction between GPT-6 Sol and Luna is how much work it takes to get an acceptable result. I’d start Luna on tasks with cheap, reliable validation and use Sol for coding and agent workflows where a failed attempt creates substantial downstream work.

OpenAI’s September 22, 2026 release added both models below Astra in the GPT-6 lineup. Astra remains the highest-capability option for the hardest end-to-end tasks; Sol targets demanding reasoning at a lower token rate; Luna targets focused, repeatable work at volume.

That positioning gives me a routing hypothesis. Production evaluations still have to establish whether it holds.

Shared limits, different workloads

The API changelog and model comparison list these specifications:

Property GPT-6 Sol GPT-6 Luna
Model ID gpt-6-sol gpt-6-luna
Input Text and images Text and images
Output Text Text
Context window 1,050,000 tokens 1,050,000 tokens
Maximum output 128,000 tokens 128,000 tokens
APIs Responses, Chat Completions Responses, Chat Completions
Intended workload Complex coding and agentic workflows Focused, high-volume tasks

Context capacity therefore gives me no reason to choose one over the other.

Sol supports reasoning effort from none through max. Its intended workloads include repository analysis, difficult debugging, tool decisions, and agent-driven software changes. I’d evaluate it where sustained reasoning might eliminate retries or reduce review time.

Luna fits extraction, classification, routing, support triage, template-driven responses, structured summaries, and first-pass transformations. Its economics are attractive when I can define success precisely and check the output cheaply.

The qualification matters: a low token bill can coexist with an expensive workflow if rejected outputs keep reaching reviewers.

The pricing threshold I’d check first

OpenAI separates Standard short-context and long-context pricing. Once a prompt exceeds 272,000 input tokens, long-context rates apply to the entire request, including the portion below that threshold.

All figures below are USD per million tokens from the official pricing table.

Token category Sol: short / long Luna: short / long
Input $2.00 / $4.00 $0.10 / $0.20
Cached input $0.20 / $0.40 $0.01 / $0.02
Cache writes $2.50 / $5.00 $0.125 / $0.25
Output $10.00 / $15.00 $0.50 / $0.75

For Standard short-context requests, that puts Sol at $2 input and $10 output versus Luna’s $0.10 and $0.50. I’d require evidence that Sol’s extra capability saves enough retries, failures, or reviewer time to justify that difference for a particular task.

Batch, Flex, Fast mode, and eligible regional processing have separate rates. A useful estimate needs the service tier, full input length, output length, cache behavior, tool fees, and retry rate.

A gateway quote needs its own verification

For a unified multi-model API, CometAPI’s September 2026 catalog snapshot lists both models at 20% below the corresponding OpenAI short-context Standard rates, including discounted cache reads and writes.

Model Listed input / million tokens Calculated output / million tokens
Luna $0.08 $0.40
Sol $1.60 $8.00

The input prices are publicly displayed; the output figures are the corresponding 20%-off calculations. I’d recheck the catalog, pricing page, and account dashboard before using them in a budget. This is a dated snapshot, and neither pricing nor account access is guaranteed.

The gateway documents an OpenAI-compatible base URL usable with the OpenAI SDK and a gateway key. Its public catalog makes Luna the safer initial example; substituting gpt-6-sol requires confirming that the account’s Sol route is enabled.

What the benchmarks actually tell me

OpenAI’s launch evaluations provide several useful comparisons, provided the reasoning settings stay attached to the scores.

Evaluation Model and setting Reported result
AutomationBench Sol, xhigh 33.2%, estimated $0.27 per task
AutomationBench Astra, low effort 30.3%
DeepSWE v1.1 Sol, max 68.8%
DeepSWE v1.1 Luna, max 66.6%
OSWorld 2.0, offline evaluation Sol, xhigh 60.5%

AutomationBench covers workflows across applications. DeepSWE tests long-horizon software-engineering tasks.

I’d use these results to choose evaluation candidates and effort settings. They do not establish a universal ranking: Sol’s AutomationBench result and Astra’s result come from different effort settings, and production outcomes also depend on tool use and task design.

OpenAI also reported roughly half as many factual mistakes for Sol as for its GPT-5.6 predecessor on an internal conversation set. Those conversations were selected because users had flagged factual errors, so the result does not describe typical traffic.

For deployment, I care about the failure modes my application can encounter:

  • Incorrect tool selection or incomplete multi-step execution.
  • Schema failures and rejected outputs.
  • Factual errors that survive validation.
  • Latency, token consumption, retries, and human-review time.

A small representative evaluation set gives me more actionable evidence than a broad leaderboard position.

My starting routing policy

I’d assign work by its boundaries, validation cost, and consequences.

Route Work I’d send there What I’d measure
Luna Narrow, high-volume tasks protected by schemas, rules, or sampling Whether rejection, retry, and review costs preserve the token savings
Sol Demanding code, debugging, repository analysis, and multi-step agents Whether stronger execution reduces total workflow cost
Astra The hardest ambiguous, tool-heavy, high-consequence, or expensive-to-review tasks Whether the highest capability ceiling reduces downstream failures and intervention

Astra has the highest token price. That can still be economical when difficult failures cost more than inference.

My default would be routine work to Luna, demanding cases to Sol, and the hardest end-to-end work to Astra. Failed validation, tool requirements, low confidence, unusually long prompts, or a high cost of error can trigger escalation.

I’d treat those signals as hypotheses to test. Confidence needs validation too, and a long prompt alone does not distinguish Sol from Luna: their published context limits are identical.

API details that affect the implementation

Both model IDs are live through OpenAI’s Responses and Chat Completions APIs.

For reasoning with built-in tools and function calling, I’d use Responses. Chat Completions supports ordinary requests, but function calling with either Sol or Luna requires reasoning_effort="none". That constraint belongs in the integration decision before any model comparison.

Gateway support needs separate testing. If the account dashboard documents Responses and chat support for a selected route, I’d exercise each required path before rollout.

For every evaluation request, I’d record the provider, route, model ID, reasoning setting, latency, tokens, validation result, and billed cost. Keeping direct-provider and gateway measurements distinguishable makes routing and billing discrepancies easier to investigate.

Terra is not an available planning assumption

As of September 23, 2026, OpenAI had not announced GPT-6 Terra or explained its absence. The public GPT-6 catalog listed Astra, Sol, and Luna; the Sol and Luna changelog entry included no Terra model ID.

GPT-5.6 included Terra, which explains the expectation. It does not establish a GPT-6 release plan. A release date, specification, benchmark, and price remain unknown; explanations involving lineup strategy, pricing, or timing are speculation.

For the models that are available, I’d make the decision with the same representative prompts and compare accepted-output rate, latency, total tokens, retries, validation failures, and review time. The number I’d optimize is cost per accepted result.


Originally published at cometapi.com

Top comments (0)