The "is GPT-6 Sol free?" question keeps circulating, and the answer is a flat no on every surface that matters. I dug through OpenAI's launch announcement, the model pages, and the pricing docs to work out what you actually get on each plan, and where the token math changes the decision.
Short version: GPT-6 Sol is not on the Free API rate-limit tier — the official comparison lists it as unsupported, with API access starting at Tier 1. On the consumer side, Free and Go accounts get GPT-6 Luna in the desktop app instead. Sol landed in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu, and neither Sol nor Luna was in the standard Chat experience at launch.
The part people conflate: free Luna on a desktop app is an interactive entitlement, not an API allowance. It doesn't hand you a key, and it doesn't give you a gpt-6-sol quota. The same boundary applies to GPT-6.1 Sol — its model page also lists the Free tier as unsupported. ChatGPT subscription access and API-key billing are two separate ledgers.
Access routes, mapped
| Route | GPT-6 Sol? | Requirement | Billing |
|---|---|---|---|
| ChatGPT Free | No | Free account | $0 |
| ChatGPT Go | No | Go subscription | Subscription |
| ChatGPT Pro | Yes | Pro subscription ($100 / $200 / $500) | Included usage, plan limits |
| ChatGPT Work | Yes | Plus, Pro, Business, Enterprise, Edu | Included usage + limits |
| Codex | Yes | Eligible paid plan | Included usage or credits |
| OpenAI API | Yes | Supported usage tier | Per token |
Sol's own positioning is complex coding and agentic workflows. It takes text and image input, returns text, and supports function calling, structured outputs, and built-in tools through the Responses API. Chat Completions function calling works only when reasoning_effort is set to none — worth knowing before you wire up a client.
What the $500 Pro tier does and doesn't buy
Pro starts at $100/month, with $200 and $500 tiers above it. All of them route you to the GPT-6 family through Work and Codex, subject to rollout and workspace settings. GPT-6.1 Sol is rolling out to eligible paid accounts starting with Pro.
The $500 tier is not a shortcut to 6.1 Sol as an exclusive model. Its actual model advantage is GPT-6 Astra in Ultrafast mode, plus bigger allowances. So if the goal is simply using Sol or 6.1 Sol, a supported Pro plan is enough — don't buy the top tier for that.
The three models, side by side
| GPT-6.1 Sol | GPT-6 Sol | GPT-6 Luna | |
|---|---|---|---|
| Positioning | Complex coding, computer use, professional work | Complex coding, agentic workflows | High-volume, bounded work |
| Context | 1.05M | 1.05M | 1.05M |
| Max output | 128K | 128K | 128K |
| Reasoning effort | low → max; no none/minimal | none → max | none → max |
| Input / 1M | $2.00 | $2.00 | $0.10 |
| Output / 1M | $10.00 | $10.00 | $0.50 |
| Free ChatGPT | No at launch | No | Yes, desktop |
The 6.1 Sol column is a distinct model profile — don't carry legacy Sol/Luna benchmark numbers onto it. Architecture details aren't publicly disclosed for any of the three.
Picking a price route
| Route | Input | Cached | Output | Use when |
|---|---|---|---|---|
| OpenAI Standard | $2.00 | $0.20 | $10.00 | Interactive API |
| OpenAI Batch/Flex | $1.00 | $0.10 | $5.00 | Async, latency-tolerant |
| CometAPI short context | $1.60 | $0.16 | $8.00 | Sync work where unit price matters |
| OpenAI Fast mode | $4.00 | $0.40 | $20.00 | Latency-critical |
Batch and Flex are confirmed at 50% of Standard; Fast mode runs at twice the applicable rate. My rule of thumb after running the numbers: batchable jobs go to Batch/Flex, ordinary synchronous traffic goes to the cheaper synchronous route, and Fast mode needs a measured reduction in total task cost before its 2x unit price is defensible.
Cache reads are where the arithmetic gets interesting
Sol's cached input is $0.20/M against $2.00/M uncached — a 90% cut on reads. 6.1 Sol pushes that to $0.10/M. Cache writes cost $2.50 on either Sol model, which is the part people forget when they model savings off the read discount alone.
Take a repeated prompt: 100K reusable input tokens, 10K fresh input tokens, 5K output.
no cache: 0.11 * $2.00 + 0.005 * $10 = $0.27
100K cached (Sol): 0.10 * $0.20 + 0.01 * $2.00 + 0.005 * $10 = $0.09
100K cached (6.1): 0.10 * $0.10 + 0.01 * $2.00 + 0.005 * $10 = $0.08
That's roughly 67% off for this request shape, and 6.1 Sol comes in about 11% under Sol on the same already-cached prefix. Real numbers will move with the reusable share of the prompt, cache eligibility, and your actual hit rate. These figures exclude write charges, tool fees, and long-context premiums.
Routing instead of picking one model
The useful pattern isn't choosing a winner. It's a Luna-first, Sol-fallback pipeline: send routine requests to Luna, escalate low-confidence output, failed validations, and genuinely hard coding to Sol or 6.1 Sol. Luna lists at $0.10/$0.50 per 1M on OpenAI's side and from $0.08/M input on the gateway route, and it reaches 66.6% on DeepSWE 1.1 at max reasoning effort — plenty for classification, extraction, structured summarization, routing, and first-pass codegen where you can verify output with tests or rules.
Where Luna falls down is long-horizon agents, ambiguous specs, and debugging loops, where one failed run costs more than the token savings.
For a different behavioral profile rather than a lower price, Claude Sonnet 5 is worth an A/B test — it's around $1.60/$8 per 1M for short-context workloads, similar to discounted Sol access. Same prompt, different planning style, tool-call patterns, and failure modes. Compare task success rate, retries, latency, tool-call efficiency, and cost per completed task, not sticker price.
Astra is an escalation tier, not a cheaper Sol. Keep it for requests already flagged as unusually hard — failed validation, confidence under threshold, repeated failed attempts, or explicitly high-risk work. Per-token price is the wrong metric; cost per successful task is the right one.
Migrating from Sol to 6.1 Sol is not automatic
Two documented differences bite: 6.1 Sol rejects none and minimal reasoning effort, and tool calling requires the Responses API. If your current deployment depends on reasoning_effort=none with Chat Completions, that's real migration work, not a config flip.
The cache-read discount only pays off if your prompts actually reuse large prefixes — system instructions, repo context, tool schemas, reference docs. Persistent coding agents hit that profile; one-shot requests don't. Benchmark both models on representative tasks and measure completion rate, retry frequency, latency, cached-token usage, and total cost per successful task before you touch production. Keep Sol if compatibility and proven results outweigh the cache saving.
Loose ends worth knowing before you deploy
A subscription does not convert into API credits — Pro, including the $500 tier, doesn't give you an unrestricted API-key allowance for either Sol model. Plan access and API billing stay separate.
Batch or Flex makes sense when completion can slip: offline enrichment, eval runs, scheduled processing. Standard is for synchronous interactive traffic.
And watch the context boundary: once a prompt exceeds 272K input tokens, OpenAI applies the long-context rate to the entire request — input and cache rates double, output rises 50%. Retrieval and context trimming can cut the bill more than model routing will.
Availability moves: as of September 30, 2026, the launch announcement states Free/Go desktop Luna access with no end date, but that's an entitlement snapshot, not a promise. Recheck plan pages and token rates before you commit a budget or an architecture to them.
Originally published at cometapi.com
Top comments (0)