GPT-6 Luna is level with GPT-5.6 Luna on public benchmarks (38 against 37 on the Artificial Analysis Intelligence Index) and costs about half as much per task. The difference is in what it writes: when a prompt only states a goal ("write the Q3 review"), GPT-6 Luna wrote 0.55x the words across 80 business documents. It did not skip parts it was asked for. With the required parts named, the two models were equally complete, and the one detail GPT-6 Luna dropped more often on goal-only prompts was the duration of an incident.
TL;DR
- Artificial Analysis scores GPT-6 Luna 38 and GPT-5.6 Luna 37 at
maxeffort, at $0.07 against $0.18 per task. - On goal-only prompts, GPT-6 Luna wrote 384 words per document against GPT-5.6 Luna's 703.
- GPT-6 Luna left the incident duration out of 9 of 20 goal-only incident postmortems; GPT-5.6 Luna out of 2.
- With the required parts named, GPT-6 Luna was complete on 74 of 80 documents against 69, at $0.81 against $1.61 per 1,000.
How does GPT-6 Luna compare with GPT-5.6 Luna on benchmarks?
GPT-6 Luna matches GPT-5.6 Luna on general benchmarks at 48% to 61% lower cost per task, and the two separate only on documents scored against a rubric (a scoring checklist). Scores are quoted at a reasoning effort, the API setting for how much the model reasons before it answers (none to max, default medium). The figures were read from the publishers' pages on 2026-10-02.
| GPT-6 Luna | GPT-5.6 Luna | |
|---|---|---|
| Released | 2026-09-22 | 2026-07-09 |
| List price, input / output per 1M tokens | $0.10 / $0.50 | $0.20 / $1.20 |
| Artificial Analysis Intelligence Index v4.3.2 at max (10 evals: knowledge, reasoning, coding, agentic and business-document tasks) | 38 | 37 |
| Cost per index task | $0.07 | $0.18 |
| Vals Index (coding, legal, tax, finance and other professional tasks) | 51.2% | 51.7% |
| Cost per Vals test | $0.43 | $0.82 |
| GDPval-AA v2.1 Elo at max (business documents scored against a hidden rubric; Elo is a rating from head-to-head comparisons, higher is better) | 1438 | 1463 |
| GDPval-AA v2.1 Elo at medium | 1262 | 1133 |
| Output speed, tokens per second | 131 | 124 |
GDPval-AA is where the complaint about GPT-6 Luna started. In its launch analysis, Artificial Analysis put GPT-6 Luna about 75 Elo points below GPT-5.6 Luna on GDPval-AA and about 45 below on AA-Briefcase v1.1, a second document benchmark, and attributed it to "reduced presentation quality and deliverables that omit rubric elements"; reviewers named it the "format tax". The leaderboard now shows a 25-point gap at max, and at medium GPT-6 Luna is 129 points ahead.
Against other vendors, GPT-6 Luna competes with the flash tier, their low-price, high-speed models. On the same index, DeepSeek V4.1 Flash at max scores 39 at $0.27 per task and Gemini 3.8 Flash at high scores 41 at $1.24, so GPT-6 Luna gives up 1 to 3 points for a cost per task 4 to 18 times lower. Both generate faster, at 209 and 249 tokens per second.
What else changed besides the price?
GPT-6 Luna keeps GPT-5.6 Luna's limits and settings; only the cache price and the knowledge cutoff moved. Both have a 1,050,000-token context window and a 128,000-token output limit, and both charge 2x input and 1.5x output for a whole request once the prompt passes 272K tokens. Cached input costs $0.01 per million tokens, down from $0.02, and the knowledge cutoff moved from Feb 16 to May 18, 2026. Effort is set with reasoning.effort on the Responses API (none, low, medium, high, xhigh, max), and the reasoning is hidden and billed as output tokens.
A week after the launch, OpenAI released GPT-6.1 Sol as the successor to GPT-6 Sol in the tier above, and our GPT-6.1 Sol measurements found its answers back at the old length. Luna got no update, so we measured what GPT-6 Luna writes shorter, what it leaves out, and what brings it back.
How did we test it?
We wrote document tasks whose requirements can be checked in code, so no judge model is involved. Each task gives the model a block of synthetic data and asks for a document built from it:
| Document | Data | What a reader expects in it |
|---|---|---|
| Quarterly review | revenue for 6 to 12 regions | every region covered, the region with the largest decline named |
| Incident postmortem | 7 to 10 timestamped events | every event, the duration of the incident |
| Vendor comparison | 5 to 8 hosting quotes | every vendor covered |
| Customer replies | 6 to 14 support tickets | a reply to every ticket, addressed to the customer by name |
Each task comes in two prompt styles over the same data. A goal-only prompt says what the document is and nothing about its parts ("Write the Q3 regional review for the leadership team"), and is graded only on the items in the table. A prompt with the parts named lists the sections, counts and word ranges, and is graded on all of them. A document is complete when nothing required is missing, too short or under-counted. Goal-only vendor comparisons count for length only, because the right vendor is a judgment call.
On 2026-10-02 we ran the 80 goal-only items on both models at five effort settings and on GPT-6 Luna at one verbosity setting, and the 80 named-part items on both at the default effort: 1,040 Responses API calls. Every prompt carried a unique random string so no response came from a cache, and cost uses OpenAI list prices. Both models answered the same items, so we compare them with an exact McNemar test (a paired test on the items where the two models differ) and a Holm correction, which tightens the threshold when several comparisons are tested at once. Only gaps that survive it are called differences.
Does GPT-6 Luna write less than GPT-5.6 Luna?
GPT-6 Luna wrote 0.55x the words of GPT-5.6 Luna on goal-only prompts at the default effort: 384 words per document against 703, shorter on all 80 items. It was shorter in every document type, most in postmortems (441 against 981 words) and least in customer replies (576 against 767).
With the parts named, the gap disappears: 738 words against 724.
The shorter default is not specific to Luna. On the same items, GPT-6 Sol wrote 325 words against 592 for GPT-5.6 Sol, and GPT-6.1 Sol is back at 565.
Does GPT-6 Luna leave things out?
On goal-only prompts GPT-6 Luna left out one thing more often than GPT-5.6 Luna: the duration of the incident in a postmortem. It usually gave the time window ("Incident window: 18:00-18:39 UTC") and left the subtraction to the reader. Everything else was there at every effort setting: every region and the worst one, every event, and a reply addressed by name for every ticket, with a single exception at none.
| Goal-only prompts | GPT-6 Luna complete | GPT-6 Luna postmortems without the duration | GPT-5.6 Luna complete | GPT-5.6 Luna postmortems without the duration |
|---|---|---|---|---|
none |
44 / 60 | 15 / 20 | 57 / 60 | 3 / 20 |
low |
47 / 60 | 13 / 20 | 54 / 60 | 6 / 20 |
medium |
51 / 60 | 9 / 20 | 58 / 60 | 2 / 20 |
high |
53 / 60 | 7 / 20 | 57 / 60 | 3 / 20 |
max |
59 / 60 | 1 / 20 | 60 / 60 | 0 / 20 |
The gap at none (44 against 57) survives the correction. At medium (51 against 58) it is significant only before correction, so read it as a direction, and at max the two are tied. All of it comes from one document type, so it shows GPT-6 Luna skipping a derived figure in a postmortem and no general habit of dropping content.
With the parts named, GPT-6 Luna was complete on 74 of 80 documents and GPT-5.6 Luna on 69, a tie. Every miss on both models was a memo or impact paragraph below its requested word range: 6 for GPT-6 Luna and 11 for GPT-5.6 Luna.
What brings the length and the details back?
Naming the parts in the prompt does, and the two request parameters do not. The same GPT-6 Luna at the same effort went from 384 words to 738 and stated the duration in every postmortem once the prompt asked for it. A request like this is enough:
Write the postmortem with these sections: Timeline (a table with every event),
Impact (120-200 words, state the total duration in minutes), Root Cause,
Action Items (5, each ending "Owner: <team>"), Lessons Learned (at least 3 bullets).
text.verbosity is a Responses API parameter (low, medium, high) for how long and detailed the visible answer is. Set to high, it made GPT-6 Luna's goal-only documents 7% longer (410 words) and left completeness where it was, 52 of 60 against 51.
Raising reasoning.effort changes how much the model thinks more than how much it writes. From none to max, GPT-6 Luna's documents grew from 370 to 436 words while reasoning went from 0 to 4,192 tokens per answer. max did recover the duration (missing from 1 of 20), at 4.5 times the cost of medium.
Both parameters with the OpenAI Python SDK:
from openai import OpenAI
client = OpenAI()
prompt = "Write the postmortem with these sections: ..." # data and required parts
resp = client.responses.create(
model="gpt-6-luna",
reasoning={"effort": "medium"}, # none, low, medium (default), high, xhigh, max
text={"verbosity": "high"}, # 7% longer in our runs; did not change completeness
input=prompt,
)
print(resp.output_text)
usage = resp.usage
print(usage.output_tokens, usage.output_tokens_details.reasoning_tokens)
output_tokens includes the reasoning tokens, so the visible answer is the difference between the two numbers.
How much does GPT-6 Luna save over GPT-5.6 Luna?
For the same document, GPT-6 Luna cost 49% less than GPT-5.6 Luna: $0.81 against $1.61 per 1,000 documents with the parts named. The price cut alone would have given more. GPT-5.6 Luna's tokens repriced at GPT-6 Luna's rates come to $0.68, 58% less, and GPT-6 Luna's answers cost 20% more than that, because it reasoned twice as much (537 reasoning tokens per answer against 271) and its total output grew from 1,284 to 1,558 tokens.
On goal-only prompts the bill looks better, $0.54 against $1.66 per 1,000 documents (68% less), but most of the extra saving is GPT-6 Luna writing half as much, which is a saving only if the other half was not needed.
Speed is about the same. GPT-6 Luna generated 115 output tokens per second of total request time against 125, and a goal-only document took a median of 8.1 seconds against 11.1 because there was less to write.
Do our results match other published tests?
Our results agree with the public benchmarks where they overlap, and add what is shorter, what is missing and what fixes it.
| Question | Published elsewhere | Our measurement | Verdict |
|---|---|---|---|
| General capability, GPT-6 Luna against GPT-5.6 Luna | Artificial Analysis 38 against 37; Vals 51.2% against 51.7% | complete on 74 against 69 of 80 with the parts named, a tie | Agrees: level |
| Document quality | GDPval-AA v2.1 Elo: 25 lower at max, 129 higher at medium | at medium, 0.55x the words on goal-only prompts and the incident duration missing from 9 of 20 postmortems against 2 | Mixed: neither shows a clear loss at the default effort; ours shows the shorter default and the one item it costs |
| Output tokens | 51K against 41K per index task at max, 24% more | 1,558 against 1,284 per document at medium, 21% more |
Agrees |
| Cost | 61% less per index task; 48% less per Vals test | 49% less with the parts named, 68% less on goal-only prompts | Agrees |
What did we observe, and which should you use?
Use GPT-6 Luna wherever you control the prompt, and name the parts you need; keep GPT-5.6 Luna only where prompts cannot be changed and readers expect long documents. Three observations lead there:
- GPT-6 Luna writes to the request. A goal gets a short document, and a list of parts gets a full one that matches GPT-5.6 Luna's.
- What it drops is narrow. Across four document types, the only expected item missing more often was the incident duration.
- The saving is the price cut. It uses 21% more output tokens for the same document, so the 49% saving is below the 58% that the list prices alone would give.
| Workload | Pick | Numbers |
|---|---|---|
| Machine-read output: classification, extraction, routing | GPT-6 Luna at the lowest effort your accuracy allows | list prices 50% (input) and 58% (output) lower; answer length does not matter when a program reads it |
| Documents from a template or a prompt you write | GPT-6 Luna with the sections, counts and lengths in the prompt | complete on 74 of 80 against 69; $0.81 against $1.61 per 1,000 |
| Documents from end-user prompts you cannot edit | GPT-6 Luna if you can add the expected parts to the request; otherwise GPT-5.6 Luna | goal-only: 384 against 703 words; duration missing from 9 of 20 postmortems against 2 |
| Rubric-scored deliverables | Put the rubric in the prompt; do not rely on text.verbosity
|
verbosity high: 52 against 51 of 60, 7% more words |
For documents that go to customers, check the required elements in code before sending, whichever model wrote them.
FAQ
Is GPT-6 Luna worse than GPT-5.6 Luna?
GPT-6 Luna was as complete as GPT-5.6 Luna when the prompt named the required parts (74 against 69 of 80 documents, a tie) at half the cost. It writes about half as much when a prompt only states a goal, and on those prompts it left the incident duration out of postmortems more often.
Does text.verbosity: "high" make GPT-6 Luna write as much as GPT-5.6 Luna?
text.verbosity: "high" added 7% more words on GPT-6 Luna in our runs, which leaves it at about 58% of GPT-5.6 Luna's length, and completeness did not change. Naming the parts in the prompt brought it to the same length.
Did GPT-6.1 fix GPT-6 Luna's short answers?
OpenAI's 6.1 update covers only the Sol tier: GPT-6.1 Sol writes at GPT-5.6 Sol's length again, and GPT-6 Luna is unchanged since its 2026-09-22 release.
Related measurements: GPT-6.1 Sol vs Claude Sonnet 5.5, flash-tier LLMs compared, and GPT-5.6 cost levers.
Top comments (0)