OpenAI has released GPT-6 Sol and GPT-6 Luna, two cheaper siblings to GPT-6 Astra that trade small benchmark shifts for a large price cut. Sol drops from $4/$20 to $2/$10 per million input/output tokens, while Luna drops from $0.20/$1.20 to $0.10/$0.50. Independent testing shows scores that are mostly flat or slightly down versus GPT-5.6, with a few clear wins in coding and automation, which makes this launch a cost-efficiency play rather than a capability leap.
What OpenAI Launched and What It Costs
OpenAI introduced GPT-6 Sol and Luna on September 22, 2026, as the everyday-work tiers of the GPT-6 family. Sol targets complex coding and professional agent work, while Luna targets fast, high-volume tasks such as summarization and extraction. Both were trained with similar methods to Astra and inherit its factuality, coding, computer-use, and alignment improvements.
The headline is pricing. OpenAI describes it as 50% cheaper than GPT-5.6 promotional pricing, and the company says improved caching and inference let it pass savings directly to users:
- GPT-6 Sol: $2 input / $10 output per 1M tokens, down from $4 / $20.
- GPT-6 Luna: $0.10 input / $0.50 output per 1M tokens, down from $0.20 / $1.20.
- Prompt caching keeps a 90% discount on cached input reads, with new dashboards and diagnostics.
- API names are gpt-6-sol and gpt-6-luna, available in ChatGPT Work and Codex, with Luna also in the desktop app for Free and Go users.
One detail worth flagging: Luna output falls from $1.20 to $0.50, which is a 58.3% cut, not 50%. OpenAI labels the row 50% cheaper, but the math favors buyers even more. Sol is exactly 50% down in both directions.
GPT-6 Sol vs GPT-5.6 Sol: Small Moves, Lower Cost
On the Artificial Analysis intelligence and coding agent testing, Sol looks level overall. The Intelligence Index v4.3.2 moves from 47 to 48, while domain indexes are mixed: Finance and Accounting 49 to 48, Strategy and Ops 54 to 52, Legal 50 to 51, Healthcare and Medical 45 to 43, Engineering flat at 49, and Economics 53 to 54.
| Benchmark | GPT-5.6 Sol | GPT-6 Sol | Change |
|---|---|---|---|
| AA-Briefcase v1.1 (score / Elo) | 49% / 1487 | 49% / 1483 | Flat score, minus 4 Elo |
| GDPval-AA v2.1 (score / Elo) | 54% / 1588 | 49% / 1487 | Down 5 percentage points, down 101 Elo |
| AutomationBench-AA | 60.1% | 61.6% | Up 1.5 percentage points |
| Terminal-Bench 4.0 | 39.9% | 43.9% | Up 4.0 percentage points |
| SciCode | 57.1% | 57.6% | Up 0.5 percentage points |
| Humanity’s Last Exam | 49.5% | 47.9% | Down 1.6 percentage points |
| AA-Omniscience Accuracy | 59% | 54% | Down 5 percentage points, fewer hallucinations |
| MMMU-Pro | 83% | 83% | Flat |
| MLCR-AA | 19.4% | 26.1% | Up 6.7 percentage points, biggest Sol gain |
Artificial Analysis notes Sol costs about $1.06 per Intelligence Index task versus $1.99 for its predecessor, with slightly higher token use offset by the price cut. In its Coding Agent Index, Sol rises from 55 to 57 and sits on the cost-efficiency frontier at roughly half the cost per task.
GPT-6 Luna vs. GPT-5.6 Luna: Cheaper, Mostly Level
Luna shows the same pattern at the low end. The Intelligence Index holds at 37, with small gains in Finance and Accounting (39 to 40), Strategy and Ops (44 to 45), Legal (37 to 39), Healthcare (36 to 37), Engineering (36 to 37), and Economics (43 to 45).
| Benchmark | GPT-5.6 Luna | GPT-6 Luna | Change |
|---|---|---|---|
| AA-Briefcase v1.1 (score / Elo) | 42% / 1345 | 40% / 1299 | Down 2 percentage points, down 46 Elo |
| GDPval-AA v2.1 (score / Elo) | 47% / 1443 | 43% / 1367 | Down 4 percentage points, down 76 Elo |
| AutomationBench-AA | 50.2% | 53.2% | Up 3.0 percentage points |
| Terminal-Bench 4.0 | 11.6% | 12.6% | Up 1.0 percentage point |
| SciCode | 53.6% | 54.6% | Up 1.0 percentage point |
| Humanity’s Last Exam | 39.5% | 38.5% | Down 1.0 percentage point |
| AA-Omniscience Accuracy | 43% | 44% | Up 1 percentage point |
| MMMU-Pro | 79% | 76% | Down 3 percentage points |
| MLCR-AA | 16.1% | 16.1% | Flat |
Luna costs about $0.07 per Intelligence Index task versus $0.18 before, a roughly 60% drop. Its Coding Agent Index slips from 43 to 41, driven by lower SWE-Atlas-QnA and DeepSWE scores, but again at much lower cost per task.

Vals.ai Results and Scores That Need a Second Look
The Vals.ai professional suites show the same mixed picture. For Sol, the overall Vals Index slips from 63.71% to 62.57%, with notable drops in Finance Agent v2 (53.76% to 49.05%), Legal Research Bench (48.08% to 28.85%), Public Benefits Bench (66.51% to 56.63%), and Tax Agent Bench (67.95% to 52.05%). Gains appear in Vibe Code Bench v1.1 (80.50% to 87.82%), Code Migration (52.92% to 57.20%), and MLCR-style long-context work.
For Luna, the Vals Index slips from 59.88% to 58.45%, with Finance Agent v2 down from 55.04% to 49.87% and Terminal-Bench 2.1 down from 79.03% to 73.03%, partly offset by gains in Vibe Code Bench, Tax Agent Bench, and SAGE.
Three results deserve caution before repeating them as fact:
- Legal Research Bench for Sol drops nearly 20 percentage points while Luna rises slightly, an unusual split that suggests a harness, retrieval, or rubric change rather than pure model quality.
- GDP.pdf and CritPt decline for both models even as AutomationBench and Terminal-Bench improve, which matches Artificial Analysis observations about shorter deliverables that omit rubric elements.
- Harvey Legal Agent scores near 1 to 3%, and ProgramBench near 0 to 2% for all four models, so those near-floor scores should not be used to rank the models.
Treat the Vals.ai finance, legal, tax, and public-benefits drops as provisional until retests confirm them, especially since OpenAI reports its own factuality error rate falling by about half for Sol.
What OpenAI Claims vs. What Independents Found
OpenAI emphasizes cost per task rather than raw scores. On AutomationBench 1.0.6, Sol at xhigh effort scores 33.2% at $0.27 per task, above Opus 5 at max effort and low-effort Astra at a fraction of the cost. On Agents Last Exam, Sol at max effort scores 56.4%, above the best Opus 5 score in that evaluation at 60% lower cost. On DeepSWE v1.1, Sol scores 68.8% and Luna 66.6%, close to Claude Fable 5 and Opus 5 levels at 80 to 96% lower cost per task.
Independent results partly support that framing but add context. Artificial Analysis confirms the cost-per-task halving and the hallucination reduction, with Sol cutting its AA-Omniscience hallucination rate from 92% to 60% by answering fewer questions. It also confirms regressions in GDPval-AA and AA-Briefcase for Luna, attributing them to weaker presentation quality and missing deliverable elements after manual review of hundreds of outputs.
The timing matters. Anthropic released Opus 5.5 a day earlier and took the lead on most benchmarks, a story already covered separately. OpenAI answered with two cheaper models but no GPT-6 Terra, which fuels speculation, clearly labeled as speculation, that this release was accelerated to blunt the Opus news cycle and that Sol and Luna could receive further tuning in the coming days or weeks.
Frequently Asked Questions
How much cheaper are GPT-6 Sol and Luna?
Sol is 50% cheaper on input and output. Luna is 50% cheaper on input and about 58% cheaper on output, even though OpenAI summarizes both as 50% cheaper.
Are GPT-6 Sol and Luna smarter than GPT-5.6?
Not consistently. Most independent scores are flat or slightly lower, with gains concentrated in automation, terminal coding, SciCode, and long-context MLCR work for Sol. The upgrade is cost efficiency, plus lower hallucination rates and better caching.
Which should developers pick, Sol or Luna?
Sol fits complex coding and agent workflows where quality matters most. Luna fits high-volume drafting, extraction, and classification where cost per task dominates. Astra remains the top choice for the most demanding work.
Will benchmarks change after launch?
Possibly. Early releases are often followed by silent tuning, and the missing Terra tier plus the rapid response to Opus 5.5 suggest more updates may arrive. Recheck provider and Artificial Analysis leaderboards before making long-term routing decisions.
Bottom Line
GPT-6 Sol and Luna do not reclaim the benchmark lead from Opus 5.5, but they do not need to. At half price, flat to slightly lower scores still move the cost-efficiency frontier, especially for coding agents and business automation. Watch retests of the finance, legal, and tax suites before treating the largest drops as final, and expect possible follow-up tuning given how quickly this release followed Anthropic.
