GPT-6 Sol and Luna: Half-Price Models With Mixed Benchmarks

Date:

OpenAI has released GPT-6 Sol and GPT-6 Luna, two cheaper siblings to GPT-6 Astra that trade small benchmark shifts for a large price cut. Sol drops from $4/$20 to $2/$10 per million input/output tokens, while Luna drops from $0.20/$1.20 to $0.10/$0.50. Independent testing shows scores that are mostly flat or slightly down versus GPT-5.6, with a few clear wins in coding and automation, which makes this launch a cost-efficiency play rather than a capability leap.

Source: OpenAI

What OpenAI Launched and What It Costs

OpenAI introduced GPT-6 Sol and Luna on September 22, 2026, as the everyday-work tiers of the GPT-6 family. Sol targets complex coding and professional agent work, while Luna targets fast, high-volume tasks such as summarization and extraction. Both were trained with similar methods to Astra and inherit its factuality, coding, computer-use, and alignment improvements.

The headline is pricing. OpenAI describes it as 50% cheaper than GPT-5.6 promotional pricing, and the company says improved caching and inference let it pass savings directly to users:

  1. GPT-6 Sol: $2 input / $10 output per 1M tokens, down from $4 / $20.
  2. GPT-6 Luna: $0.10 input / $0.50 output per 1M tokens, down from $0.20 / $1.20.
  3. Prompt caching keeps a 90% discount on cached input reads, with new dashboards and diagnostics.
  4. API names are gpt-6-sol and gpt-6-luna, available in ChatGPT Work and Codex, with Luna also in the desktop app for Free and Go users.

One detail worth flagging: Luna output falls from $1.20 to $0.50, which is a 58.3% cut, not 50%. OpenAI labels the row 50% cheaper, but the math favors buyers even more. Sol is exactly 50% down in both directions.

GPT-6 Sol vs GPT-5.6 Sol: Small Moves, Lower Cost

On the Artificial Analysis intelligence and coding agent testing, Sol looks level overall. The Intelligence Index v4.3.2 moves from 47 to 48, while domain indexes are mixed: Finance and Accounting 49 to 48, Strategy and Ops 54 to 52, Legal 50 to 51, Healthcare and Medical 45 to 43, Engineering flat at 49, and Economics 53 to 54.

Benchmark GPT-5.6 Sol GPT-6 Sol Change
AA-Briefcase v1.1 (score / Elo) 49% / 1487 49% / 1483 Flat score, minus 4 Elo
GDPval-AA v2.1 (score / Elo) 54% / 1588 49% / 1487 Down 5 percentage points, down 101 Elo
AutomationBench-AA 60.1% 61.6% Up 1.5 percentage points
Terminal-Bench 4.0 39.9% 43.9% Up 4.0 percentage points
SciCode 57.1% 57.6% Up 0.5 percentage points
Humanity’s Last Exam 49.5% 47.9% Down 1.6 percentage points
AA-Omniscience Accuracy 59% 54% Down 5 percentage points, fewer hallucinations
MMMU-Pro 83% 83% Flat
MLCR-AA 19.4% 26.1% Up 6.7 percentage points, biggest Sol gain

Artificial Analysis notes Sol costs about $1.06 per Intelligence Index task versus $1.99 for its predecessor, with slightly higher token use offset by the price cut. In its Coding Agent Index, Sol rises from 55 to 57 and sits on the cost-efficiency frontier at roughly half the cost per task.

GPT-6 Luna vs. GPT-5.6 Luna: Cheaper, Mostly Level

Luna shows the same pattern at the low end. The Intelligence Index holds at 37, with small gains in Finance and Accounting (39 to 40), Strategy and Ops (44 to 45), Legal (37 to 39), Healthcare (36 to 37), Engineering (36 to 37), and Economics (43 to 45).

Benchmark GPT-5.6 Luna GPT-6 Luna Change
AA-Briefcase v1.1 (score / Elo) 42% / 1345 40% / 1299 Down 2 percentage points, down 46 Elo
GDPval-AA v2.1 (score / Elo) 47% / 1443 43% / 1367 Down 4 percentage points, down 76 Elo
AutomationBench-AA 50.2% 53.2% Up 3.0 percentage points
Terminal-Bench 4.0 11.6% 12.6% Up 1.0 percentage point
SciCode 53.6% 54.6% Up 1.0 percentage point
Humanity’s Last Exam 39.5% 38.5% Down 1.0 percentage point
AA-Omniscience Accuracy 43% 44% Up 1 percentage point
MMMU-Pro 79% 76% Down 3 percentage points
MLCR-AA 16.1% 16.1% Flat

Luna costs about $0.07 per Intelligence Index task versus $0.18 before, a roughly 60% drop. Its Coding Agent Index slips from 43 to 41, driven by lower SWE-Atlas-QnA and DeepSWE scores, but again at much lower cost per task.

Bar chart concept showing GPT-6 Sol and Luna benchmark scores beside falling price tags
Independent scores are mostly flat, so the price cut carries this launch. (Credit: Intelligent Living)

Vals.ai Results and Scores That Need a Second Look

The Vals.ai professional suites show the same mixed picture. For Sol, the overall Vals Index slips from 63.71% to 62.57%, with notable drops in Finance Agent v2 (53.76% to 49.05%), Legal Research Bench (48.08% to 28.85%), Public Benefits Bench (66.51% to 56.63%), and Tax Agent Bench (67.95% to 52.05%). Gains appear in Vibe Code Bench v1.1 (80.50% to 87.82%), Code Migration (52.92% to 57.20%), and MLCR-style long-context work.

For Luna, the Vals Index slips from 59.88% to 58.45%, with Finance Agent v2 down from 55.04% to 49.87% and Terminal-Bench 2.1 down from 79.03% to 73.03%, partly offset by gains in Vibe Code Bench, Tax Agent Bench, and SAGE.

Three results deserve caution before repeating them as fact:

  • Legal Research Bench for Sol drops nearly 20 percentage points while Luna rises slightly, an unusual split that suggests a harness, retrieval, or rubric change rather than pure model quality.
  • GDP.pdf and CritPt decline for both models even as AutomationBench and Terminal-Bench improve, which matches Artificial Analysis observations about shorter deliverables that omit rubric elements.
  • Harvey Legal Agent scores near 1 to 3%, and ProgramBench near 0 to 2% for all four models, so those near-floor scores should not be used to rank the models.

Treat the Vals.ai finance, legal, tax, and public-benefits drops as provisional until retests confirm them, especially since OpenAI reports its own factuality error rate falling by about half for Sol.

What OpenAI Claims vs. What Independents Found

OpenAI emphasizes cost per task rather than raw scores. On AutomationBench 1.0.6, Sol at xhigh effort scores 33.2% at $0.27 per task, above Opus 5 at max effort and low-effort Astra at a fraction of the cost. On Agents Last Exam, Sol at max effort scores 56.4%, above the best Opus 5 score in that evaluation at 60% lower cost. On DeepSWE v1.1, Sol scores 68.8% and Luna 66.6%, close to Claude Fable 5 and Opus 5 levels at 80 to 96% lower cost per task.

Independent results partly support that framing but add context. Artificial Analysis confirms the cost-per-task halving and the hallucination reduction, with Sol cutting its AA-Omniscience hallucination rate from 92% to 60% by answering fewer questions. It also confirms regressions in GDPval-AA and AA-Briefcase for Luna, attributing them to weaker presentation quality and missing deliverable elements after manual review of hundreds of outputs.

The timing matters. Anthropic released Opus 5.5 a day earlier and took the lead on most benchmarks, a story already covered separately. OpenAI answered with two cheaper models but no GPT-6 Terra, which fuels speculation, clearly labeled as speculation, that this release was accelerated to blunt the Opus news cycle and that Sol and Luna could receive further tuning in the coming days or weeks.

Frequently Asked Questions

How much cheaper are GPT-6 Sol and Luna?

Sol is 50% cheaper on input and output. Luna is 50% cheaper on input and about 58% cheaper on output, even though OpenAI summarizes both as 50% cheaper.

Are GPT-6 Sol and Luna smarter than GPT-5.6?

Not consistently. Most independent scores are flat or slightly lower, with gains concentrated in automation, terminal coding, SciCode, and long-context MLCR work for Sol. The upgrade is cost efficiency, plus lower hallucination rates and better caching.

Which should developers pick, Sol or Luna?

Sol fits complex coding and agent workflows where quality matters most. Luna fits high-volume drafting, extraction, and classification where cost per task dominates. Astra remains the top choice for the most demanding work.

Will benchmarks change after launch?

Possibly. Early releases are often followed by silent tuning, and the missing Terra tier plus the rapid response to Opus 5.5 suggest more updates may arrive. Recheck provider and Artificial Analysis leaderboards before making long-term routing decisions.

Bottom Line

GPT-6 Sol and Luna do not reclaim the benchmark lead from Opus 5.5, but they do not need to. At half price, flat to slightly lower scores still move the cost-efficiency frontier, especially for coding agents and business automation. Watch retests of the finance, legal, and tax suites before treating the largest drops as final, and expect possible follow-up tuning given how quickly this release followed Anthropic.

Aaron Jackson
Aaron Jackson
With a decade of hands-on experience in publishing and social media, and a B.Eng in Robotics from UWE, I'm passionate about turning challenges into opportunities. My focus is on creating solutions rather than merely highlighting problems.

Share post:

Popular

Text to Speech in 2026: How AI Voice Generation Got Its Human Edge Back

For most of the last decade, the phrase "text...

Light Origins Open-Sources Light-O1-Preview 6B Whole-Body Model

On September 21, 2026, Shenzhen-based startup Light Origins published...

UBTECH Started Delivering the UWorld U1 Companion Humanoid to Homes

UBTECH Robotics began delivering the UWorld U1 to consumers'...

Hy4 Preview vs. MiMo V2.6 Pro: Every Known Benchmark and Price

Two open-weight flagships landed within four weeks of each...