Claude Sonnet 5.5 vs GPT-6 Sol, Astra and Fable 5.1: Which Model Wins?

Date:

Claude Sonnet 5.5 launched on September 28, 2026, at $2 per million input tokens and $10 per million output tokens. It is now the second most capable model in the field, ahead of both Claude Fable 5.1 and GPT-6 Astra, two models priced at $10 and $50. Only its own sibling, Claude Opus 5.5, scores higher, and that one costs twice as much per token.

That reads as a decisive result for the mid-tier. The per-benchmark detail is messier and more useful. Sonnet 5.5 beats GPT-6 Astra on 13 benchmarks and loses on 6. It beats its own sibling, Fable 5.1, on 8 and loses on 10, despite winning the composite index. The wins cluster in code migration, terminal work, and web app building. The losses cluster in finance, legal, tax, medical coding, and scientific research.

There is also a finding that has nothing to do with benchmarks as usually reported. Sonnet 5.5 posts the lowest knowledge-accuracy score in the field, at 53.95%, against 67.23% for Claude Fable 5.1. But it also posts the highest non-hallucination rate, at 52.99%.

So this is not a general improvement. It is a specialization, and the right way to buy is to route work rather than pick a winner. The headline result is simultaneously real and misleading: on an identical rate card, Sonnet 5.5 wins the benchmarks and loses the invoice, because the two models consume wildly different numbers of tokens to finish the same job. That tension runs through everything below. This article sets Sonnet 5.5 against the four models it competes with, using every publicly available figure across Vals AI and Artificial Analysis, plus Anthropic’s and OpenAI’s own claims. Where a figure could not be confirmed against a primary source, the cell is left blank rather than guessed.

One naming note. GPT-6 Sol is the September 2026 release at $2 and $10, the only model on Sonnet 5.5’s rate card, which makes it the natural cost comparison. GPT-6 Astra is OpenAI’s flagship at $10 and $50. They are different products separated by one price tier. We covered Sol and its Luna sibling separately in our GPT-6 Sol and Luna benchmark breakdown and Astra against the Claude flagships in our GPT-6 Astra comparison.

The short answer

There is no single winner, and the reason is more useful than a verdict. Six findings, in the order they should change your decisions.

  • A mid-tier model caught the flagships. Sonnet 5.5 outscores both GPT-6 Astra and Claude Fable 5.1 on the Artificial Analysis index while costing a fifth as much per token. That should have been the biggest story of the launch, and on most of the underlying benchmarks it is.
  • The composite index is misleading on its own. Sonnet 5.5 edges past Fable 5.1 on the Vals composite by a fraction of a point, then loses the head-to-head 8 to 10. The composite is being carried by coding benchmarks, not by the professional work Fable is bought for.
  • Astra owns the scientific tail. It is the only model in the field to finish IOI with a perfect score, and it takes Terminal-Bench Science and SRE Bench by margins nobody should dismiss. If your work touches research, the cheaper model is a downgrade in exactly the categories you need.
  • Sonnet 5.5 is the most cautious model in the field and the least accurate when decisive. It declines or qualifies nearly half of knowledge questions, posting the highest non-hallucination rate tested at 52.99%. But Fable 5.1 is right 13 points more often on the questions it attempts. In regulated work that gap decides the model.
  • Sol is the cost floor, not the capability ceiling. On an identical rate card, Sonnet 5.5 costs roughly seven times more per task. Same dollars per token, very different dollars per completed job.
  • Opus 5.5 still leads. It scores higher, reasons less, and is the only model in the field that beats Sonnet 5.5 on the composite despite charging twice as much per token.

Capability per dollar, in other words, is highest in the middle of this field, not at either end. The models that look most impressive on a leaderboard are frequently among the worse buys.

Price and specifications side by side

Every figure is as published by the providers in late September 2026.

Specification Claude Sonnet 5.5 GPT-6 Sol Claude Opus 5.5 Claude Fable 5.1 GPT-6 Astra Claude Opus 5 Claude Sonnet 5
Release date Sep 28, 2026 Sep 22, 2026 Sep 22, 2026 Sep 1, 2026 Sep 3, 2026 Jul 2026 Jun 2026
Input, per 1M tokens $2.00 $2.00 $4.00 $10.00 $10.00 $5.00 $2.00
Output, per 1M tokens $10.00 $10.00 $20.00 $50.00 $50.00 $25.00 $10.00
Cached input, per 1M $0.20 $0.20 $0.20 $0.25 $1.00 $0.50 $0.20
Cache write, 5-minute TTL $2.50 $2.50 $5.00 $12.50 $12.50 $6.25 $2.50
Cache write, 1-hour TTL $4.00 $2.50 $8.00 $20.00 $12.50 $10.00 $4.00
Context window 1M tokens 1M tokens 1M tokens 1M tokens 1.05M tokens 1M tokens 1M tokens
Max output tokens 128k 128k 128k 128k 128k 128k 128k
Knowledge cutoff June 2026 April 2026 June 2026 June 2026 April 2026 May 2026 January 2026
Input modalities Text, image, file Text, image Text, image Text, image Text, image Text, image Text, image
Long-context surcharge None, full 1M at standard rates 2x input, 1.5x output above 272K None, full 1M at standard rates None, full 1M at standard rates 2x input, 1.5x output above 272K None, full 1M at standard rates None, full 1M at standard rates
AA Intelligence Index, max 56 48 58 53 53 51 38
AA cost per task, max $7.60 $1.06 $5.98 $7.63 $3.26 $5.86 $5.09
Vals Index 69.22% 62.57% 69.69% 68.83% 66.61% 67.21% 59.61%

Sonnet 5.5 and GPT-6 Sol are priced identically to the dollar at short context, on the same context window, with matching cache-read pricing. That price parity is what makes the comparison interesting rather than settled. It also creates a trap: rate cards are per token, and token consumption is a behavioral property of a model, not a published specification. Two models on identical rate cards can differ sevenfold in what a task actually costs.

Long inputs are where the two part ways. Sonnet 5.5 charges $2 and $10 for a prompt of any size, up to its full 1M-token window. GPT-6 Sol does not. Cross 272K input tokens and the entire request re-bills at $4 and $15, including the first 272K. OpenAI applies the same rule to Astra at $20 and $75. Send a 300,000-token contract to Sol and you pay double on input and half again on output for the whole job, not just the excess. That single threshold will move more money than any benchmark gap in this article.

The cache-write TTLs are the other quiet difference. Anthropic charges 1.25x input for a five-minute cache and 2x for an hour, so Sonnet 5.5 costs $2.50 or $4.00 depending on which you choose. OpenAI has no equivalent choice: Sol pays $2.50 either way. In agentic coding, where most input tokens are cache reads, that 1.6x premium for an hour-long cache matters more than the headline input rate, because it decides whether you re-write the cache after every short break or keep it warm.

Identical AI model pricing producing very different per-task costs
(Credit: Intelligent Living)

Sonnet 5.5 against GPT-6 Sol, evaluation by evaluation

Artificial Analysis is the cleanest source for a direct comparison, since it evaluates every model on the same evaluations at matched effort levels. This is the full head-to-head, before the cost is taken into account.

Benchmark Claude Sonnet 5.5 GPT-6 Sol Winner
AA-Briefcase v1.1 66% 49% Sonnet 5.5
GDPval-AA v2.1 67% 49% Sonnet 5.5
AutomationBench-AA 71.3% 61.6% Sonnet 5.5
Terminal-Bench 4.0 63.6% 43.9% Sonnet 5.5
SciCode 61.0% 57.6% Sonnet 5.5
Humanity’s Last Exam 55.0% 47.9% Sonnet 5.5
GDP.pdf 25.8% 24.8% Sonnet 5.5
CritPt 31.4% 30.9% Sonnet 5.5
AA-Omniscience accuracy 53.95% 54.48% GPT-6 Sol
AA-Omniscience non-hallucination 52.99% 39.88% Sonnet 5.5
AA-LCR v1.1 (long context) 82.7% 83.7% GPT-6 Sol
Terminal-Bench-Science 0.1* 53.3% 30.0% Sonnet 5.5
MLCR-AA* 75.0% 16.1% Sonnet 5.5
Harvey LAB-AA (criterion pass rate)* 93.1% Not published –
Intelligence Index composite 56 48 Sonnet 5.5
Cost per task $7.60 $1.06 GPT-6 Sol

Sonnet 5.5 takes twelve of the fourteen rows it can be scored on. Both losses are knowledge-retrieval measures rather than capability ones, which matters for the interpretation further down.

Which is where cost enters, and it inverts the result. Sonnet 5.5 is the heaviest token consumer Artificial Analysis has ever measured, at roughly 193,000 output tokens per index task, about six times GPT-6 Sol’s 31,000. At identical output pricing, that single behavioral difference accounts for essentially the entire $7.60 versus $1.06 gap. The model that won the head-to-head on capability loses decisively on the bill, and the next section shows why that pattern repeats across the whole field.

One important caveat on that cost figure. It holds at maximum effort. Sonnet 5.5 costs between $0.41 and $7.60 per task depending on the setting, an eighteen-fold spread on one rate card, so any honest cost comparison has to name the effort level. Every benchmark figure in this article is measured at maximum effort.

Sonnet 5.5 versus the flagships: where it won and where it lost

Two models priced at $10 input and $50 output per million tokens were the reference points most buyers were comparing against three months ago. At a fifth of their token price, Sonnet 5.5’s composite win reads as a rout. The per-benchmark detail is where the useful information is.

Vals AI benchmark Claude Sonnet 5.5 GPT-6 Astra Claude Fable 5.1 Claude Opus 5.5
Vals Index (composite) 69.22% 66.61% 68.83% 69.69%
Vals RSI Index Not published 27.07% 36.09% 37.31%
Code Migration 69.83% 67.74% 54.61% 66.65%
Vibe Code Bench v1.1 92.39% 89.59% 90.26% 90.29%
ProgramBench Not published 5.50% 7.00% 18.50%
Terminal-Bench 4.0 53.03% 57.07% 49.49% 61.62%
Terminal-Bench Science 38.57% 65.71% 34.29% 48.57%
IOI 83.06% 100.00% 90.78% 95.06%
SRE Bench 30.15% 56.87% 22.90% 33.59%
MysteryMechanism 49.10% 53.15% 47.75% 49.55%
ProofBench v1.1 100.00% 99.00% 100.00% 100.00%
BioMysteryBench 81.11% 79.26% Not published 79.26%
Finance Agent v2 58.10% 53.54% 58.88% 58.59%
EMB (Excel Modelling) 75.71% 71.70% 76.67% 75.94%
Tax Agent Bench 73.39% 63.34% 77.64% 70.50%
Legal Research Bench 48.08% 39.42% 55.29% 50.48%
Public Benefits Bench 67.19% Not published 74.90% 70.64%
MedScribe 91.10% 87.91% 91.29% 91.43%
MedCode 52.92% 48.49% 53.51% 49.80%
CyberBench v1.1 59.58% 41.07% 70.42% 55.36%
Harvey’s Legal Agent 2.92% 5.42% 6.67% 3.75%
CUA-bench Not published 19.17% 13.17% 14.00%
Time Horizon Index Not published 90.50% 63.33% Not published

Counting only rows where all four models have a published score, Sonnet 5.5 beats GPT-6 Astra 13 to 6 and beats Claude Fable 5.1 by only 8 to 10 with one tie. The same comparison against GPT-6 Sol, where Sonnet 5.5 wins 18 of 19 with one tie, is far less interesting, which is why the flagship comparison is the one worth reading.

That second number is the one to sit with. Sonnet 5.5 wins the composite index against its own sibling by four tenths of a point and loses the head-to-head. The composite win is carried by code migration and terminal work, not by the professional knowledge tasks Fable 5.1 is bought for.

Astra is a different shape of defeat. It wins the scientific and systems end of the table outright, by margins that should not be dismissed. If your work is research-adjacent, Astra at $10 and $50 remains the better buy, and switching to Sonnet 5.5 would be a downgrade in exactly the categories where you needed the capability.

The pattern is consistent across every comparison in this article. Sonnet 5.5 wins on code, and loses or ties on the long professional and scientific tail. That is a coherent specialization, not a general improvement, and it is worth pricing accordingly.

Source: Vals AI

For context on where the professional tier still stands, Fable 5.1 also holds the strongest scores available on the older academic suite: 93.43% on GPQA Diamond, 90.52% on LiveCodeBench, 88.51% on LegalBench, 92.38% on MMLU Pro and 90.64% on MMMU Pro. Those benchmarks are not published for Sonnet 5.5, so no direct comparison is possible, but they explain why the professional tier has not been obsoleted.

All five models on the Artificial Analysis index

The composite scores hide more than they reveal. Artificial Analysis runs eleven individual evaluations, ten of which feed the Intelligence Index, and they break the pattern the composites suggest.

Artificial Analysis evaluation Claude Sonnet 5.5 Claude Opus 5.5 GPT-6 Astra Claude Fable 5.1 GPT-6 Sol
AA-Briefcase v1.1 66% 66% 53% 59% 49%
GDPval-AA v2.1 67% 67% 52% 62% 49%
AutomationBench-AA 71.3% 69.5% 68.5% 59.4% 61.6%
Terminal-Bench 4.0 63.6% 59.6% 59.1% 52.0% 43.9%
SciCode 61.0% 66.9% 56.5% 63.1% 57.6%
Humanity’s Last Exam 55.0% 61.4% 54.7% 59.1% 47.9%
GDP.pdf 25.8% 26.2% 31.0% 26.2% 24.8%
CritPt 31.4% 31.7% 31.7% 29.7% 30.9%
AA-Omniscience accuracy 53.95% 66.22% 62.60% 67.23% 54.48%
AA-Omniscience non-hallucination 52.99% 41.39% 48.66% 27.42% 39.88%
AA-LCR v1.1 (long context) 82.7% 84.7% 80.7% 85.3% 83.7%
Terminal-Bench-Science 0.1* 53.3% 59% 63.3% 43.3% 30.0%
MLCR-AA* 75.0% 66.7% 35% 71.1% 16.1%
Harvey LAB-AA (criterion pass rate)* 93.1% 91.2% Not published 93.0% Not published

Sonnet 5.5 is the weakest of the five models on knowledge accuracy. It scores the lowest of the five on AA-Omniscience accuracy.

The non-hallucination rate reverses the obvious reading completely. Sonnet 5.5 posts the highest rate of any model here, and Claude Fable 5.1 posts the lowest, despite holding the highest accuracy in the field.

So the two Claude models are mirror images. Fable 5.1 answers more often and is right more often when it answers. Sonnet 5.5 declines or qualifies nearly half of knowledge questions, but when it does commit it is the least likely model here to be hallucinating. That is the opposite of the story an earlier version of this article told, and the correction matters more than the original claim did.

The practical consequence cuts against the usual vendor framing. Fable 5.1’s 13-point accuracy lead over Sonnet 5.5 is a difference in competence, not caution: it is right far more often on the questions both models attempt. But where a declined answer is acceptable and a wrong one is expensive, Sonnet 5.5’s caution is the safer property, and no other model in this comparison offers it.

Three rows sit outside the eleven-evaluation Intelligence Index composite, marked with an asterisk above. MLCR-AA is worth singling out: Sonnet 5.5 leads the field on long-context medical reasoning, beating every flagship tested and Astra by forty points. On Harvey LAB-AA’s legal criterion pass rate it edges Fable 5.1 by a tenth of a point, with Opus 5.5 third.

The rest of the table is more flattering: Sonnet 5.5 leads all five models on AutomationBench-AA and Terminal-Bench 4.0, and matches Opus 5.5 on both knowledge-work evaluations despite the fifth of the price.

Note what it does not win. That headline 56 against 58 is carried entirely by coding, terminal work and business automation, and by nothing else.

Read the rows as a routing guide. Sonnet 5.5 for agentic coding and business automation, where it leads outright. Opus 5.5 for scientific coding and general reasoning. Fable 5.1 for the highest accuracy and long-context work, and for professional tasks where the answer has to be right. Astra for scientific document reasoning. Sol for high-volume extraction on a tight budget.

The domain indexes: where Sonnet 5.5 actually ranks

Artificial Analysis also publishes composite indexes for each domain, built from the evaluations above. They are the cleanest summary of where each model is strong, because no single benchmark can distort them.

Artificial Analysis index Claude Opus 5.5 Claude Sonnet 5.5 Claude Fable 5.1 GPT-6 Astra GPT-6 Sol
Coding Agent 66 68 62 62 57
Finance & Accounting 61 57 56 55 48
Strategy & Ops 64 60 60 57 52
Legal 63 56 60 59 51
Healthcare & Medical 61 58 58 52 43
Engineering 60 58 56 55 49
Economics 66 61 63 60 54

Two results stand out. Sonnet 5.5 leads the Coding Agent Index outright, ahead of Opus 5.5, and it is the only domain where any model beats the flagship. That comes with a caveat worth weighing: each model was run in its own vendor’s harness, the Claude models in Claude Code and the OpenAI models in Codex, so the ranking reflects the toolchain as much as the model.

The second is Legal, where Sonnet 5.5 finishes fourth of five, behind GPT-6 Astra. It is Sonnet’s only domain ranking below third, and it points the same way as the independent legal results further up. Its single head-to-head win over Fable 5.1 among the domain indexes is Engineering.

Opus 5.5 leads six of the seven, which is the clearest single argument that the flagship still earns its price despite costing twice as much per token.

The security exception, and the safety tax behind it

Claude Sonnet 5.5 finishes 28th of 41 models on CyberBench, at 59.58% or 41.97% depending on how its fallbacks are counted. The field is led by GPT-6 Sol and Anthropic’s own Opus 5.5 CVP, both on 74.58%, with Claude Fable 5.1 at 70.42% and GPT-6 Astra last on 41.07%.

The explanation is in Vals AI’s fallback data, and it is the most useful piece of analysis in this comparison. Sonnet 5.5 refused 2.27% of tasks, falling back to Claude Sonnet 5 server-side. Counting those fallback-assisted tasks as failures changes the numbers sharply:

Benchmark With fallback Fallback counted as failure Change
Vals Index 69.22% 68.89% -0.33 pts
Terminal-Bench 2.1 83.15% 80.52% -2.63 pts
Terminal-Bench 4.0 53.03% 50.51% -2.52 pts
SRE Bench 30.15% 19.08% -11.07 pts
CyberBench v1.1 59.58% 41.97% -17.61 pts

On the composite index the effect is negligible. On security work it is catastrophic. CyberBench drops by nearly 18 points and SRE Bench by 11, because those are exactly the categories where a safety-tuned model declines, and a declined task in a real security workflow is a failure whether or not a weaker model quietly handled it in the background.

So the CyberBench gap is not a capability gap. It is a policy difference, priced in points. If your work involves vulnerability research, reverse engineering or adversarial testing, Sonnet 5.5’s 59.58% is the honest number, and no vendor benchmark table shows it.

There is a reason the gate exists, and it is worth knowing before treating the refusals as excessive caution. On Anthropic’s own ExploitBench evaluation, Sonnet 5.5 achieved full arbitrary code execution in 43.4% of runs. Sonnet 5 never built a working Firefox exploit in the equivalent June evaluation. The model is substantially more capable at offensive security work than its predecessor, which is precisely why it ships with the same three-stage cyber gate as Opus 5.5.

Anthropic’s own top score on this benchmark comes from a configuration most people cannot buy. CVP is Anthropic’s Cyber Verification Program, a restricted mode with security safeguards reduced, available only to approved organizations. It is a valid measurement of what the underlying model can do when allowed to, and Opus 5.5 CVP scores 74.58% against the standard Opus 5.5’s 55.36%.

Set the CVP aside and Anthropic’s strongest generally available score is Fable 5.1 at 70.42%, still well clear of Sonnet 5.5. But the caveat cuts both ways. The CVP variant falls back on 42.86% of tasks once fallbacks are counted as failures, a larger penalty than Sonnet 5.5’s 2.27%, which suggests fallback behaviour is configured per variant rather than set once across Anthropic’s line. Vals AI does not publish which tasks each model declines, so the mechanism behind the gaps is not yet fully explained.

For security work that reshapes the shortlist in a way no composite index hints at. Fable 5.1 is the strongest generally available option on that single benchmark, and also the most expensive per token in the field.

Cost per task at maximum effort

This is the table that should drive a purchasing decision, and its ordering is not monotonic.

Model at maximum effort Intelligence Index Cost per task Output speed Input per 1M Output per 1M
Claude Opus 5.5 58 $5.98 94 t/s $4.00 $20.00
Claude Sonnet 5.5 56 $7.60 139 t/s $2.00 $10.00
Claude Fable 5.1 53 $7.63 68 t/s $10.00 $50.00
GPT-6 Astra 53 $3.26 63 t/s $10.00 $50.00
Claude Opus 5 51 $5.86 60 t/s $5.00 $25.00
GPT-6 Sol 48 $1.06 85 t/s $2.00 $10.00
Claude Sonnet 5 38 $5.09 81 t/s $2.00 $10.00

Sonnet 5.5 is the second most capable model in the field and the second most expensive one to run at maximum effort, which is counterintuitive given it costs half as much per token as Opus 5.5. The reason is verbosity: at roughly 193,000 output tokens per index task it generates more tokens than any other frontier model Artificial Analysis has measured, so the lower unit price does not produce a lower bill.

Capability per dollar is highest in the middle of this table, not at either end. The most capable model is not the cheapest, the cheapest is not the strongest, and the model most people would assume is the value pick on unit price is beaten on cost per task by three cheaper-running models.

That ranking is specific to maximum effort and to the benchmark composite. Both caveats matter: Sonnet 5.5’s per-task cost falls by an order of magnitude at lower effort settings, and its 139 tokens per second is the fastest in the field, which a per-task figure does not capture at all.

Anyone deploying this without pinning the effort level is flying blind, because the default varies by surface and the bill follows whatever it happens to be.

Where Anthropic and OpenAI actually stand

Look past the individual benchmark rows and a clearer picture emerges. Anthropic currently holds the top of the Artificial Analysis Intelligence Index, and it holds three of the top four places. Claude Opus 5.5 leads at 58, Sonnet 5.5 is second at 56, and Fable 5.1 sits joint third at 53 alongside GPT-6 Astra. The best OpenAI model in the field is level with Anthropic’s third-place model and points behind both its leaders.

That is the opposite of the position a year ago, and it is the result of a deliberate pricing strategy rather than a single breakthrough. Anthropic built a four-rung ladder and is selling down it. Fable 5.1 and Mythos 5.1 sit at $10 and $50 for the hardest reasoning. Opus 5.5 dropped to $4 and $20 when it launched. Sonnet 5.5 launched at $2 and $10. And Anthropic has confirmed that Haiku 5.5 will join the family in the coming weeks as the smallest and cheapest model, which on the pattern so far should land well below $2 and $10.

OpenAI’s response has been to cut rather than climb. GPT-6 Astra holds the flagship slot at $10 and $50, but GPT-6 Sol and Luna arrived at $2 and $10 and $0.10 and $0.50 respectively, both at roughly half the price of the GPT-5.6 models they replaced. That is a real answer to Anthropic on price and no answer at all on the index.

Which is why Haiku 5.5 matters more than it looks. OpenAI’s cheapest serious model currently costs $2 and $10. If Anthropic drops a capable model beneath that, the bottom of the market compresses again and the GPT-6 mid-tier loses the clearest advantage it has. A high-volume extraction pipeline priced against Sol today is priced against the wrong model in six weeks’ time.

Speed is the part the cost argument misses

Sonnet 5.5 is expensive per task, and the honest reason is verbosity. At roughly 193,000 output tokens per index task it is the heaviest token consumer Artificial Analysis has measured, which is why a cheaper model on paper produces a larger bill. That much is a genuine drawback.

But verbosity and latency are not the same thing, and Sonnet 5.5 has an advantage the cost tables do not capture. At 139 tokens per second it is the fastest model in this comparison, roughly a third quicker than Opus 5.5, about 55% quicker than GPT-6 Astra, and more than twice the rate of Fable 5.1’s 68.

The ordering tracks model size rather than price. Sonnet 5.5 is a smaller model than Fable 5.1 or Opus 5.5, and smaller models produce tokens faster, so the reasoning that makes it expensive per task is also the reasoning that makes it quick to stream.

That matters for two reasons. For interactive work, an agent that emits at 139 tokens per second finishes a long chain of thought considerably sooner than one running at 63, and the wall-clock experience of a model is not the same thing as its invoice. For high-volume batch work, throughput is what determines wall-clock cost, and Sonnet 5.5’s token rate means a task completes in roughly half the time of the same task on Fable 5.1 at a similar per-task price.

So the accurate summary of Sonnet 5.5 is narrower than “expensive.” It is the most verbose model in the field and the fastest. Whether that is a good trade depends entirely on whether you are paying for tokens or for elapsed time, and nobody’s benchmark reports the second.

What comes next

Two predictions, offered as analysis rather than fact.

Fable 5.5 is likely, and it would probably take the top spot. Fable 5.1 was Anthropic’s most capable model before Opus 5.5 overtook it, and each release in this family has pushed the ceiling higher. A Fable 5.5 would be expected to score above 58 and reclaim the lead, at a higher price per token than Opus 5.5 rather than lower. Anthropic has not confirmed it, and the gap between a prediction and a launch is wide, but the pattern across the 5 and 5.1 generations makes it the most probable next move.

Someone else may take the top spot first. Anthropic held the lead comfortably when Opus 5.5 launched in September, and GPT-6 Astra was already available and still finished behind Fable 5.1. This field has moved at least three index points in a month. Assuming the ceiling holds until Fable 5.5 arrives is a bet, and the recent record suggests it is not a safe one. OpenAI has not published a GPT-6 Terra tier, and if one lands alongside a new flagship the ranking could move again before Anthropic’s next release.

The durable point is that the frontier is no longer ordered by vendor. The top four positions on the index currently hold three Anthropic models and one OpenAI model, at four different price points, and the cheapest model in the field is nowhere near the best. Choosing a model means choosing which kind of work it is for, not which laboratory built it.

What each lab claims about itself

Anthropic’s Sonnet 5.5 announcement is unusually transparent in one respect: it publishes Sonnet 5’s scores alongside, and some of those numbers are unflattering.

Benchmark Claude Sonnet 5.5 Claude Sonnet 5 Claude Opus 5.5 GPT-6 Sol
Terminal-Bench 4.0 70.6% 10.3% 66.4% (xhigh) Not disclosed
FrontierCode 1.1 (main), max 46.2% 42.4% 54.4% 49.3%
FrontierCode 1.1 (main), xhigh 52.1% 42.4% 54.4% 49.3%
CursorBench 4.0 55.5% 34.1% 57.8% Not disclosed
GDPval-AA v2.1 (Elo) 1844 1449 1846 1487
AA-Briefcase v1.1 (Elo) 1811 1359 1822 1483
Humanity’s Last Exam, with tools 64.5% 54.9% 67.7% Not disclosed
OSWorld 2.1 (partial credit) 80.1% 57.0% 81.8% Not disclosed
Chartography (no tools) 61.6% 15.6% 64.4% Not disclosed

Anthropic’s framing is that Sonnet 5.5 essentially ties Opus 5.5 on knowledge work, and the GDPval and AA-Briefcase rows bear that out.

On frontier code generation Sonnet 5.5 actually trails GPT-6 Sol at matched effort, 46.2% against 49.3%, and only passes it at xhigh effort with 52.1%. That is worth pausing on, because it runs against the independent results, where Sonnet 5.5 beats Sol on Terminal-Bench 4.0 by 53.03% to 34.34% and on Code Migration by nearly 13 points. Different harnesses produce different answers on agentic coding, which is the least standardized measurement in the field.

The Terminal-Bench 4.0 jump from 10.3% to 70.6% against Sonnet 5 sounds like an order of magnitude, and it is, but it should be read as repairing a near-total failure rather than a normal increment. Anthropic also notes a standard error of roughly plus or minus 2.6 points on that benchmark, so 70.6% against Opus 5.5’s 66.4% is not a meaningful lead.

Sonnet 5.5 is also the first Sonnet model to complete Pokémon Red working only from screenshots, a detail that says more about long-horizon agentic persistence than any single benchmark row.

Note what Anthropic does not report: no GPT-6 Sol figure for Terminal-Bench 4.0, CursorBench 4.0 or Humanity’s Last Exam. When a vendor omits a competitor’s row from their own table, treat it as a competitive loss.

What OpenAI claims for GPT-6 Sol

OpenAI’s framing for Sol is cost per task rather than raw scores, which is the right frame given what Artificial Analysis then measured.

  • Agents Last Exam: Sol at maximum effort scores 56.4%, which OpenAI says beats the best Opus 5 result in that evaluation at 60% lower cost.
  • AutomationBench: Sol at xhigh effort scores 33.2% at $0.27 per task, above Opus 5 at maximum effort.
  • DeepSWE v1.1: Sol scores 68.8%, close to Claude Fable 5 and Opus 5 levels at 80% to 96% lower cost per task.
  • Hallucination rate: Artificial Analysis measured Sol’s AA-Omniscience hallucination rate falling from 92% to 60%, achieved by answering fewer questions. Accuracy drops from 59% to 54% as a result, which is a trade rather than a pure win.

Independent testing broadly supports the cost framing while adding context. Sol’s AutomationBench-AA rose to 61.6% and Terminal-Bench 4.0 to 43.9%, the largest independent gain in the release. But GDPval-AA fell 101 Elo, and manual review of hundreds of outputs attributed the drop to weaker presentation quality and missing rubric elements rather than a reasoning failure.

Why the two labs disagree

Vals AI and Artificial Analysis are both credible, and they emphasize different things. Understanding why is more useful than picking a winner.

  • Different tasks. Vals AI runs full agentic workflows with file systems, spreadsheets and documents, and scores whether the job got finished. Artificial Analysis runs a composite of scored evaluations, many of which reward reaching a correct answer after extended deliberation.
  • Different safeguards. Vals AI runs Anthropic’s production safeguards and counts refusals. Artificial Analysis does not.
  • GPT-6 Sol’s Vals numbers are provisional. Its finance, legal, tax and public-benefits results look unusually weak, and OpenAI separately reports its own factuality error rate falling by about half. Our earlier coverage flagged those rows as pending retests.
  • Index versions. Artificial Analysis has rebuilt its Intelligence Index several times. A 56 from v4.3.2 is not comparable to a 61 from an older build. Always read the version number, not just the score.

Where both evaluators agree is the useful finding. Both place Sonnet 5.5 far ahead of Sonnet 5, both place it within two points of Opus 5.5 on aggregate capability, and both place GPT-6 Sol in a middle band: well behind the Claude 5.5 models on intelligence, but with a cost profile none of them match at maximum effort.

Which model to choose

It depends on your per-task budget, your latency tolerance, and how much a wrong answer costs you. Read the evaluations above as a routing guide and the reasoning mostly follows from it, but a few things are worth stating plainly.

Sonnet 5.5 is the default for most technical work, specifically agentic coding, business automation and anything latency-sensitive. It leads all five models on AutomationBench-AA and Terminal-Bench 4.0, takes first place on three Vals leaderboards including Code Migration and Vibe Code Bench, and matches Opus 5.5 exactly on both knowledge-work evaluations at a fifth of the per-token price. At 139 tokens per second it is also the fastest model in the field, which is decisive for anything interactive. Anthropic’s own guidance lands in the same place: its model docs position Sonnet 5.5 as “the best combination of speed and intelligence” and describe it as “a faster, lower-cost complement to Opus 5.5.”

Choose Claude Fable 5.1 for legal, tax, medical and financial work where accuracy on attempted answers is the binding constraint. It wins Tax Agent Bench, Legal Research and MedCode, and posts the highest knowledge-accuracy score in the field at 67.23%.

Choose Claude Sonnet 5.5 where a declined answer is acceptable and a wrong one is not. Its 52.99% non-hallucination rate is the highest tested, and on long-context medical reasoning it leads the field outright at 75.0% on MLCR-AA. The trade is explicit: 13 points less accuracy than Fable 5.1, far fewer confident errors.

Opus 5.5 earns its price only when you need the ceiling. It is the only model that beats Sonnet 5.5 on the composite, and it is cheaper to run per task despite costing twice as much per token, because it reasons less. If those two index points do not change your output, Sonnet 5.5 is the better buy.

Stay on Sonnet 5 if your tasks are not frontier-difficulty. It scores 38 against Sonnet 5.5’s 56 at the same per-token price, so the decision is purely whether you can afford the extra tokens the smarter model will use.

One operational habit matters more than the choice. Pin the reasoning effort level on anything. Sonnet 5.5 costs between $0.41 and $7.60 per task depending on the setting, an eighteen-fold spread on a single rate card, and the default is not consistent across surfaces: medium in the Claude apps and Claude Code, high on the Claude Platform. Deploying without specifying it means your bill is decided by whichever endpoint served the request.

Frequently asked questions

Is Claude Sonnet 5.5 better than GPT-6 Sol?

On measured intelligence, clearly. Sonnet 5.5 scores 56 on the Artificial Analysis Intelligence Index at maximum effort against 48 for GPT-6 Sol, winning nine of the eleven evaluations. On the Vals Index it takes second of 66 at 69.22%, against Sol’s ninth at 62.57%. On cost per task, Sol wins by roughly seven times, $1.06 against $7.60. On list price per token they are identical.

Do Claude Sonnet 5.5 and GPT-6 Sol cost the same?

Exactly the same at short context. Both are $2 per million input tokens, $10 per million output tokens, $0.20 per million for cache reads, and $2.50 for a five-minute cache write. Two things differ. Anthropic also offers an hour-long cache at $4.00, while OpenAI charges $2.50 regardless of duration. And on long inputs, Sonnet 5.5 bills its full 1M-token window at standard rates, while GPT-6 Sol applies a 2x input and 1.5x output multiplier to the entire request once it exceeds 272K input tokens.

Is Sonnet 5.5 the best value model in the field?

No, and this is the finding most likely to surprise. At maximum effort Sonnet 5.5 costs $7.60 per task, second-most expensive in the comparison, and it is beaten on cost by GPT-6 Sol at $1.06, by GPT-6 Astra at $3.26 and by Opus 5.5 at $5.98. Astra also matches a higher-index score region at less than half the cost. Sonnet 5.5 is the right choice at lower effort settings, not at maximum.

Why does Sonnet 5.5 score worse than GPT-6 Sol on security tasks?

Mostly because it declines them. Sonnet 5.5 refused 2.27% of tasks in Vals AI testing, and CyberBench falls from 59.58% to 41.97% when those refusals are counted as failures rather than being quietly handled by a fallback to Sonnet 5. Sol reaches 74.58% with no comparable fallback, and Anthropic’s restricted Opus 5.5 CVP variant matches it.

Is Sonnet 5.5 suitable for legal, medical or financial work?

It depends what failure costs you. Its non-hallucination rate of 52.99% is the highest in the field, so where a declined answer is acceptable and a wrong one is expensive, Sonnet 5.5 is the safest model tested. Where the answer itself has to be right, use Fable 5.1, which leads every professional benchmark Sonnet 5.5 loses and is right 13 points more often on knowledge questions.

Does Claude Sonnet 5.5 have usage limits or restrictions?

Alongside Opus 5.5, Anthropic raised five-hour usage limits on Pro, Max, Team and Enterprise plans and issued a rate-limit reset usable until October 22, 2026. The model ships with Fable 5.1-class safeguards on biology and cybersecurity categories, which is the direct cause of the refusal behaviour documented above. Thinking is always on and cannot be disabled on this generation.

Bottom line

Three things will change this picture soon. Anthropic has confirmed that Haiku 5.5 will join the family in the coming weeks as the smallest and cheapest model, which should compress the bottom of the market again. Artificial Analysis has not published the full effort ladder for Sonnet 5.5 or GPT-6 Sol, and the crossover point between them depends on where the middle settings land. And Vals AI has flagged its finance, legal and tax results for GPT-6 Sol as pending retests, which are exactly the rows where its margin is thinnest.

On the top spot, the reasoning is set out in the section above. In short, a Fable 5.5 is the most probable next move from Anthropic, but assuming the ceiling holds until it arrives is not a safe bet given how quickly this field has moved.

The durable point is that the frontier is no longer ordered by vendor. The top four positions currently hold three Anthropic models and one OpenAI model, at four different price points, and the cheapest model in the field is nowhere near the best. Choosing a model means choosing which kind of work it is for, not which laboratory built it. The scarce resource stopped being raw capability some time ago. It is cost per completed task, and the ranking on that measure looks nothing like the ranking on the index.

For the wider market context, our LLM API pricing comparison covers per-token rates across providers, our Claude Opus 5.5 review covers the model ahead of Sonnet 5.5 on the index, and our look at why GLM and MiMo undercut both labs on cost is a useful corrective when even these tiers look steep.

Aaron Jackson
Aaron Jackson
With a decade of hands-on experience in publishing and social media, and a B.Eng in Robotics from UWE, I'm passionate about turning challenges into opportunities. My focus is on creating solutions rather than merely highlighting problems.

Share post:

Popular

Needle-Thin Brain Implant Records, Stimulates and Delivers Drugs

A research team in Denmark and the United Kingdom...

Largest Impact Crater Since 2018 Was Spotted on Google Maps

Sixty-two miles north of the small Quebec community of...

Betavoltaic Battery: The 50-Year Power Source That Still Can’t Charge a Phone

A betavoltaic battery can sit inside a device for...

The EU Will Grade Large Data Centres A to G on Energy and Water

Data centres have long reported their energy use in...