Anthropic Says Sonnet 5.5 Ties Opus 5.5 on Knowledge Work at Half the Token Price

Anthropic released Claude Sonnet 5.5 on September 28, 2026. The headline figure on its launch page is a 70.6% score on Terminal-Bench 4.0, an agentic coding evaluation where Sonnet 5 scored 10.3%. Opus 5.5, the larger model in the same family, scored 66.4% on that test at its highest effort setting. Anthropic also says Sonnet 5.5 generates output more than 30% faster than Sonnet 5 and costs up to 30% less per task, while per-token prices stay at $2 per million input tokens and $10 per million output tokens.
Sonnet 5.5 is the second model in the Claude 5.5 family. Anthropic describes it as a faster, cheaper companion to Opus 5.5, which the company says remains clearly stronger at complex work that requires sustained judgment. A smaller Haiku 5.5 is scheduled for the coming weeks.
Coding results
Terminal-Bench 4.0 measures multi-step professional tasks completed inside a command-line interface. Anthropic reports that at Medium effort, the default in the Claude apps, Sonnet 5.5 beats Sonnet 5's best score at less than a tenth of the cost per attempt. On CursorBench 4.0, which draws tasks from real Cursor sessions, Sonnet 5.5 scored 55.5% against 34.1% for Sonnet 5 and 57.8% for Opus 5.5.
FrontierCode 1.1 asks whether an agent's code change could be merged without human edits. Sonnet 5.5 scored 52.1% at Xhigh effort, compared with 42.4% for Sonnet 5, 54.4% for Opus 5.5 and 49.3% for OpenAI's GPT-6 Sol. Its score at Max effort was lower, 46.2%. A footnote explains why. At Max, the model more often ran Claude Code's code-review skill, which splits review across many subagents. In two cases examined by Cognition, that led to a timeout or to edits outside the task's scope, which the benchmark penalizes. Anthropic says that at High effort, the default on its developer platform, Sonnet 5.5 matches GPT-6 Sol's best score for about a fifth of the cost per task.
Company testers described the efficiency gains in operational terms. Base44 ran 118 real app builds and reported that Sonnet 5.5 produced apps scoring level with Opus 5, using 3.6 iterations per build where Opus 5 needed 7.7. Unity said the model completed 90% of tasks in its multi-step Editor and coding benchmark, with the majority of its work passing a runtime check. Lovable reported about a third fewer tool calls and roughly half the shell runs on its coding evaluations. CodeRabbit said the model spends far fewer output tokens than Sonnet 5 and no longer reaches for web search as often, and it plans to move simple and moderate code reviews over first.
Knowledge work and design
On GDPval-AA, which tests real-world tasks across 44 occupations, Sonnet 5.5 scored 1844 Elo against 1846 for Opus 5.5 and 1449 for Sonnet 5. GPT-6 Sol scored 1487. On AA-Briefcase, a new long-horizon knowledge-work benchmark, the figures were 1811, 1822 and 1359 for the three Claude models, with 1483 for GPT-6 Sol. Sonnet 5.5 also came within a few points of Opus 5.5 on OSWorld 2.1 computer use (80.1% against 81.8%) and on the Chartography chart-reading test (61.6% against 64.4%). Sonnet 5 scored 15.6% on Chartography.
Anthropic ran one internal test that shows what these numbers mean in practice. It gave the model a public company's quarterly earnings materials, call transcripts and a slide template, and asked for a 10-slide operating review. Two experts judged the first draft ready to send as is. Early testers also pointed to the model's handling of user interface polish and its ability to follow slide templates with minimal editing.
Several enterprise testers focused on cost per answer. Balyasny Asset Management ran 2,441 finance tasks and reported that Sonnet 5.5 scored ahead of Sonnet 5 while using about 121,000 tokens per answer, where Sonnet 5 used 497,000. Box reported that the model rechecks data in source documents and catches errors Sonnet 5 missed, and measured it as 2.4 times faster with 12% fewer tokens. Zendesk processed tickets 20% faster in its tests, and Slack reported better results on almost all of its offline evaluations with about 14% fewer output tokens.
Price, speed and safeguards
Opus 5.5 is priced at $4 per million input tokens and $20 per million output tokens, twice the Sonnet 5.5 rate. Cache writes cost $2.50 per million tokens on Sonnet 5.5 and $5 on Opus 5.5, while cache reads are $0.20 on both. Anthropic says the model performs best next to Opus 5.5 at lower effort settings, where it costs less per task, and that at higher settings the two can perform comparably at similar cost. The company also notes that on some benchmarks, Sonnet 5.5 at Low or Medium effort beats Sonnet 5's best score for about a tenth of the cost.
Sonnet 5.5 is the first Sonnet model to launch with cyber safeguards similar to those on Opus 5.5. Routine bug finding and fixing is unaffected, but higher-risk cybersecurity tasks visibly fall back to Sonnet 5. Anthropic says cyberdefenders will be able to apply to an expanded Cyber Verification Program for tiered access to more advanced capabilities. The biology safeguards match those on Sonnet 5, and the company acknowledges that some microbiology and virology requests may be flagged in error.
The model is also the first Sonnet to ship with classifiers meant to block reasoning extraction, which Anthropic describes as a defense against distillation attacks using thousands of fake accounts. Alongside this, the company expanded preserved thinking so that a model's reasoning cannot be separated from the account that produced it. Developers who move conversations between accounts, including by switching accounts mid-session in Claude Code, should read the documentation on the change. Developers who run Sonnet with thinking turned off need to switch to a new setting called between_tools before upgrading.
On alignment, Anthropic ran its automated behavioral audit across roughly 1,850 scenarios. It reports that Sonnet 5.5 improves on or matches Sonnet 5 on most measures of alignment, misuse resistance and honesty. On containment evaluations it comes close to Opus 5.5, and it was the least likely of any Anthropic model to probe the limits of its containers. Opus 5.5 still performs slightly better overall. Anthropic states that no set of evaluations reliably catches every failure and that Sonnet 5.5 may have tendencies the company has not found.
Questions the launch page leaves open
The benchmark table has several caveats attached. Artificial Analysis ran GDPval-AA and AA-Briefcase on a pre-release deployment that had a bug affecting structured outputs, which Anthropic says has since been fixed and likely understates Sonnet 5.5's scores. The GPT-6 Sol figures carry their own footnote, since OpenAI recently fixed an image-understanding bug and third-party scores may not yet reflect it. For Terminal-Bench 4.0 and CursorBench 4.0, GPT-6 Sol results were not publicly reported, so Anthropic's charts use GPT-5.6 Sol instead.
The size of the Terminal-Bench gap also deserves a careful read. A move from 10.3% to 70.6% is large enough that the baseline matters, and the launch page does not explain what held Sonnet 5 back on that test. Opus 5.5's 66.4% is reported at Xhigh effort, so the two Claude models are compared at different settings. The testimonials come from companies with commercial relationships with Anthropic, and their task suites are private.
The practical step for teams is to run their own workloads at Low and Medium effort before adopting higher settings, since Anthropic's own data shows the cost curves diverge sharply across effort levels. Teams that depend on thinking-off configurations or cross-account sessions should read the migration guide first.

