Anthropic has introduced Claude Sonnet 5.5 as a faster, more efficient successor to Sonnet 5. In a post on X, @claudeai says the model generates output more than 30% faster and can cost up to 30% less per task, even though its listed token prices are unchanged. Anthropic positions it between routine, well-scoped work and the more demanding open-ended tasks intended for Claude Opus 5.5.
The reported improvements cover agentic coding—where a model uses tools and takes multiple steps to complete a task—as well as knowledge work, computer use, visual understanding, and document creation. Anthropic's announcement refers to external testers, but the benchmark, speed, cost-per-task, and tester comparisons below remain results reported by Anthropic rather than independent measurements.
What Sonnet 5.5 changes from Sonnet 5
Anthropic describes Sonnet 5.5 as a complement to Opus 5.5 rather than a universal replacement for it. Sonnet 5.5 is designed for well-scoped everyday tasks, fixing bugs, fast iteration, and producing polished documents, slides, and spreadsheets. Opus 5.5 remains stronger in Anthropic’s testing and that of external testers for complex, open-ended work that requires sustained judgment.
The model’s clearest reported change is efficiency. Anthropic says Sonnet 5.5 typically uses fewer tokens to complete the same work as Sonnet 5. Since token usage affects API costs and can also affect response time, that efficiency is separate from the models’ identical list prices. The company also says Sonnet 5.5 batches tool calls more often in coding tasks, which can reduce the number of steps needed.
Anthropic’s headline Terminal-Bench 4.0 result illustrates the size of the reported coding improvement: Sonnet 5.5 scored 70.6%, compared with 10.3% for Sonnet 5. Terminal-Bench measures complex, multi-step professional tasks completed through a command-line interface. Anthropic reports that Sonnet 5.5’s score was also above the 66.4% result shown for Opus 5.5, although the Opus result used Xhigh effort, its highest reported setting.

Image credit: @claudeai on X
That result should not be read as a general ranking across every task. Anthropic’s announcement says benchmark scores measure only one facet of capability. Its broader table reports Sonnet 5.5 ahead of Sonnet 5 on CursorBench 4.0, GDPval-AA, AA-Briefcase, Humanity’s Last Exam with tools, OSWorld, and Chartography. On several of those evaluations, Opus 5.5 still scores higher.
The benchmark footnotes also matter. Artificial Analysis ran GDPval-AA and AA-Briefcase on a pre-release Sonnet 5.5 deployment that had a structured-output bug. Anthropic says the bug could have degraded responses and has since been fixed, but it does not provide independent confirmation of how much the scores changed. The GPT-6 Sol comparisons also carry version and availability caveats in Anthropic’s table.
Coding and knowledge-work performance
Coding is the area where Anthropic reports the most dramatic difference from Sonnet 5. On FrontierCode 1.1, Sonnet 5.5 scored 46.2% at Max effort and 52.1% at Xhigh effort, compared with 42.4% for Sonnet 5. On CursorBench 4.0, which uses tasks from real Cursor coding sessions, Sonnet 5.5 reached 55.5%, compared with 34.1% for Sonnet 5 and 57.8% for Opus 5.5.

Image credit: @claudeai on X
Those figures also show why the effort setting matters. Sonnet 5.5’s FrontierCode score was lower at Max than at Xhigh. Anthropic attributes the result partly to more frequent use of a code-review skill that split work among subagents; in two examined cases, this led to a timeout or to extra edits beyond the task’s scope. FrontierCode penalizes out-of-scope changes even when they are otherwise useful.

Image credit: @claudeai on X
Anthropic reports similar gains in knowledge work. On GDPval-AA v2.1, which evaluates real-world tasks across 44 occupations and nine industries, Sonnet 5.5 scored 1844 versus 1449 for Sonnet 5 and 1846 for Opus 5.5. On AA-Briefcase v1.1, Sonnet 5.5 scored 1811, compared with 1359 for Sonnet 5 and 1822 for Opus 5.5.
The company also highlights improvements that are harder to reduce to a single score. It says early testers found Sonnet 5.5 more natural for collaboration, better at understanding codebases quickly, and capable of adding polish to interfaces and slide decks. One Anthropic internal test involved creating a 10-slide operating review from earnings materials, call transcripts, and a slide template; two experts judged the first draft ready to send. That is an Anthropic-reported internal test, not independent evidence of typical output quality.
Effort settings change the cost-and-capability trade-off
Effort settings control how long the model reasons and checks its work. Lower settings generally produce faster responses with fewer tokens, while higher settings give the model more time to work through a task. Anthropic says Claude Code and its apps use Medium effort by default, while the Claude Platform defaults to High.
This means the same model can have different cost, speed, and benchmark results depending on how it is configured. Anthropic says Sonnet 5.5 at Low or Medium effort beats Sonnet 5’s best score on several evaluations for about one-tenth of the cost per task. The exact savings vary by benchmark and workload, so that statement should not be treated as a universal per-request price.

Image credit: @claudeai on X
The practical choice is therefore not simply Sonnet 5.5 versus Sonnet 5. Developers need to decide how much effort a task warrants. Low or Medium may suit routine bug fixes, bounded transformations, and quick drafting. Higher settings may be more appropriate when the model must inspect a larger codebase, use tools over many steps, or check a complicated answer. Anthropic’s results suggest that Sonnet 5.5 can approach Opus 5.5 on some tasks at higher settings, but the company still describes Opus as the stronger option for sustained, open-ended judgment.

Image credit: @claudeai on X
Pricing, speed, and availability
Sonnet 5.5 keeps Sonnet 5’s listed token prices: $2 per million input tokens, $10 per million output tokens, and $0.20 per million tokens for cache reads. Anthropic lists cache writes at $2.50 per million tokens for Sonnet 5.5 and $5 for Opus 5.5. Sonnet 5.5’s lower reported cost per task comes from using fewer tokens and taking fewer steps, not from a lower input or output rate.
Anthropic says Sonnet 5.5 generates output more than 30% faster than Sonnet 5, making it the fastest Sonnet model in the company’s comparison. The announcement does not provide a universal latency figure, so the actual difference will depend on the task, effort setting, tool calls, and platform.
The model is available through the Claude Platform, Claude products, Amazon Web Services, Google Cloud, and Microsoft Azure, according to Anthropic. The announcement does not specify regional availability, quotas, plan restrictions, or the exact rollout conditions for each third-party platform.
For developers migrating from a setup that used Sonnet with thinking disabled, Anthropic says Sonnet 5.5 requires the new between_tools setting. That setting keeps up-front thinking disabled, and Anthropic directs users to its migration guide for the details. The announcement does not provide a complete compatibility matrix or migration procedure.
Safety changes and their limits
Anthropic says Sonnet 5.5 matches or improves on Sonnet 5 across most measures in its automated behavioral audit. The company also says the model has cybersecurity capabilities comparable to Opus 5.5, so it is the first Sonnet model to launch with cyber safeguards and fallbacks like those used for Anthropic’s most capable models.
Anthropic describes those cyber safeguards as targeting a narrow set of higher-risk requests. Routine software development, including finding and fixing ordinary bugs, is intended to remain available. Higher-risk cybersecurity tasks may fall back to Sonnet 5, according to the announcement.
The model uses the same biology safeguards as Sonnet 5. Anthropic says most research, education, and clinical work is unaffected, although some microbiology and virology requests may be flagged incorrectly. Its separate Life Sciences Verification Program has its own eligibility, monitoring, retention, and access rules; it should not be treated as a standard Sonnet 5.5 product feature or as evidence that all biology-related requests are unrestricted.
Anthropic also says that no set of evaluations reliably catches every failure, so its reported alignment results are not a guarantee of behavior in every deployment or use case.
Who should consider Sonnet 5.5?
Sonnet 5.5 looks most relevant for teams that need faster iteration on bounded coding and knowledge-work tasks, especially when token usage and response speed matter. It may be a practical upgrade for Sonnet 5 users whose workloads involve bug fixing, codebase understanding, document generation, slide creation, spreadsheet work, or repeated tool-assisted tasks.
The choice is less obvious for highly open-ended work. Anthropic’s own announcement says Opus 5.5 remains clearly stronger for complex tasks requiring sustained judgment. Developers should also test their own prompts, tools, effort settings, output checks, and safety workflows before switching production traffic. The reported benchmark and cost-per-task results are useful signals, but they do not establish that Sonnet 5.5 will be better or cheaper for every application.
For many teams, the most meaningful change may be the combination of unchanged token pricing, lower reported token use, and faster output. Whether that makes Sonnet 5.5 the better choice depends on the task’s complexity, the acceptable error rate, and how much additional reasoning the application needs.
Read Anthropic’s announcement and benchmark notes.





0 comments
No approved comments yet. You can start the conversation.
Leave a comment