Scores that catch Opus at half the price
Start with the most striking set of official numbers: on the agentic coding benchmark Terminal-Bench 4.0, Sonnet 5.5 scores 70.6% against Opus 5.5's 66.4%; on GDPval-AA, which spans 44 occupations, it is 1844 to 1846; on the computer-use benchmark OSWorld 2.1, 80.1% to 81.8%. Opus 5.5 currently costs $4 per million input tokens and $20 per million output; Sonnet 5.5 costs $2 and $10. **On these tests, the half-price model matches or beats the flagship.** The gap is wider on FrontierCode 1.1 (46.2% vs 54.4%), where Opus still leads clearly on harder coding tasks. That fits Anthropic's own division of labor: well-scoped work goes to Sonnet, judgment-heavy work to Opus.
"Up to 30% less" depends on which effort level you use
The price hasn't changed; the claimed savings come from "doing the same work with fewer tokens." That claim is qualified: it applies to "most work," and savings are "up to" 30%. Artificial Analysis's independent testing shows the other side. **At max effort, Sonnet 5.5 output 410 million tokens across its Intelligence Index, for an average cost of $7.60 per task, about 50% more than Sonnet 5's $5.09 at max effort.** The score is genuinely higher (56, just 2 points behind Opus 5.5), but it was bought with more tokens. At high effort, the score drops to 47 and the cost per task falls to $1.08. The two sets of numbers don't contradict each other: Anthropic compares everyday tasks, while the third party measured a harder benchmark suite at the top effort setting. **For users, the takeaway is that whether you save money depends mainly on the reasoning effort you choose, not the model's name.**
Cyber safeguards come to Sonnet for the first time
The other change is safety policy. Sonnet 5.5 is the first Sonnet to launch with cyber safeguards similar to those on Opus 5.5: higher-risk cybersecurity tasks visibly fall back to Sonnet 5, and security teams that need full capability can apply to the Cyber Verification Program. Anthropic also mentions new classifiers that prevent reasoning extraction, with thinking bound to the account that created it. That means users doing security research may hit fallbacks or refusals and should go through verification in advance.
What this means for readers
- If you already use Sonnet 5, you can swap the ID and try it at the same price. - If you pay by usage, run your real tasks at high or default effort first and check token usage before turning on max effort. - If you currently use Opus 5.5 for coding agents, it is worth benchmarking Sonnet 5.5 on Terminal-Bench-style tasks; it could halve the cost. Limits of what is known: the official benchmarks are Anthropic's own; Artificial Analysis's cost data reflects its own evaluation suite, not your actual task costs.
via: Anthropic announcement, Artificial Analysis model page, VentureBeat