Claude Sonnet 5.5 Launches: Anthropic Says It Costs Up to 30% Less for Most Work, While a Third-Party Test at Max Effort Found About 50% Higher Cost per Task

Anthropic released Claude Sonnet 5.5 on September 28, the second model in the Claude 5.5 family after Opus 5.5 on September 22, positioned as a faster, lower-cost complement to Opus 5.5: Opus 5.5 is built for complex work requiring careful judgment, while Sonnet 5.5 is strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides and spreadsheets. Anthropic says it runs 30%+ faster than Sonnet 5, typically needs far fewer tokens to do the same work, and costs up to 30% less for most work. Pricing matches Sonnet 5: $2 per million input tokens, $10 per million output tokens, $0.20 for cache reads and $2.50 for cache writes. The model ID is claude-sonnet-5-5, available through the Claude Platform, AWS, Google Cloud and Microsoft Azure. In Anthropic's benchmarks it scores 70.6% on Terminal-Bench 4.0 (Sonnet 5: 10.3%, Opus 5.5: 66.4%), 80.1% on OSWorld 2.1 (Opus 5.5: 81.8%) and 1844 on GDPval-AA v2.1 (Opus 5.5: 1846). It is the first Sonnet model to launch with cyber safeguards similar to Opus 5.5's; higher-risk cybersecurity tasks visibly fall back to Sonnet 5, and expanded capabilities are available through a Cyber Verification Program. Independent evaluator Artificial Analysis found that at max effort it scores 56 on its Intelligence Index but output 410 million tokens across the index, for an average cost of $7.60 per task, about 50% more than Sonnet 5's $5.09 at max effort; at high effort it scores 47 at $1.08 per task. Haiku 5.5 will join the family in the coming weeks.

Scores that catch Opus at half the price

Start with the most striking set of official numbers: on the agentic coding benchmark Terminal-Bench 4.0, Sonnet 5.5 scores 70.6% against Opus 5.5's 66.4%; on GDPval-AA, which spans 44 occupations, it is 1844 to 1846; on the computer-use benchmark OSWorld 2.1, 80.1% to 81.8%. Opus 5.5 currently costs $4 per million input tokens and $20 per million output; Sonnet 5.5 costs $2 and $10. **On these tests, the half-price model matches or beats the flagship.** The gap is wider on FrontierCode 1.1 (46.2% vs 54.4%), where Opus still leads clearly on harder coding tasks. That fits Anthropic's own division of labor: well-scoped work goes to Sonnet, judgment-heavy work to Opus.

"Up to 30% less" depends on which effort level you use

The price hasn't changed; the claimed savings come from "doing the same work with fewer tokens." That claim is qualified: it applies to "most work," and savings are "up to" 30%. Artificial Analysis's independent testing shows the other side. **At max effort, Sonnet 5.5 output 410 million tokens across its Intelligence Index, for an average cost of $7.60 per task, about 50% more than Sonnet 5's $5.09 at max effort.** The score is genuinely higher (56, just 2 points behind Opus 5.5), but it was bought with more tokens. At high effort, the score drops to 47 and the cost per task falls to $1.08. The two sets of numbers don't contradict each other: Anthropic compares everyday tasks, while the third party measured a harder benchmark suite at the top effort setting. **For users, the takeaway is that whether you save money depends mainly on the reasoning effort you choose, not the model's name.**

Cyber safeguards come to Sonnet for the first time

The other change is safety policy. Sonnet 5.5 is the first Sonnet to launch with cyber safeguards similar to those on Opus 5.5: higher-risk cybersecurity tasks visibly fall back to Sonnet 5, and security teams that need full capability can apply to the Cyber Verification Program. Anthropic also mentions new classifiers that prevent reasoning extraction, with thinking bound to the account that created it. That means users doing security research may hit fallbacks or refusals and should go through verification in advance.

What this means for readers

- If you already use Sonnet 5, you can swap the ID and try it at the same price. - If you pay by usage, run your real tasks at high or default effort first and check token usage before turning on max effort. - If you currently use Opus 5.5 for coding agents, it is worth benchmarking Sonnet 5.5 on Terminal-Bench-style tasks; it could halve the cost. Limits of what is known: the official benchmarks are Anthropic's own; Artificial Analysis's cost data reflects its own evaluation suite, not your actual task costs.

via: Anthropic announcement, Artificial Analysis model page, VentureBeat