Anthropic shipped Claude Sonnet 5.5 on September 28. The headline most coverage led with was the benchmark table. The number that actually matters if you run AI inside a business is quieter: the same work now costs less, because the model uses fewer tokens to finish it — not because the price per token went down.
The price per token did not go down. Sonnet 5.5 costs exactly what Sonnet 5 cost: $2 per million input tokens, $10 per million output. Cache reads stay at $0.20, cache writes at $2.50. Nothing on the invoice line changed.
What changed is how much of it a task consumes.
The efficiency story, in one chart
Balyasny Asset Management ran 2,441 finance tasks through both models before launch. Same tasks, same prompts. The difference was in how much thinking each model spent getting to an answer.
View the data
| Model | Tokens per answer |
|---|---|
| Sonnet 5 | 497,000 |
| Sonnet 5.5 | 121,000 |
Balyasny is the extreme case. The rest of the early testers landed in a narrower, more believable band, and they are worth reading as a range rather than a promise:
- Slack measured roughly 14% fewer output tokens with no prompt changes at all — a straight model swap.
- Box reported 2.4× faster responses with 12% fewer tokens.
- Base44 needed 3.6 iterations per app build, against 7.7 for Opus 5, across 118 real builds.
- Zendesk closed support tickets about 20% faster.
- Atlassian saw its Rovo agents run up to 30% faster.
Note what those five have in common. None of them is a benchmark. They are all the same shape of work: high volume, well-scoped, repeated thousands of times a day. That is exactly the shape of work sitting inside a brick-and-mortar operation — quoting, routing, following up, summarising, classifying.
Where it lands against Opus
The benchmark picture is genuinely strong, and in one place surprising: on Terminal-Bench 4.0, Sonnet 5.5 beats Opus 5.5. Not by a rounding error, and from a Sonnet 5 baseline of 10.3%.
- Sonnet 5
- Sonnet 5.5
- Opus 5.5
View the data
| Benchmark | Sonnet 5 | Sonnet 5.5 | Opus 5.5 |
|---|---|---|---|
| Terminal-Bench 4.0 | 10.3% | 70.6% | 66.4% |
| FrontierCode 1.1 (Xhigh) | 42.4% | 52.1% | 54.4% |
| CursorBench 4.0 | 34.1% | 55.5% | 57.8% |
| GDPval-AA v2.1 (Elo) | 1449 | 1844 | 1846 |
| AA-Briefcase v1.1 (Elo) | 1359 | 1811 | 1822 |
Read the last two rows of that table again. On GDPval-AA, Sonnet 5.5 scores 1844 against Opus 5.5's 1846. On AA-Briefcase, 1811 against 1822. Those are rounding errors between a mid-tier model and a flagship.
Anthropic is unusually direct about the limit, though, and it is worth repeating rather than burying: Opus 5.5 is still clearly stronger on complex, open-ended work that needs sustained judgment over a long horizon. Sonnet 5.5 closed the gap on bounded tasks. It did not close the gap on ambiguous ones.
Where to actually plug it in
Model routing is the whole game, and most teams get it wrong in the same direction: they pick one model and send everything to it. That means either overpaying for simple work or underperforming on hard work. Here is the split the launch data supports.
CodeRabbit put this into practice publicly: they moved their simple and moderate code reviews over to Sonnet 5.5 and kept the harder ones where they were. That is the pattern. Not a migration — a reallocation.
What to do this week
If you already have AI running in production, the migration is genuinely cheap. Sonnet 5.5 is a drop-in swap at the same price, the model ID is claude-sonnet-5-5, and it is available on AWS Bedrock, Google Cloud Vertex and Azure from day one. Slack got their 14% by changing the model string and nothing else.
Three things worth doing before you swap anything:
- Measure your current token spend per task type, not per month. A monthly total tells you nothing about which workflow is expensive. You cannot see a 14% improvement if you were never tracking the baseline.
- Swap one workflow, not all of them. Pick the highest-volume, most-repetitive one. That is where an efficiency gain compounds and where a regression is easiest to spot.
- Leave your judgment-heavy work on Opus. The temptation after a launch like this is to move everything down a tier to save money. The benchmarks say that is a mistake on open-ended tasks, and Anthropic says so too.
One honest caveat on the numbers in this piece: every early-tester figure here is vendor-reported, measured by the customer and published by Anthropic. They are specific and attributed, which is better than most launch claims, but they are not independent. Your own workload is the only benchmark that predicts your own bill.
- Anthropic launch materials and published benchmark scores, September 28, 2026.
- Customer-measured figures reported by Anthropic: Balyasny Asset Management, Slack, Box, Base44, Zendesk, Atlassian, CodeRabbit.
- Launch coverage: TechCrunch, VentureBeat, SiliconANGLE, The New Stack, unite.ai.
