Editorial illustration for Anthropic's Claude Sonnet 5.5 Scores 70.6% on Terminal-Bench 4.0
Claude Sonnet 5.5 Hits 70.6% on Terminal-Bench
Anthropic put out Claude Sonnet 5.5 this week, the second release in its Claude 5.5 lineup after Opus 5.5. Where Opus is built for the hardest problems, Sonnet 5.5 is pitched as the cheaper, quicker option for the stuff people actually do all day: fixing bugs, writing documents, building slides and spreadsheets. It's live now on the Claude Platform under the name claude-sonnet-5-5, with availability on AWS, Google Cloud, and Microsoft Azure. Anthropic hasn't opened the weights, so running it yourself isn't an option.
The model comes with a 1-million-token context window, a 128K max output, and a knowledge cutoff Anthropic calls reliable through June 2026. Adaptive thinking is switched on by default, with five effort settings ranging from low to max. Anthropic says the jump from Sonnet 5 shows up in four places: faster output generation, lower cost per task from needing fewer tokens and tool calls, cleaner writing that early testers preferred, and stronger performance on vision and long-horizon work. That last claim gets tested in an unusual place: Pokémon Red.
Anthropic just released Claude Sonnet 5.5. It is the second model in the Claude 5.5 family, following Claude Opus 5.5. Anthropic positions it as a faster, lower-cost complement to Opus 5.5.
Why this matters
The Terminal-Bench jump from 10.3% to 70.6% is the number worth sitting with. That's not an incremental gain, and it's happening in the cheaper, faster tier of Anthropic's lineup, not the flagship. For developers picking models for CI pipelines, bug triage, or agentic coding tasks, Sonnet 5.5 beating Opus 5.5 Xhigh on Terminal-Bench (66.4%) while trailing only slightly on CursorBench (55.5% vs 57.8%) suggests the price-to-capability gap between Anthropic's two tiers has narrowed to almost nothing for terminal-heavy workflows.
That's a real pricing signal: at the same $2/$10 rate as its predecessor, Sonnet 5.5 makes it harder to justify defaulting to Opus for everyday coding work. Keep in mind these are vendor-reported numbers pulled from Anthropic's own system card, so independent verification on FrontierCode and CursorBench matters before teams rewrite their model-routing logic. Still, for founders running cost-sensitive agent fleets, this is the kind of release that changes which model sits behind your default API calls next quarter, not just a leaderboard footnote.
Common Questions Answered
What is Claude Sonnet 5.5 and how does it differ from Claude Opus 5.5?
Claude Sonnet 5.5 is the second model in Anthropic's Claude 5.5 family, positioned as a faster and lower-cost alternative to Claude Opus 5.5. While Opus 5.5 is built for solving the hardest problems, Sonnet 5.5 is designed for everyday tasks like fixing bugs, writing documents, and building slides and spreadsheets. Both models are available on the Claude Platform, AWS, Google Cloud, and Microsoft Azure.
What is Claude Sonnet 5.5's Terminal-Bench 4.0 score and why does it matter?
Claude Sonnet 5.5 achieved a 70.6% score on Terminal-Bench 4.0, representing a dramatic jump from previous performance metrics. This significant improvement is noteworthy because it occurs in the cheaper, faster tier of Anthropic's lineup rather than the flagship model, demonstrating that the price-to-capability gap between the two tiers has narrowed considerably.
How does Claude Sonnet 5.5 compare to Claude Opus 5.5 on coding benchmarks?
Claude Sonnet 5.5 beats Claude Opus 5.5 Xhigh on Terminal-Bench with a score of 70.6% versus 66.4%, while trailing only slightly on CursorBench at 55.5% compared to Opus 5.5's 57.8%. This performance parity between the two models at different price points makes Sonnet 5.5 particularly attractive for developers working on CI pipelines, bug triage, and agentic coding tasks.
What are the primary use cases for Claude Sonnet 5.5?
Claude Sonnet 5.5 is optimized for everyday development tasks including fixing bugs, writing documents, building slides and spreadsheets, and handling CI pipelines and bug triage. Its faster speed and lower cost compared to Opus 5.5 make it particularly well-suited for agentic coding tasks and other routine development work that developers perform daily.
Where can developers access Claude Sonnet 5.5?
Claude Sonnet 5.5 is available on the Claude Platform under the model name claude-sonnet-5-5, with availability across major cloud providers including AWS, Google Cloud, and Microsoft Azure. This wide availability ensures developers can integrate the model into their existing infrastructure and workflows.
Further Reading
- Anthropic launches Claude Sonnet 5 as a cheaper way to run agents - TechCrunch
- Anthropic releases Opus 5.5 with lower prices and Fable-level performance - TechCrunch
- Introducing Claude Sonnet 5.5 - Anthropic
- Terminal-Bench 4.0 Benchmark Leaderboard - Vals AI
- Claude Sonnet 5.5 Benchmarks Explained: Coding, Speed & Cost - Vellum AI