Skip to content
← All news
4 min read

Anthropic released Sonnet 5.5, six days after Opus 5.5

Sonnet 5.5 runs 30% faster and up to 30% cheaper per task than Sonnet 5, at the same list price, and beats Opus 5.5 on one agentic coding benchmark.

Anthropic's second model release in a week scores higher than Opus 5.5 on one benchmark, at Sonnet pricing.

Anthropic released Claude Sonnet 5.5 on September 28, its second model release in six days. It runs more than 30 percent faster than Sonnet 5 and cuts effective cost per task by up to 30 percent, at the same $2 per million input tokens and $10 per million output tokens Sonnet 5 already charged.

One benchmark it wins outright

Sonnet 5.5 scores 70.6 percent on Terminal-Bench 4.0, an agentic coding evaluation, against Sonnet 5's 10.3 percent and Opus 5.5's own 66.4 percent. That is not a typo: the cheaper model beats the more expensive one on this specific test. On GDPval-AA v2.1, a professional knowledge-work benchmark, Sonnet 5.5 scores 1844 Elo, two points behind Opus 5.5 and roughly 360 points ahead of OpenAI's GPT-6 Sol.

Where Anthropic says to use which model

Anthropic positions Opus 5.5 for complex work that needs careful judgment and Sonnet 5.5 for well-scoped everyday tasks: fixing a bug, drafting a document, building a spreadsheet. The Terminal-Bench win complicates that pitch a little, since agentic coding is exactly the kind of task Opus is supposed to be for.

What the benchmark does not cover

Terminal-Bench 4.0 measures one class of agentic coding task, not general reasoning or judgment under ambiguity, so a model that wins it can still lose on the kind of open-ended work Opus is built for. Anthropic has not published where that line actually falls for a given task, only the two-tier pricing that assumes one exists.

Why a build studio cares

We pick a model per task, not per project, and a model that beats a pricier sibling on agentic coding at a third of the reasoning-tier price changes which tasks get routed where. The open question for us is exactly the one Anthropic left unanswered: which of our own agent workflows are Terminal-Bench-shaped enough to move down a tier without a quality drop, and which ones are actually judgment calls wearing a coding task's clothes.

Next step: read Anthropic's own announcement for the full benchmark suite and pricing.

AIAnthropicDev ToolsClaudeSonnet55AnthropicClaudeCodeAIModelsAgenticCodingLLMTerminalBenchAIBenchmarksOpusAndSonnetAIWorkflows

Ready to know?

Send what you want checked or built. Fixed scope, price, and date in writing inside 24 hours, or the website or audit fee on your first project is refunded in full.

24 clock hours. Weekends included.
Book a call