Source-led article

Claude Sonnet 5.5 nearly matches Opus 5.5 on benchmarks, costs up to 30 percent less

AI News India//4 min read
Claude Sonnet 5.5 interface showing pricing and benchmark scores compared to Opus 5.5
Claude Sonnet 5.5 interface showing pricing and benchmark scores compared to Opus 5.5
Featured image from the source article

Anthropic has released Claude Sonnet 5.5, the second model in its Claude 5.5 family, offering output speeds more than 30 percent faster and per-task costs up to 30 percent lower than the flagship Opus 5.5, while matching it on several key benchmarks. The model is available immediately across major cloud platforms including Amazon Web Services, Google Cloud and Microsoft Azure.

Sonnet 5.5 is designed for well-defined everyday tasks such as fixing bugs, writing documentation, building presentations and creating spreadsheets. Opus 5.5 remains the choice for complex tasks requiring careful judgment. Anthropic also confirmed that Claude Haiku 5.5 will launch in the coming weeks, targeting high-throughput, low-cost use cases.

With Fable, Opus and Sonnet already in the market, Anthropic now has counterparts to OpenAI’s three GPT-6 models — Astra, Sol and Luna. Roughly, Opus sits a bit above Sol, Sonnet above Luna, and Fable above Astra, though Anthropic charges more across the board. Performance differences between matched tiers are small enough that cost may become the deciding factor for many teams.

Performance leap in coding and knowledge work

The performance gap between Sonnet 5.5 and its predecessor is most significant in coding. On Terminal-Bench 4.0, a test for agentic coding, Sonnet 5.5 scores 70.6 percent compared to Sonnet 5’s 10.3 percent. On CursorBench 4.0, which recreates real coding sessions from the Cursor editor, Sonnet 5.5 scores 55.5 percent versus 34.1 percent — just two points below Opus 5.5 at 57.8 percent.

On FrontierCode 1.1 at the “High” setting, Sonnet 5.5 scores ten points above Sonnet 5 at roughly one-fifteenth the cost per task, Anthropic says. Early testers noted how quickly the model grasps a codebase. Sonnet 5.5 also batches tool calls more often than its predecessor, reducing the number of steps needed.

One odd wrinkle appears at the highest reasoning effort level, “Max.” Sonnet 5.5 actually scores worse on FrontierCode at “Max” than at “Xhigh.” Anthropic says that at maximum effort, the model more frequently triggers a code-review function that splits work across multiple sub-agents. In some cases, this led to timeouts or changes outside the task scope, both of which FrontierCode penalises.

On GDPval-AA, an OpenAI-developed knowledge-work benchmark covering tasks from 44 professions and nine industries, Sonnet 5.5 scores 1,844 points — nearly matching Opus 5.5 at 1,846 and sitting about 400 points above Sonnet 5 at 1,449. OpenAI’s GPT-6 Sol lands at 1,487 by Anthropic’s numbers. On Chartography, a visual chart recognition test, Sonnet 5.5 jumps from 15.6 to 61.6 percent.

Pricing, speed and effort settings

Sonnet 5.5 costs the same per million tokens as Sonnet 5: USD 2 for input tokens, USD 10 for output tokens and USD 0.20 for cache reads. Because the model uses fewer tokens per task, effective costs drop by up to 30 percent, Anthropic says. Output generation is also more than 30 percent faster. Independent testing still needs to confirm these claims.

Like Anthropic’s other models and OpenAI’s counterparts, Sonnet 5.5 offers an adjustable effort setting that lets users trade cost and speed against output quality. At low or medium settings, Sonnet 5.5 already beats Sonnet 5’s best scores on several benchmarks at about one-tenth the per-task cost, Anthropic says. Finding the right effort level for each task, however, remains more art than science.

Safety measures for cybersecurity use

Because Sonnet 5.5 is far more capable in cybersecurity than its predecessor, Anthropic is adding safeguards to a Sonnet model for the first time. Requests involving high-risk cybersecurity tasks get visibly rerouted to Sonnet 5. Through an expanded Cyber Verification Program, qualified professionals can apply for tiered access.

Anthropic has also added safety classifiers against distillation attacks, the same kind used in its most powerful models. Whether these measures actually work should become clear over the coming months. If Chinese labs have been benefiting from such attacks, closing that vector could widen the gap with Western labs again.

Datos clave

Metric Claude Sonnet 5.5 Claude Opus 5.5 Claude Sonnet 5
Terminal-Bench 4.0 score 6% – 3%
GDPval-AA score 1,844 1,846 1,449
Input/output cost per million tokens USD 2 / USD 10 – USD 2 / USD 10
Effective cost reduction vs Sonnet 5 Up to 30% – –

What this means for Indian developers and enterprises

For Indian startups, SaaS companies and enterprise teams running AI workloads on AWS, Google Cloud or Azure, Sonnet 5.5 offers a practical middle ground. The model delivers near-flagship quality on knowledge work and coding at significantly lower cost and higher speed. Teams that previously reserved Opus for critical tasks can now shift routine work to Sonnet without sacrificing output quality.

The adjustable effort setting also gives developers flexibility to optimise for cost or accuracy depending on the task. However, the unusual behaviour at the “Max” effort level on coding benchmarks means teams should test thoroughly before deploying at that setting.

Availability

Claude Sonnet 5.5 is available now on Amazon Web Services, Google Cloud and Microsoft Azure. Like Opus 5.5 and Sonnet 5, Anthropic offers the model with zero data retention. Developers can access it through the Claude Platform using the model ID “claude-sonnet-5-5.”

Source: The Decoder – https://the-decoder.com/anthropics-claude-sonnet-5-5-nearly-matches-opus-5-5-on-benchmarks-while-costing-up-to-30-percent-less-per-task/