AI

Anthropic Unveils Claude Sonnet 5.5, Boosting Speed and Cutting Work Costs

TECHWORLD ·

Claude Sonnet 5.5. [Photo: Anthropic]

✦ AI Summary

Anthropic unveiled Claude Sonnet 5.5 on the 28th (local time).

Claude Sonnet 5.5 is a complementary model for everyday work, coding, and document tasks, with output generation speed more than 30% faster than Sonnet 5 and pricing unchanged.

It scored 70.6% on Terminal-Bench 4.0, and Anthropic said Opus 5.5 retains the edge in complex, open-ended tasks that require long periods of judgment.

Anthropic unveiled Claude Sonnet 5.5 on the 28th (local time). Claude Sonnet 5.5 is the second model in the Claude 5.5 lineup. Anthropic had previously unveiled Claude Opus 5.5, which focuses on complex long-term tasks, advanced coding, and agent work.

Claude Sonnet 5.5 has been positioned as a complementary model for everyday work, coding, and document tasks. It is characterized by improved speed, cost efficiency, and agentic coding performance. Anthropic said Claude Sonnet 5.5's output generation speed is more than 30% faster than Sonnet 5, its predecessor.

Pricing remains unchanged. It costs USD 2 per 1 million input tokens and USD 10 per 1 million output tokens. Anthropic said the number of tokens required for the same task has decreased, cutting per-task costs by up to 30%.

Anthropic outlined role distinctions by model. Opus 5.5 is focused on complex open-ended tasks that require sustained judgment, while Sonnet 5.5 is aimed at bug fixes and producing documents, slides, and spreadsheets, as well as work with relatively clear scope. Claude Haiku 5.5 is specialized for high-volume processing and cost efficiency, and Claude Haiku 5.5 is expected to be added in the coming weeks.

The most notable performance gains came in agentic coding. In Anthropic's own evaluations, Sonnet 5.5 scored 70.6% on Terminal-Bench 4.0, while Sonnet 5 scored 10.3% on Terminal-Bench 4.0. As a result, Sonnet 5.5's Terminal-Bench 4.0 result was presented as a sharp increase from Sonnet 5. Sonnet 5.5 also scored 55.5% on CursorBench 4.0, while Opus 5.5 scored 57.8% on CursorBench 4.0, indicating that Sonnet 5.5's CursorBench 4.0 performance was close to Opus 5.5.

Anthropic attributed this to improved codebase understanding, more efficient tool calling, and better performance on multi-step development tasks. In other words, the gains in Sonnet 5.5's metrics and its benchmark results near Opus 5.5 were presented as an example of improved codebase understanding and tool-use efficiency translating into better development task performance.

The performance evaluation across several benchmarks shows a narrowing gap with top models under knowledge work evaluation standards. On GDPval-AA v2.1, it scored 1844 points, while Opus 5.5 scored 1846 points on GDPval-AA v2.1, indicating a similar level to Opus 5.5 in that result. It also scored 64.5% on Humanity's Last Exam under tool-use criteria and 80.1% on OSWorld 2.1, an evaluation of computer-use capabilities.

Anthropic said there are differences by item. Anthropic said Sonnet 5.5 comes close to Opus 5.5 in some evaluations. On the other hand, Anthropic said Opus 5.5 retains the edge in complex, open-ended tasks that require long periods of judgment.

It also supports an Effort setting for adjusting reasoning resources. The Effort feature is a setting that controls reasoning resources, with the default set to Medium for Claude Code and general apps, and High by default on the Claude platform. Lowering the setting reduces response speed and token usage, while raising it increases the amount of reasoning performed.

Anthropic checked the safety of Claude Sonnet 5.5 through an automated behavioral audit based on about 1,850 scenarios. Based on the evaluation results, Anthropic said it was equal to or better than Sonnet 5 on most metrics, including alignment, misuse resistance, and honesty. It also said it applied a protection system similar to the Opus family to reflect stronger cyber security capabilities.

Claude Sonnet 5.5 is available on the Claude platform, AWS, Google Cloud, and Microsoft Azure. The model name used on the Claude platform is 'claude-sonnet-5-5'.

Source: TECHWORLD · Kim Seung-gi
Original: https://www.epnc.co.kr/news/articleView.html?idxno=407495

References

This article was produced with the help of an automated content generation algorithm.


Source: TECHWORLD

View original

This article was summarized and organized by BizCrush based on the original article from TECHWORLD. For exact quotations and full details, please refer to the original article.