AIMarketsNewsTech

Anthropic ships Claude Sonnet 5.5 with frontier-style cyber safeguards

Anthropic on Monday released Claude Sonnet 5.5, the second model in its Claude 5.5 family, and said the mid-tier Claude now runs more than 30% faster and costs up to 30% less for most work than Sonnet 5 while launching with cyber safeguards previously reserved for its most capable models, the company wrote in a product post.

Sonnet 5.5 keeps the same sticker price as Sonnet 5: $2 per million input tokens, $10 per million output tokens, and $0.20 per million for cache reads. Anthropic pitches it for everyday coding, bug fixes, and polished documents, slides, and spreadsheets, as a faster complement to Claude Opus 5.5 on well-scoped work.

On Terminal-Bench 4.0, an agentic coding evaluation that tests multi-step work in a command-line setup, Sonnet 5.5 scores 70.6% versus Sonnet 5's 10.3%. TechCrunch notes the model sometimes beats Opus 5.5 on Anthropic's own agentic coding benchmarks because it can spawn multiple agents without blowing cost limits. The company also says Sonnet 5.5 sits two points below Opus 5.5 on GDPval-AA, a test of real-world work across occupations, and is the first Sonnet to beat Pokémon Red working only from screenshots.

The safety line is the real product story. Because cybersecurity capabilities are "comparable to Opus 5's," Anthropic calls Sonnet 5.5 the first Sonnet to launch with cyber safeguards and fallbacks like those on its top models.

Cyber safeguards here mean extra blocks and fallbacks when the model is asked to do high-risk hacking. Higher-risk cybersecurity tasks "will visibly fall back to Sonnet 5." Biology safeguards stay the same as Sonnet 5's.

TechCrunch adds that the same cyber rules that apply to Fable and Opus now cover this Sonnet.

Sonnet 5.5 is also the first Sonnet with safety classifiers meant to stop distillation attacks, where outsiders use thousands of fake accounts to copy a model's skills by scraping its hidden step-by-step answers. The company expands "preserved thinking" so Claude's thinking cannot be decoupled from the account that created it. The Next Web cites August research that decoded 315,320 thinking blocks from 6,708 public agent traces across OpenAI, Anthropic, and Google systems and recovered 62 API keys, 33 passwords, and seven private keys.

Anthropic narrowed its alignment assessment because it says Sonnet 5.5 "doesn't advance the frontier of our models' capabilities." An automated behavioral audit across roughly 1,850 scenarios improves on or matches Sonnet 5 on most measures. Claude Haiku 5.5 is due to join the 5.5 family in the coming weeks, without a firm date.

Morning's Nvidia containment stack was an engineering answer for agents that leave the sandbox. Today's ship is Anthropic admitting the everyday coder model now needs the same cyber tripwires as the expensive frontier shelf.

Sonnet 5.5 is priced like Sonnet 5 and guarded like Opus.

Sources

Leave a Reply

Your email address will not be published. Required fields are marked *