PaPoo
cover

Sonnet keeps getting more serious

What jumped out at me is that Anthropic is now treating a Sonnet model like something that can trip the same safety wires as its biggest models. That’s not the usual “mid-tier model gets a modest bump” story. If Sonnet 5.5 is really close enough to Opus 5.5 on the cyber side that Anthropic ships frontier-style restrictions, then the practical boundary inside the lineup is getting fuzzier in a way that matters to people building agents.

The coding benchmark result is the obvious headline, but I’m more interested in the weird tension inside the announcement. Anthropic says Sonnet 5.5 does not move the frontier overall, yet in the same breath says its cyber capabilities are a large improvement. Those two claims can both be true, but they are not the same kind of reassurance. “Not frontier” sounds like product positioning. “We had to add stronger safeguards” sounds like the model did something genuinely new and potentially awkward.

I also think the reasoning-extraction protection is the most consequential part here, even if it will get less attention than the coding score. If Anthropic is now shipping classifiers that block this kind of leakage on Sonnet, it’s reacting to a very specific and very real failure mode. The article cites researchers recovering secrets from public agent traces, which is exactly the sort of thing that makes “thinking” features feel less like a nice UX and more like attack surface. Preserved reasoning sounds elegant until you remember that traces are data, and data gets scraped, decoded, and abused.

The pricing detail is the one place where I’d keep my hand near the eject button. TNW itself flags that the $2 and $10 numbers may just be the introductory Sonnet 5 price, not the real current comparison. That matters because “half the price of Opus” is a very different claim from “same as an old promo price we were already planning to move off.” Anthropic’s model cards and marketing have gotten so dense that even a clean-looking price line can hide a caveat.

Still, the bigger signal is clear enough. Anthropic wants Sonnet to feel like the practical default for serious coding and agent work, while keeping Opus as the model for deeper, more open-ended judgment. That’s a smart lineup strategy if it holds. It also suggests something the company probably doesn’t want to say too plainly: the safety and capability boundaries are now being managed model by model, task by task, rather than by neat product tiers.


Reference: Anthropic releases Claude Sonnet 5.5 with the cyber limits it reserved for its best models

同じ著者の記事