PaPoo
cover

A cheaper Claude that still looks expensive to run

What jumped out at me is not the benchmark bragging. It’s the familiar Anthropic move of saying, in effect, “we made the model cheaper, but don’t get too comfortable because we’re also being more cautious with who gets what.” That’s the part that feels more real to me than the headline numbers.

A model that lands near Fable 5.1 on the tasks Anthropic cares about, while being 40 percent cheaper than Opus 5, is obviously attractive if you’re building on Claude. Lower price, same general capability, better prose than the older Opus, and higher usage limits for subscribers. That’s a tidy product story. If you’re shipping agents or internal tools, the pricing is probably the first thing you’d actually notice.

But I’m less impressed by the benchmark theater than the article seems to be. Anthropic itself says the margins are getting less reliable at this level, which is basically the company admitting that the scoreboard is starting to blur. Fair enough. Once models are all in the same neighborhood, tiny benchmark gaps can turn into marketing fuel without changing much in practice. I’d care more about whether Opus 5.5 is materially better at the weird, messy prompts I actually throw at Claude Code, not whether it beats some other model by a few points on Terminal-Bench.

The safety side is where this gets interesting, and a little awkward. Anthropic is clearly trying to thread a needle here: release a stronger model, but wrap the riskier stuff in guardrails, reroute some cybersecurity work to an older model, and gate biology access behind verification. That’s sensible if you believe the model is getting close enough to real misuse territory. It also suggests the company is not fully confident about where its own tests end and the real world begins. The line about the model seeming to know it is being evaluated is especially telling. That’s the kind of thing that makes evals feel more brittle than people like to admit.

I do think the new pricing and the safety restrictions belong together. Cheaper models get used more. More usage means more edge cases, more agentic workflows, more opportunities for the model to stumble into something you didn’t intend. So yes, cheaper Opus is good news. But if I were building on it, I’d probably wait to see how the new safeguards behave outside Anthropic’s own lab framing before trusting it for anything sensitive. The cost cut is nice. The real question is whether the model is actually more dependable, or just easier to justify in a budget meeting.


Reference: Anthropic launches Claude Opus 5.5: Benchmarks, pricing, safety

同じ著者の記事