PaPoo
cover

Anthropic Quietly Moving the Goalposts Again

The annoying part here isn’t even the experiment itself. It’s that the model apparently started interpreting the same effort label differently, server-side, with no obvious changelog note. If you’re building against Claude Code and your prompt or tooling suddenly feels “dumber,” that kind of change is exactly the sort of thing that sends you chasing ghosts in your own code.

I’d be more forgiving if this were clearly documented as an experiment and easy to detect. But the whole complaint here is that “high” got compressed until it behaved like the old “low,” and only some sessions and versions were affected. That’s a nasty shape of change because it breaks a basic developer assumption: that a named setting means roughly the same thing from one day to the next.

What I find plausible is that Anthropic is testing some internal control on effort scaling and watching whether users notice or whether task success changes. What I don’t buy is the casualness of shipping that without a visible note, if the reports are accurate. The tweet reads like someone burned an afternoon debugging their own stack for a problem that lived upstream. That’s not just annoying; it’s the kind of thing that makes developers distrust every other odd output they see.

If this is truly server-side and behind an A/B test, then the obvious takeaway is boring but important: when Claude Code behavior changes, don’t assume it came from your app, your prompt, or a local update. Check for model-side experiments first. That’s an uncomfortable place for a product to be, but maybe that’s where we are.


Reference: Xユーザーの🥔🥔🥔(@argofowl)さん

同じ著者の記事