PaPoo
cover

The model upgrade that looked better on paper than in my terminal

What jumps out to me is not that Claude Fable 5.1 “won.” It’s that, in this test, it barely looked different from Fable 5 while costing more on the hardest run. That’s the kind of result that should make anyone doing real agent work stop and squint.

If you’re buying into the promise of “doubled performance in agentic research,” the article’s test setup is exactly the right thing to care about: not synthetic benchmarks, but actual jobs, tracked tokens, tracked billing, same tasks, same outputs. And the uncomfortable part is that both models apparently got everything right. Perfect scores sound good until you realize they erase the very signal you hoped the upgrade would reveal. If two models both finish the assignment cleanly, the question becomes less “which is smarter?” and more “which one is cheaper, faster, or less annoying to operate?” On that front, the reported result is not flattering to the newer model.

I also think this is a useful reminder that “agentic” gains are slippery. A vendor can say a model is much better at research, planning, or tool use, but if your actual workflow is already inside the easy zone, you may see almost nothing. Or worse, you may see the same answer with a higher bill. That doesn’t mean the claim is false; it means the claim may be real only in the kinds of messy, branching tasks the article didn’t hit, or in aggregate over lots of runs. But if that’s the case, I’d want Anthropic to say it plainly, because “doubled performance” reads a lot bigger than “sometimes better on harder agent traces.”

The detail I’d trust most here is the author’s insistence on measuring tokens and billing. People talk about model quality as if it floats above cost. It doesn’t. For Claude Code users and anyone wiring Claude into agent loops, cost is part of quality. A model that behaves the same and bills more is not an upgrade in any practical sense, at least not for that workload.

If I were testing this myself, I’d want to push past a four-task sample and look for the edge cases where 5.1 is supposed to separate: longer tool chains, more retries, more ambiguous retrieval, more opportunities to go off the rails. Because the article leaves me with a pretty specific suspicion: this may be an upgrade for benchmark graphs and a wash for a lot of real developer work.


Reference: Claude Fable 5.1 vs. Fable 5: On real work, I couldn’t tell them apart.

同じ著者の記事