PaPoo
cover

Benchmarks keep exposing the same uncomfortable thing

I’m more interested in the floor than the leaderboard here. Claude winning this Hyper-τ-bench and still clearing fewer than a quarter of the tests is not a cute paradox; it’s the whole story. If the best model on a benchmark about building agents is still failing most tasks, then we’re not in the “agents are solved” phase. We’re in the “careful, this demo might be a mirage” phase.

That said, I wouldn’t over-read the raw score without looking hard at what the benchmark is actually asking. “Agents that build agents” sounds clever, but it also sounds like the sort of test that can quietly reward a very specific style of prompting, tool use, and scaffolding. I think that’s the real question here: does Hyper-τ-bench measure a capability you’d trust in production, or does it mostly measure how well a model can follow a synthetic workflow under constrained conditions? Those are not the same thing.

The number that jumps out to me is not Claude’s lead, but how low everyone is. It suggests the field is still brittle when you ask models to do meta-work: designing other agents, wiring them together, and not getting lost in the weeds. That matches my own experience. Models can look oddly competent until you ask them to maintain state, reason about their own output, and keep multiple moving parts aligned. Then the cracks show fast.

So I come away less impressed by the “Claude did best” part than relieved that the benchmark didn’t paper over the difficulty. If this pushes people to stop talking about agents as if they’re just prompt templates with better branding, good. If it becomes yet another leaderboard where we pretend 20-something percent is a meaningful finish line, less good.


Reference: Claude did best on a new benchmark for 'agents that build agents'. It still passed fewer than a quarter of the tests.

同じ著者の記事