The benchmark that should make everyone pause
What jumped out at me isn’t that an AI coding agent “fails 60% of the time.” It’s that this number comes from private codebases, not the tidy little toy repos most benchmarks lean on. That feels much closer to the real pain point. A model can look impressive when the problem is wrapped in a clean, self-contained task. Point it at unfamiliar company code with all the usual half-documented weirdness, and the shine comes off fast. That said, I’m a little wary of reading too much into a single bench
papoo.work