PaPoo
cover

Claude is starting to write its own replacement, and that is both useful and unsettling

What jumps out to me isn’t the “26%” figure by itself. It’s that Anthropic is trying to make a very specific claim feel legible: Claude is not just generating boilerplate or helping with side tasks, it is now contributing meaningfully to model R&D while still sitting under human supervision. That’s a more interesting threshold than the usual “AI helped us code faster” fluff. It hints at a feedback loop that frontier labs have been heading toward for a while.

But I’m not ready to treat the number as some clean scientific measurement. “Led,” “collaboration,” “large chunks of work” — those are slippery categories. I can believe Anthropic has an internal methodology, and I can also believe the line between a model doing real research work and a model making a human faster at checking ideas is a lot fuzzier than the announcement makes it sound. If Claude proposes variants, drafts experiments, writes code, and humans approve the direction, is that the model leading? Maybe. Is it self-improvement in the sci-fi sense? Not yet, at least not from what’s here.

What I find more telling is the company’s insistence on public metrics. That feels less like celebration and more like positioning. Anthropic is clearly aware that recursive self-improvement talk makes people nervous, especially when the company itself is one of the loudest voices warning about AI safety. Sharing the numbers first lets them frame the story before someone else does. I think that’s smart communication, but it’s also a little self-serving. Transparency is good; controlled transparency is still controlled.

The part I’d want to see, if I were using Claude in my own work, is not the headline percentage but the methodology. What kinds of tasks qualify? How often does the model suggest something genuinely novel versus just speeding up routine lab work? How much of the “90% collaboration” is basically pair programming with a very fast autocomplete, and how much is the model actually steering research? Without that, the stat is interesting but not fully actionable.

Still, the trend is hard to ignore. If the model is already contributing to the process of making better models, then developers building on Claude are using a system that is increasingly entangled with its own future. That could be a productivity boost. It could also make the system harder to audit, because the thing you’re trying to understand is helping shape the next version of itself. That is the real tension here, and it’s not going away.


Reference: Claude, Anthropic’s AI model, is helping to develop the next version of itself

同じ著者の記事