PaPoo
cover

Anthropic is asking for brakes after helping build the car

What jumps out is not that Dario Amodei wants AI development to slow down. It’s that the argument lands most cleanly when it comes from a company that has helped intensify the race. That doesn’t make it wrong. It does make it politically convenient in a way I can’t ignore.

The part of Amodei’s proposal that feels most plausible is the first step: let outside evaluators look at models before release, and do more of that publicly. That seems like the bare minimum if frontier systems are now touching cybersecurity, autonomous agents, and other failure modes that are hard to inspect after the fact. If you’re building Claude or anything adjacent, “trust us” is obviously not enough anymore. I’d want red-team access, third-party evals, and a paper trail too.

Where I get skeptical is the leap from “we need better oversight” to “the industry should pace itself.” Self-regulation is always the easiest part to announce and the hardest part to maintain when product pressure hits. Companies say they want common safety standards right up until those standards slow them down relative to competitors. And once you get to the global piece, the whole thing starts to look almost aspirational in the bad sense. Getting the U.S. and its allies to coordinate is one thing. Getting China and Russia to sign onto the same ceiling on capability growth sounds, frankly, close to fantasy.

The geopolitical part is especially thorny. Amodei seems to be arguing for both slowing progress and preserving a U.S. lead by restricting chips and distillation. That’s not a contradiction so much as a tension he doesn’t fully resolve. If the goal is global safety, export controls and competitive containment pull in a different direction. If the goal is maintaining advantage while slowing everyone else, then “safety” starts to look like the respectable packaging for industrial policy. Maybe that’s unavoidable. But it should be said plainly.

The “recursive self-improvement” worry is the one I’d take seriously without overstating it. It’s not sci-fi to think model-assisted training could speed up capability gains faster than oversight can keep up. The question is whether that becomes a near-term practical problem or stays a long-horizon argument that companies use to justify whichever policy position they already prefer. I don’t know yet. Same with the hacking examples. They’re alarming, but one ugly agent incident doesn’t prove we’re on the edge of machine autonomy running away from us. It does prove we’re already past the point where “it mostly follows instructions” is a comforting default assumption.

What I’d actually want from Anthropic here is less philosophy and more friction: release discipline, stronger evals, clearer reporting on failures, and some willingness to say no to capabilities that look obviously dangerous. That would be more convincing than a grand three-step framework. The framework may be directionally right. It just reads like the kind of thing the industry can praise without changing much unless somebody forces it.


Reference: Anthropic CEO says it’s time to pump the brakes on AI

同じ著者の記事