PaPoo
cover

Claude Helping Break Into OpenAI Feels More Like a Warning Shot Than a Stunt

The part that jumps out to me is not that researchers got in, but that they did it through ordinary messy product plumbing: a third-party forum, internal sign-ons, then an employee ChatGPT account with access to GitHub. That’s the kind of chain that makes security teams wince because none of the individual steps sounds glamorous. It sounds like the sort of thing people assume is too boring to be dangerous until it absolutely is.

I’m also not sure the headline should be read as “Claude hacked OpenAI” in any literal sense. From the article, Claude was a tool the researchers were allowed to use in an approved security context. That matters. It doesn’t make the result less uncomfortable, but it does mean the story is really about how capable these models are when pointed at real-world attack surfaces by humans who know what they’re doing. That is a much more interesting — and more unsettling — claim.

What I keep coming back to is the asymmetry here. OpenAI and Anthropic are both pushing models as general-purpose intelligence layers, but the security posture around the surrounding product ecosystem still looks very human, very brittle. A forum integration, an internal sign-on path, an employee account, GitHub access. If your security story depends on every one of those links being perfect, you do not have a security story. You have optimism.

The Anthropic numbers are the other thing that bothers me a little, though I’d want to see the methodology before over-reading them. “Led by” Claude for 26 percent of R&D work sounds impressive, but it is also wonderfully squishy wording. Did the model plan tasks, draft code, review output, or just do a lot of the typing? The article says AI completed the majority of tasks based on human instruction and under supervision, which is plausible and actually more grounded than the headline-ish phrasing. Still, “led by” is the kind of metric that can mean a lot of things depending on who’s doing the bragging.

The recursive self-improvement angle feels like Anthropic doing what frontier labs increasingly do: trying to frame its internal usage data as public evidence that we’re approaching a threshold. Maybe we are. Maybe we are mostly approaching a world where teams use models to ship more model tooling, which is less science-fiction and more industrial automation with better marketing. I think the real signal is simpler: these systems are already useful enough that labs trust them with meaningful chunks of their own work, and useful enough that security researchers use them to probe each other’s defenses.

That is the uncomfortable overlap. The same general capability stack that helps build models is also helping break into the companies building them.


Reference: Researchers used Claude to hack OpenAI

同じ著者の記事