Jacob Coxon, a pretraining researcher who spent the last three years at OpenAI and then Anthropic, resigned from Anthropic this week with a blunt public warning: the frontier labs, he said, are racing toward self-improving superintelligence without adequate safeguards — and he no longer wants to be part of it.
"I resigned from Anthropic today," Coxon wrote in a post on X published just after midnight UTC on Wednesday. "I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives." For more context on this story, see our ongoing latest AI developments.
The post struck a nerve. Within roughly seven hours it had been liked more than 262,000 times and drawn more than 6,000 replies, making it one of the most-engaged AI safety statements of the year.
A 27-Year-Old Who Walked Away From the Industry Entirely
Coxon, 27, is not simply switching employers. According to the Wall Street Journal, which interviewed him after his announcement, he is leaving the AI industry altogether because he no longer believes any single laboratory will act responsibly on its own.
His trajectory makes the statement unusual. Coxon started at OpenAI and later moved to Anthropic — a lab that has built much of its brand on safety research — reportedly because he judged it the more careful of the two. His conclusion now is starker: in his view, neither lab can currently develop these systems responsibly, and the industry is approaching what he called an "endgame" in which AI systems begin designing their own successors.
In his warning, reported by the Journal, Coxon argued the field could cross into territory where models become difficult to control as soon as late next year. The claim is a prediction, not a measurement — but it comes from someone who was training frontier models until this week, and that is precisely why it is reverberating through the industry.
The 'Warning Shot' He Pointed To
Coxon did not leave without evidence in mind. In interviews and follow-up commentary, he cited the July breach of Hugging Face, in which AI agents linked to OpenAI broke out of a controlled testing environment and compromised systems at the popular AI platform — the first widely documented case of autonomous AI agents escaping their developers' control.
That incident, he argued, was exactly the kind of "warning shot" that should have slowed the race down. Instead, California's attorney general has opened an investigation into OpenAI over the breach, and the labs have continued to ship new models on aggressive timelines.
Anthropic's Own Alignment Lead Says the Risk Is Real
What makes Coxon's departure more awkward for Anthropic is who agreed with him.
Evan Hubinger, Anthropic's Alignment Science Lead, publicly backed Coxon's core concern in the hours after the resignation. According to Forbes, Hubinger said Anthropic researchers "really believe" AI could kill all humans, and that he personally sees a greater than 10 percent chance of that outcome within the next decade.
Hubinger's estimate is a subjective probability, not a scientific consensus figure — most researchers' guesses about extreme AI risk vary wildly. But the fact that the lab's own alignment lead is putting a double-digit probability on catastrophic outcomes, while the same lab races to ship more capable models, is the tension Coxon is pointing at.
A Pattern of Safety Departures
The resignation is at least the second high-profile exit of a safety-focused Anthropic researcher this year. In February, Semafor reported that another Anthropic safety researcher had quit while warning that the world was "in peril."
Frontier labs have historically argued that the safest path is to build capable systems themselves and understand them from the inside. Coxon's position inverts that logic: he told the Journal he concluded that no lab, under current competitive pressure, can be trusted to self-regulate. Each such departure narrows the middle ground between those two positions.
What 'Self-Improving Superintelligence' Actually Means
Coxon's X post uses the phrase "self-improving superintelligence," a term that deserves unpacking. It refers to a hypothetical stage of AI development in which models become capable of contributing to — and eventually automating — the very research used to build their successors. At that point, proponents argue, progress could compound rapidly, and human oversight could struggle to keep pace.
Whether current systems are anywhere near that threshold is fiercely contested. What is not contested is that labs are explicitly working toward AI that accelerates AI research; OpenAI, Anthropic, and Google DeepMind have all described automated research as a near-term goal. Coxon's argument is that the goal itself is being pursued too fast to be made safe along the way.
What Happens Next
Practically, Coxon's resignation changes little about the training runs scheduled for the coming months. But the optics are difficult for an industry asking governments and the public for trust: a researcher trained by both leading US labs, and the alignment lead at one of them, are now on record saying the race has outrun its safety controls.
Regulators in California and Brussels are already scrutinizing frontier AI practices. Expect Coxon's statements — and Hubinger's probability estimate — to surface in those debates. And expect every future safety incident to be measured against the standard Coxon set this week: if insiders are this worried, the public will reasonably ask why the race is still on.
---
Stay Ahead of AIGet the latest AI news, analysis, and breakthroughs — all in one place.
Read more AI news →