Anthropic researcher's departure turns into a public existential warning on self-improving AI
Jacob Coxon's exit from Anthropic landed differently Tuesday night. The AI researcher did not announce a startup or lodge a policy complaint. He used his departure to state publicly that frontier AI companies are…
Key takeaways
- Anthropic researcher Jacob Coxon announced his departure Tuesday night with a public warning that frontier AI companies are "gambling with our lives."
- Coxon warned about future "self-improving superintelligence" that could hack anything, revolutionize fields overnight, and acquire real power and resources, not about today's models.
- Evan Hubinger, Anthropic's Alignment Science lead, publicly agreed and put his personal probability of AI causing human extinction within the next decade at above 10%.
- Hubinger's estimate was personal and offered in support of Coxon, not an official Anthropic company policy statement.
- A formal Anthropic response to Hubinger's comment would mark the next confirmable shift from individual disclosure to institutional position.
Jacob Coxon's exit from Anthropic landed differently Tuesday night. The AI researcher did not announce a startup or lodge a policy complaint. He used his departure to state publicly that frontier AI companies are "gambling with our lives" with systems their teams "earnestly believe... could kill us all by the end of the decade."
What the resignation said
The risk Coxon identified sits in the future, not in today's models. He pointed to the prospect of "self-improving superintelligence" producing "superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources." On why labs continue building despite that view, Coxon wrote that researchers have either not "internalized the civilizational stakes" or believe they must "speedrun" the race to superintelligence first, before a less responsible actor gets there.
The confirmation carries more weight than the resignation itself. Evan Hubinger, Anthropic's Alignment Science lead, responded on social media and agreed with Coxon. Hubinger stated that Anthropic "really do earnestly believe AI could kill all humans" and offered his personal probability at above 10% within the next decade.
A public, on-record figure of greater than 10% odds of human extinction, from the person running alignment science at a frontier lab, is an uncomfortable number to sit alongside the current pace of model releases. The consensus position has been that alignment risk is real but tractable, something being worked on in parallel with capability development. Hubinger did not step outside that framework entirely. His estimate is personal, offered in support of a researcher who had already walked out, and not a company policy statement.
That distinction is exactly what to watch. If Anthropic issues a formal response to Hubinger's public comment, this moves from individual disclosure to institutional position. That is the next confirmable moment.
Filed via arstechnica.com