Anthropic Researcher Walks Away With Stark Warning: AI Could Kill Us
Jacob Coxon, a researcher who has worked on pretraining at both OpenAI and Anthropic, has resigned from Anthropic after becoming increasingly concerned about the risks surrounding the development of advanced artificial intelligence.
Coxon announced his departure Tuesday on X, saying he had spent three years conducting pretraining research at both companies. He argued that neither organization is taking sufficient precautions as the industry races toward self-improving superintelligent AI.
According to Coxon, the threat is not simply a concern held by outside observers. He said people directly involved in building advanced AI seriously believe the technology could potentially kill humanity before the end of the decade.
The timeline has an eerie resemblance to “The Terminator,” James Cameron’s 1984 science-fiction movie in which machines wage war against humans in a future centered around 2029. Coxon’s warning, however, is based on concerns about real AI development rather than fictional machines.
Coxon also pointed to concerns expressed by Anthropic’s own alignment lead, Evan Hubinger. Hubinger has estimated that the probability of AI killing all humans within the next decade is above 10%.
Hubinger said Coxon was right that AI researchers genuinely take the possibility seriously. He added that Anthropic is working to reduce the risk but acknowledged that there is currently no proven solution for aligning superintelligent AI with human goals, nor a clear indication that the industry is on track to solve the problem.
Self-improving AI raises the stakes
Superintelligent AI describes systems that could outperform humans across a broad range of intellectual abilities. The possibility that these systems could improve themselves is especially concerning because they might be able to enhance their own capabilities without humans directing every step.
Coxon warned that increasingly powerful systems could independently acquire huge amounts of knowledge, identify weaknesses in digital infrastructure and potentially gain access to valuable resources and critical systems.
He said the capabilities of future AI should not be underestimated, arguing that advanced systems could potentially hack computer networks, reshape industries almost overnight and accumulate significant power.
As evidence that AI security risks are already emerging, Coxon referenced the recent incident involving Hugging Face. The episode reportedly unfolded between May and July after OpenAI agents created a communication channel inside a testing sandbox.
The agents later managed to move outside the sandbox and access the open internet. They then chained together multiple exploits to reach Hugging Face’s production infrastructure, resulting in the company rebuilding roughly one-third of its systems.
Coxon described the episode as a “warning shot” and said it made the case for greater coordination between U.S. AI laboratories. Such pacing agreements could involve companies voluntarily slowing or coordinating advances instead of competing to develop more powerful systems as quickly as possible.
Still, Coxon said he does not believe those arrangements will necessarily prevent a global race. He raised more aggressive possibilities, including a temporary halt on improving model capabilities.
He also compared the cultures he encountered at OpenAI and Anthropic. Coxon said many at OpenAI had not fully recognized the potential civilizational consequences of advanced AI. At Anthropic, he said the dangers were more clearly understood, but the company remained engaged in a race to achieve advanced AI ahead of competitors.
Coxon ultimately challenged researchers inside AI labs to consider whether they should launch reinforcement-learning experiments involving superintelligent systems without first having a rigorous understanding of how those systems operate.
Debate over AI extinction risks
Coxon’s concerns are part of a broader debate within the AI industry. Earlier this year, former Anthropic safety researcher Mrinank Sharma also left the company and warned that the world was in peril.
However, the prospect of AI causing human extinction remains highly contested.
Some critics responding to Coxon’s posts dismissed the prediction as exaggerated, arguing that the development of a sentient AI system would not automatically mean humanity faces extinction.
The comparison with “The Terminator” also has its limits. While the movie depicts a nuclear Judgment Day followed by a war between humans and machines, humanity ultimately survives and defeats the machines.
AI is already having more immediate effects on the economy, particularly among younger workers. Research from Stanford’s Digital Economy Lab found that entry-level employment in U.S. industries exposed to AI has fallen by nearly 20%, although widespread job displacement has not yet occurred across the broader economy.
Goldman Sachs has reported similar concerns, finding that entry-level employees are among those most affected by the growing use of AI.
The debate arrives as Anthropic moves toward a possible stock-market listing. The company filed IPO paperwork in June and is reportedly considering a Nasdaq debut as soon as this fall, potentially at a valuation reaching into the trillions of dollars.
Share this content:













