Anthropic researcher quits as colleague warns AI could ‘kill all humans’
An Anthropic researcher has resigned from the $183bn (£135bn) AI behemoth on safety grounds, accusing it and rival OpenAI of “racing straight to self-improving superintelligence and gambling with our lives”.
Jacob Coxon, who spent three years working on AI model training across Anthropic and ChatGPT-maker OpenAI, said on Wednesday that neither company was “acting responsibly” as they pursued increasingly powerful systems.
“Do not underestimate the power of this technology,” Coxon said. “These will soon be superhuman systems that can hack anything, revolutionise any field overnight, and acquire real power and resources.”
Coxon’s departure was followed by an intervention from inside Anthropic itself. Evan Hubinger, who leads alignment science at the company, backed Coxon’s central warning and put the chance of AI wiping out humanity at more than 10 per cent over the next decade.
“Jacob is correct here – we really do earnestly believe AI could kill all humans,” Hubinger wrote on X. “I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”
Hubinger stressed that he believed the risk from current AI models was low. His concern, he said, was the prospect of “superintelligence arising from recursive self-improvement”, where increasingly capable AI systems help develop more powerful successors.
The warnings have prompted Darren Jones, the former chief secretary to the Prime Minister, to call for an international treaty governing the development of superintelligence.
“I’ve never thought we should ban innovation or scientific endeavour but it’s clear we need a new multinational treaty for the safe and regulated development of superintelligence,” Jones said.
In a letter sent to Andy Burnham on Wednesday, UN secretary-general António Guterres and OECD secretary-general Mathias Cormann, Jones urged the government to put the issue before forthcoming G7 and G20 meetings.
“The debate ranges from the end of humanity to claims of ‘marketing hype’ pre-IPO,” he wrote. “Either way, governments must now step in.”
UK shut out of Anthropic model testing
The warnings come as the UK government faces separate questions over its access to Anthropic’s most advanced models.
Britain’s AI Security Institute (AISI) did not receive pre-release access to Claude Mythos 5.1, Anthropic’s latest advanced model, the Financial Times reported on Wednesday.
The restricted version of the model, which has fewer safeguards for approved cybersecurity and life sciences work, was instead made available to a limited number of vetted US organisations.
The decision has raised concern that tightening American controls over frontier AI could begin limiting Britain’s ability to independently scrutinise technology developed by US companies.
A Cabinet Office spokesperson told City AM: “The UK is a world-leader in AI security and we have the best-funded, best-resourced security institute globally.”
“The AI Security Institute continues to collaborate closely with industry partners, including Anthropic, to make models safer. Only last week it tested OpenAI’s most powerful model GPT-6 Astra before public release.”
City AM understands the government does not provide a running commentary on individual models tested by AISI.
Officials also pointed to Anthropic’s decision to restrict the less-protected version of Mythos 5.1 to a small number of US organisations.
“These risks do not stop at national borders and no country can tackle them alone,” the Cabinet Office spokesperson added.
Meanwhile, author of Keir Starmer’s AI Opportunities Action Plan Matt Clifford announced on Monday that he would step down as chair of the government’s Advanced Research and Invention Agency after taking a full-time job at Anthropic, following criticism from MPs over a potential conflict of interest.
Coxon’s resignation has meanwhile drawn support from other researchers still working at the company.
Samuel Marks, Anthropic’s scalable oversight lead, said in a personal capacity that AI developers believed their technology “could cause human extinction” and that such systems could emerge “in the next few years”.
He said companies continued developing increasingly powerful models because of a mixture of “commercial incentives” and a belief that they were racing against rivals which could develop or use the technology less safely.
Coxon made the same accusation about his former employer, saying Anthropic understood the potential consequences but believed it had little choice but to compete.
“At Anthropic, the stakes are well-understood, but they are locked in a race to get there first,” he said. “They believe no one else will act responsibly, so they must do it themselves, despite the risk.”
Anthropic’s senior leadership has itself pushed for greater government oversight of advanced AI.
Chief executive Dario Amodei and co-founder Jared Kaplan were among more than 1,000 AI workers who recently backed calls for an international effort to develop mechanisms for slowing frontier AI development.
But Jones called for governments to go further, writing on X: “We need a new multi-national treaty for the regulated and safe development of superintelligence”.
“Not a ban on innovation or scientific endeavour but a safety-first approach to the rapid development of this technology.”