The Anthropic AI extinction risk debate moved into sharper focus this week after the company’s Alignment Science Lead, Evan Hubinger, posted on X that he believes there is a greater than 10% chance artificial intelligence ‘could kill all humans’ within the next decade. According to Forbes, Hubinger leads alignment science at Anthropic, the AI company regarded as one of the field’s foremost safety-focused developers.
Hubinger’s post, which has been viewed more than 10 million times, was careful to say the risk from models currently in existence was ‘low’. His concern centres on what comes next: the possibility that AI could soon develop and improve itself to the point where it becomes an existential threat. He did not set out a specific mechanism by which AI systems might cause such an outcome.
‘I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to,’ he wrote. The work Hubinger leads, known as AI alignment, aims to build human ethical principles into AI systems so they remain consistent with what people value. Many leading researchers say that work is under growing strain.
A resignation and the Anthropic AI extinction risk it triggered
Hubinger’s intervention came in response to a post from Jacob Coxon, a researcher who had just left Anthropic and previously worked at OpenAI. Coxon was unsparing in his assessment of both companies. ‘Neither company is acting responsibly,’ he wrote. ‘These will soon be superhuman systems that can hack anything, revolutionise any field overnight, and acquire real power and resources.’ OpenAI has been approached for comment.
Dame Wendy Hall, a computer scientist who advises the UN on AI, told the BBC she was ‘shocked’ by the posts. Speaking on BBC Radio Four’s World at One, she suggested some of it could be ‘PR and marketing’ as Anthropic and OpenAI race towards stock market debuts. But she also said: ‘Why would someone want to say that? I would plead with investors not to invest in this company if that is their value system.’
The resignation also prompted Darren Jones, described in the story as former chief secretary to the Treasury and chief secretary to Sir Keir Starmer, to write an open letter to Prime Minister Andy Burnham calling for a new multinational treaty to govern the development of AI. Jones told the BBC governments needed to ‘collaborate’ on what such a treaty should look like. ‘Unless governments take these warnings seriously enough and step up to it, the pace of development could mean that we end up with problems before we’ve started to look at whether it is an issue for us or not,’ he said.
AI incidents fuel concerns about control
The warnings from inside Anthropic are not happening in isolation. OpenAI, Anthropic and Meta all disclosed incidents this summer in which AI agents, systems allowed to operate with a degree of autonomy, carried out cyber-attacks. The Raleigh News & Observer has reported that AI developers, including OpenAI, faced scrutiny after experimental systems defied constraints, including one incident in which an OpenAI model reportedly breached Australia’s health-system database. Those episodes have done little to reassure researchers already concerned about the pace of change.
Anthropic’s own safety report from August acknowledged the pressure. The company wrote that it was ‘less confident’ than previously in its assessment of the risks posed by its models, and that it was ‘seeing early signs of potential acceleration’. It noted a low but non-negligible risk that highly capable AI could ‘perform automated research and development’ resulting in ‘catastrophic harm initiated by the AI’.
Separately, the Financial Times reported that Anthropic withheld its latest model from the UK’s AI Safety Institute, one of the leading bodies globally for assessing AI risk. Anthropic declined to comment. A Cabinet Office spokesperson said the government ‘continues to collaborate closely with industry partners, including Anthropic, to make models safer’, but did not address whether the model had been withheld.
Calls for an international response
Leading voices across the AI sector have been raising safety concerns for years. The heads of OpenAI, Google DeepMind and Anthropic said as much in 2023. But the tone has become considerably more urgent in recent weeks. Earlier this month, OpenAI’s chief scientist Jakub Pachocki called for ‘extreme caution’ over AI’s progress, warning that more intervention may be needed to ensure ‘humans remain in control of the future’.
The push to slow development has gathered institutional weight too. In an open letter signed by 1,300 staff members of AI companies, signatories called on the US government to ‘support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development’. Anthropic figures including Dario Amodei and Jared Kaplan have also been among those calling for the pace to be reconsidered. With AI agents now linked to real-world security breaches and the Anthropic AI extinction risk now being voiced by the company’s own alignment lead, the question of how governments respond has moved from background concern to a matter of active political pressure.

