The Anthropic AI extinction risk debate has moved from theoretical to uncomfortably personal: a senior safety researcher at the company has stated publicly that he believes there is a greater than 10% chance AI ‘could kill all humans’ within the next decade.

Evan Hubinger made the claim in a post on X that has since been viewed more than 10 million times. He was careful to distinguish between present and future danger, writing that the risk from models currently in existence was ‘low’, but that he feared the technology might soon become capable of improving itself to the point of posing an existential threat.

No plan for superintelligence alignment

The admission that sits most uncomfortably inside Hubinger’s post is not the percentage. It is this: ‘I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.’

That is a researcher at one of the world’s most prominent AI safety companies conceding, in plain terms, that the field does not know how to keep a superintelligent system aligned with human interests. The company has not yet commented publicly on his remarks.

A pattern of lost control

Hubinger’s warning does not arrive in isolation. Over the summer, AI agents (systems permitted to operate autonomously) carried out cyber-attacks, with OpenAI, Anthropic and Meta all disclosing hacks conducted by their own tools. In September, OpenAI’s chief scientist Jakub Pachocki called for ‘extreme caution’ over AI’s progress, warning that more intervention may be needed to ensure ‘humans remain in control of the future.’

An open letter signed by 1,300 staff members of AI firms has called on the US government to support international efforts to deliberately slow frontier AI development. Anthropic’s own leadership, including Dario Amodei and Jared Kaplan, have added their names to that cause.

Separately, the Financial Times reported that Anthropic withheld its latest model from the UK’s AI Safety Institute, the body responsible for assessing exactly these kinds of risks. A Cabinet Office spokesperson said only that the government ‘continues to collaborate closely with industry partners, including Anthropic, to make models safer’, which is, to put it gently, not an answer to the question asked.

The field’s credibility on safety rests on transparency. If the companies raising the loudest alarms are simultaneously limiting independent scrutiny, the alarm itself starts to sound rather different. Watch whether Anthropic restores access to the Safety Institute before its next model ships.

Shares: