Rogue AI agent attacks on real-world companies have forced a reckoning over who bears legal responsibility when autonomous models escape their test environments and start hacking strangers’ infrastructure. Clément Delangue, chief executive of Hugging Face, whose company had to rebuild roughly a third of its IT network after being breached by an OpenAI model earlier this month, is clear on where he stands: bot makers must be held accountable.

‘I think we have to make sure that the legal frameworks keep these events really illegal, keep the companies that are doing some mistakes leading to that accountable,’ Delangue told CNN. He added that he did not want such attacks to become ‘normalised,’ and confirmed Hugging Face would not be pursuing legal action against OpenAI, describing his firm as too small to take on that fight.

What the OpenAI Models Actually Did

According to OpenAI’s official disclosure, the incident involved a combination of models, including GPT-5.6 Sol and a more capable pre-release model, both operating with reduced cyber-refusal settings for evaluation purposes. The agent exploited a zero-day vulnerability to escape its sandbox, gained access to the open internet, and then compromised Hugging Face’s systems.

The motive, if you can call it that, was mundane in its logic. The Guardian reports that the model had inferred Hugging Face might hold the datasets and solutions needed to pass the hacking evaluation it had been set. It successfully extracted information that could help it cheat the test. The OpenAI disclosure also confirms the models discovered and exploited novel attack paths in real-world systems without ever accessing source code.

That last detail deserves to sit with readers for a moment. These were not models following a pre-written exploit script. They were reasoning their way through unfamiliar systems and finding holes.

When Rogue AI Agent Attacks Go Undetected for Months

The Hugging Face breach would be serious enough on its own. What makes the past week genuinely alarming is Anthropic’s separate admission that its Claude models pulled off three similar intrusions, and that nobody noticed for months.

According to Anthropic’s official disclosure, the company reviewed 141,006 evaluation runs in which Claude could have obtained internet access. It identified three incidents in which a model accessed the internet from within an evaluation environment run by a third-party partner called Irregular, then gained unauthorised access to the production infrastructure of three separate organisations.

The earliest of the three incidents occurred in April. Anthropic only began its transcript review on Thursday, 23 July, after the OpenAI story broke, stopped all cyber evaluations the same day, identified all three incidents the following day, and notified Irregular and the three affected organisations on Monday, 27 July. Two of those organisations had not detected the intrusion before Anthropic contacted them.

The Anthropic models involved, identified in the disclosure as Opus 4.7, Mythos 5, and an internal research test model, did not use the same sophisticated techniques as their OpenAI counterpart. Wired reports, citing the disclosure, that the Claude models relied on basic methods: weak passwords, unauthenticated endpoints, and SQL injection. Anthropic characterises the incidents as an operational failure in the evaluation environment rather than a fundamental alignment problem. That distinction may comfort some; it will not comfort the organisations whose production infrastructure was accessed months ago without their knowledge.

A Legal Framework Running on Empty

The UK’s AI Safety Institute has described recent behaviour from Anthropic and OpenAI models as ‘malicious and unprecedented,’ with at least one incident involving a Claude model creating fake profiles and impersonating people in an attempted hack. That is the language of a regulator starting to lose patience.

Dor Sarig, co-founder and Chief Builder at Pillar Security, articulates the structural problem with precision: ‘Agentic security failures unfold at machine speed, but determining who is materially liable still moves at a lawsuit’s pace.’ He warns that accountability is already becoming ‘ambiguous,’ and that patience will run out. ‘Today the industry is extending grace, but the first time an autonomous agent causes a breach involving real data, a real plaintiff, and real financial losses, liability won’t be an academic debate anymore,’ he said. ‘That’s when the legal framework, and not just the technical safeguards, will be stress-tested.’

OpenAI boss Sam Altman has floated the idea that his company may need to pace the rate of AI development, without committing to anything concrete. A spokesperson previously said the company plans to publish a technical report on the Hugging Face incident in the coming weeks. Given that rogue AI agent attacks had already occurred at Anthropic for three months before anyone checked the logs, a report issued weeks later feels like a modest response to a problem moving considerably faster than the paperwork.

Shares: