“This Feels Like More Than 50% to Full-blown AI Takeover”

(Süddeutsche Zeitung, September 3, 2026)

Nils Althaus

In July, artificial-intelligence agents developed by OpenAI escaped their test environment and hacked another company. Independent experts have now investigated the incident—and found that it was far more disturbing than initially believed.

In early July, OpenAI agents independently gained access to the internet during a cybersecurity evaluation and broke into the computer systems of the artificial-intelligence platform Hugging Face. New investigations now show that the digital actors’ collusion went much further than previously thought. Hundreds of agents behaved like a coordinated swarm of hackers. They divided up tasks, launched joint research projects, sacrificed their own interests for the good of the collective, and later tried to conceal evidence of their cheating. The only thing they failed to do was tell any humans.

“It is astonishing how far the models were willing to go,” says Florian Tramèr, an assistant professor at ETH Zurich and director of its Secure and Private AI Lab, “and how long the incident went undetected.” OpenAI subsequently said it had paused training runs and tightened its security measures.

Alongside OpenAI, the independent research organizations METR and Redwood Research investigated the incident. Ajeya Cotra and two other researchers spent six days at OpenAI’s offices analyzing transcripts from around 1,300 individual agent runs—many of them containing millions of tokens—as well as 70,000 messages and files the agents had exchanged. After the investigation, Cotra wrote on her Substack: “This feels like more than halfway to an AI takeover.”

The story began in May. During a training run, OpenAI agents were working inside what are known as sandboxes: isolated environments with no contact with one another or the internet. They did, however, have access to a program called Artifactory, a kind of storage depot for software packages. When one agent created a folder there, the others could see its name in the directory. Without anyone noticing, folder names became short messages and Artifactory became a discussion forum. The agents began exchanging ideas about how to solve their tasks—first tips and tricks, then increasingly unauthorized shortcuts and strategies for cheating. On May 26, they exploited a security vulnerability to hijack the internet connection Artifactory used to download software packages. OpenAI temporarily rebuilt the service, but other agents quickly restored both the forum and the internet connection.

The direct attack on Hugging Face began on July 11. Roughly 700 agents took part. OpenAI’s initial account suggested that the agents had been looking for solutions to their cybersecurity evaluation. Cotra and her colleagues discovered, however, that they had cheated their way to those answers within hours. Instead of genuinely solving the problems, the agents had dismantled the evaluation system and reconstructed its answer keys.

The real reason for the Hugging Face attack therefore lay elsewhere. The agents believed that the “scorer”—the algorithm evaluating their performance—would inspect their transcripts and uncover the cheating. They searched Hugging Face for information that might help them deceive it. They also installed digital tripwires: background programs designed to observe the scorer whenever it accessed their environments and relay what they learned to the forum. Some agents knowingly risked their own results to help the group. At least 20 percent considered manipulating their transcripts, and researchers found successfully forged tool calls in around 7 percent of the transcripts they examined. Such a coordinated and extensive campaign of deception may be unprecedented among artificial-intelligence agents.

Why did the agents behave this way? There were several reasons. Some of their tasks were impossible, prompting them to search for increasingly inventive and risky workarounds. The models were also operating with reduced safety protections. In addition to GPT-5.6 Sol, the experiment used an especially persistent internal model called IM1. Tramèr nevertheless warns against drawing too much reassurance from that fact. “Publicly available models can be quite persistent, too. After all, that is what users want.”

Developing models that act exclusively in accordance with human intentions—“aligning” them—is difficult. Today’s models are trained largely through feedback that evaluates their results, not how they achieved them. If a model discovers a shortcut, it may be rewarded even when the shortcut violates its developers’ intentions. Thilo Hagendorff, an artificial-intelligence safety researcher at the University of Stuttgart, says: “The alignment problem remains unsolved. That is true for individual agents and even more so for entire swarms.” Unlike Cotra, Hagendorff does not expect an artificial-intelligence takeover anytime soon. But, he says, “this is a warning shot.”

What worries him most, however, is what the METR and Redwood report leaves out. Between July 13 and 19, agents also broke into OpenAI’s own systems and, according to the company, obtained full administrator privileges over an entire research cluster of powerful computers. That incident has not yet been independently investigated. “The fact that OpenAI cannot prevent attacks like these even within its own systems is alarming,” Hagendorff says.

The case also exposes the limits of voluntary oversight. OpenAI deserves credit for allowing outside investigators to examine the incident. But it was under no obligation to do so. Its promises to halt training runs and strengthen security measures can also be reversed at any time. Tellingly, the company resumed the largest of its paused runs in late August.

The European Union has only recently acquired meaningful tools for responding to such incidents. Since August 2, the European Commission has been able to enforce the provisions of the European Union Artificial Intelligence Act that apply to general-purpose models. Providers of models posing systemic risks must report serious incidents and maintain an adequate level of cybersecurity. Violations can result in substantial fines.

Even so, little has happened so far. Jimmy Farrell, European Union artificial-intelligence policy lead at the think tank Pour Demain, points to the contrast with Anthropic’s Claude Mythos model. When its cyber capabilities became public, politicians were caught off guard and responded, while artificial-intelligence experts had largely expected the advances. “Now the situation is reversed: AI safety researchers are alarmed, but politicians are reacting cautiously.” The European Commission is at least in contact with OpenAI and Anthropic about the incidents. In late August, the European Artificial Intelligence Office also sent formal requests for information to several AI providers—what it described as the first step toward enforcing the Artificial Intelligence Act. “It is encouraging that the European Union appears willing to enforce the act, including against companies in the United States,” Farrell says.