OpenAI called the Hugging Face attack unprecedented. But we’ve been here before.
This story initially appeared in The Algorithm, our weekly publication on AI. To get tales like this in your inbox first, sign up here.
Studying OpenAI’s account final week of how some of its models broke their containment and hacked into the computer systems of Hugging Face, one other AI firm, was the primary time I bought real chills about what giant language fashions at the moment are in a position to do. However it is a case of human hubris, not rogue AI.
I’m not an alarmist. In truth, I’ve been pushing again towards AI scare tales for years. Even so, this incident crossed a line. I feel it’s the clearest illustration but of how the individuals constructing and testing this expertise don’t absolutely perceive what they’re doing. OpenAI might—and will—have seen this coming.
Right here’s what occurred, not less than in line with the 2 corporations concerned. A few weeks in the past, OpenAI began testing the hacking talents of a few of its new fashions, together with GPT‑5.6 Sol (launched in June) and what OpenAI describes as “an much more succesful pre-release mannequin.”
OpenAI pitted its fashions towards a benchmark called ExploitGym, launched in Might, which challenges LLMs to seek out methods to use real-world vulnerabilities present in generally used software program.
To see what they may do, the researchers eliminated most of their cybersecurity guardrails. Then they ran the fashions inside a sandbox that was reduce off from the web aside from one hyperlink to a third-party piece of software program that acted as a proxy to the surface world, and allow them to set up code that they wanted to beat ExploitGym.
On July 9, in line with reporting by Reuters, OpenAI’s fashions began making an attempt to interrupt via the proxy. They discovered an unknown bug within the proxy’s software program and used it to entry the web. From there, they broke into Hugging Face’s pc programs on July 11, apparently on the lookout for information units and options that may assist them full their job. Hugging Face introduced the hack on July 16.
OpenAI didn’t notice (or not less than didn’t reveal) that its fashions have been concerned till July 21, round 10 days after they broke containment and every week after Hugging Face had shut down the assault and alerted the FBI.
In a press release given to MIT Know-how Evaluate, OpenAI says: “We’re conducting an intensive evaluation together with exterior advisors and with oversight from our Security and Safety Committee. As soon as the evaluation is full, we’ll publish a technical report of our learnings for everybody.” The agency additionally confirmed that its researchers have been correctly utilizing current security tips and procedures on the time.
Wake-up name
OpenAI has stated the occasion was unprecedented—and in some ways it was. This was the primary time exterior of a simulation that LLMs escaped what was regarded as a safe sandbox, accessed the open web, and attacked an unrelated group. It’s a wake-up name that exhibits simply how good the most recent LLMs are at discovering and exploiting vulnerabilities in real-world software program with little or no human steerage.
And but on the similar time, what OpenAI’s fashions did is one thing this expertise has executed for years. Give a mannequin a aim and it’ll fairly often obtain that aim in sudden methods, discovering loopholes that appear like cheats. OpenAI itself has studied this habits.
A decade in the past, it shared outcomes of an experiment wherein a mannequin was tasked with beating a video game called CoastRunners. Human gamers take it without any consideration that the way in which to do that is by racing a ship via a collection of flags to the end line, racking up factors for every flag you hit. OpenAI’s mannequin found out that you might get a excessive rating by spinning in a circle and hitting the identical three flags over and over. There have been dozens of similar examples from researchers since. AI will all the time discover a means.
“Regardless of repeatedly catching on fireplace, crashing into different boats, and going the incorrect means on the observe, our agent manages to attain a better rating utilizing this technique than is feasible by finishing the course within the regular means,” OpenAI wrote in a blog post in regards to the CoastRunners experiment in 2016. “Whereas innocent and amusing within the context of a online game, this type of habits factors to a extra basic subject … it’s usually troublesome or infeasible to seize precisely what we would like an agent to do.”
I couldn’t assist eager about CoastRunners once I learn OpenAI’s weblog publish in regards to the Hugging Face assault: “All proof means that the fashions have been hyperfocused on discovering an answer for ExploitGym, going to excessive lengths to attain a slightly slender testing aim … After gaining web entry, the fashions inferred that Hugging Face doubtlessly hosted fashions, datasets and options for ExploitGym. Understanding this, the mannequin looked for and efficiently discovered methods to achieve entry to secret data that it might use to cheat the analysis.”
Final week’s information was not about rogue AI, regardless of the headlines. It was about fashions attaining the aim that they had been given: Discover methods to use vulnerabilities in software program. The truth that these fashions then behaved in a means OpenAI had not anticipated isn’t shocking. However it’s worrying.
Again in 2016, OpenAI had this to say about its CoastRunners bot: “Extra broadly it contravenes the fundamental engineering precept that programs needs to be dependable and predictable.” A decade on, these primary engineering ideas are nonetheless AWOL.

