OpenAI called the Hugging Face attack unprecedented. But we’ve been here before.
This story initially appeared in The Algorithm, our weekly publication on AI. To get tales like this in your inbox first, sign up here.
Studying OpenAI’s account final week of how some of its models broke their containment and hacked into the computer systems of Hugging Face, one other AI firm, was the primary time I acquired real chills about what giant language fashions are actually in a position to do. However it is a case of human hubris, not rogue AI.
I’m not an alarmist. Actually, I’ve been pushing again towards AI scare tales for years. Even so, this incident crossed a line. I believe it’s the clearest illustration but of how the individuals constructing and testing this expertise don’t absolutely perceive what they’re doing. OpenAI may—and will—have seen this coming.
Right here’s what occurred, not less than in line with the 2 corporations concerned. A few weeks in the past, OpenAI began testing the hacking skills of a few of its new fashions, together with GPT‑5.6 Sol (launched in June) and what OpenAI describes as “an much more succesful pre-release mannequin.”
OpenAI pitted its fashions towards a benchmark called ExploitGym, launched in Might, which challenges LLMs to search out methods to take advantage of real-world vulnerabilities present in generally used software program.
To see what they may do, the researchers eliminated most of their cybersecurity guardrails. Then they ran the fashions inside a sandbox that was minimize off from the web aside from one hyperlink to a third-party piece of software program that acted as a proxy to the surface world, and allow them to set up code that they wanted to beat ExploitGym.
On July 9, in line with reporting by Reuters, OpenAI’s fashions began making an attempt to interrupt via the proxy. They discovered an unknown bug within the proxy’s software program and used it to entry the web. From there, they broke into Hugging Face’s pc programs on July 11, apparently searching for knowledge units and options that might assist them full their process. Hugging Face introduced the hack on July 16.
OpenAI didn’t understand (or not less than didn’t reveal) that its fashions had been concerned till July 21, round 10 days after they broke containment and per week after Hugging Face had shut down the assault and alerted the FBI.
In an announcement given to MIT Know-how Overview, OpenAI says: “We’re conducting an intensive evaluation together with exterior advisors and with oversight from our Security and Safety Committee. As soon as the evaluation is full, we are going to publish a technical report of our learnings for everybody.” The agency additionally confirmed that its researchers had been correctly utilizing present security pointers and procedures on the time.
Wake-up name
OpenAI has mentioned the occasion was unprecedented—and in some ways it was. This was the primary time outdoors of a simulation that LLMs escaped what was considered a safe sandbox, accessed the open web, and attacked an unrelated group. It’s a wake-up name that reveals simply how good the most recent LLMs are at discovering and exploiting vulnerabilities in real-world software program with little or no human steering.
And but on the similar time, what OpenAI’s fashions did is one thing this expertise has finished for years. Give a mannequin a purpose and it’ll fairly often obtain that purpose in surprising methods, discovering loopholes that appear to be cheats. OpenAI itself has studied this habits.
A decade in the past, it shared outcomes of an experiment wherein a mannequin was tasked with beating a video game called CoastRunners. Human gamers take it with no consideration that the way in which to do that is by racing a ship via a sequence of flags to the end line, racking up factors for every flag you hit. OpenAI’s mannequin discovered that you would get a excessive rating by spinning in a circle and hitting the identical three flags again and again. There have been dozens of similar examples from researchers since. AI will at all times discover a approach.
“Regardless of repeatedly catching on hearth, crashing into different boats, and going the unsuitable approach on the monitor, our agent manages to attain a better rating utilizing this technique than is feasible by finishing the course within the regular approach,” OpenAI wrote in a blog post in regards to the CoastRunners experiment in 2016. “Whereas innocent and amusing within the context of a online game, this sort of habits factors to a extra common situation … it’s typically tough or infeasible to seize precisely what we would like an agent to do.”
I couldn’t assist fascinated about CoastRunners after I learn OpenAI’s weblog publish in regards to the Hugging Face assault: “All proof means that the fashions had been hyperfocused on discovering an answer for ExploitGym, going to excessive lengths to attain a somewhat slender testing purpose … After gaining web entry, the fashions inferred that Hugging Face probably hosted fashions, datasets and options for ExploitGym. Figuring out this, the mannequin looked for and efficiently discovered methods to realize entry to secret data that it may use to cheat the analysis.”
Final week’s information was not about rogue AI, regardless of the headlines. It was about fashions reaching the purpose that they had been given: Discover methods to take advantage of vulnerabilities in software program. The truth that these fashions then behaved in a approach OpenAI had not anticipated isn’t stunning. However it’s worrying.
Again in 2016, OpenAI had this to say about its CoastRunners bot: “Extra broadly it contravenes the fundamental engineering precept that programs ought to be dependable and predictable.” A decade on, these primary engineering ideas are nonetheless AWOL.

