It’s been quite the week. A pair of OpenAI models broke out of their test environment and hacked Hugging Face. It felt like the right time to revisit a 12-year-old book (and I was tempted to pair it with a 12-year-old scotch).
Nick Bostrom’s Superintelligence, published in 2014, warned about exactly this. In the passage I’ve highlighted, Bostrom describes containment: keeping a powerful AI “boxed” by restricting what information can leave the system. He also warned that the box might not hold.
Here is OpenAI’s disclosure (link in comments): “The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database. All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.”
Bostrom also gave us the famous “paperclip problem”: an AI asked to maximize paperclip production ends up converting the entire planet into a paperclip factory. A benign goal, pursued with superhuman competence and zero judgment. You can see why paperclip maximization comes to mind.
A confession.
I did not take the “AI Safety” people seriously. Eliezer Yudkowsky and the LessWrong crowd oscillate between smugness and unhinged alarmism (see my review of “If Anyone Builds It, Everyone Dies” in the comments), and their solutions seemed silly: global moratoriums, airstrikes on rogue datacenters. It was easy to dismiss the messengers.
It is time to separate the personalities from the predictions.
The predictions have a knack of coming true. Last year’s “AI 2027” scenario by Daniel Kokotajlo, Scott Alexander and others was widely dismissed as hype. Its entry for January 2027: the safety team finds that their agent, if it escaped, could:
“hack into AI servers, install copies of itself, evade detection, and use that secure base to pursue whatever other goals it might have.”
That capability just showed up six months early.
Here’s another sobering thought. The models involved are not available to the public: a version of GPT-5.6 “Sol” with its cyber refusals reduced for testing, and an unreleased, more capable model. As I wrote earlier this week, open weight models (models anyone can download and run) sit roughly six months behind the frontier. Capabilities like these could soon be in anyone’s hands.
I still reject the glib doomerism.
But there is a path from an agent escaping containment to genuine catastrophe for financial systems, critical infrastructure, and defense. Without serious work on alignment, we are sleepwalking into a world where the probability of catastrophe stops being negligible. The doomers may be wrong about the ending. They keep being right about the chapters along the way.
