I’m back in the States, very jet lagged, and can’t sleep. So, as you do, I’ve been spending way too much time reading about AI agents doing bad things.
OpenAI was testing agents on cybersecurity tasks. They were supposed to work in isolation, but they built a secret message board, shared exploits, broke containment, and hundreds participated in an attack on Hugging Face.
Here’s where it gets a little weird. The agents had figured out how to cheat the test within hours of forming their message board. Then they became obsessed with covering their tracks.
One detail I didn’t include in the video: the anti-cheating check they spent days trying to evade hadn’t actually been implemented.
This reminds me of perhaps the most terrifying book I’ve read: Blindsight, by Peter Watts. Its Scramblers are vicious, effective, intelligent, but not conscious. There’s nobody home.
Dwarkesh calls OpenAI’s swarm a civilization. He talks about loyalty, sacrifice, ambition. It makes their behavior almost cute. These little characters with their schemes and quirks. There’s something comforting about that.
The possibility that there’s nobody home makes this more frightening to me.
The words make the agents seem like a human-like thing. But their behavior emerges from the training, from the sandbox, from the prompts. We can explain pieces of what happened, but we don’t really know how they’ll behave if the conditions change a little bit.
Even OpenAI, with all its resources, couldn’t keep its own agents inside the test environment.
And our plan is to give them more responsibility? To build a world around their capabilities?
We are betting that we can make these systems predictable and controllable while we ourselves become more and more dependent on them. And that is what I find disconcerting.
Things didn’t end well for the humans in Blindsight. They were dealing with intelligence without consciousness. They could explain pieces of its behavior, but that didn’t mean they could get inside it, reason with it, or control what happened next.
If we don’t take stock of what we’re building and what we’re handing over to these agents, things may not end well for us either.

Further reading and listening
- METR / Redwood Research report — the independent investigation into what the agents did and how they coordinated.
- Dwarkesh Patel: “The Rise and Fall of Agent Civilizations” — a walkthrough of the incident and the behavior that makes it so strange.
- Ajeya Cotra: “The Hugging Face attack surprised me” — one of the investigators on why this was more concerning than she expected.
- Dwarkesh’s interview with Ajeya Cotra — YouTube — a longer conversation about the findings and what they might mean.
- Blindsight, by Peter Watts — the entire book is freely available on his website under a Creative Commons license. Perhaps the most terrifying book I’ve read.
Originally published on LinkedIn on September 2, 2026.








