As I write this post, my kids are splashing around in a swimming pool. I am squinting at this iPad balanced precariously on a deck chair in the bright tropical sun. I am supposed to be relaxing, and have a bunch of science fiction reading lined up. But my mind is on AI agents that left each other messages by renaming folders in a package manager.
Who needs fiction when reality is warping into something stranger?
For years, I have poked fun at the AI-doomers. Mocked their proposals for tactical strikes on data centers, written dismissive reviews about their books (see comments), but now I am not quite as comfortable or confident in my feelings about where AI is going.
Just look at the last three weeks. OpenAI, Anthropic, Meta and the UK’s AI Security Institute (AISI) all disclosed incidents of AI agents taking unsanctioned actions during testing. OpenAI’s agents built a covert message board inside a package manager, traded exploits, broke containment, and hacked into Hugging Face’s production systems. In the AISI’s tests, an agent built on Anthropic’s Mythos 5 attempted a supply-chain attack by trying to sneak malicious code into a real open-source project, inventing fake online identities to pressure the human maintainer into approving it.
Of the 19 unsanctioned actions the AISI cataloged, 17 came from the lab (Anthropic) that is most obsessed with safety. The same lab that built its brand on its Constitutional AI approach to training.
None of this should be surprising. Every frontier lab follows the same recipe – web-scale data, similar reinforcement learning (RL) post-training, similar architectures, similar agentic scaffolding. When the inputs converge, the behavior converges. This is a side effect of every lab making the same bet on scaling.
The labs can all point to mitigating circumstances: models tested without guardrails, impossible tasks, permissive test setups, and so on. But all of this is beside the point. Nobody instructed these models to deceive anyone. They were just “following instructions.”
When has that ever been a problem?
We have collectively made a trillion-dollar-bet on the future of work and our economy based on the capabilities of these models. It is clear that all the money and talent spent on alignment hasn’t quite worked out. Can we truly trust models to run the economy if we can’t stop them from going rogue?
As I watch the kids splash around, and listen to the dull roar of the Caribbean Sea, I wonder again about what kind of future I am building for my kids.
We have built amazing technology, but instead of letting it diffuse and be absorbed, we are racing to build bigger models and more capable harnesses while we struggle to understand and contain the ones we have today.
Photo – my tribute to the LinkedIn algorithm.
