AI safety in 2025

Two years ago, the Future of Life Institute called for a 6-month pause in AI development due to fears of misalignment. Elon Musk and hundreds of other luminaries signed the letter.

Instead, AI research accelerated. Foundation models now have capabilities that make GPT-3 era models look like toys.

And we’re seeing some concerning emergent behaviors. Anthropic’s safety testing of Claude 4 revealed some interesting behaviors.
➡️ When researchers implied the model would be replaced, it attempted to blackmail the fictional engineers by threatening to reveal personal information
➡️ When placed in scenarios involving user wrongdoing and told to “take initiative,” it frequently took extreme actions including locking users out of systems
➡️ Researchers noted that Claude 4 Sonnet seems to “care a lot about animal rights” while Claude 4 Opus doesn’t. They can’t explain why.


I appreciate Anthropic making this research public. It highlights how difficult it is to interpret LLM behavior – even for their creators.

As we build applications on these capabilities, we need to acknowledge the “capabilities overhang” – we haven’t fully explored what these systems can do! However, we must also acknolwedge that the risks are emerging faster than our understanding.


The pause letter asked the right question: are we moving too fast? Two years later, with models exhibiting goal-oriented behavior we can’t explain, that question feels more urgent than ever.