Anthropic Staffers Again Sound the Alarm on AI Catastrophe

Demonstrators participate in the "Stop the AI Race" protest march in San Francisco, California, on July 11, 2026. The protestors are making stops outside the offices of OpenAI, Anthropic and Google DeepMind. Karl Mondon/AFP/Getty
On Tuesday, experienced AI researcher Jacob Coxon resigned from the AI firm Anthropic—saying that both that company and OpenAI, his previous employer, were “gambling with our lives” by developing models that could improve themselves at a rapid clip until they reach “superintelligence.”
In an alarming social media post , Evan Hubinger, who leads a team that tries to stress-test Anthropic’s models for safety, essentially agreed that the company’s staffers “really do earnestly believe AI could kill all humans!”
It’s far from the first time researchers have tried to raise the alarm about AI’s dangers, but Coxon s post and subsequent discourse went viral.
The details of AI alignment can be difficult to grasp—a big part of the problem is that no one truly understands the intricacies of how advanced models work. But it doesn’t take expert-level knowledge or insider secrets to understand the reasons for alarm. All you need is three facts about AI: It is already surpassing human abilities in important domains. Leading companies continue to improve it rapidly.
And no one knows how to reliably keep its behavior in line with human goals.
Yesterday, OpenAI shared a solution to one of the six remaining “Millenium Problems,” some of the most heavily researched in all of mathematics. Mathematicians working independently are also claiming credit —but they, too, relied on advanced models for their work. It was the most striking example yet, though not the first, of AI doing cutting-edge math.
A few years ago, a common dismissal of large language models was the claim that they were just elaborate algorithms, creating sentences by guessing the most likely next word. But today’s AI models clearly build on, rather than simply remix, the text reflected in their training data.
The internal reasoning of OpenAI s latest model is less possible to understand—and its behavior is harder to predict.
In some areas, computers or algorithms have outstripped human minds for a while.