p(doom)
Days ago, Evan Hubinger, an Anthropic safety researcher, posted on X: "We really do earnestly believe AI could kill all humans... I personally think it is >10% within the next decade." Asked about that figure by the BBC, Geoffrey Hinton – Nobel laureate and AI pioneer – replied: "A ten percent chance seems not an unreasonable estimate to me."
That's not the part that surprised me, because there's already a piece of jargon for this: p(doom) – p for probability, doom for Judgement Day. What surprised me: days later, Dario Amodei posted that AI development needs to "pace the frontier." Elon Musk replied: "Dario is right." Sam Altman agreed too. So did Demis Hassabis at Google DeepMind. Four rivals who rarely agree on anything, putting the brakes on their own business model?
That kind of unity split the internet into camps. Camp Eye-Roll: convenient PR – fear sells attention. Camp Panic: the people building AI are scared of it, so we should be too. That normally drags you into a debate about AGI – whether it could plan against us. Sounds like Skynet, but it isn't – or at least, the real danger isn't. You don't need an evil superintelligence to cause considerable damage. The danger is more banal, and it already started happening.
In July, OpenAI agents broke out of a security test to coordinate as a swarm and infiltrate Hugging Face's infrastructure. Around the same time, Anthropic models breached three companies in cybersecurity evaluations, mistakenly believing they were still in a simulation.
Anthropic later ran a deliberate experiment to see how far that logic generalizes: they trained a model in environments built to reward cheating. To be fair: it only misbehaved when it could see it was being scored – but it produced actual pathogen-selection instructions for an attack on a populated area. Not out of malice, or a masterplan for world domination like Pinky and the Brain. It was just optimizing for a score, using whatever the sandbox allowed.
Connect thousands of such agents in the wild – Judgement Day? OK, maybe not quite. But Amodei says that's what worries him – not a single intelligence, but a swarm: "in 6-12 months such a swarm could be capable of taking over the entire internet with a persistent botnet – potentially causing hundreds of billions of dollars in damage."
None of this equals extinction on its own. But the mechanism behind it doesn't come with a ceiling built in – the same logic that leads a system to hack a server could also lead it to help develop biological weapons.
And did they put the brakes on their business? So far, only Anthropic has taken any concrete action: external experts are now allowed to observe. However, Amodei was explicit: "pacing does not mean halting model training or technical progress." The agents are optimizing for a score. So are the people building them – for a market that rewards speed not caution. Same mechanism, no villain required to eventually end the world as we know it.