On 9 September, a 27-year-old researcher named Jacob Coxon resigned from Anthropic and said the quiet part out loud: the people building frontier AI genuinely believe it might kill everyone by the end of the decade, and they are speeding up anyway. Coxon had spent three years on pretraining, first at OpenAI, then at Anthropic, which is to say he sat inside the two labs at the centre of the race. His verdict, given to the Wall Street Journal and expanded in a public thread, is that the industry is “gambling with our lives” and sprinting toward self-improving superintelligence with no idea how to control it.
The easy move is to file this under “disgruntled ex-employee has dramatic exit” and scroll on. Resist it, for one reason: the most senior alignment researcher still at Anthropic publicly agreed with him. That is the part worth sitting up for, and it is where this piece is going to spend its time, because the interesting question is not whether Coxon is sincere (he plainly is) but whether he is right, and the honest answer requires actually looking at the numbers rather than the vibes.
What Coxon actually said
Coxon’s specific claim is sharper than the headlines suggest. He says his colleagues have adopted the vocabulary of “crunchtime” and “endgame” to describe the current sprint toward self-improving models, systems that help build the next, more capable systems. His timeline is startlingly short: “We’re on track for a lot of the most aggressive of these scenarios where by the end of next year things could be out of control already.”
His indictment of the two labs is different for each. At OpenAI, he says, many staff simply have not absorbed the civilisational stakes of what they are building. At Anthropic, by contrast, the stakes are widely understood, and the company races anyway, because its leaders are convinced that if they slow down, a less careful rival reaches the frontier first. That is the trap in one sentence: everyone speeds up because everyone else is speeding up, and “we are the responsible ones, so we must win the race to be safe” becomes the justification for the exact behaviour that makes it dangerous.
The admission that should have been the headline
Coxon resigning is a story. Evan Hubinger, Anthropic’s alignment science lead, replying in public to say he was basically correct is a bigger one. Hubinger confirmed that many at the company earnestly believe AI could kill all humans, put his own estimate above 10% within the next decade, and added, remarkably, that Anthropic is trying hard but “does not yet have a plan for aligning superintelligence and is not clearly on track to produce one.”
Read that again slowly. The person whose job is to keep Anthropic’s AI aligned with human interests is saying, on the record, that he assigns better-than-one-in-ten odds to the technology killing everyone, and that his own employer has no plan to prevent it. If a lead safety engineer at Boeing said there was a 10% chance the new aircraft kills all its passengers and no plan yet to stop it, the fleet would be grounded by lunchtime. In AI, it was a Tuesday.
So is he right? Here is where the honesty starts
This is the point where a lesser piece would either scream that we are all doomed or sneer that it is all hype. Both are cop-outs. The truth is that the people who have thought hardest about this disagree with each other by a staggering margin, and any writer pretending otherwise is selling you something.
Consider the actual spread of “p(doom)”, the informal shorthand for the probability that AI causes human extinction or something close to it. Yann LeCun, Meta’s chief AI scientist and a Turing Award winner, puts it at effectively zero, under 1%. A large 2023 survey of AI researchers landed on a median around 5% and a mean of roughly 14% over a century. Geoffrey Hinton, who won the same Turing Award as LeCun for the same underlying work, says 10 to 20%. Anthropic’s own Dario Amodei has floated 10 to 25%. Yoshua Bengio, the third of that Turing trio, says around 20%. And then, at the far end, Eliezer Yudkowsky says north of 95%, and researcher Roman Yampolskiy says 99.9%.
That range spans roughly six orders of magnitude. Three men who shared the same prize for the same invention land at “basically zero”, “one in five”, and the survey middle. That is not a rounding error; it is a chasm, and it tells you something crucial: nobody actually knows. As Bengio put it when swatting down LeCun’s confidence, “he doesn’t know, nobody knows”, and it is reckless to claim certainty in either direction. The single most defensible position on AI extinction risk is deep uncertainty, which is also the least shareable, which is why you rarely hear it.
The case for calm: the constraints nobody markets
If you only listen to the doom side, you would think uncontrollable superintelligence is a matter of months. It is worth deferring to the sceptics on why it might not be, because their arguments are technical rather than emotional.
The first constraint is architectural. LeCun and Gary Marcus have argued for years that today’s transformer-based models, for all their fluency, cannot reach genuine general intelligence by scaling alone. On this view, hallucinations and brittle reasoning are not teething problems but symptoms of a design that lacks a world model, planning, and grounded understanding. If they are right, the current paradigm hits a ceiling well short of the god-machine, and the “endgame” is further off than the insiders fear.
The second is physical. The nightmare scenario relies on recursive self-improvement, an AI rewriting itself into ever-greater intelligence in a runaway loop. But intelligence does not delete the physical world. Training runs need chips, chips need fabs, fabs need years and nations, and everything needs enormous amounts of power. As one careful analysis put it, intelligence multiplies what is parallelisable but does not delete serial or physical constraints. You cannot think your way past a power grid, a supply chain, or the laws of thermodynamics on a timescale of months.
The third is data and signal. Several researchers argue that data quality sets a hard ceiling on how far self-improvement can run: a model training on its own output tends to degrade rather than explode, and the specific capabilities that would drive a true takeoff, research taste and knowing which problems are worth solving, are exactly the ones current systems are weakest at. The scary loop assumes the machine already has the judgment it is supposedly trying to acquire.
Ranking the nightmares by plausibility
Not all doom scenarios are equal, and lumping them together is how the debate goes stupid. Roughly, from most to least plausible:
Human misuse is the near-certain one and gets the least airtime because it is not cinematic. An AI that helps a bad actor build a bioweapon or run a civilisation-scale disinformation campaign does not need to be superintelligent or “want” anything; it needs to be capable and available. The AI researcher Melanie Mitchell, a noted sceptic of the sci-fi scenarios, argues this is the one genuinely plausible existential pathway. It is also the most tractable, which is the good news buried in the bad.
Gradual disempowerment is the middle case: no single robot uprising, just humans handing over more and more decisions to systems we understand less and less, until we cannot meaningfully take the wheel back. Less a bang than a slow surrender of agency, and arguably already underway in small ways.
Fast recursive takeover, the Terminator scenario the headlines love, is the one the constraints above hit hardest. Not impossible, but the most speculative and the most dependent on assumptions about self-improvement that many serious researchers doubt. It deserves to be taken seriously and it does not deserve to dominate the conversation the way it does.
Who actually deserves the criticism
Here is where we stop being diplomatic. The scandal is not that some researchers hold high p(doom) numbers. The scandal is the behaviour those numbers coexist with. If you genuinely believe, as Hubinger says he does, that there is a better-than-10% chance your industry ends the human story, and your response is to keep racing because a competitor might be worse, you have talked yourself into a moral position that would embarrass a Bond villain. “We must build the thing that might kill everyone, so that we are the ones holding it” is not a safety strategy. It is a hostage note written to yourself.
The glib dismissers earn a share too. Anyone declaring the risk is flatly zero, as if a technology no one fully understands could be pronounced safe by assertion, is making a claim they cannot support, and Bengio is right to call it dangerous. But the doom absolutists deserve the same scepticism. When Yampolskiy says 99.9%, he is dressing a hunch in the costume of precision. A number that specific implies a rigour that does not exist, and it discredits the legitimate concern underneath it. False certainty is the disease here, and it is bipartisan.
The honest villains, if we must have them, are the incentives. A race with no brakes, run by people who privately admit the car might go off a cliff, funded by investors who treat “might end civilisation” as a quarterly risk factor. That is worth being angry about. The uncertainty is not an excuse for panic, but it is absolutely an argument for slowing down, and the people closest to the technology saying so, and then not slowing down, is the actual story.
Where this leaves you
Do not let anyone bully you into either camp with false confidence. The reasonable position, boring as it is, goes like this: extinction by rogue superintelligence in the next few years is unlikely and heavily constrained by physics, data and architecture; catastrophic misuse of powerful-but-dumb AI is a real and present danger that deserves far more attention than the robot-apocalypse fantasy; and the correct response to genuine uncertainty about a technology this powerful is caution, not a sprint.
Jacob Coxon may or may not be right about his timeline. What is not in dispute is that a serious person walked away from a lot of money and status because he could not stomach the gamble, and the person paid to reassure us instead confirmed his fears. You do not have to believe the world ends in 2030 to think that is worth stopping to look at. (None of this is investment or safety advice; it is an attempt to describe an argument honestly, which is rarer than it should be.)
This piece touches on existential risk and the mental toll of working in the field. If it is weighing on you personally, that is worth talking through with someone you trust. Related on Top Tool Stack: the Ban Artificial Superintelligence Act and OpenAI declaring the AGI era.