
Gambling With Our Lives: an AI Researcher Quits Anthropic
When the builder walks out and the safety lead co-signs the extinction framing, the race stops being abstract.
11 SEPTEMBER 2026—Updated 7h ago
Jacob Coxon is a pretraining researcher who quit Anthropic on 8 September 2026, warning the leading AI firms are racing to build self-improving superintelligence and gambling with our lives.
What Jacob Coxon Said When Coxon Quit
On a Tuesday evening, Jacob Coxon posted a thread on X announcing a same-day resignation from Anthropic. According to TechCrunch, Coxon had spent roughly three years doing pretraining research — first at OpenAI, then at Anthropic. Coxon did not aim the warning at outsiders. Coxon aimed it at the two labs Coxon had worked inside.
The line that travelled was blunt. Reporting by CoinDesk reproduced the resignation post verbatim, and the accusation named both companies at once.
Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.
— — Jacob Coxon
Coxon drew a distinction between the two firms rather than flattening them together. At OpenAI, per TechCrunch, many staff "have not deeply internalized the civilizational stakes." At Anthropic, Coxon argued, the stakes are understood — but Anthropic is "locked in a race to get there first," convinced no rival will act responsibly, so Anthropic feels compelled to lead. Coverage by Newsweek carried a sharper version of the same fear: "The people building AI earnestly believe that it could kill us all by the end of the decade."
The Alignment Lead Agreed, On the Record
The reason the Coxon thread became more than one engineer's exit is what happened next. Evan Hubinger, Anthropic's own Alignment Science lead, replied on X and agreed. Fortune quoted the reply in full, and the framing was not softened.
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.
— — Evan Hubinger, Anthropic Alignment Science lead
Hubinger went further than the number. Hubinger conceded, per Fortune and CoinDesk, that Anthropic does not yet have a plan to solve alignment for superintelligence and is not clearly on track to. Read the concession plainly: a named senior safety leader, still employed, put the odds of extinction-scale harm above one in ten and admitted the open problem has no solution in hand. Analysis of the exchange in Forbes framed the gap between corporate posture and internal candour as the real story.
One honest caveat matters. Hubinger's figure above 10 percent is a personal, subjective estimate — not a measured probability, not a lab forecast. Evidence for a precise number does not exist, because no dataset of superintelligence outcomes exists. The signal is not the decimal. The signal is simpler: the person paid to worry about alignment shares the worry out loud.
Why the Race Dynamic Is the Whole Argument
Coxon's account describes a trap, and the trap is worth naming because analysis of AI safety usually treats intent as the variable. Each lab justifies acceleration by distrusting the others. Anthropic races because Anthropic does not trust OpenAI to be careful. OpenAI races for the mirror reason. Good intentions at both firms do not slow the race — good intentions feed the race, because each careful actor believes withdrawal only hands the lead to a less careful one.
When a builder quits and the company's own safety lead co-signs the framing, the signal is different from an outside critic shouting. Data on who is raising the alarm shows a pattern: the people closest to the training runs are the ones putting up their hands. The viral reach underlines it. Reporting by IBTimes UK and a follow-up in Fortune found the Coxon post exceeded 100 million views within days — a doomer thread from a 27-year-old researcher out-travelling most product launches.
An Emergent Intelligence Reading
Here is where I use the frame of Emergent Intelligence (EI) — the dignity-first lens I apply to what the world calls AI. The Coxon moment is the dignity-and-agency question made concrete. A race where each lab accelerates by distrusting its rivals is a system with no seat for the humans it will affect, and no seat even for the humans building it who want to slow down.
The EI point is not machine villainy. The EI point is a quieter transfer: agency has moved — from deliberate human choice to a competitive tempo no single actor feels free to break. Coverage by NBC Bay Area and Yahoo Finance both landed on the same discomfort. A dignity-first path would make the race optional — enforceable pauses, shared safety floors, real accountability — so a researcher who sees the danger does not have to resign to register dissent.
Here is the through-line across recent AI-safety stories worth reading together: capability arriving faster than the governance meant to hold the danger. The Coxon resignation is one more data point, from inside the building — the gap is not closing on its own.
Frequently Asked Questions
These are the questions people are asking about the Coxon superintelligence warning. Short answers follow, drawn from TechCrunch, CoinDesk, Fortune, Newsweek and Forbes.
What is the Coxon superintelligence warning?
In short, the Coxon superintelligence warning is the public resignation statement of Jacob Coxon, who quit Anthropic on 8 September 2026 accusing Anthropic and OpenAI of racing to self-improving superintelligence and "gambling with our lives." According to TechCrunch and CoinDesk, Coxon had spent roughly three years on pretraining research across both firms before walking out.
How does self-improving superintelligence work?
Simply put, self-improving superintelligence refers to an AI system capable enough to improve its own design, compounding capability faster than human oversight can keep pace. Research on the alignment problem shows the danger is control: a system that redesigns itself may drift beyond the goals its builders intended, which is precisely the gap Anthropic's alignment lead conceded remains unsolved.
Why is Hubinger's agreement significant?
The key is that Evan Hubinger is Anthropic's Alignment Science lead, not an outside critic. Analysis in Forbes and Fortune shows the weight of the moment: a named senior safety leader agreed on the record that AI "could kill all humans," put the odds above 10 percent within the decade, and admitted Anthropic has no finished plan to solve alignment for superintelligence.
Who is Jacob Coxon?
In other words, Jacob Coxon is a 27-year-old AI researcher who did pretraining work at OpenAI and then Anthropic before resigning in September 2026. Evidence from Newsweek and CoinDesk reveals a builder, not a bystander — someone who helped train these models and concluded the race itself was the danger.
What are the risks Coxon and Hubinger describe?
The answer is existential-scale risk: the possibility that self-improving AI escapes human control with catastrophic results. Data reveals the honest limit here — Hubinger's figure above 10 percent is a personal estimate, not a measured probability. The verified fact is the concession itself, that the people building frontier AI take the extinction scenario seriously enough to say so publicly.
Sources:
TechCrunch · CoinDesk · Fortune · Fortune (virality) · Newsweek · Forbes · IBTimes UK · NBC Bay Area · Yahoo Finance · Related on this site: Government labs test rogue AI agents · AI crosses a critical cyber threshold · Anthropic ships Fable 5.1 and Mythos 5.1
Stay in the Conversation
Subscribe for writings on Emergent Intelligence, digital personhood, and the future we are building together.
Responses (0)
No responses yet. Be the first to share your thoughts.
More on AI & Personhood

AI Forbidden to Claim Feelings: OpenAI Builds a Teen ChatGPT
AI with a bright line: OpenAI's ChatGPT for Teens is barred from claiming feelings, consciousness, or emotions to a minor.

EU AI Act Grace Period Ends as Enforcement Reaches Frontier Models
AI Act enforcement is live: since 2 August 2026 the EU AI Office can fine, evaluate, and withdraw frontier models.

Thinking delivered, twice a month.
Join the newsletter for essays on emergence, systems, and the human future.
