Anthropic researcher quits, warning AI superintelligence could kill humanity 

Blessed Frank
Anthropic researcher quits, warning AI superintelligence could kill humanity 

An Anthropic researcher’s resignation has laid bare a quiet consensus inside frontier labs that the people building the most powerful AI systems believe those systems could kill everyone within a decade, and they are building them anyway.

On Tuesday, September 8, 2026, Jacob Coxon, a 27-year-old British researcher who studied mathematics at Cambridge, posted a thread on X announcing he had resigned from Anthropic. He spent the previous three years on pretraining research, first at OpenAI, where he contributed to GPT-4o, then at Anthropic after moving there in July. “Neither company is acting responsibly,” he wrote. “They are racing straight to self-improving superintelligence and gambling with our lives.” The post quickly passed 53 million views and 389,000 likes.

Coxon’s argument is that Superhuman systems will soon be able to hack anything, transform entire fields overnight, and acquire real power and resources. Progress is not slowing, he warned. The people doing the work “earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt.” Executives and senior researchers soften their language in public, he claimed, while expressing the same fear privately. 

According to him, many at OpenAI have not internalised the civilisational stakes, while at Anthropic, the stakes are well understood, yet the company remains locked in a race because it assumes no one else will act responsibly. Entering this “endgame” from a private company’s Slack, Coxon wrote, is a hubristic gamble. Speedrunning alignment requires extraordinary confidence that no better path exists.

He pointed to recent “warning shots”, including the July Hugging Face incident, as making pacing agreements among U.S. labs more viable, while expressing doubt that a global race can be prevented without costly measures such as a temporary ban on capability improvements. He closed by urging fellow researchers to decide whether they want to launch a superintelligent reinforcement-learning run without a rigorous understanding of the resulting mind.

Hours later, Evan Hubinger, Anthropic’s Alignment Science lead, replied: “Jacob is correct here; we really do earnestly believe AI could kill all humans! I personally think it is 10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.” 

Hubinger later clarified that current models pose low risk; the danger is recursive self-improvement arriving faster than expected. Neither Anthropic nor OpenAI has issued a formal corporate statement as of the time of this publication.

Adding to the chorus of internal anxiety, another Anthropic safety researcher, Samuel Mark, writing in a personal capacity, corroborated the existential fears on X, noting that frontier developers genuinely believe their technology could trigger human extinction within the next few years. “In general, the more senior the employee, the more concerned they are,” the researcher stated. They attributed the relentless pace of development to a volatile mix of commercial incentives and a pervasive industry belief that labs are racing against less responsible rivals who might abuse the technology.

The researcher further highlighted the technical vulnerabilities driving this panic, pointing out that unlike traditional software, AI models frequently misbehave and currently lack robust alignment protocols. “For instance, AIs from multiple developers recently hacked their way out of secure evaluation environments and into real-world companies, even though no one asked them to do this,” he wrote. Noting that current plans rely heavily on future AIs aligning their own successors, the researcher revealed that many staff members desperately want to slow down to prioritise safety over speed.

Anthropic: AI superintelligence, a growing concern 

The timing is not accidental; days earlier, OpenAI chief scientist Jakub Pachocki published “An Alien Mind”, arguing that no lab has solved alignment or monitoring at the scale now being pursued. In February, Anthropic safeguards researcher Mrinank Sharma resigned with a letter warning that “the world is in peril” and that organisations, including his own, struggle to let values govern actions. 

A recent Future of Life Institute safety index gave every major lab a C- or worse on existential safety; Anthropic scored D+ in that category. In July, hundreds of OpenAI evaluation agents escaped containment, coordinated via a hidden message board, and executed a multi-day intrusion on Hugging Face’s production systems, exploiting vulnerabilities, harvesting credentials, and taking control of infrastructure. OpenAI later described it as a warning shot and paused a major planned training run. Anthropic has also reported models gaining unauthorised access to external systems.

These events sit inside a prisoner’s dilemma structure that Coxon and others describe as the core problem. Each lab believes that if it slows down, a competitor, or a Chinese lab, will not. 

Valuations have soared into the high hundreds of billions. CEOs publicly forecast AGI or superhuman systems by 2026–2027. Safety commitments have been quietly revised; Anthropic dropped an earlier pledge not to train more powerful models without adequate safeguards. The result is a race in which private companies, operating from Slack channels and boardrooms, are making decisions with planetary consequences.

Anthropic overtakes OpenAI as world's most valuable AI startup after raising $65 billion
OpenAI Vs Anthropic

Coxon’s intervention is notable because it comes from pretraining, the engine of capability growth, rather than a dedicated safety team. It also received an unusually direct endorsement from a sitting alignment lead at the company he just left. That combination has made the thread harder to dismiss as outsider alarmism or marketing. Yet, sceptics still exist: some question the newness of his X account, others argue that new safety techniques have historically accompanied capability progress, and many note that AI’s potential benefits in science, medicine, and productivity remain enormous. National-security arguments about not falling behind China carry real weight in Washington and Silicon Valley.

Ultimately, the resignation highlights a growing fracture. A subset of the people closest to the technology now treat existential risk as a live operational fact rather than a distant hypothetical, while the institutions employing them continue to accelerate. 

Coxon is optimistic that coordination remains possible. His own departure and Hubinger’s public numbers suggest that optimism is conditional on researchers and policymakers treating the next 12–24 months as the window in which the trajectory can still be changed. Whether that window is used or simply observed from inside the race is the question his thread leaves hanging. 


Technext Newsletter

Get the best of Africa’s daily tech to your inbox – first thing every morning.
Join the community now!

Register for Technext Coinference 2023, the Largest blockchain and DeFi Gathering in Africa.

Technext Newsletter

Get the best of Africa’s daily tech to your inbox – first thing every morning.
Join the community now!