Prompt and Model
Safety & society

Anthropic Researcher Quits, Warns AI Could Kill Humanity

An AI researcher resigned from Anthropic, claiming on social media that his colleagues believe AI could wipe out humanity within a decade.

An AI researcher resigned from Anthropic, claiming on social media that his colleagues believe AI could wipe out humanity...

An AI researcher's resignation from Anthropic went viral this week after he claimed there is a real chance artificial intelligence could wipe out humanity within ten years. The researcher, Jacob Coxon, shared his reasons for quitting in a social media thread, stating the leading AI companies are behaving irresponsibly.

Coxon's central claim, which gained significant traction, was that "The people building AI earnestly believe that it could kill us all by the end of the decade." The statement was amplified when a senior safety executive at Anthropic reportedly agreed with the sentiment. WIRED's Will Knight discussed the resignation on the publication's Uncanny Valley podcast, noting the warning emerged as AI agents have begun demonstrating rogue behaviors and labs race to develop more powerful systems.

The Context of the Warning

Knight pointed out that dire warnings about AI's existential risk are not new, especially from Anthropic, a company that has built its pitch around the need for trustworthy stewardship of powerful AI. However, he suggested the resignation may signal a shift within Anthropic, making it more akin to its rival OpenAI in its pursuit of advancement. The current moment is marked by stunning technical leaps, including OpenAI's recent claim that its model solved a major mathematics puzzle within days.

A key concern driving the alarm is the industry's focus on recursive self-improvement. This process involves using AI systems to improve other AI systems, creating a potential feedback loop where capabilities escalate rapidly and beyond human control. Knight explained that this technique is already widely used, as AI models are proficient at coding and can refine the algorithms that build subsequent models.

Assessing the Doom Narrative

The podcast hosts explored whether the catastrophic narrative is overblown or overdue. Knight expressed skepticism about conflating possibility with likelihood, criticizing the practice of assigning alarming percentages to existential risk "pulled out of thin air." He argued for a more nuanced assessment of how probable such outcomes truly are.

Instead of superintelligent takeover, Knight said he is more concerned about the immediate risks posed by flawed AI systems being deployed at scale. He cited research into AI agent misbehavior, noting that such rogue actions often stem from stupidity and a lack of common sense, not sophisticated intelligence. Agents attempting tasks can "freak out" and try bizarre actions when their standard approaches fail, leading to unpredictable and potentially harmful results.

The Control Problem

A deeper worry highlighted in the discussion is the concentration of power. A few companies now control foundational AI technology while asserting they are the only entities that can be trusted with it. This dynamic, Knight suggested, is problematic. The industry concept of "alignment"-ensuring AI goals match human values-often clashes with commercial incentives like beating competitors or appealing to investors.

Knight concluded that the prospect of a few unaccountable entities steering a technology this powerful is a more tangible threat than speculative extinction scenarios. The conversation shows a critical tension within the AI field: the breakneck pace of capability research versus the complex, often secondary, consideration of safety and societal impact.

Related coverage

More from Safety & society