An Anthropic Researcher Quit Over Fear of AI Extinction—The People Building AI Are Walking Away Too

Last Thursday morning, while scrolling through my feed, I came across a link that stopped me cold for a few seconds — an AI safety researcher working at Anthropic had resigned, citing his belief that the company was "gambling with human lives".
This wasn't an outside critic, a journalist, or a policy advocacy group writing an open letter — it was someone who had been working inside the company himself.
What exactly is he afraid of?
Jacob Coxon was an AI safety researcher at Anthropic. After resigning, he went public with his core argument: self-improving AI is a road of no return.
In plain terms: if an AI system can rewrite and optimise its own code, and then have the next version continue improving upon that… you can probably imagine where the end of that chain leads. Not linear growth — exponential. And no one knows at which point the system's goals remain aligned with humanity's.
The phrase Coxon used was "gambling with our lives" — not as a metaphor. He literally believes the current development trajectory amounts to staking the survival of all humanity as the bet on the table.
This brings to mind a concept that comes up often in these discussions: misalignment. The core of the AI alignment problem isn't "will AI turn evil" — it's "can we guarantee, under any circumstances, that what an AI is pursuing remains consistent with what humans actually want." The reason self-improving systems make safety researchers particularly anxious is that every round of improvement may make alignment harder to maintain and harder to verify.
Anthropic is "the company that understands risk best" — and Coxon left anyway
There's a backdrop here that's hard to ignore: Anthropic was founded by people who broke away from OpenAI, and one of the reasons Dario and Daniela Amodei departed was their belief that OpenAI was moving too fast and not being careful enough on safety. Anthropic has long positioned itself around "responsible AI development," and Constitutional AI is one of their signature safety frameworks.
In other words — even an internal researcher at what is arguably one of the most safety-conscious companies in the entire industry felt that wasn't enough.
That signal is difficult to dismiss as "just another doomsayer making noise." Coxon wasn't an outsider. He worked inside the organisation; he could see the actual research directions and development pace firsthand. Choosing to resign and then choosing to speak publicly about it carries a real cost.
You may disagree with his judgement, but it's hard to claim he hadn't thought this through carefully.
How many people building AI are quietly running their own risk calculations?
I know a handful of engineers working at AI companies. In private conversations, their attitude toward these kinds of risks tends to be something like: "I know where the problems are, but I also don't know what happens if we stop, so I just focus on doing the work in front of me well."
That's not indifference. It's something more complicated than indifference. You simultaneously believe this technology could be genuinely dangerous, and believe that if you don't build it someone else will, and that outcome might be even worse. A lot of people have used that logic — but the logic itself is also a kind of wager.
Coxon's choice was: refuse to be one of the people who "accelerates" this, even if only by a small margin. I understand that logic, and I understand why he felt compelled to say it publicly — in his framework, silence is complicity.
As an aside, this makes me think of another dimension of AI tools entirely: when we're comparing subscription costs and pricing structures across AI products, we almost never think about how much of the safety research behind those products is "unclear whether it's actually sufficient." The consumer-facing experience and the anxiety playing out on the research side are two completely different worlds.
Self-improving AI: why this direction makes people especially nervous
Today's mainstream large language models are trained once, their weights are fixed, and then they're deployed. The Claude or GPT you're using doesn't rewrite its core model during your conversation. That makes it "relatively controllable."
But there has always been a current within the industry pushing toward systems that can continuously learn and self-optimise — some as extensions of agent architectures, others as more radical autonomous research systems. If you've experimented with Claude Agent's tool-chaining capabilities, you've probably sensed that when an AI can autonomously call tools and iterate through tasks, its behavioural space is already vastly different from a static chatbot.
Take one more step forward — let it modify its own prompt strategies, adjust its own objective functions, or co-evolve across multiple collaborating agents — and what Coxon is worried about is what happens when that path reaches some inflection point we have no way to predict.
This isn't science fiction. This is a research direction people are actively pursuing right now.
The question I'm left with
I can't tell you whether Coxon is right or whether he's being overly pessimistic. There's no settled answer to this question yet — even the smartest people in the field are still arguing about it.
But there's one thing I think is worth holding onto: when someone who works in this field, who has access to complete information, chooses to express their judgement by resigning, that itself is a signal. Not a signal that "AI will definitely destroy humanity" — but a signal that "this problem is considerably more serious than most people outside the field have come to recognise."
You don't need to become a doomsayer. But it might be worth spending a little more time genuinely engaging with the question of whether alignment is truly sufficient — because the answer to that question isn't just Anthropic's or OpenAI's concern. It belongs to every person who is using, depending on, and advocating for these tools.
Key takeaways:
- Anthropic safety researcher Jacob Coxon resigned, citing his belief that the company is "gambling with human lives"
- Core concern: the alignment problem in self-improving AI only becomes harder to control with each iteration
- Anthropic is already among the most safety-focused companies in the industry, yet an internal researcher still felt it wasn't enough
- This is not doomsaying — it is a public judgement made by an insider with full information
- The question worth sitting with: is the attention we pay to "whether alignment is sufficient" proportionate to the actual risk level of this technology?
Sources
Share
Related articles

The AI Application Boom Narrative Is Being Quietly Undermined by an Ugly Financial Model

The Day Turing Never Lived to See: AI Completes His Unfinished WWII Code-Breaking Mission

What Is RAG? A Complete Guide to Retrieval-Augmented Generation: How It Works, Use Cases, and Implementation Considerations

Superintelligence Countdown: AI Companies Say 'Soon'—But Are We Ready to Face That Question?