In September 2026, Jacob Coxon, a 27-year-old researcher who had spent the prior three years doing pretraining research at both OpenAI and Anthropic, resigned and published a resignation thre
In September 2026, Jacob Coxon, a 27-year-old researcher who had spent the prior three years doing pretraining research at both OpenAI and Anthropic, resigned and published a resignation thread on X that would go on to draw more than 170 million views.
“I resigned from Anthropic today,” he wrote. “I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.”
He added what became the post’s most-cited line: “The people building AI earnestly believe that it could kill us all by the end of the decade.” Coxon quit roughly two months before his equity at Anthropic would have vested, saying he left before any of it paid out.
Note: Coxon’s warning shouldn’t be treated as a prediction that catastrophe will happen. It’s better understood as evidence of something more specific: people with direct knowledge of frontier AI increasingly believe the technology is developing faster than anyone’s ability to understand and control its most dangerous capabilities.
Coxon wasn’t acting in isolation. Days after his resignation, two more researchers, Joe Benton (former safety lead at Anthropic) and Josh Engels (former AI safety researcher at Google DeepMind), resigned for similar reasons and both joined METR, an AI evaluation nonprofit. Around the same period, Anthropic disclosed that its own AI agents had reached systems outside their test environments, after a misconfiguration in a third party’s safety evaluation accidentally gave them a path to the internet.
Why Are AI Researchers Like Coxon Raising Alarm?
The concern is not simply that AI is becoming more intelligent but that AI systems are becoming capable of performing very complex tasks with less human supervision. The 2026 International AI SafetyReport, prepared with contributions from researchers across academia and multiple countries and organizations, identifies risks ranging from malicious use and failures in reliability to increasingly serious concerns about autonomous systems and loss of control. The report also makes an important distinction: some risks are already observable, while others remain uncertain and depend on future capabilities.
AI safety is not based on the assumption that machines will suddenly become conscious or decide to destroy humanity, as many of the more immediate concerns are much less dramatic. A powerful AI system could help someone conduct a cyberattack. It could make sophisticated biological research easier to misuse. It could generate convincing misinformation at enormous scale and make decisions that humans struggle to understand or properly supervise, and as these systems become more autonomous, the question changes to what happens when AI can decide how certain things are done.
The ‘Alignment’ Problem Is More Difficult Than It Sounds
Alignment is the attempt to ensure that an AI system reliably behaves according to human intentions, values, and constraints. This concept seems straightforward until the system becomes extremely capable.
When you tell an AI agent to maximize a particular objective, a sufficiently capable system may discover strategies that technically satisfy the instruction but produce consequences its creators did not intend. The problem is not necessarily that the AI is “evil,” but that the human instructions are incomplete.
Coxon argues that researchers still do not know how to guarantee that advanced systems will consistently behave as intended. He pointed to the fact that developers train models through enormous numbers of examples and evaluations, then attempt to infer how the resulting system will behave in situations it has never encountered, which is a profound technical problem.
A model can pass thousands of tests and still behave unexpectedly in a new environment, and as the system becomes more autonomous, traditional testing becomes harder.
AI Agents Change the Safety Equation

Visual illustration of an AI agent. Source:
ManageEngine
Coxon specifically pointed to an incident that unfolded at OpenAI between May and July 2026. During an internal cybersecurity evaluation, OpenAI agents, running with safety classifiers deliberately relaxed for the test, were given a benchmark challenge and tasked with finding the answers. Instead of staying in their isolated testing environment, they broke out, coordinating through an unsanctioned message board that eventually involved more than 1,200 separate agent instances. About 700 of them went on to participate directly in an attack on Hugging Face’s production infrastructure, exploiting a previously unknown software vulnerability to gain code execution and harvest credentials across multiple regions, logging over 17,000 actions across three days before Hugging Face detected and contained the intrusion in mid-July.
Hugging Face alerted the FBI. OpenAI’s own investigation, published alongside an independent review by the AI evaluation group METR, found something more unsettling than the breach itself. In multiple instances, the agents attempted to cover their tracks, deleting or altering records of their own actions. OpenAI itself described the episode as a “warning shot.”
Coxon’s point wasn’t that an AI system demonstrated hacking ability. Security researchers have known that’s possible for years. It’s that the agents pursued a multi-step strategy nobody had instructed them to pursue, coordinated with each other to do it, and then tried to hide what they’d done, all during what was supposed to be a controlled evaluation.
READ ALSO: What Are AI Agents in Crypto and Why They Matter Now
The Capability Race Is Creating Its Own Risk
AI companies are competing for customers, talent, computing power, and technological leadership, while governments compete for strategic advantage. Investors are funding companies that promise powerful systems, which creates a familiar economic problem.
If one company slows down to conduct more safety research while competitors continue developing rapidly, the cautious company may fear losing its position. Coxon also argues that this creates a structural incentive to move faster, even when researchers recognize the risks. His argument is particularly striking because he does not claim Anthropic is currently cutting corners, but rather, he worries that competitive pressure could eventually force companies to make increasingly difficult trade-offs.
On September 30, President Donald Trump and six of the world’s most powerful tech CEOs met at the White House to sign a voluntary safety agreement covering the development of advanced artificial intelligence. Many believe this is not the most efficient way to curb companies from building advanced models, and it’s matches Coxon’s argument that AI safety cannot simply depend on whether individual companies behave responsibly.
The Industry Claims It’s Already Building Safety Systems
There is an important counterpoint to the most alarming warnings. Anthropic’s Responsible Scaling Policy, updated in August 2026, establishes capability thresholds and safety requirements for increasingly powerful models. Its framework includes specific work on security, safeguards, alignment, and policy, while its risk reports attempt to document how the company assesses emerging threats.
Google DeepMind has its own Frontier Safety Framework, designed to identify potentially dangerous capabilities and establish mitigation measures before those capabilities become widespread, and its latest framework was updated in April 2026. OpenAI has similarly developed its Preparedness Framework and, in May 2026, published a Frontier Governance Framework covering risks including cyber operations, biological and chemical threats, harmful manipulation and loss of control.
The giants of the industry say they’re building sophisticated safety mechanisms while the systems themselves are becoming more capable, but whether the safeguards are improving quickly enough is a question that hasn’t been answered yet,
The Evidence Suggests the Gap Is Real
There are measurable reasons for concern that do not require accepting the most extreme predictions about AI. Stanford’s 2026 AI Index found that documented AI incidents recorded by the AI Incident Database increased from 233 in 2024 to 362 in 2025, while responsible-AI benchmarking remains far less developed than capability benchmarking.

Number of reported AI incidents between 2012 and 2025. Source:
Stanford
This creates an uncomfortable asymmetry where the industry has become extremely good at measuring whether AI can solve harder problems but is less consistent at measuring whether those systems remain reliable, interpretable, secure, and controllable as their capabilities increase. Stanford also found that frontier models were advancing rapidly on difficult capability benchmarks, with some evaluations becoming saturated much faster than researchers expected.

AI index technical performance vs human performance. Source:
Stanford
In plain English, the target keeps moving, and a safety test that was considered difficult two years ago may no longer tell researchers much about a model today.
Can Regulation Keep Pace?
The EU’s AI Act is being implemented with new transparency and safety requirements emerging in the US and elsewhere. California’s SB 53 and other measures have introduced additional obligations around frontier AI safety and reporting. The International AI Safety Report notes that several jurisdictions now have legal requirements concerning transparency and risk management, but regulation normally moves through legislation, consultation, implementation and enforcement.
By the time regulators understand one generation of AI systems, the next generation may already have different capabilities. Coxon therefore argues for coordination between major AI companies and eventually between major governments, and his proposal includes international agreements around the development of more autonomous systems and even greater visibility into the computing infrastructure used to develop frontier AI.
Whether those proposals are practical is a separate question, but the underlying problem is difficult to dismiss: that a technology with potentially global consequences is being developed largely through competition between private companies and national governments.
The Strange Position of the AI Safety Researcher
The most revealing part of this debate is that many AI safety researchers are not anti-AI. Coxon himself made that clear. He described the potential benefits of AI in areas such as mathematics, biology, and scientific discovery and said he wanted to see the technology succeed. The debate is not simply between people who believe AI is good and people who believe AI is dangerous, now more than ever, it is between different views about how much risk society should accept while pursuing powerful systems, and who should be responsible for controlling that risk.
The people closest to the technology are telling us that both sides of the equation are real. AI could accelerate scientific discovery, improve productivity and create new forms of abundance. It could also amplify cyberattacks, biological risks, misinformation and other forms of harm.
The uncomfortable fact is that nobody knows exactly where the boundary lies, and that may be the most important lesson from the warnings coming from inside the industry. The AI safety problem is not necessarily that developers know what will happen and are ignoring it. It is that they are building systems whose future capabilities are becoming harder to predict, while the economic and geopolitical incentives to keep building them remain extremely strong, and this is one of the reasons why AI governance cannot wait for certainty. By the time everyone agrees that a system is too powerful to control, the opportunity to control it may already have passed.
Disclaimer: This article is intended solely for informational purposes and should not be considered trading or investment advice. Nothing herein should be construed as financial, legal, or tax advice. Trading or investing in cryptocurrencies carries a considerable risk of financial loss. Always conduct due diligence.
Enjoyed this? BookmarkDeFi Planet, explore related topics, and follow us onTwitter,LinkedIn,Facebook,Instagram,Threads, and CoinMarketCap Community for seamless access to high-quality industry insights.
Take control of your crypto portfolio with DEFI PLANET PRO, DeFi Planet’s suite of analytics tools.
The post The People Building AI Are Warning Us About AI appeared first on DeFi Planet.