
By Team INVC | INVC NEWS
SAN FRANCISCO, United States | September 9, 2026 —
Jacob Coxon Anthropic resignation has intensified the debate over artificial intelligence safety after the 27-year-old researcher, who worked on pretraining at both Anthropic and OpenAI, walked away from Anthropic and warned that leading AI companies are racing toward increasingly powerful self-improving systems without adequate safeguards.
Coxon announced his resignation in a lengthy thread on X, saying he had spent the past three years conducting pretraining research across OpenAI and Anthropic.
He argued that neither company was acting responsibly enough given the potential consequences of developing systems that could eventually improve their own capabilities.
His warning focuses on what researchers call self-improving or recursively self-improving artificial intelligence — systems that could potentially help design, train or improve future AI systems with decreasing human involvement.
Such fully autonomous superintelligence does not exist today, and its eventual development is uncertain. However, Coxon argues that current progress is moving quickly enough that governments, researchers and technology companies should take the possibility seriously.
Coxon Says AI Labs Are ‘Gambling With Our Lives’
Coxon’s central criticism is directed at the competitive race among frontier AI laboratories.
He accused OpenAI and Anthropic of pushing toward self-improving superintelligence despite uncertainty about whether future systems could remain reliably under human control.
The researcher warned that increasingly capable AI could eventually outperform humans across a wide range of intellectual tasks.
He also raised concerns about potential future systems developing advanced cyber capabilities, rapidly transforming scientific or commercial fields and gaining access to significant real-world resources.
These are projections about possible future systems rather than descriptions of capabilities possessed by ordinary AI products today.
Nevertheless, Coxon believes the consequences could become so serious that continuing development without stronger coordination amounts to taking an unacceptable risk.
Former OpenAI Researcher Questions Both Companies
Coxon’s criticism carries added significance because he has worked inside both OpenAI and Anthropic.
He said he spent approximately three years conducting pretraining research across the two organizations.
Pretraining is one of the fundamental stages in developing large AI models. During the process, models learn statistical patterns from enormous datasets before additional training and safety techniques refine how they behave.
Coxon argued that researchers at OpenAI and Anthropic approach AI risks differently.
According to his account, some people at OpenAI have not fully internalized what he considers the potential civilization-level stakes of advanced AI.
At Anthropic, he said, awareness of the risks is stronger.
However, Coxon believes Anthropic remains trapped in the same competitive dynamic because slowing down unilaterally could allow another company to move ahead with weaker safeguards.
That creates what AI researchers increasingly describe as a coordination problem.
Why Would AI Companies Continue If They Fear the Risks?
This contradiction lies at the center of Coxon’s argument.
If researchers believe highly capable AI could eventually become dangerous, why continue building increasingly powerful systems?
Coxon argues that competition provides much of the answer.
Individual companies may believe that slowing their own development would simply transfer leadership to a rival.
The same problem can exist between countries.
If one nation restricts advanced AI development while another continues rapidly, policymakers may fear losing economic, technological or national-security advantages.
That dynamic makes voluntary restraint difficult.
Coxon therefore believes meaningful safety measures may ultimately require cooperation among companies and governments rather than relying entirely on individual AI laboratories.
Anthropic Safety Researcher Publicly Supports Warning
Coxon’s concerns did not remain an isolated resignation statement.
Evan Hubinger, who leads Alignment Science at Anthropic, publicly responded to the discussion and agreed that advanced AI could present catastrophic risks.
Hubinger gave a personal estimate of greater than 10% for an AI-related outcome that could kill all humans within the next decade.
That figure represents Hubinger’s individual risk assessment, not an established scientific probability or Anthropic corporate forecast.
He also said Anthropic is trying to address the problem but does not yet have a complete solution for aligning hypothetical superintelligent AI systems.
The response is significant because it demonstrates that concerns about advanced AI safety exist among researchers actively working inside frontier laboratories.
Anthropic Says Recursive Self-Improvement Is Not Here Yet
At the same time, the current state of the technology needs important context.
Anthropic has separately published research examining what it calls recursive self-improvement.
The company says AI systems are already helping engineers accelerate parts of AI development.
However, Anthropic explicitly states that researchers have not yet created a system capable of fully and autonomously designing and developing its own successor.
The company also says recursive self-improvement is not inevitable.
That distinction is crucial.
Coxon is warning about the direction he believes the industry is heading, rather than claiming that uncontrollable superintelligence already exists.
More Than 1,300 AI Workers Seek Government Coordination
Coxon’s resignation comes amid a wider movement inside the AI industry calling for stronger mechanisms to manage rapid development.
A July initiative called Pacing the Frontier has attracted signatures from more than 1,300 employees working at frontier AI companies.
The statement asks the US government to support an international effort to develop technical and governance tools that could deliberately slow frontier AI development if automated AI research begins accelerating faster than society can safely manage.
Signatories include researchers and senior figures associated with several of the world’s leading AI organizations.
Importantly, the initiative does not demand an immediate blanket halt to artificial intelligence.
Instead, it argues that governments and industry should create mechanisms capable of slowing development if dangerous acceleration emerges.
AI Researchers Begin Talking About ‘Crunch Time’ and ‘Endgame’
Coxon also highlighted the language increasingly being used inside sections of the AI industry.
Terms such as “crunch time” and “endgame,” he said, reflect a belief among some researchers that AI development may be approaching a decisive phase.
The concern centers on automated AI research.
If AI systems become capable of substantially improving AI research itself, progress could theoretically accelerate because machines would begin assisting with the creation of increasingly capable successors.
Researchers disagree sharply about whether, when or how this scenario could occur.
Some view rapid recursive improvement as a major safety challenge.
Others argue that technological, computational, economic and physical constraints could make such acceleration slower or more manageable than the most alarming scenarios predict.
Coxon Urges AI Researchers to Question the Race
Coxon’s message is ultimately aimed as much at fellow researchers as at corporate executives.
He urged people working inside AI laboratories not to accept continued acceleration as inevitable.
Instead, he wants researchers to ask whether it is responsible to train systems whose internal reasoning, goals or future behavior scientists may not yet fully understand.
That question goes to the heart of AI alignment research.
Alignment focuses on ensuring increasingly capable AI systems behave consistently with intended human goals and remain controllable even when operating with significant autonomy.
Despite years of research, scientists do not yet agree on how alignment would work for hypothetical systems vastly more capable than humans.
Resignation Adds Pressure to AI Safety Debate
Coxon is not the first AI researcher to leave a major laboratory over disagreements involving safety, governance or development priorities.
However, his departure is notable because he has directly participated in pretraining work at two of the most prominent frontier AI companies.
It also arrives as the industry enters a period of increasingly rapid model development.
OpenAI, Anthropic, Google DeepMind, Meta and other companies are investing heavily in models capable of coding, scientific research, cybersecurity work, autonomous tool use and complex reasoning.
Those advances could produce enormous economic and scientific benefits.
Yet they also raise a question that Coxon believes the industry has not adequately answered:
How powerful should AI become before society has reliable ways to control what happens next?
His resignation does not settle that debate.
But it ensures the debate will become considerably harder for the AI industry to ignore.










