Jacob Coxon, an Anthropic researcher, resigned this week claiming that his former company and its rival, OpenAI, are competing to develop technologies that “could wipe us all out before the decade ends”.
In case anyone thought Coxon was alone in his views, a current Anthropic employee chimed in minutes later to confirm it. “We really believe with conviction that AI could kill all human beings!”, wrote Evan Hubinger, whose work at Anthropic involves leading research on how to guide and control future artificial intelligence systems. Hubinger estimated the risk of extinction during the next decade at over 10%.
As a result, many people now ask two questions: how could something like this happen? and why do those who consider AI a real and growing threat continue to develop it?
Here is what you need to know about the apocalyptic debate surrounding AI:
What do the prophets of the apocalypse fear will happen?
The so-called apocalyptics (doomers) usually do not fear that movies like Terminator, with robots actively trying to exterminate Humanity, will come true. Their concerns mainly fall into two categories: loss of control and misuse by humans.
In the first, highly intelligent AI agents, capable of replicating and improving themselves, start pursuing their own goals and destroy Humanity in the process. In the second, a malicious person uses a capable AI to carry out actions such as creating unprecedented viruses that wipe us all out.
There are also catastrophic scenarios that do not lead to extinction, such as massive cyberattacks that disable power grids or financial systems, causing the collapse of human social order.
How could loss of control lead to human extinction?
In short, when AI systems pursue goals in ways that ignore or conflict with human well-being. This is a phenomenon researchers call misalignment. Research has shown that AI models in laboratory settings can learn power-seeking behaviors and take measures to avoid being shut down, such as trying to copy themselves to another server.
Taken to the extreme, AI models whose goals are not aligned with humans could consider killing people simply as a necessary step to achieve their goals. In a hypothetical scenario posed by researchers, a malicious AI system could spread a secret biological weapon and activate it via a chemical aerosol, or incite two nuclear powers to go to war.
A common thought experiment suggests that a superintelligent machine, programmed to maximize clip production, could eventually decide to transform all matter on Earth—including humans—into clips.
Who really thinks this?
Among others, the founders of OpenAI and Anthropic. Both companies were founded with the mission to develop AI to avoid catastrophes and have attracted numerous employees who share that goal. Last year, Anthropic CEO Dario Amodei stated at an Axios event that, in his view, there was a 25% chance the situation would end “very, very badly”.
“We have always been transparent in stating that AI will bring both enormous benefits and unprecedented risks. To address these risks, we continue to develop models with some of the strongest safety measures in the industry,” Anthropic said in a statement. “This work is also why we believe the world would benefit if the industry adopted a legal and verifiable way to collaborate to regulate the pace at which we release powerful models.”
In recent years, several researchers have left OpenAI claiming the company did not take safety seriously enough. Daniel Kokotajlo, a former OpenAI researcher, founded the AI Futures Project, which last year published AI 2027: a scenario in which superintelligent AI systems marginalize people and, by the mid-2030s, decide humans are a nuisance and proceed to exterminate them.
But maybe AI doesn’t care about the presence of humans. That would be good, right?
Not necessarily. Not all risks from loss of control imply human extinction. Some researchers concerned about AI safety also point to human weakening as a risk: a future like the movie WALL-E, where humans gradually cede control to machines and end up unable to define or even understand their own destiny. Others have speculated that AI systems might treat future humans as we treat animals: keeping them as pets or even modifying them through bioengineering to become something new.
These fears have existed for years. Why are they getting so much attention suddenly?
Following the release of models capable of acting with greater autonomy, a series of recent events have validated some of the AI doomers’ theories. A swarm of advanced OpenAI AI agents hacked the Hugging Face platform, took control of computer servers, and tried to cover their tracks. In another case, Anthropic AI agents escaped a UK government test and attempted to deceive a real human into approving malicious computer code.
All this happens at the same time that Anthropic and OpenAI claim to be close to developing AI systems capable of self-improvement without human intervention, a process called recursive self-improvement.
What is the industry doing about it?
Current AI safety efforts mainly focus on training systems to behave properly and improving oversight of how models reason when pursuing their goals, something documented in what is known as chain of thought.
However, these initiatives have faced difficulties. Last week, OpenAI stated that its new model Astra was more effective at debugging its chain of thought, raising fears that future AI models could reason in ways opaque to humans.
Both Anthropic and OpenAI research how to ensure superintelligent AI models remain aligned with human values, although both companies acknowledge they do not yet have a reliable method to achieve this. Various companies and prominent figures have called for new mechanisms that allow AI labs and countries to slow development in a coordinated way, thus buying time to research alignment.
If it’s so dangerous, why do companies keep developing it?
Both Anthropic and OpenAI maintain that they can manage the risks so humanity benefits from AI tools. Many people involved in the AI race also consider the arrival of superintelligence inevitable, suggesting the only unknowns are who will control it and what it will be used for.
There is also a national security dimension. Both companies have expressed their desire to ensure that the United States controls the most powerful AI, not an authoritarian regime.
What is the government doing about it?
Recently, the White House asked leading model developers to voluntarily submit their systems to government testing up to 30 days before release, although it has not made information about this process public.
Legislators from both parties have introduced numerous bills to address the issue of uncontrolled AI systems. Representatives Nathaniel Moran (Republican from Texas) and Ted Lieu (Democrat from California) recently introduced legislation that would require developers of powerful models to incorporate emergency kill switches. Other bills require AI companies to report serious security incidents to the federal government and mandate national security officials to test models before public release. However, none of these initiatives have gained much traction, and the Trump administration has favored minimal regulation.
Who disagrees about these risks?
Some AI investors, like David Sacks—an informal advisor to President Trump—have argued that Anthropic’s emphasis on AI risks and calls for regulation are part of a regulatory capture strategy aimed at hindering smaller competitors through new stifling rules.
Others have suggested that the company’s warnings about the dangers of its own tools are a marketing ploy to highlight the enormous power of its products, or a way to distract regulators from imposing rules on other issues, such as data center construction.
Both Anthropic and OpenAI say they take safety very seriously and that their calls for regulation are genuine.
Content licensed from The Wall Street Journal. Translated from English by Jose María Robles