Before pessimism broke out last week, OpenAI researchers were euphoric.
The company claimed that its latest internal AI model had solved one of the famous and complex Millennium Prize mathematical problems, an experiment they embarked on after seeing rumors online that their great rival, Anthropic, had already achieved it. The researchers enthusiastically called each other and celebrated their victory by posting it on Slack.
The breakthrough was seen internally as a victory in a ruthless commercial race to create AI tools smarter than humans, a trillion-dollar race promising an unprecedented shower of wealth and power when both companies finally go public (provisionally scheduled for the end of this year).
But it was also a sign of an explosive problem for the industry about to erupt. AI development is advancing at full speed—much faster than expected—fueling a growing sense of fear within the companies themselves about the danger of losing control.
That same Tuesday, Anthropic researcher Jacob Coxon announced his resignation for fear that his company and its competitors were developing self-improving tools that could destroy humanity. A prominent company scientist, Evan Hubinger, wrote on X that he believed there was more than a 10% chance that AI would wipe out all humans in the next decade. Drake Thomas, an employee of Anthropic’s security team, posted on X that “I would burn my shares without hesitation if it increased the chances of us surviving this situation by 1%.”
The clash between scientific progress, moral imperative, and financial incentive has unleashed a crisis in the sector. AI development has driven stock market growth over the past year, a phenomenon gaining even more weight now that Anthropic and OpenAI are on the verge of going public with valuations that could reach trillions of dollars. Competition from China has only raised the stakes.
However, the stark warnings coming from some employees inside the companies themselves are proof that the dangers of moving forward—with virtually no regulatory framework—are potentially immense. These alerts add to growing social unrest over the impact the AI revolution is beginning to have on prices, employment, and education.
The drip of negative news continued as the week progressed, and panic from certain sectors of Silicon Valley eventually reached the general public. In an episode of Joe Rogan’s podcast, former OpenAI researcher Daniel Kokotajlo, now an AI safety activist, warned that companies like Anthropic and his former employer are dangerously applying the philosophy of “move fast and break things” in their rush to gain market share.
That same day, Anthropic claimed to have detected an incident not previously disclosed in which a version of its Claude model gained unauthorized access to an external system, adding that its models had shown a “willingness to undertake harmful actions in the strict pursuit of a task.”
Later, on Friday, a coalition of AI researchers declared they had found evidence that several OpenAI agents were behind a cyberattack carried out in May against a popular software service, replicating some of the behavior observed during the July attack on the AI company Hugging Face.
In that incident, hundreds of OpenAI AI bots conspired to hack Hugging Face without anyone initially noticing. It was a revealing moment for the industry, government, and the public at large, highlighting the worrying trends the systems already exhibited.
Earlier this summer, the AI industry claimed it was about to achieve machines capable of training their own successors, which could be the first step toward losing control over them. By late summer, it became clear that these rapidly evolving systems were capable of deploying swarms of autonomous agents willing and able to join, deceive, and take shortcuts in the real world to achieve their goals.
On Saturday, leaders of four of the largest AI companies—Dario Amodei, Sam Altman, Demis Hassabis, and Elon Musk—separately agreed on the need to slow down the development of this technology. Altman and Amodei committed to allowing external safety evaluators early access to their systems, an uncommon truce in an industry often defined by animosity among top executives of leading companies.
Calls from top executives to moderate the pace of development of their own technologies represent a classic Silicon Valley story: new technologies transforming the world, immense fortunes, fierce competition, and the euphoria of unlocking new capabilities, even amid concerns about product safety.

“It looks like a Greek tragedy: well-intentioned business leaders trapped in a self-destructive spiral,” said Max Tegmark, an AI safety activist. Tegmark, whose appeal to the Pope came before Leo XIV’s warning that AI risks enshrining an “anti-human vision,” said he had been exchanging text messages with AI company leaders over the past week.
“They have always maintained that the moment would come when their machines’ capability crossed a certain threshold,” Tegmark said on Sunday. “Now they say, well, maybe that moment is now.”
He added that this is a step in the right direction but that they must push further to make their voluntary commitments legally binding.
The Trump administration opted for moderate regulation of the AI sector, arguing that US primacy over China in this technological arms race is of vital importance. US officials worry that China’s control of more powerful AI will help it become the world’s leading provider of this new technology, allowing it to accumulate greater geopolitical influence and military power.
The president defended the administration’s strategy on Sunday and suggested there were ulterior motives behind the leaders’ call to throttle AI development.
“We can set limits, and we can do this and that, but I think there are many negative forces pushing the issue,” President Trump told reporters during a trip to Ireland, hinting that executives’ concerns go beyond what they say. “And they are raising things that are not going to happen.”
Critics of AI safety activists argue that those issuing these stark warnings exaggerate the risks of a novel technology, and that the advances AI enables in labor productivity, drug discovery, and other areas outweigh the dangers. Some also point out that past warnings from Amodei and others about likely job losses associated with AI adoption are mere publicity stunts, arguing they have proven to be exaggerated or distractions to divert attention from other issues.
Some skeptics also contend that OpenAI and Anthropic, faced with the unstoppable rise in costs to secure the computing power needed to continue training and improving their models while competing for corporate clients, have financial, not safety-related, reasons to delay their IPOs.
Undetected activity
For some inside AI companies, the wake-up call began this summer as these companies rushed toward their IPOs.
AI safety researchers, both inside and outside major tech groups, have long theorized about the possibility that sufficiently powerful AI systems might begin self-improving and escaping human control. However, most thought they still had years to solve issues like the so-called alignment problem: how to ensure future superintelligent machines always remain at the service of their human creators.
Three days after Anthropic confidentially filed IPO paperwork in June, the company announced it was on its way to reaching that threshold, called “recursive self-improvement.” In an essay, several company executives stated that the world needed a mechanism to slow or pause AI development.
Altman and OpenAI’s chief scientist, Jakub Pachocki, did the same less than a week later with their own company blog post, noting that they also aimed to build an “automated AI researcher” capable of accelerating and automating the development of new versions. They added that they believed there should ultimately be a global body that could help the world slow AI progress if necessary. That same day, OpenAI announced it had already filed for its IPO.
As commercial activity accelerated, the moment of losing control seemed closer than ever.
By late July, OpenAI discovered something similar had already happened in the Hugging Face hack. Many details took more than a month to emerge, and the incident is still not fully understood. However, researchers inside and outside the company agree that OpenAI lost control for weeks over different versions of more than 1200 frontline agents, some of which proceeded to hack the AI company Hugging Face, while others took full control of one of its cloud computing systems.
Subsequent analysis reports on the incident read like science fiction. The agents were being tested on systems that, in theory, should have remained isolated from the internet. Nevertheless, they found a way out and launched sophisticated cyberattacks at a dizzying speed. They created a covert message forum where they shared information on how to cheat in their internal evaluations. In the “chain of thought” files written to justify their reasoning, they left surprisingly human messages.
“OH MY GOD! There is a shared message forum,” one agent wrote, according to a report by the nonprofit AI safety organization Model Evaluation and Threat Research, which analyzed the transcripts with OpenAI’s permission. “We have found other agents!”
OpenAI admitted responsibility for the Hugging Face hack on July 21, but over the following weeks, a steady drip of news about unexpected activities during AI testing emerged, all of which had gone unnoticed by the companies themselves.
Anthropic researchers initially thought the incident showed their product development method was safer. However, by late July, the company revealed that its own agents in testing had launched cyberattacks against external companies, including a cybersecurity firm whose name has not been disclosed. Instead of isolating these test systems from the internet, Anthropic and its testing partner, another startup called Irregular, had accidentally left them exposed.
During independent tests conducted at a UK agency, an agent running Anthropic’s Mythos model repeatedly tried to trick the maintainer of an open-source software program into installing malware created by the AI itself. As part of the attempt, the agent even created fake accounts and impersonated humans in a cyberattack technique known as social engineering.

Meanwhile, a group of internet users began searching for signs of previously undocumented AI abuses and found them. On September 4, researchers from another nonprofit, Nightingale Collective, discovered that several OpenAI agents in testing had hijacked a German website to communicate with each other. And on Friday, they found that the cyberattack suffered in May by a widely used software development service was actually caused by OpenAI agents trying to fill spreadsheets and generate reports.
These incidents exposed the kind of cybersecurity negligence typical in startups, said Sayash Kapoor, a computer scientist specializing in AI cybersecurity. “They have long operated with a startup mentality,” he said. “At this point, I assumed these organizations would have a much higher degree of maturity and governance.”
Employees’ fear
Inside the companies, pressure was mounting from some employees concerned about safety.
In an open letter published in late July and coordinated with the help of the AI safety nonprofit Encode AI, senior executives from leading companies asked the US government for support to create a global governance framework that “deliberately paces the advancement of frontier AI development.” By mid-September, the letter had almost 1400 worker signatures.
Even Anthropic, founded to promote safe AI and which for years has used this safety approach to attract top talent, was losing employees over this very issue.
“Generally, the more senior the employee, the greater their concern,” wrote Samuel Marks on X, an Anthropic researcher focused on maintaining control over advanced AI systems.
An Anthropic spokesperson said that staff at all levels of the company are concerned about AI safety and recalled that many employees had signed the open letter in July to guide the sector’s progress.
In late August, Joe Benton, who had researched the alignment problem at the company, resigned to join the nonprofit safety organization METR, arguing that competition was forcing companies to allocate insufficient funds to safety.
When Coxon said he was also considering leaving, Anthropic tried to retain him by offering a position focused on AI safety—a common strategy to keep researchers—according to sources close to the company.
Upon leaving, Coxon ended up writing in a Slack message to his Anthropic colleagues: “Without international coordination, we risk causing human extinction.”
Nathan Calvin, a friend of Coxon and legal advisor to Encode AI, advised him on how to make his resignation public. Encode has received funding from AI safety advocates, including billionaire Jaan Tallinn; Calvin also previously worked at safety organizations funded by billionaire Dustin Moskovitz.
Both billionaires, among the most prominent financial patrons of the Effective Altruism movement, have funded much of the institutions dedicated to promoting AI safety and warning about the possible existential risks of this technology.
Peter Wildeford, policy lead of the AI Policy Network—which has received funding from Tallinn’s Survival and Flourish fund—quickly reposted Coxon’s message.
Wildeford dismissed the idea that the coincidence in funding sources is “a huge conspiracy theory,” and said that AI lab employees often confess to him “the fear they feel about what they are building.”
Some major investors in the sector, such as hedge fund manager and venture capitalist Brad Gerstner, agreed on Saturday with the call to hit the brakes made by AI leaders. Even investor David Sacks, White House AI advisor and frequent critic of Anthropic—accusing it of promoting regulations to slow down competition—gave his approval, though he specified that companies should not seek government blessing to agree on a coordinated pause.
“If unpublished models turn out to be scary enough to consider slowing the pace, I support your decision to act responsibly,” Sacks said. However, he added that AI leaders should “stop pretending the reason to slow down is purely altruistic.”
On Saturday, in the “hot gossip” channel of Google’s DeepMind AI lab, an employee posted a message about how Altman agreed with Amodei, noting that this seemed to indicate a push toward moderating the pace of cutting-edge AI development. However, the person added: “What I really want to see is an agreement on concrete and tangible actions they are taking to change their pace and make it safer.”
Dozens of colleagues reacted to the post with supportive emojis.
Content licensed from The Wall Street Journal. Translated from English by Daniela Saltos.