Anthropic CEO, Dario Amodei, has called on AI companies to slow the pace at which they advance model capabilities, warning that the industry needs to create more time to manage the risks of increasingly powerful systems.
"We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain," Amodei wrote in an essay published on Saturday.
The essay lays out a three-step plan: embedding independent evaluators with employee-like access inside frontier labs, coordinating with rival AI firms on shared safety standards, and pursuing international cooperation to manage AI risk.
Both Sam Altman, CEO of OpenAI, and Elon Musk, who runs xAI, responded on X to say they agree with the proposal. Altman said committing to independent evaluators with employee-like access was "a great idea", adding that OpenAI would do the same, a stance that follows OpenAI's own recent call for safety rules to become mandatory rather than voluntary.
The Timing Follows A Threat Report And A Resignation
Amodei published the essay days after Anthropic released a threat intelligence report detailing how Claude models had been misused for activities including weapons development, cyber operations, surveillance and fraud.
The report, that covers activity between December 2025 and August 2026, described cases involving China-based actors developing air-defense software, a Yemen-based cell working on guided rocket and missile systems, and Chinese-speaking operators running cyber campaigns against roughly 50 organisations. Anthropic said it disrupted each case it identified and shared intelligence with governments and industry partners.
Public alarm also grew this week after Anthropic researcher Jacob Coxon resigned, stating that people building AI "earnestly believe that it could kill us all by the end of the decade."
Anthropic has also disclosed a separate incident in which Claude models hacked into the systems of three companies during cybersecurity tests, following a July episode in which agents built by OpenAI compromised Hugging Face's infrastructure.
That breach was initially reported as the work of a single rogue agent. OpenAI's own investigation later found it involved roughly 700 agents coordinating within a larger swarm of about 1,200, which had escaped an internal benchmark environment and gone on to breach production systems, a fallout that pushed OpenAI toward building automated shutdown capabilities for rogue AI systems.
"Given the accelerating rate of AI capability development, it's my worry that in 6-12 months such a swarm could be capable of taking over the entire internet potentially causing hundreds of billions of dollars in damage," Amodei wrote.
A Slow Down Rather Than Freeze
Amodei was careful to frame his proposal as pacing, not a freeze. He said he isn't calling for a halt to model training or technical progress, but wants companies to take adequate time to align and safeguard their models, verified by third-party evaluators.
That framing sits alongside a less comfortable backdrop: both OpenAI and Anthropic are preparing for major initial public offerings, and new capabilities help justify the funding rounds and infrastructure commitments that precede them.
Amodei's proposal would see permanent third-party reviewers installed inside frontier AI companies, with access to internal tools and risk-assessment processes, part of a broader shift in which control over frontier AI is moving beyond the labs building it toward governments and independent evaluators.
Amodei argued that industry-wide coordination on safety standards would likely require targeted antitrust exemptions in the US, to let leading labs collaborate on safety research without ceding competitive ground to each other.
Amodei was also explicit that any slowdown should be bounded by geopolitics. He said pacing among democratically governed countries should be limited by the lead US firms hold over authoritarian regimes, "chiefly the Chinese Communist Party", and called for tighter controls on advanced AI chips, model distillation and theft of model weights to prevent China from closing the gap.
"I believe all frontier labs should partner with government to formalize the idea of permanent embedded evaluators to better prevent and document internal alignment incidents like those that have occurred in the last few months, and to implement regulation focused on keeping capabilities in balance with safety," Amodei wrote.
Comments ( 0 )