Anthropic says it has disrupted several attempts to misuse its Claude AI models, including a suspected Russia-linked cyber-espionage campaign as well as efforts by Chinese AI firms to extract and replicate the models' capabilities.
Reuters reports that Anthropic identified several uses of Claude over the past eight months that they suspected were malicious, as cybercriminals and state-backed actors increasingly used AI to automate and execute significant portions of cyber-related operations.
According to Anthropic's latest Threat Intelligence report, the activity went beyond typical chatbot interactions. Threat actors used multi-agent structures to carry out task sequences, with humans increasingly acting as overseers rather than performing the operations themselves.
AI Used Across Russian-Linked Cyber Campaign
One campaign was linked by Anthropic to activity consistent with Midnight Blizzard, a known Russia-based threat actor. Previously, the US government had linked this very same hacker to the SVR, which is the Foreign Intelligence Service of the Russian Federation.
This group has allegedly been targeting Ukrainian government, military and diplomatic organisations, using Claude across several stages of organised criminal operations. These included phishing, hotel Wi-Fi hijacking and WhatsApp account takeovers.
Additionally, Anthropic said the attackers used AI to build a system that could automatically detect when its malware had been identified by security defences and then take action by rewriting the code to evade detection.
It was said that the campaign demonstrated how AI can move beyond merely providing assistance to becoming an active component in carrying out a cyberattack.
Chinese AI Firms Accused of Model Distillation
Anthropic also said it had disrupted activity involving seven China-based AI labs that it accused of attempting to extract capabilities from Claude. This is known as model distillation, where a smaller, more efficient AI model is used to mimic the behaviour and outputs of a larger, more powerful model.
According to Anthropic, some of the companies linked to this included Alibaba, Moonshot, DeepSeek and Xiaomi. Alibaba, in particular, carried out what it described as the largest illicit distillation attack it had observed. The activity allegedly sought to extract Claude's capabilities and use them to improve Alibaba's internal Qwen models.
Anthropic said it observed an astonishing number of exchanges (more than 151 million) that were attributed to Alibaba between May and July 2026. These exchanges involved more than 3,500 accounts that it described as fraudulent. Activity peaked at almost three million exchanges every single day.
It was also alleged that Moonshot and DeepSeek routed live customer conversations through Claude and used those responses as a means to train data. The company said some of those conversations could have contained sensitive information.
The companies named in Anthropic's report did not immediately respond to requests for comments by Reuters.
AI Misuse Extends Beyond Cyberattacks
Anthropic also identified what it described as new categories of threat actors misusing Claude. These included operators using the model to develop software associated with conventional weapons, including firearms, missiles, armed drones and other munitions.
Additional activity involving weapons design, intelligence gathering and procurement were linked to programmes in China, Russia and Yemen has been detected. The company also said it had disrupted activity associated with affiliates of the ShinyHunters cybercrime collective, which has been linked to attacks against major corporations.
Jacob Klein, Anthropic's head of threat intelligence, told Reuters that the increasing capabilities of AI models were creating new, increasingly worrying risks. As models become more capable of handling complex technical tasks, the potential consequences of their misuse also increase, he said.
AI Security Is Becoming an Infrastructure Problem
The incidents highlight how the security challenge around AI is changing as models become more capable and increasingly connected to other systems.
For enterprises, the issue is no longer simply whether employees or attackers can use an AI model to generate malicious content. AI systems can increasingly be incorporated into workflows that interact with software, data and external tools.
That creates a broader question around how organisations monitor what AI systems are doing once they move beyond simple question-and-answer interactions and whether appropriate guardrails are being put in place.
In an effort to limit the damage these campaigns had done or attempted to do, Anthropic banned the accounts associated with the activity and introduced additional safeguards and monitoring measures.
The latest incidents suggest that as AI becomes more autonomous, securing the technology will increasingly depend on understanding not only what models can generate, but what they can help orchestrate when connected to the wider enterprise environment.
Comments ( 0 )