U.S. cybersecurity and intelligence agencies have accused six China-based artificial intelligence companies of conducting coordinated, industrial-scale campaigns to extract advanced capabilities from American frontier models, escalating a long-running dispute over model distillation into a national security confrontation.
In a joint cybersecurity advisory, the Cybersecurity and Infrastructure Security Agency (CISA), National Security Agency (NSA) and Federal Bureau of Investigation (FBI) named DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI as participants in systematic efforts to obtain proprietary functionality from models developed by Anthropic, OpenAI, Google and xAI.
The agencies allege that the companies extracted billions of tokens through millions of automated interactions beginning in at least late 2024. According to the advisory, the campaigns were not limited experiments or routine research projects but formed a central part of the Chinese companies’ model-development strategies.
U.S. officials further assessed that the operations probably took place with the awareness of the Chinese government. The advisory does not publicly present evidence establishing direct government tasking, however, and the named companies have not all provided detailed responses to the specific allegations.
China has rejected the accusations. Its Commerce Ministry described distillation as a widely used and technically neutral development method, accused Washington of applying double standards and said some U.S. companies had also distilled Chinese models. Beijing warned that it could respond if the allegations were used to justify additional restrictions against Chinese AI businesses, according to Reuters.
The dispute draws an important distinction between conventional knowledge distillation—a legitimate and widely used machine-learning technique—and campaigns designed to evade access controls, violate service conditions and reproduce restricted capabilities through enormous volumes of model interactions.
What model distillation means
Knowledge distillation normally involves using a highly capable “teacher” model to help train a smaller or more efficient “student” system. Developers can submit prompts to the teacher, collect its answers and use those responses as training examples.
The technique is not inherently malicious. AI researchers and vendors use authorized distillation to reduce operating costs, improve smaller models and transfer selected capabilities to systems designed for specific applications or devices.
The controversy concerns how access is obtained, how much data is collected and whether the operator is deliberately circumventing contractual, geographic and technical restrictions.
A normal user might submit occasional prompts to an AI assistant for research, writing or coding. A distillation campaign can instead generate millions of carefully structured queries intended to map and reproduce the model’s behaviour. These queries may target reasoning methods, programming abilities, tool use, mathematical problem-solving, specialised professional knowledge and agentic functions.
By repeatedly requesting outputs, evaluating their quality and feeding successful examples into another training pipeline, an operator can potentially shorten development cycles and reduce some of the cost associated with independently creating comparable capabilities.
The U.S. agencies allege that the Chinese campaigns went substantially further than ordinary use. The operators allegedly built automated infrastructure, distributed requests across large pools of accounts and access providers, and sought to conceal the origin and coordinated nature of the traffic.
Billions of tokens collected through distributed infrastructure
CISA, the NSA and FBI said the campaigns reached U.S. systems through several pathways, including official application programming interfaces, remote cloud providers and third-party model aggregators.
The operators allegedly used proxy services known as “transfer stations” to route requests through infrastructure outside mainland China. These intermediary services can conceal customers’ locations, provide access to products unavailable in their home markets and make activity appear to originate from multiple unrelated users.
Some transfer stations reportedly acquire large numbers of legitimate subscriptions or API accounts and resell access through a common interface. This arrangement allows a customer to distribute automated requests across different accounts, services and countries while making central coordination more difficult for an individual AI provider to detect.
The campaigns also allegedly relied on bulk purchases of premium subscriptions shared among teams of developers. Accounts could be operated continuously until they reached usage limits, after which an automated routing system would move traffic to another subscription, provider or access channel.
That approach presents a detection problem. Each account may generate only a fraction of the overall collection activity, while the complete pattern becomes visible only when events are correlated across accounts, cloud platforms, payment systems, API aggregators and model providers.
The advisory says more advanced operations used automatic failover mechanisms that redirected queries when accounts were suspended or pathways blocked. Operators also removed or altered identifying metadata and employed evaluation systems designed to identify whether providers had changed responses as a defensive measure.
DeepSeek allegedly targeted reasoning and agentic capabilities
The U.S. agencies alleged that DeepSeek carried out organised distillation activity between late 2024 and the middle of 2025 to assist development of its R1 and V3 models.
The activity reportedly targeted reasoning, writing, question answering, agentic functions and specialised capabilities, including tasks associated with legal work. Rather than attempting to copy a model as a single undifferentiated system, operators could construct datasets focused on capabilities where leading models perform particularly well.
DeepSeek previously attracted widespread attention by reporting that the final training run for one of its models cost approximately $5.6 million. The U.S. advisory argues that such a figure does not represent the full economic cost of developing the system because it excludes precursor research, infrastructure and the value of data and capabilities obtained through external models.
That does not mean every capability in a distilled system was copied, or that reported training-run costs are necessarily false. Training-cost disclosures frequently cover a defined phase rather than the entire research programme. The agencies’ broader argument is that extensive access to high-quality outputs can provide substantial development advantages that are not reflected in the cost of the final training run.
Moonshot AI linked to Claude and GPT extraction
Moonshot AI is accused of conducting distillation operations from at least the middle of 2025 to improve its Kimi model family.
According to the advisory, the company extracted data from Anthropic’s Claude Fable 5 for development of Kimi-K3 and used GPT-4o outputs in work associated with Kimi-K2. The targeted areas reportedly included mathematics, software engineering, supervised fine-tuning and reinforcement-learning capabilities.
Supervised fine-tuning uses labelled examples to teach a model how to respond to particular inputs, while reinforcement learning uses feedback signals to encourage more desirable behaviour. High-quality responses obtained from a capable external model can be used in both processes: directly as training examples or indirectly as a benchmark for ranking and improving outputs.
Alibaba accused of targeting coding and customer-service functions
The advisory alleges that Alibaba distilled capabilities from several Claude and GPT variants during late 2025.
The operation reportedly sought improvements in software engineering, customer-service dialogue, image and character creation, as well as the integration of reinforcement learning, supervised fine-tuning and model-distillation processes.
Separate allegations against Alibaba surfaced earlier in 2026. Anthropic accused operators affiliated with Alibaba and its Qwen research organisation of conducting what it characterised as the largest distillation campaign the company had identified at that time.
That campaign allegedly generated more than 28.8 million exchanges through nearly 25,000 fraudulent accounts between April 22 and June 5, according to a company letter reported by Reuters.
MiniMax, StepFun and Z.AI also named
MiniMax allegedly collected chain-of-thought reasoning, reinforcement-learning, fine-tuning and software-engineering data from Claude Code, Claude Sonnet 4, Claude Opus and multiple generations of Google’s Gemini models. The advisory links this activity to improvements in the company’s M2 model in late 2025.
StepFun is accused of drawing on numerous Claude and GPT products between late 2025 and early 2026 to strengthen coding and agentic capabilities in its Step 4 model.
The models identified by the agencies include Claude Opus 4.1 and 4.5, Claude Sonnet 4.5, Claude Haiku 4.5, GPT-5 Mini, GPT-5 Pro and several GPT-5.x and Codex variants.
Z.AI, formerly widely known as Zhipu AI, allegedly distilled billions of tokens from GPT-5.5 and Claude Opus 4.8 by the middle of 2026. The collection was reportedly focused on developing chain-of-thought reasoning capabilities.
Chain-of-thought extraction is particularly sensitive because the objective is not simply to reproduce a final answer. Carefully designed prompts may attempt to obtain intermediate reasoning, problem-solving structures and reusable methods that can help another model perform complex tasks.
Providers have increasingly restricted access to internal reasoning traces, instead returning answers or condensed explanations. Determined operators may nevertheless use prompt variation, multi-stage questions and large-scale evaluation to approximate useful reasoning patterns from observable outputs.
Behaviour that could expose a distillation campaign
The advisory describes several signals that AI providers and cloud operators can use to distinguish large-scale collection from ordinary customer activity.
One is continuous usage with few or no natural pauses. Human activity typically changes according to working hours, sleep patterns, weekends and local time zones. An account producing high-volume traffic 24 hours a day may be automated or shared across a large group.
Another warning sign is a subscription account generating throughput that more closely resembles an enterprise API deployment. Newly created subscriptions that rapidly consume their full quotas can also indicate accounts being created for short-lived collection.
Providers should investigate groups of accounts that submit similar prompts, use related infrastructure or change identifiers in a coordinated way. Abrupt alterations in user-agent strings, network locations, payment information or other metadata can indicate an attempt to evade detection.
The difficulty is that no single indicator proves malicious distillation. Researchers, software developers and legitimate enterprises can generate high-volume or highly repetitive traffic. Defenders therefore need to combine technical telemetry, account history, identity information, payment data and behavioural analysis rather than relying on simple request thresholds.
Agencies recommend stronger identity and account controls
The U.S. government is urging AI companies to strengthen identity verification, particularly for accounts seeking high-volume access or displaying enterprise-scale behaviour.
Providers can restrict the sharing of premium subscriptions, impose more granular rate limits and require additional verification when usage suddenly increases. Cloud companies and API aggregators can also analyse whether multiple nominally independent customers are sending coordinated requests through shared infrastructure.
Cross-platform cooperation is likely to be critical. An operator blocked by one model provider can move to another access channel, use a third-party aggregator or rotate through fresh accounts. Isolated enforcement decisions may therefore displace the activity without disrupting the wider campaign.
The agencies recommend correlating signals between frontier-model developers, cloud providers and intermediaries while observing privacy, contractual and legal requirements. Shared indicators could include known proxy infrastructure, coordinated account characteristics and distinctive automation patterns.
Defensive responses may alter what suspected operators receive
The advisory also proposes measures intended to reduce the value of collected outputs.
A provider could limit reasoning depth, vary the form of otherwise correct answers or use different reasoning routes for suspicious accounts. It could also quietly route suspected collectors to a less capable model, making the resulting training dataset less consistent or less valuable.
For activity assessed to be malicious, a company may alter responses without notifying the operator. Such measures must be deployed cautiously: undisclosed model changes can affect legitimate research, safety testing and third-party evaluation. The advisory consequently recommends informing authorised AI-safety researchers and evaluators when defensive changes could influence their results.
Differential privacy is another possible defence. It introduces controlled statistical noise to reduce the amount of information that can be extracted about a system or its training data. Stronger privacy settings, however, can reduce accuracy or usefulness, meaning providers may have to combine them with monitoring, response controls and rate limits.
CISA Acting Director Nick Andersen urged AI companies to act quickly, warning that industrial-scale distillation could narrow the capability gap between American developers and overseas competitors. The NSA confirmed that it jointly issued the advisory with CISA and the FBI on September 8.
From commercial abuse to national security concern
The advisory frames distillation as more than a contractual or intellectual-property dispute. U.S. officials argue that copied reasoning, coding and agentic capabilities could strengthen military systems, intelligence programmes and offensive cyber operations.
Advanced models can assist with vulnerability research, malware analysis, software development, intelligence processing and the coordination of automated tools. A country that accelerates model development through large-scale extraction could, in the agencies’ assessment, obtain strategically important capabilities while avoiding part of the cost and time required for independent research.
The government has not publicly demonstrated that every distilled capability cited in the advisory has been deployed in Chinese military or cyber operations. Its warning instead describes a risk pathway: commercial model extraction can accelerate domestic AI development, and increasingly powerful domestic systems may then become available to state security and defence organisations.
The accusations also arrive during a sensitive period in U.S.-China relations. Washington and Beijing are preparing for discussions about AI safety and a planned meeting between President Donald Trump and Chinese President Xi Jinping. Reuters reported that Chinese officials regard the advisory as part of a broader U.S. effort to constrain China’s technological development.
A difficult enforcement boundary
The case exposes a fundamental problem for the AI industry. Model outputs are products delivered to users, but they can also become training material for competitors. Unlike traditional source-code theft, distillation does not necessarily require breaching the provider’s internal network or stealing model weights.
Instead, an operator can abuse the interface the company intentionally provides.
That makes the legal and technical boundary complicated. Distillation may be authorised, prohibited by service conditions or conducted through fraudulent access depending on the circumstances. Questions remain over which elements of a model’s behaviour qualify for intellectual-property protection and how existing laws apply when millions of individually generated outputs are assembled into a competing training dataset.
China maintains that the United States is attempting to transform a common research technique into a political weapon. The U.S. position is that the scale, concealment, access-control evasion and targeted extraction described in the advisory place the activity far outside legitimate research.
Regardless of how the dispute develops diplomatically, the security implications extend beyond the six companies named. API credentials, premium subscriptions and model access accounts now have strategic value because they can provide entry to expensive, difficult-to-reproduce capabilities.
For AI developers, preventing model extraction will require the same kind of layered approach already used against payment fraud, credential stuffing and automated abuse: stronger identity controls, behavioural monitoring, infrastructure intelligence, cross-provider cooperation and carefully designed technical countermeasures.
The government’s warning effectively recasts frontier-model security as a form of critical technology protection. The central challenge is no longer limited to keeping model weights and source code inside secured environments. Providers must also defend the public-facing behaviour of their systems against adversaries capable of collecting, analysing and reproducing it at industrial scale.
Click Here For The Original Source.
