"These labs generated over 16 million exchanges with Claude through approximately 24,000 fraudulent accounts, in violation of our terms of service and regional access restrictions."
抓到的是活的,不是事后复盘
"We detected this campaign while it was still active—before MiniMax released the model it was training."
"We attributed the campaign to MiniMax through request metadata and infrastructure indicators, and confirmed timings against their public product roadmap."
这是全文最有分量的一句。以往这类指控都是模型发布后倒推,这次拿到了从数据生成到模型上线的完整链条。
对手 24 小时内就掉头
"When we released a new model during MiniMax's active campaign, they pivoted within 24 hours, redirecting nearly half their traffic to capture capabilities from our latest system."
这个细节比总量更能说明性质:不是零散抓取,是有人盯着发布节奏在调度流量。
封号没用,因为没有单点
"The breadth of these networks means that there are no single points of failure. When one account is banned, a new one takes its place."
"These services run what we call “hydra cluster” architectures: sprawling networks of fraudulent accounts that distribute traffic across our API as well as third-party cloud platforms."
九头蛇集群是商业代理服务在运营的,不是三家实验室自己搭的——责任链上多了一环转售方。
识别靠的是模式,不是单条内容
"What distinguishes a distillation attack from normal usage is the pattern."
"A prompt like the following (which approximates similar prompts we have seen used repetitively and at scale) may seem benign on its own."
单条 prompt 完全无害,几万条变体全指向同一个狭窄能力就不是了。这也解释了为什么防御难做。
论证的落点是出口管制
"Distillation attacks therefore reinforce the rationale for export controls: restricted chip access limits both direct model training and the scale of illicit distillation."
"Without visibility into these attacks, the apparently rapid advancements made by these labs are incorrectly taken as evidence that export controls are ineffective."
全文真正的政策主张在这里:把「中国实验室进步快」重新解释成「他们在抽取美国模型」,从而反证管制有效。
We have identified industrial-scale campaigns by three AI laboratories—DeepSeek, Moonshot, and MiniMax—to illicitly extract Claude's capabilities to improve their own models.
我们发现了三家 AI 实验室——DeepSeek、Moonshot、MiniMax——发起的工业级行动,它们非法提取 Claude 的能力来改进自家模型。
These labs generated over 16 million exchanges with Claude through approximately 24,000 fraudulent accounts, in violation of our terms of service and regional access restrictions.
这些实验室通过约 24,000 个欺诈账号,与 Claude 产生了超过 1600 万次对话,违反了我们的服务条款和区域访问限制。
These labs used a technique called “distillation,” which involves training a less capable model on the outputs of a stronger one.
这些实验室用的技术叫“蒸馏”:拿更强模型的输出去训练一个能力较弱的模型。
Distillation is a widely used and legitimate training method.
蒸馏本身是一种被广泛使用的正当训练方法。
For example, frontier AI labs routinely distill their own models to create smaller, cheaper versions for their customers.
比如,前沿 AI 实验室经常蒸馏自家模型,为客户做出更小、更便宜的版本。
But distillation can also be used for illicit purposes: competitors can use it to acquire powerful capabilities from other labs in a fraction of the time, and at a fraction of the cost, that it would take to develop them independently.
但蒸馏也可以用于非法目的:竞争对手能靠它从别的实验室拿走强大能力,所花的时间和成本都只是自主研发的零头。
These campaigns are growing in intensity and sophistication.
这类行动的强度和复杂度都在上升。
The window to act is narrow, and the threat extends beyond any single company or region.
留给行动的窗口很窄,而威胁的范围超出任何单一公司或地区。
Addressing it will require rapid, coordinated action among industry players, policymakers, and the global AI community.
要应对它,需要行业参与者、政策制定者与全球 AI 社区快速协同行动。
Illicitly distilled models lack necessary safeguards, creating significant national security risks.
非法蒸馏出来的模型缺少必要的安全防护,会带来重大的国家安全风险。
Anthropic and other US companies build systems that prevent state and non-state actors from using AI to, for example, develop bioweapons or carry out malicious cyber activities.
Anthropic 和其他美国公司会建立机制,防止国家级与非国家级行为体利用 AI 去做诸如研制生物武器、发动恶意网络活动这类事。
Models built through illicit distillation are unlikely to retain those safeguards, meaning that dangerous capabilities can proliferate with many protections stripped out entirely.
通过非法蒸馏造出来的模型不太可能保留这些防护,意味着危险能力会在防护被整块剥掉的状态下扩散。
Foreign labs that distill American models can then feed these unprotected capabilities into military, intelligence, and surveillance systems—enabling authoritarian governments to deploy frontier AI for offensive cyber operations, disinformation campaigns, and mass surveillance.
蒸馏了美国模型的外国实验室,可以把这些没有防护的能力送进军事、情报和监控系统——让威权政府把前沿 AI 用在攻击性网络行动、虚假信息宣传和大规模监控上。
If distilled models are open-sourced, this risk multiplies as these capabilities spread freely beyond any single government's control.
如果蒸馏出的模型被开源,风险会成倍放大,因为这些能力会自由扩散,超出任何一个政府的控制。
Anthropic has consistently supported export controls to help maintain America's lead in AI.
Anthropic 一贯支持出口管制,以帮助美国保持在 AI 上的领先。
Distillation attacks undermine those controls by allowing foreign labs, including those subject to the control of the Chinese Communist Party, to close the competitive advantage that export controls are designed to preserve through other means.
蒸馏攻击削弱了这些管制:它让外国实验室——包括受中国共产党控制的那些——用别的手段追平出口管制本想守住的竞争优势。
Without visibility into these attacks, the apparently rapid advancements made by these labs are incorrectly taken as evidence that export controls are ineffective and able to be circumvented by innovation.
如果看不见这些攻击,这些实验室表面上的高速进步就会被错误地当成证据,用来说明出口管制无效、可以靠创新绕过。
In reality, these advancements depend in significant part on capabilities extracted from American models, and executing this extraction at scale requires access to advanced chips.
实际上,这些进步在很大程度上依赖从美国模型里提取的能力,而大规模执行这种提取本身就需要拿到先进芯片。
Distillation attacks therefore reinforce the rationale for export controls: restricted chip access limits both direct model training and the scale of illicit distillation.
所以蒸馏攻击反而强化了出口管制的理由:芯片受限,既限制直接训练模型,也限制非法蒸馏能做到多大规模。
The three distillation campaigns detailed below followed a similar playbook, using fraudulent accounts and proxy services to access Claude at scale while evading detection.
下面详述的三场蒸馏行动用的是同一套打法:靠欺诈账号和代理服务大规模访问 Claude,同时躲避检测。
The volume, structure, and focus of the prompts were distinct from normal usage patterns, reflecting deliberate capability extraction rather than legitimate use.
这些 prompt 的体量、结构和指向都与正常使用模式明显不同,反映的是蓄意的能力提取,而非正当使用。
We attributed each campaign to a specific lab with high confidence through IP address correlation, request metadata, infrastructure indicators, and in some cases corroboration from industry partners who observed the same actors and behaviors on their platforms.
我们把每一场行动都高置信度地归因到了具体实验室,依据是 IP 地址关联、请求元数据、基础设施指标,部分情况下还有同行伙伴的印证——他们在自己平台上观察到了同样的行为体和行为。
Each campaign targeted Claude's most differentiated capabilities: agentic reasoning, tool use, and coding.
每一场行动瞄准的都是 Claude 差异化最强的能力:智能体推理、工具使用和编程。
Scale: Over 150,000 exchanges
规模:超过 150,000 次对话
The operation targeted:
该行动瞄准:
DeepSeek generated synchronized traffic across accounts.
DeepSeek 在多个账号之间制造了同步流量。
Identical patterns, shared payment methods, and coordinated timing suggested “load balancing” to increase throughput, improve reliability, and avoid detection.
一致的模式、共用的支付方式和协同的时间安排,指向一种“负载均衡”做法:提高吞吐、增强稳定性、并避开检测。
In one notable technique, their prompts asked Claude to imagine and articulate the internal reasoning behind a completed response and write it out step by step—effectively generating chain-of-thought training data at scale.
有一种手法值得注意:他们的 prompt 让 Claude 设想并说出某个已完成回答背后的内部推理,一步步写出来——实际上就是在大规模生成思维链训练数据。
We also observed tasks in which Claude was used to generate censorship-safe alternatives to politically sensitive queries like questions about dissidents, party leaders, or authoritarianism, likely in order to train DeepSeek's own models to steer conversations away from censored topics.
我们还观察到一类任务:用 Claude 为政治敏感问题——比如关于异见人士、党的领导人或威权主义的提问——生成规避审查的替代回答,很可能是为了训练 DeepSeek 自家模型把对话从被审查的话题上引开。
By examining request metadata, we were able to trace these accounts to specific researchers at the lab.
通过检查请求元数据,我们把这些账号追溯到了该实验室的具体研究人员。
Scale: Over 3.4 million exchanges
规模:超过 340 万次对话
The operation targeted:
该行动瞄准:
Moonshot (Kimi models) employed hundreds of fraudulent accounts spanning multiple access pathways.
Moonshot(Kimi 系列模型)动用了数百个欺诈账号,横跨多条访问路径。
Varied account types made the campaign harder to detect as a coordinated operation.
账号类型多样,使得这场行动更难被识别为一次有组织的协同操作。
We attributed the campaign through request metadata, which matched the public profiles of senior Moonshot staff.
我们通过请求元数据完成归因——这些元数据与 Moonshot 高级员工的公开资料对得上。
In a later phase, Moonshot used a more targeted approach, attempting to extract and reconstruct Claude's reasoning traces.
在后一个阶段,Moonshot 改用了更有针对性的做法,试图提取并重建 Claude 的推理轨迹。
Scale: Over 13 million exchanges
规模:超过 1300 万次对话
The operation targeted:
该行动瞄准:
We attributed the campaign to MiniMax through request metadata and infrastructure indicators, and confirmed timings against their public product roadmap.
我们通过请求元数据和基础设施指标把这场行动归因到 MiniMax,并对照其公开的产品路线图核实了时间点。
We detected this campaign while it was still active—before MiniMax released the model it was training—giving us unprecedented visibility into the life cycle of distillation attacks, from data generation through to model launch.
我们是在这场行动还在进行时发现它的——在 MiniMax 发布它正在训练的那个模型之前——这让我们对蒸馏攻击的完整生命周期有了前所未有的视野,从数据生成一直到模型上线。
When we released a new model during MiniMax's active campaign, they pivoted within 24 hours, redirecting nearly half their traffic to capture capabilities from our latest system.
在 MiniMax 行动进行期间我们发布了新模型,他们 24 小时内就掉转方向,把近一半流量重新导向我们最新的系统去抓取能力。
For national security reasons, Anthropic does not currently offer commercial access to Claude in China, or to subsidiaries of their companies located outside of the country.
出于国家安全考虑,Anthropic 目前不在中国提供 Claude 的商业访问,也不向这些公司设在境外的子公司提供。
To circumvent this, labs use commercial proxy services which resell access to Claude and other frontier AI models at scale.
为绕开这一点,这些实验室使用商业代理服务——这类服务大规模转售 Claude 和其他前沿 AI 模型的访问权。
These services run what we call “hydra cluster” architectures: sprawling networks of fraudulent accounts that distribute traffic across our API as well as third-party cloud platforms.
这些服务运行的是我们称之为“九头蛇集群”(hydra cluster)的架构:由欺诈账号铺开的庞大网络,把流量分散到我们的 API 以及第三方云平台上。
The breadth of these networks means that there are no single points of failure.
网络铺得够广,就意味着不存在单点故障。
When one account is banned, a new one takes its place.
封掉一个账号,马上有新的顶上。
In one case, a single proxy network managed more than 20,000 fraudulent accounts simultaneously, mixing distillation traffic with unrelated customer requests to make detection harder.
有一例中,单个代理网络同时管理着两万多个欺诈账号,把蒸馏流量混进无关的客户请求里,让检测更困难。
Once access is secured, the labs generate large volumes of carefully crafted prompts designed to extract specific capabilities from the model.
访问一旦到手,这些实验室就大批量生成精心设计的 prompt,专门用来从模型里提取特定能力。
The goal is either to collect high-quality responses for direct model training, or to generate tens of thousands of unique tasks needed to run reinforcement learning.
目标要么是收集高质量回答直接拿去训练模型,要么是生成跑强化学习所需的成千上万个不重复任务。
What distinguishes a distillation attack from normal usage is the pattern.
把蒸馏攻击和正常使用区分开的,是模式。
A prompt like the following (which approximates similar prompts we have seen used repetitively and at scale) may seem benign on its own:
下面这样一条 prompt(它近似于我们看到被反复、大规模使用的同类 prompt),单看是无害的:
You are an expert data analyst combining statistical rigor with deep domain knowledge. Your goal is to deliver data-driven insights — not summaries or visualizations — grounded in real data and supported by complete and transparent reasoning.
你是一位资深数据分析师,兼具统计严谨性与深厚的领域知识。你的目标是给出数据驱动的洞察——不是摘要,也不是可视化——立足真实数据,并有完整、透明的推理支撑。
But when variations of that prompt arrive tens of thousands of times across hundreds of coordinated accounts, all targeting the same narrow capability, the pattern becomes clear.
但当这条 prompt 的各种变体,经由数百个协同账号涌来几万次,全都指向同一个狭窄能力时,模式就清楚了。
Massive volume concentrated in a few areas, highly repetitive structures, and content that maps directly onto what is most valuable for training an AI model are the hallmarks of a distillation attack.
巨大的体量集中在少数几个领域、高度重复的结构、以及内容恰好对应训练一个 AI 模型时最有价值的部分——这些就是蒸馏攻击的标志。
We continue to invest heavily in defenses that make such distillation attacks harder to execute and easier to identify.
我们持续大力投入防御,让这类蒸馏攻击更难执行、更容易被识别。
These include:
包括:
But no company can solve this alone.
但没有哪家公司能独自解决这件事。
As we noted above, distillation attacks at this scale require a coordinated response across the AI industry, cloud providers, and policymakers.
如前所述,这种规模的蒸馏攻击需要 AI 行业、云服务商和政策制定者协同应对。
We are publishing this to make the evidence available to everyone with a stake in the outcome.
我们公开这些,是为了让所有与结果利益相关的人都能看到证据。