三家中国实验室被点名,工业规模蒸馏 Claude
"We have identified industrial-scale campaigns by three AI laboratories—DeepSeek, Moonshot, and MiniMax—to illicitly extract Claude's capabilities to improve their own models."
"These labs generated over 16 million exchanges with Claude through approximately 24,000 fraudulent accounts, in violation of our terms of service and regional access restrictions."
这是 Anthropic 第一次公开"点名 + 给数据"地指控同行实验室——DeepSeek 15 万 / Moonshot 340 万 / MiniMax 1300 万 次对话,合计 1600 万+。
蒸馏本身合法,但用来抄竞争对手不合法
"Distillation is a widely used and legitimate training method… But distillation can also be used for illicit purposes: competitors can use it to acquire powerful capabilities from other labs in a fraction of the time, and at a fraction of the cost."
"Frontier AI labs routinely distill their own models to create smaller, cheaper versions for their customers."
蒸馏 = 用强模型输出训弱模型,自己蒸自己的模型(为了便宜版)完全 OK;别人蒸你的模型(为了少花训练成本拿到能力)= illicit。法律边界从"技术行为"挪到了"对象 + 意图"。
这件事被定性为"国家安全风险",不是"商业纠纷"
"Illicitly distilled models lack necessary safeguards, creating significant national security risks."
"Foreign labs that distill American models can then feed these unprotected capabilities into military, intelligence, and surveillance systems."
论点:Anthropic 在防"生物武器 / 网络攻击"的 safeguard 蒸馏过程中会被剥掉;裸能力流入军队 / 监控 / 攻击场景;开源版扩散后再也收不回。把技术议题"国家安全化"。
蒸馏攻击 = 出口管制有效性的"反向证据"
"In reality, these advancements depend in significant part on capabilities extracted from American models, and executing this extraction at scale requires access to advanced chips."
"Distillation attacks therefore reinforce the rationale for export controls: restricted chip access limits both direct model training and the scale of illicit distillation."
巧妙的论证逆转:外界看"中国 lab 快速追赶 → 出口管制无效",Anthropic 说"恰恰相反 —— 他们靠蒸馏追赶,而蒸馏本身需要芯片才能跑得动,所以管芯片其实双重打击"。
"Hydra cluster"代理网络:一个 proxy 跑 2 万账号
"In one case, a single proxy network managed more than 20,000 fraudulent accounts simultaneously, mixing distillation traffic with unrelated customer requests to make detection harder."
"When one account is banned, a new one takes its place."
攻击手法被公开:不是直接调 API,而是通过商业代理服务"分发流量到几万个伪造账号",混进真实客户流量里。封一个补一个,无单点失败。
MiniMax 在 Anthropic 发新模型后 24 小时 切流量
"When we released a new model during MiniMax's active campaign, they pivoted within 24 hours, redirecting nearly half their traffic to capture capabilities from our latest system."
"We detected this campaign while it was still active—before MiniMax released the model it was training—giving us unprecedented visibility into the life cycle of distillation attacks."
最惊人的运营细节:MiniMax 的蒸馏 pipeline 是实时响应式的——Anthropic 一上新,他们半天内就把一半算力转去抓新能力。攻防对抗的节奏暴露。
We have identified industrial-scale campaigns by three AI laboratories—DeepSeek, Moonshot, and MiniMax—to illicitly extract Claude's capabilities to improve their own models.
我们已经识别出三家 AI 实验室——DeepSeek、Moonshot、MiniMax——以工业规模非法提取 Claude 能力来改进自家模型的行动。
These labs generated over 16 million exchanges with Claude through approximately 24,000 fraudulent accounts, in violation of our terms of service and regional access restrictions.
这些实验室通过大约 24,000 个伪造账号与 Claude 产生了超过 1,600 万次对话,违反了我们的服务条款和区域访问限制。
These labs used a technique called "distillation," which involves training a less capable model on the outputs of a stronger one.
这些实验室使用了一种叫"蒸馏"(distillation)的技术——用一个更强的模型的输出去训练一个能力较弱的模型。
Distillation is a widely used and legitimate training method.
蒸馏本身是一种被广泛使用的合法训练方法。
For example, frontier AI labs routinely distill their own models to create smaller, cheaper versions for their customers.
比如,前沿 AI 实验室经常蒸馏自己的模型,做出更小、更便宜的版本提供给客户。
But distillation can also be used for illicit purposes: competitors can use it to acquire powerful capabilities from other labs in a fraction of the time, and at a fraction of the cost, that it would take to develop them independently.
但蒸馏也可以被用于非法目的:竞争对手可以借此从其他实验室获取强大的能力,只花独立研发所需时间和成本的很小一部分。
These campaigns are growing in intensity and sophistication.
这些攻击行动在强度和精密程度上都在持续上升。
The window to act is narrow, and the threat extends beyond any single company or region.
应对的时间窗口很窄,而这个威胁也已经超出了任何一家公司或地区的范畴。
Addressing it will require rapid, coordinated action among industry players, policymakers, and the global AI community.
要应对它,需要业界、政策制定者、以及全球 AI 社区采取快速、协调的行动。
Illicitly distilled models lack necessary safeguards, creating significant national security risks.
非法蒸馏出来的模型缺乏必要的 safeguard,带来显著的国家安全风险。
Anthropic and other US companies build systems that prevent state and non-state actors from using AI to, for example, develop bioweapons or carry out malicious cyber activities.
Anthropic 和其它美国公司构建了一套系统,防止国家或非国家行为体用 AI 来——比如说——研发生物武器、或者发起恶意网络行动。
Models built through illicit distillation are unlikely to retain those safeguards, meaning that dangerous capabilities can proliferate with many protections stripped out entirely.
通过非法蒸馏构建出来的模型,基本不会保留这些 safeguard——这意味着危险能力会带着"被剥光保护"的状态扩散出去。
Foreign labs that distill American models can then feed these unprotected capabilities into military, intelligence, and surveillance systems—enabling authoritarian governments to deploy frontier AI for offensive cyber operations, disinformation campaigns, and mass surveillance.
外国实验室蒸馏美国模型之后,可以把这些没有保护的能力直接喂给军事、情报、监控系统——让威权政府能够把前沿 AI 用于进攻性网络行动、虚假信息运动、大规模监控。
If distilled models are open-sourced, this risk multiplies as these capabilities spread freely beyond any single government's control.
如果蒸馏出的模型还被开源,这种风险会成倍放大——这些能力将自由扩散,任何单一政府都无法控制。
Anthropic has consistently supported export controls to help maintain America's lead in AI.
Anthropic 一贯 支持出口管制,以帮助美国保持 AI 领先地位。
Distillation attacks undermine those controls by allowing foreign labs, including those subject to the control of the Chinese Communist Party, to close the competitive advantage that export controls are designed to preserve through other means.
蒸馏攻击会削弱这些管制——它让外国实验室(包括那些受中国共产党管控的实验室)通过别的途径,把出口管制本来想保留的竞争优势给追平。
Without visibility into these attacks, the apparently rapid advancements made by these labs are incorrectly taken as evidence that export controls are ineffective and able to be circumvented by innovation.
在看不到这些攻击的时候,这些实验室表面上的快速进展会被错误地理解为"出口管制无效,可以被创新绕开"的证据。
In reality, these advancements depend in significant part on capabilities extracted from American models, and executing this extraction at scale requires access to advanced chips.
而事实上,这些进展很大程度上依赖于从美国模型里提取出来的能力——而要在规模上完成这种提取,本身就需要先进芯片的算力支撑。
Distillation attacks therefore reinforce the rationale for export controls: restricted chip access limits both direct model training and the scale of illicit distillation.
所以蒸馏攻击反过来强化了出口管制的理据:限制芯片访问,既限制了直接训模型的规模,也限制了非法蒸馏能跑到的规模。
"Distillation attacks therefore reinforce the rationale for export controls."
"所以,蒸馏攻击反而强化了出口管制的理据。"
The three distillation campaigns detailed below followed a similar playbook, using fraudulent accounts and proxy services to access Claude at scale while evading detection.
下面详述的三个蒸馏行动遵循类似的操作手册——使用伪造账号和代理服务,在规模上访问 Claude 的同时规避检测。
The volume, structure, and focus of the prompts were distinct from normal usage patterns, reflecting deliberate capability extraction rather than legitimate use.
这些 prompt 的数量、结构、关注点都跟正常使用模式截然不同,反映的是有意为之的能力提取,而非合法使用。
We attributed each campaign to a specific lab with high confidence through IP address correlation, request metadata, infrastructure indicators, and in some cases corroboration from industry partners who observed the same actors and behaviors on their platforms.
我们通过 IP 地址关联、请求元数据、基础设施指纹、以及在某些案例中来自行业伙伴(他们在自家平台上观察到相同行为者与行为模式)的佐证,以高置信度把每一场行动归因到具体的实验室。
Each campaign targeted Claude's most differentiated capabilities: agentic reasoning, tool use, and coding.
每一场行动针对的都是 Claude 最具差异化的能力:agentic reasoning、tool use、coding。
Scale: Over 150,000 exchanges
规模:超过 15 万次对话
The operation targeted:
行动针对的方向:
Reasoning capabilities across diverse tasks
多任务场景下的推理能力
Rubric-based grading tasks that made Claude function as a reward model for reinforcement learning
基于评分细则(rubric)的打分任务,让 Claude 充当强化学习的 reward model
Creating censorship-safe alternatives to policy sensitive queries
为政策敏感问题生成"过审版"的替代答案
DeepSeek generated synchronized traffic across accounts.
DeepSeek 在不同账号之间生成了同步的流量。
Identical patterns, shared payment methods, and coordinated timing suggested "load balancing" to increase throughput, improve reliability, and avoid detection.
相同的模式、共享的支付方式、协调一致的时序,提示这是一种"负载均衡"——为了提高吞吐、可靠性,同时规避检测。
In one notable technique, their prompts asked Claude to imagine and articulate the internal reasoning behind a completed response and write it out step by step—effectively generating chain-of-thought training data at scale.
其中一种值得注意的手法:他们的 prompt 让 Claude 想象并表达一个已完成回答背后的"内部推理过程",然后一步步写出来——本质上就是在规模化生成 chain-of-thought 训练数据。
We also observed tasks in which Claude was used to generate censorship-safe alternatives to politically sensitive queries like questions about dissidents, party leaders, or authoritarianism, likely in order to train DeepSeek's own models to steer conversations away from censored topics.
我们也观察到一类任务:Claude 被用来给政治敏感问题——比如关于异见人士、党的领导人、威权主义的问题——生成"过审版"替代答案,目的很可能是训练 DeepSeek 自家模型把对话引离被审查的话题。
By examining request metadata, we were able to trace these accounts to specific researchers at the lab.
通过分析请求元数据,我们能够把这些账号追溯到该实验室的具体研究人员。
Scale: Over 3.4 million exchanges
规模:超过 340 万次对话
The operation targeted:
行动针对的方向:
Agentic reasoning and tool use
Agentic 推理与工具调用
Coding and data analysis
编程与数据分析
Computer-use agent development
Computer-use agent 开发
Computer vision
计算机视觉
Moonshot (Kimi models) employed hundreds of fraudulent accounts spanning multiple access pathways.
Moonshot(Kimi 系列模型)使用了数百个伪造账号,横跨多种访问路径。
Varied account types made the campaign harder to detect as a coordinated operation.
账号类型多样化,让这场行动作为一个协调操作时更难被识别出来。
We attributed the campaign through request metadata, which matched the public profiles of senior Moonshot staff.
我们通过请求元数据完成归因,这些元数据与 Moonshot 高级员工的公开履历相吻合。
In a later phase, Moonshot used a more targeted approach, attempting to extract and reconstruct Claude's reasoning traces.
在后期阶段,Moonshot 改用了更有针对性的手法——尝试提取并重建 Claude 的推理轨迹。
Scale: Over 13 million exchanges
规模:超过 1300 万次对话
The operation targeted:
行动针对的方向:
Agentic coding
Agentic coding
Tool use and orchestration
工具调用与编排
We attributed the campaign to MiniMax through request metadata and infrastructure indicators, and confirmed timings against their public product roadmap.
我们通过请求元数据和基础设施指纹把这场行动归因到 MiniMax,并将时间线与他们的公开产品路线图相互印证。
We detected this campaign while it was still active—before MiniMax released the model it was training—giving us unprecedented visibility into the life cycle of distillation attacks, from data generation through to model launch.
我们在这场行动仍在进行中的时候就检测到了——在 MiniMax 发布它正在训练的那个模型之前——这让我们对蒸馏攻击的完整生命周期(从数据生成到模型发布)拥有前所未有的可见度。
When we released a new model during MiniMax's active campaign, they pivoted within 24 hours, redirecting nearly half their traffic to capture capabilities from our latest system.
在 MiniMax 行动活跃期间,我们发布了一个新模型——他们 24 小时内就转向了,把将近一半的流量重新导向去抓取我们最新系统的能力。
"They pivoted within 24 hours, redirecting nearly half their traffic to capture capabilities from our latest system."
"他们 24 小时内就切了,把将近一半的流量转去抓我们最新系统的能力。"
For national security reasons, Anthropic does not currently offer commercial access to Claude in China, or to subsidiaries of their companies located outside of the country.
出于国家安全考虑,Anthropic 目前在中国境内,以及中国公司在境外的 子公司,都不提供 Claude 的商用访问。
To circumvent this, labs use commercial proxy services which resell access to Claude and other frontier AI models at scale.
为了绕过这一限制,这些实验室使用商业代理服务——这些服务以规模化的方式转售对 Claude 和其它前沿 AI 模型的访问权。
These services run what we call "hydra cluster" architectures: sprawling networks of fraudulent accounts that distribute traffic across our API as well as third-party cloud platforms.
这些服务运行的是我们称之为 "hydra cluster"(九头蛇集群)的架构:由海量伪造账号组成的庞大网络,把流量分散到我们的 API 以及第三方云平台上。
The breadth of these networks means that there are no single points of failure.
这个网络的广度意味着——没有单点失败。
When one account is banned, a new one takes its place.
封一个账号,马上有一个新账号顶上。
In one case, a single proxy network managed more than 20,000 fraudulent accounts simultaneously, mixing distillation traffic with unrelated customer requests to make detection harder.
在其中一个案例里,单个代理网络同时管理了超过 20,000 个伪造账号——把蒸馏流量跟无关的真实客户请求混在一起,让检测变得更难。
Once access is secured, the labs generate large volumes of carefully crafted prompts designed to extract specific capabilities from the model.
访问权一旦到手,这些实验室就开始生成大量精心设计的 prompt,用来从模型中提取特定能力。
The goal is either to collect high-quality responses for direct model training, or to generate tens of thousands of unique tasks needed to run reinforcement learning.
目标要么是收集高质量回答用于直接模型训练,要么是生成跑强化学习所需的几万个独立任务。
What distinguishes a distillation attack from normal usage is the pattern.
让蒸馏攻击和正常使用区分开的,是"模式"。
A prompt like the following (which approximates similar prompts we have seen used repetitively and at scale) may seem benign on its own:
下面这样的 prompt(近似我们在规模化、重复性使用中看到的真实 prompt),单看一条似乎是无害的:
"You are an expert data analyst combining statistical rigor with deep domain knowledge. Your goal is to deliver data-driven insights — not summaries or visualizations — grounded in real data and supported by complete and transparent reasoning."
"你是一位专家级数据分析师,兼具统计严谨与深厚的领域知识。你的目标是给出基于真实数据的数据驱动洞察——不是摘要、也不是可视化——并由完整、透明的推理过程支撑。"
But when variations of that prompt arrive tens of thousands of times across hundreds of coordinated accounts, all targeting the same narrow capability, the pattern becomes clear.
但当这条 prompt 的各种变体出现几万次、跨数百个协调账号、全部针对同一个狭窄能力时,模式就显形了。
Massive volume concentrated in a few areas, highly repetitive structures, and content that maps directly onto what is most valuable for training an AI model are the hallmarks of a distillation attack.
在少数领域集中的巨大体量、高度重复的结构、以及内容上直接对应于训练 AI 模型最有价值的部分——这些就是蒸馏攻击的标志。
We continue to invest heavily in defenses that make such distillation attacks harder to execute and easier to identify.
我们持续重投于"让此类蒸馏攻击更难执行、更易识别"的防御。
These include:
这些包括:
Detection. We have built several classifiers and behavioral fingerprinting systems designed to identify distillation attack patterns in API traffic. This includes detection of chain-of-thought elicitation used to construct reasoning training data. We have also built detection tools for identifying coordinated activity across large numbers of accounts.
检测。我们已经构建了若干分类器和行为指纹系统,用于在 API 流量中识别蒸馏攻击模式——包括识别"用 chain-of-thought 引诱来构建推理训练数据"这类手法。我们也搭建了检测工具,用于识别跨大量账号的协调行为。
Intelligence sharing. We are sharing technical indicators with other AI labs, cloud providers, and relevant authorities. This provides a more holistic picture into the distillation landscape.
情报共享。我们正在与其他 AI 实验室、云服务商、相关监管机构共享技术指标,以拼出蒸馏攻击图景的更完整版本。
Access controls. We've strengthened verification for educational accounts, security research programs, and startup organizations—the pathways most commonly exploited for setting up fraudulent accounts.
访问控制。我们加强了对教育账号、安全研究项目、创业组织的身份验证——这些是最常被用来开伪造账号的入口。
Countermeasures. We are developing Product, API and model-level safeguards designed to reduce the efficacy of model outputs for illicit distillation, without degrading the experience for legitimate customers.
反制措施。我们正在开发产品层、API 层、模型层的 safeguard,目的是降低模型输出在非法蒸馏场景下的有效性,同时不损伤合法客户的使用体验。
But no company can solve this alone.
但没有一家公司能单独解决这个问题。
As we noted above, distillation attacks at this scale require a coordinated response across the AI industry, cloud providers, and policymakers.
正如上文所说,这种规模的蒸馏攻击需要 AI 行业、云服务商、政策制定者三方的协同响应。
We are publishing this to make the evidence available to everyone with a stake in the outcome.
我们公开发布此文,正是要把证据摆给所有利益相关方看。