Wink Pings

Anthropic Names 7 Chinese Companies for 'Distilling' Claude: What Did They Actually Steal?

Anthropic's report accuses seven companies, including Alibaba and DeepSeek, of distilling Claude. The numbers are alarming, but the real technical issue has been overlooked: What is distillation? Why do major tech companies dread it? How does the latest OPD method change the game?

In September 2026, Anthropic published a report naming Alibaba, DeepSeek, Moonshot AI, Zhipu AI, Xiaomi, SenseTime, and MiniMax, accusing them of using various methods to distill Claude models. The report cites striking numbers: Alibaba is accused of generating 150 million interactions between May and July; DeepSeek is accused of forwarding 12 million user requests to Claude over 14 days; Moonshot AI is accused of passing nearly 300,000 real user questions directly to Claude for answers.

These are unilateral allegations from Anthropic, and most companies have not responded. But setting aside the allegations themselves, there is a technical question worth clarifying: What exactly is distillation? Why do major companies fear it? How is proper distillation done?

![image](https://wsrv.nl/?url=https://p3-xtjj-sign.byteimg.com/tos-cn-i-73owjymdk6/b7f65734dfde4d08a8e77dfd172c105e~tplv-73owjymdk6-jj-mark-v1:0:0:0:0:5o6Y6YeR5oqA5pyv56S-5Yy6IEAg56iL5bqP5ZGY5LqO6ICB5LiD:q75.awebp?rk3s=f64ab15b%26x-expires=1791280318%26x-signature=sDe1X5TB0oCCMBuq%252FRBV7xIDntM%253D)

## Distillation: Letting a Grade-Schooler Copy the Top Student's Homework

The idea behind distillation is simple: have a small model learn the behavior of a large model. Large models are powerful but expensive—each run requires a cluster of GPUs. Small models are cheap but not as smart. The trick is to let the large model act as the teacher and the small model as the student. The teacher solves a million problems and records the answers and reasoning processes; the student learns from those answers.

The classic process: the teacher generates ground-truth answers, and the student does supervised learning.

This process has a fatal flaw.

## Once the Student Makes a Mistake, No One Corrects It

Imagine a student model doing math problems. The teacher's correct answer is "first complete the square, then solve for the roots," but the student never learned how to complete the square. It goes down a wrong path, makes a mistake in the first step, and everything after is wrong. The problem is that the teacher's notebook never contains this error. The teacher never makes mistakes, so the ground-truth answers contain no teaching examples of wrong paths.

The student has only seen correct paths. Once it starts deviating during generation, it enters territory the teacher never demonstrated, and every subsequent token drifts further off track.

This is called exposure bias: during training, the student only sees ground-truth answers, but during generation, it must face its own errors. The distributions are completely different.

Analogy: A coach only shows you videos of world champions skating and never tells you what happens if your takeoff posture is wrong. You step onto the ice, your first move is off, and everything after falls apart—the videos can't help.

## RL's Solution: Practice Itself, but Scoring Is Too Slow

Reinforcement learning takes a different approach: instead of watching videos, let the student generate its own answers, then give reward signals. This solves the "no one corrects the student's mistakes" problem, but a new issue arises: rewards are too sparse.

A problem may require 100 tokens, but at the end you receive only a single right/wrong signal. The model doesn't know which of the 99 intermediate tokens were right or wrong. It's like an exam that gives only a total score without a breakdown of mistakes.

Qwen3 used RL training, spending 17,920 GPU hours to achieve only 67.6% on the AIME math benchmark.

## OPD: Let the Student Answer, and the Teacher Correct Token by Token

From late 2025 into 2026, a method called on-policy distillation (OPD) became popular. In the Qwen3 technical report, it took only 1,800 GPU hours to reach 74.4%—nearly 10 times faster than RL.

The approach has three steps: the student generates its own answers (the RL idea), the teacher scores every token (the distillation idea), and the student revises based on the scores.

The key is the second step. The teacher puts down the ground-truth answers and acts as a token-by-token grader. Each time the student writes a token, the teacher says "this token is good" or "this token is wrong."

There is a subtlety here: the teacher scores using reverse KL divergence.

![image](https://wsrv.nl/?url=https://p3-xtjj-sign.byteimg.com/tos-cn-i-73owjymdk6/ae11c528aaf9442185c0ec0283f921e1~tplv-73owjymdk6-jj-mark-v1:0:0:0:0:5o6Y6YeR5oqA5pyv56S-5Yy6IEAg56iL5bqP5ZGY5LqO6ICB5LiD:q75.awebp?rk3s=f64ab15b%26x-expires=1791280318%26x-signature=%252BLlwXy7geA8MXQUZNDlRCyL77ds%253D)

## Why Reverse KL?

KL divergence measures the gap between two distributions, but it is asymmetric. Depending on the direction, the meaning is completely different.

Imagine a teacher teaching a student to draw. The teacher can draw 10 animals, with cats being the best.

Forward KL: The student must learn all 10 animals. Even if the teacher has drawn snakes only 1% of the time, the student must imitate them faithfully. The result is that the student knows a little about everything and is master of none—this is called mean-seeking.

Reverse KL: The student only needs to draw cats as well as the teacher. The snakes the teacher rarely draws can be skipped. This is called mode-seeking—focusing on polishing the modes the teacher is best at.

Mathematically, forward KL weights by the teacher's probability, while reverse KL weights by the student's probability. This yields a key property: when the student confidently expresses itself in regions where the teacher assigns low probability, the penalty explodes. If the student writes a wrong token, teacher probability 0.001, student probability 0.5, loss contribution 0.5 × log(0.5 / 0.001) = 3.1. If the student stays close to the teacher—teacher says 0.7 and the student also says 0.7—the loss is 0.

Reverse KL severely punishes "confident mistakes by the student" but does not punish "staying close to the teacher." This resolves exposure bias: as soon as the student strays, it gets pulled back.

Reverse KL has another property: it is cheat-proof. Under forward KL, the student can flatten its probability into a uniform distribution to push the loss down. Under reverse KL, flattening the probability is itself penalized, because it assigns low probability where the teacher assigns high probability. The only way to get a low loss is to learn to imitate the teacher.

![image](https://wsrv.nl/?url=https://p3-xtjj-sign.byteimg.com/tos-cn-i-73owjymdk6/9cb69abc51c049b59a545a07e3c1ae2e~tplv-73owjymdk6-jj-mark-v1:0:0:0:0:5o6Y6YeR5oqA5pyv56S-5Yy6IEAg56iL5bqP5ZGY5LqO6ICB5LiD:q75.awebp?rk3s=f64ab15b%26x-expires=1791280318%26x-signature=S3ExkRAaxdYizdorNv9HXvbMXD0%253D)

This can be clearly seen in a bimodal distribution diagram. The solid yellow line is the real distribution with two peaks. The forward KL fit is the blue dashed line—it tries to cover both, blurring the middle and missing the peaks. The reverse KL fit is the pink dashed line—it locks onto one peak on the left, sharp and focused, abandoning the peak on the right.

![image](https://wsrv.nl/?url=https://p3-xtjj-sign.byteimg.com/tos-cn-i-73owjymdk6/a8cea451758243a3930dd026bf137171~tplv-73owjymdk6-jj-mark-v1:0:0:0:0:5o6Y6YeR5oqA5pyv56S-5Yy6IEAg56iL5bqP5ZGY5LqO6ICB5LiD:q75.awebp?rk3s=f64ab15b%26x-expires=1791280318%26x-signature=txi91oOo9G%252BCMgPuUKBdv%252FDWE4o%253D)

## Black-Box Models Can't Do OPD

What about closed-source models like GPT and Claude? Can others run OPD on them?

No. The OPD training objective requires the teacher to provide logprobs for each token of the student. Specifically, to compute:

```

L_OPD = E_x~πθ [ log πθ(xₜ₊₁ | x₁..ₜ) − log π_teacher(xₜ₊₁ | x₁..ₜ) ]

```

where πθ is the student and π_teacher is the teacher. The key requirement is the teacher's logprobs.

But closed-source APIs like GPT-4o and Claude:

- OpenAI removed the logprobs parameter starting with GPT-4

- The Claude API has never offered logprobs

Without the teacher's logprobs, reverse KL cannot be computed. Black-box models can only fall back to using the output text for SFT:

```

L_SFT = E_x~π_teacher [ −log πθ(x) ]

```

Note that the sampling distribution for SFT is the teacher's output, not the student's own distribution. This is off-policy distillation and suffers from the exposure bias ceiling.

Open-weight models can do full OPD with each other, which is why distillation gains in open ecosystems are so significant. Closed-source black boxes can only be distilled in a degraded way—they learn wording and style, not the true distribution.

This may explain why the "stolen goods" that those 7 companies are accused of taking turn out to be of limited value.

## Back to the 7 Companies

If Anthropic's allegations are true, the named companies were doing exactly this: obtaining a large volume of high-quality answers from a teacher and running supervised learning. That is the off-policy version of classic distillation: fast initial gains, but a low ceiling due to exposure bias.

Interestingly, the latest direction of the mainstream players is exactly the opposite. OPD allowed Qwen3 to surpass RL with a tenth of the compute, shifting the focus from the teacher's answers to the student's own generation. The teacher only does the grading.

Stealing teacher answers works in the short term but has a low ceiling. Letting the teacher act as an examiner is the long-term solution.

## References

- This article is rewritten from: Programmer Yu Laoqi, 'On Large Models: Seven Chinese Companies Named for Distillation—What Did They Actually Steal?' https://juejin.cn/post/7690769804492341298

- On-Policy Distillation - Thinking Machines Lab

- On-Policy Self-Distillation for Large Language Models (arXiv 2601.18734)

- SFT, RL, and On-Policy Distillation Through a Distributional Lens

- Anthropic Accuses Multiple Chinese Companies of Distilling Claude Models - WSJ

发布时间: 2026-10-05 00:00