Comment by jstummbillig
4 hours ago
If you are distilling from other models (according to Anthropic reports they are [1]), there are probably a bunch of things that you can just do away with.
[1] https://www.anthropic.com/news/detecting-and-preventing-dist...
That's one of the reasons why you should never trust a single word from Anthropic and OpenAI (Sam Altman also blamed them back in the day of R1, in a pretty convenient moment). If you know anything about Claude, DeepSeek, jailbreaking, and distillation, you know the claims are clearly bullshit and the models are nothing alike, and forensic attempts agree, in fact we just had another one [1] [2].
Meanwhile, DeepSeek makes their models and methodology open, so Anthropic can (and likely do) grab without giving back.
[1] https://gist.github.com/wsxiaoys/e0286dc6bb624ff5fdf49e7f4c5...
[2] https://gist.github.com/wsxiaoys/102e8654c14d5d27b7b77532026...
Those are rookie numbers for "distillation" and one of moonshot or minimax used to offer tooling via these shady routing services for their harness/chat platforms which they served to chinese users.