Comment by Garpagan

1 day ago

> I don't understand how you're so caught up on the term "distillation". Distillation is using a larger model's outputs to train a (weaker) student model. Which is exactly what's happening.

But from evidence it was data generated by Opus 4.5. Is Opus 4.5 larger, stronger model and Kimi K3 a weaker, student model? I don't think so.

Also, I searched on HuggingFace "fable data", first link was a "distillation dataset", including samples from Fable, there are 2 milion code samples of Fable output. So this fully legal and sits in the open, somehow is not a distillation attack? Who knows, maybe Moonshot used this for training.