← Back to context

Comment by dotancohen

2 days ago

  > distealling

Apt typo.

Though I am of the opinion that distilling is no different than how extant frontier LLMs have also been trained on other people's data, I could actually see the word distealling becoming useful in discussion.

Its not a typo, someone coined that during the DeepSeek R1 hype period and I kept using it since then.

I totally agree with you on the fact that it's not morally any different than pre-training. IMHO we should have a legislation that force base models to be released publicly without any restrictions whatsoever as it's basically the product of the whole humanity's intelligence.

Opus/Sol/Fable are valuable because of their reasoning and coding ability, not their bedside manners.

While still not okay, I suspect the latter is what gets stolen by Chinese distillation (and some evidence suggest this happens the other way round with US models talking in Chinese)