← Back to context

Comment by Hamuko

1 day ago

Get data from Claude. That's what the Chinese (allegedly) do.

Note that this sort of distillation is NOT for pre-training data (which is tens of trillions of tokens). I think the allegations against Chinese companies by Anthropic is more so that they distill SFT data (which is good for post-training, but you still need a strong base model)