← Back to context

Comment by davyAdewoyin

14 hours ago

How is China further behind if distillation cannot stop? I think it's a reasonable strategy to follow, even if they could train from scratch.

distillation results in a worse product than the actual teacher model iirc

  • Nobody care if Chinese models are only 99%, 95% or 90% as good as SotA US models.

    Because we only have weights and able to self-host Chinese ones. Gemma 4 and GPT OSS are nice to have, but nowhere close to that.