← Back to context Comment by undeveloper 14 hours ago distillation results in a worse product than the actual teacher model iirc 1 comment undeveloper Reply SXX 12 hours ago Nobody care if Chinese models are only 99%, 95% or 90% as good as SotA US models.Because we only have weights and able to self-host Chinese ones. Gemma 4 and GPT OSS are nice to have, but nowhere close to that.
SXX 12 hours ago Nobody care if Chinese models are only 99%, 95% or 90% as good as SotA US models.Because we only have weights and able to self-host Chinese ones. Gemma 4 and GPT OSS are nice to have, but nowhere close to that.
Nobody care if Chinese models are only 99%, 95% or 90% as good as SotA US models.
Because we only have weights and able to self-host Chinese ones. Gemma 4 and GPT OSS are nice to have, but nowhere close to that.