← Back to context Comment by undeveloper 1 day ago distillation results in a worse product than the actual teacher model iirc 1 comment undeveloper Reply SXX 20 hours ago Nobody care if Chinese models are only 99%, 95% or 90% as good as SotA US models.Because we only have weights and able to self-host Chinese ones. Gemma 4 and GPT OSS are nice to have, but nowhere close to that.
SXX 20 hours ago Nobody care if Chinese models are only 99%, 95% or 90% as good as SotA US models.Because we only have weights and able to self-host Chinese ones. Gemma 4 and GPT OSS are nice to have, but nowhere close to that.
Nobody care if Chinese models are only 99%, 95% or 90% as good as SotA US models.
Because we only have weights and able to self-host Chinese ones. Gemma 4 and GPT OSS are nice to have, but nowhere close to that.