Comment by satvikpendem
2 days ago
Yes for the first question. Google of all companies didn't even get it right with Gemma for a while until recently. For some reason it doesn't seem like people can actually get these templates right.
2 days ago
Yes for the first question. Google of all companies didn't even get it right with Gemma for a while until recently. For some reason it doesn't seem like people can actually get these templates right.
Is the chat template used at all when they benchmark the model?
It’s unlikely they are benchmarking that downstream. They probably have private benchmark scripts with purpose-specific run telemetry and logging, etc.
I’m sure there is some basic testing but it might be agentic (LLM likes its own output) and maybe just some human smoke tests.
So what's needed to solve this is that someone publicly and popularly benchmarks using the chat template. So that they have an incentive to look good using it.
they likely use their internal infra to run benchmarks; aligning external releases with internal environments is always painful and somewhat underincentivized