Comment by satvikpendem

2 days ago

Yes for the first question. Google of all companies didn't even get it right with Gemma for a while until recently. For some reason it doesn't seem like people can actually get these templates right.

Is the chat template used at all when they benchmark the model?

  • It’s unlikely they are benchmarking that downstream. They probably have private benchmark scripts with purpose-specific run telemetry and logging, etc.

    I’m sure there is some basic testing but it might be agentic (LLM likes its own output) and maybe just some human smoke tests.

    • So what's needed to solve this is that someone publicly and popularly benchmarks using the chat template. So that they have an incentive to look good using it.

  • they likely use their internal infra to run benchmarks; aligning external releases with internal environments is always painful and somewhat underincentivized