← Back to context

Comment by dannyw

2 days ago

It’s unlikely they are benchmarking that downstream. They probably have private benchmark scripts with purpose-specific run telemetry and logging, etc.

I’m sure there is some basic testing but it might be agentic (LLM likes its own output) and maybe just some human smoke tests.

So what's needed to solve this is that someone publicly and popularly benchmarks using the chat template. So that they have an incentive to look good using it.