Comment by suprjami

2 days ago

You have understood correctly.

One really would think these companies (including Google) who spend many millions of dollars on compute could write a few hundred lines of Jinja correctly, so their investment works optimally or at all.

But they don't.

Then a couple of individuals on HuggingFace fix it, either a 2-person startup like Unsloth or a volunteer like froggeric.

I also don't understand how this repeatedly happens.

I did SFT / RL post-training on Qwen3 models a bit. This is an issue that dates back long ago.

My favorite theory is that they had many model variants internally, each using a slightly different chat template, so when it comes to the release they are not even sure what to use any more.