← Back to context

Comment by verdverm

7 days ago

I haven't dug into QAT deeply, better recovery is my understanding as well, and also that it is out of reach for most people because you have to train a model to back prop errors based on estimated error under quant.

Hopefully more of the lab releases are trained under QAT so we can all benefit.

I think they did Gemma 3 QAT models and there are QAT versions of essentially all the Gemma 4 models (including DiffusionGemma).