← Back to context

Comment by zxexz

3 hours ago

This whole thing reminds me of DeepMind’s Variational Bayesian Last Layers[0], which never gained much traction in the broader “AI” world, but is a remarkably useful tool. And a relatively obvious one that anyone with experience in SVI and with transformer pretraining, seems to independently rediscover (including me) before finding this paper.

[0] https://arxiv.org/pdf/2404.11599