Comment by scorpion7
5 days ago
It's fascinating how this works with such a small model. Especially given that the training is a kind of meta learning of "how to do in-context learning". I wonder, is there a good intuition of the role of the MLP in this architecture? For LLMs the consensus seems to be that they store knowledge...what would that be for tabular data?
No comments yet
Contribute on Hacker News ↗