Comment by janalsncm
6 days ago
1) Mixture of experts as an LLM architecture is not the same thing. In MoE each “expert” can activate for any given token.
2) Bitter lesson is misunderstood. Specialization and inductive biases still matter. ChatGPT isn’t the best chess player in the world just because it’s seen more math problems or read more Japanese poetry. Stockfish is, because it bakes in useful inductive biases like minimax.
No comments yet
Contribute on Hacker News ↗