← Back to context

Comment by mdp2021

10 hours ago

Sure, the "6b subset" can be more knowledgeable on its area than a whole 27b generalist (and more efficient), but where is the simulated Intelligence encoded? A 6b subset as or more intelligent than a 27b raises the question of how metacognition skills are stored.

MOEs are built by training a second "router" model to identify which parts matter inside the dense model.

Think of MOE as ignoring noise, rather than a more efficient encoding of the data we teach it which ones can be ignored. Turning down the noise actually sharpens the results sometimes.

Modern LLM's are wildly inefficient.