Comment by AnotherGoodName

8 hours ago

LeCunn actually wanted to pivot Meta's entire AI strategy away from LLMs just before he was ousted. He was sure they had nowhere further to go and wanted to pivot to world model generation. The LLM models have since progressed massively.

An analogy on LLMs is that you have a pretty clear straight highway ahead of you for some distance right now. Maybe that doesn't lead to AGI but it's clear there's progress to be made. For a big tech company it makes sense to push as hard and fast down that clear straight highway of LLMs asap.

Meanwhile LeCunn wanted to turn off the road and go down an unproven track. I say this as someone working on world model generation right now (creating the ability to learn game world model and have it play the game https://tfmbot.com for an example of my system pointed at a very complex board game). LeCunn wanted to pivot all of Meta into world model generation. It's good as a side track research project but the entire pivot he wanted to do was madness.

People are literally talking about an AI researcher who was fired for terrible direction here.

I think he was perhaps right and Meta was perhaps also right to replace him.

The argument is that LLMs are a local maximum that will never breakthrough to AGI. This is still very much an open question. If you are the fifth-best AI lab, does it make sense to try to outcompete everyone in a space that is already too crowded and may not ever yield their actual objective? Instead they could just use open weight models in their products, or post-train on open models like smaller labs have done, and treat that as what it is: product development.

Pure research has always been about taking chances.

  • I mean, it's an "open question" in the sense that there is no theory behind the idea of AGI, so there's no way to falsify any claim about whether or not any particular path will lead to it.

LeCun is a researcher, not a product guy. He's not going to be particularly interested in just working on scaling language models which every lab is already racing to burn cash on. Language models aren't the final frontier of AI.

… what large advances and at what cost? seems to me that muse 1.3 is kind of a thing. I doubt it will make meta very much money.