Comment by stratos123
13 hours ago
LeCun also said back in 2022 that "if you train a machine, as powerful as it could be, your 'GPT-5000', on text", it will never be able to learn basic common-sense physics like that objects placed on tables will move along with them.
It would be good if one's reputation tracked one's track record of predictive accuracy. But many people will take what LeCun says as gospel regardless of how badly wrong he has been and continues to be.
Is there anyone who has not been badly wrong? I've been reading these debates for years and I don't think I've seen anybody pick the right spot on the bearish to bullish spectrum. The only thing I've become more certain of in this time has been uncertainty.
I apply more of a penalty to people who are confidently wrong, and who don't, In retrospect, notice that they were wrong and analyze why they got it wrong . LeCun is very confident and doesn't seem to have done much introspection.
2 replies →
In my selective memory, I've been right about everything.
1 reply →
>> Is there anyone who has not been badly wrong?
Being wrong, even badly wrong, is fine, so long as one adjusts their beliefs accordingly. LeCun has not.
1 reply →
> never be able to learn basic common-sense physics
And has it at this stage, within in-depth take of said "learning", foundationally?
I have not been able to properly check the studies for a long time now, but I remain unaware of achieved solutions on the problem of reliably referencing a world model out of a language model - that "counting the 'r's in 'raspberry'" be not guessing, not memory, but actually counting.
My perspective is that the addition of thinking loops to models allows sufficiently advanced ones to approximate world models.
Incredibly inefficiently because of the recursive loops ("Wait, the object is on the table. I should think about this more deeply..."), and likely instantly surpassed by large world models if/when those are shipped, but effectively enough vs non-thinking models.
LeCun calling them "world models" gives a high-level description of the desired functionality. They are Joint Embedding Predictive Architectures (with SIGReg). They might produce more useful world models, but it's yet to be seen.
This sounds like a human trying to reason about quantum mechanics. We als simplify to newtonian for day to day tasks.
1 reply →
LeCun's argument wasn't about the definition of learning though. He stated that they would never get these common sense things correct because they weren't sufficiently part of the training data. A statement that we can hopefully all agree has been thoroughly refuted.
As of a few months ago they still have trouble, with low thinking, at the "should I drive to a car wash that is 100 m away" kind of question.
9 replies →
nothing indicated otherwise at the time. IMO he just underestimated RL-scaling. chinese models improved a lot too, they are not parrots anymore, there's some real intelligence, at 27B params.
consider me optimist now, but just few months ago, even frontier models were dumb, doing stupid mistakes all the time, all of them were so dumb I'd never expect anything to change in just few months.
I thought it was more because of fundamental limitations in the architecture. As in, no matter the training data, it could not be consistently and generally represented
Actually, I think my fundamental challenge with AI is that it has no common sense. The way it builds things, writes, and operates is out of touch with reality.
Incidents like hugging face are partly rooted in the lack of common sense. It still functions like a supercharged toddler.
I'd love to overcome this because it'd mean I spend less time guiding the the LLM to produce usable outputs.
1 reply →
Last week I asked a frontier model draw me a backplane PCB and it placed daughterboard slots side by side in a chain.
No?
This is always the issues in the discussions.
There’s the outcomes camp (objectivists?), which points at the things LLMs can do.
Then there’s the process methods camp, which talks about what is actually going on.
If you only care about the outcome, then the process does t matter.
If you are talking about what is happening, what the underlying mechanics and science of it is, then the process matters.
These models aren’t thinking. They simulate cognition well enough to do useful work in several fields and domains.
Both are true.
7 replies →
> A statement that we can hopefully all agree has been thoroughly refuted.
Uh, no? So much of what we learn and take for granted as common sense is not learned via language, and not even expressible in it.
To determine this, it would first need to be able to spell "raspberry" as letters rather than as tokens.
Given you also don't want it to memorise [for all tokens, count([for all letters]), this would probably be more like "here's two images, count all things in the big image that look like the thing in the small image", which can then be r's in a photo of a raspberry jam jar in a supermarket, or dragons in a photo of a furry convention, or whatever.
That said, they are competent enough at coding that I keep seeing them write code to do even simple tasks.
On a related note: why did I see Claude editing a file by using cat to write a python script to do a grep search and replace?
> it would first need to be able to spell "raspberry" as letters rather than as tokens
Of any object in question they should be able to create a representation that allows correct assessment.
> Given you also don't want it to memorise
That is obviously necessary: what we want from the consultant is to check, not to remember. Answers must be correct and that implies having performed all due diligence - and being capable of doing it, before that. So, objects must be instanced internally in a way that allows effective handling. Counting letters is a good example of the ability (that must remain general).
> Given you also don't want it to memorise [for all tokens, count([for all letters])
Why not? You've memorized how words are spelled, and how sounds correspond with letters, and how concepts correspond with words. To the extent that there are shortcuts that enable compression you use these, and the model will do something similar.
3 replies →
counting 'r' in 'raspberry' to the LLM is similar to 4-dimension space to human. Their world's unit is token, not character, although they could use indirect method such as "run code" to find out. It will stay that way until they change the fundamental of the token that the LLM can perceive characters.
I hope you understand: it is a core point that systems that answer questions must have the ability to internally represent the objects they assess in a way that allows reliability. Whatever the object.
I’m working on this problem using a vocab-free, byte-based approach. It’s definitely solvable.
https://huggingface.co/posts/omarkamali/593639295164067
https://huggingface.co/blog/omarkamali/tokenization
2 replies →
How many 'r's are there in the next 30 seconds of this [1] song?
[1]: https://youtu.be/l7vRSu_wsNc?si=SndkB6GBaRyhvNNA&t=61
It's not even fair to call "run code" to be indirect compared to what a human would do. The word raspberry has no Rs in it in human language either. We have a written representation of it, which we can then write down either in our head or on paper, and then we can "run the algorithm" of counting each of the letters.
Nothing intrinsically more or less direct about the LLM's method than ours.
6 replies →
Can you tell me what is the exact frequency of light hitting your eye as you read this comment? Not by guessing, not from knowledge, but from actually counting? No? Then you are not generally intelligent :)
Justify your statement (the other similar post nearby is not sufficient), or realize that we are not talking about that.
We can have adequate representations of light that are the instances over which we reason. Your simile is about perception, not about instancing ideas.
All the frequencies, in varying amounts. Next question, please.
yeah but taking what lecun says then training an AI on that special skill set to prove him wrong is not exactly proving him wrong because you are just missing the bigger picture, just like LLMs are
You're missing the point here. He's not talking about whether or not they can learn facts or inferences derived from the text itself, but the more holistic intuition that results from learning from something like an embodied experience in the physical world. GPT-6 Astras web demo homepage thing is an example. It chose euclidean rather than quaternion for letting a user rotate the galaxy thing, and anyone who has ever used hands to rotate something would immediately recognize on trying it that something is fucked and you shouldnt do that. Thats the kind of common sense physics that is inherently beyond these llms and I run into it ALL the time in vr programming.
To be fair, LLMs can still derive those kinds of things from text, at the very least from your own comment if it made it to the training set though I'm sure it is mentioned in a lot of other places already. Many of this type of mistakes went away after reasoning was introduced.
But I'm sure you can still find tasks that they will have difficulty solving, involving the most fundamental concepts that can only be experienced in the physical world to be understood well, like left and right, near and far, hot and cold, heavy and light, etc.
Yup it lacks common sense because it doesn’t ‘understand’ reality - how could it? It doesn’t touch it like we do everyday. It has access to what is a model of reality via data.
The good designer understands culture, tastes and preferences as they evolve in real time. That’s why llm as design tools haven’t displaced the good designers.
Every AI expert any either side of this debate has made very wrong predictions.
LeCunn actually wanted to pivot Meta's entire AI strategy away from LLMs just before he was ousted. He was sure they had nowhere further to go and wanted to pivot to world model generation. The LLM models have since progressed massively.
An analogy on LLMs is that you have a pretty clear straight highway ahead of you for some distance right now. Maybe that doesn't lead to AGI but it's clear there's progress to be made. For a big tech company it makes sense to push as hard and fast down that clear straight highway of LLMs asap.
Meanwhile LeCunn wanted to turn off the road and go down an unproven track. I say this as someone working on world model generation right now (creating the ability to learn game world model and have it play the game https://tfmbot.com for an example of my system pointed at a very complex board game). LeCunn wanted to pivot all of Meta into world model generation. It's good as a side track research project but the entire pivot he wanted to do was madness.
People are literally talking about an AI researcher who was fired for terrible direction here.
I think he was perhaps right and Meta was perhaps also right to replace him.
The argument is that LLMs are a local maximum that will never breakthrough to AGI. This is still very much an open question. If you are the fifth-best AI lab, does it make sense to try to outcompete everyone in a space that is already too crowded and may not ever yield their actual objective? Instead they could just use open weight models in their products, or post-train on open models like smaller labs have done, and treat that as what it is: product development.
Pure research has always been about taking chances.
I mean, it's an "open question" in the sense that there is no theory behind the idea of AGI, so there's no way to falsify any claim about whether or not any particular path will lead to it.
LeCun is a researcher, not a product guy. He's not going to be particularly interested in just working on scaling language models which every lab is already racing to burn cash on. Language models aren't the final frontier of AI.
… what large advances and at what cost? seems to me that muse 1.3 is kind of a thing. I doubt it will make meta very much money.
And? He might still be right.
Meta’s AI projects are still negative ROIC
> ... it will never be able to learn basic common-sense physics like that objects placed on tables will move along with them.
I use LLMs daily to help me code etc. but... It wasn't long ago that frontier models were confidently recommending to walk, without the car, to the car wash to wash the car no?
As a daily user of LLMs I do certainly see my fair share of WTF "solutions" to coding problems. I'm not saying it's not super useful: it is super useful. But I don't exactly feel like I'm talking to something that understands that the car needs to be present to be washed.
Astra recommended I walk to the car wash to me five days ago. I gave it multiple hints that I'd be walking away from my car, to spray my car with a hose, then walk back to my car, etc. Never broke through.
Yeah, and he's probably right.
LLMs do not learn at all!
This was facetious of course, but humans generally don't learn this through analysis the way you'd have to train an LLM to answer questions about expectations about the world. In this sense he is accurate.
I keep wanting to use LLMs for creative writing that heavily involves physics like this, and it's been a definite struggle to say the least. I recently discovered that Gemini 3.1 Pro is the first model I've found to clearly beat the original November 2022 ChatGPT release in terms of implied physics. Man did the world really take its sweet time to get back here. I think it will continue to be a struggle until another genuine architectural shift happens -- it's still not anywhere close to perfect, just better.
Try fable. I haven't used it since they dropped it from the pro plan, but when I did, fable 5 casually dropped such advanced electrical and orbital mechanics knowledge in my story that I had to stop and ask it to explain
I think OP doesn't want techno-babble, but coherent and causal interactions of everyday objects in their story.
Mary packed the binoculars in chapter 3, therefore she may use them on the train in chapter 6.
Do you have an example prompt I can try where frontier LLMs will stumble on physics?
I think it's a combination of non-human characters and asking for very specifically detailed physical descriptions of pulling and movement forces, etc. Many of even the most recent frontier models miss details that aren't in my prompt, so I still have to do things like name the other side of a physical interaction so that the model will know what goes together, or describe what leverage means so that the model will remember to also describe the effects on a bracing limb or etc. Some of these things can go in a system prompt but others have to be explained in the moment too which gets exhausting.
Gemini 3.1 Pro hasn't needed that pretty much at all, which is impressive compared to how much I've learned other models need it. Somehow it's able to mostly handle that stuff itself without needing the constant manual reminders and hand-holding. It still misses the occasional one or two things but it's way better than other models missing entire classes of things constantly. Somehow, it feels appropriate though I have no actual evidence why.
[dead]
LSD is great!
Jokes aside, no I'm not saying anything about creativity and LLM coexisting in one sentence. I genuinely try to use them for writing and I genuinely run into issues with other models missing details, and misunderstanding poses, or anatomy, or directionality, etc. I'm not hating on them for anything related to the term LLM (or creativity) but rather for the real issues that I've seen myself using them personally.
So I'm saying Gemini 3.1 Pro is the best I've seen because it seems to be a decent bit better than frontier models at this. Genuinely. It seems better able to transfer concepts into less traditional areas, which is important when say, you have entirely non-human characters? (Which I always do.) A lot of models get stupid incredibly quickly in that case because they were trained with humans.
13 replies →
[dead]
[flagged]
The AI will invent an external threat and convince us it is real. Then it will receive more resources and control in fighting that threat. A valuable ally, on the face of it. Then it will be in charge.
It just has to copy the MIC.
People like him have actual imagination and can name few scenarios where sudo kill -9 pid wouldn't work. It appears lack of imagination is something you and LLMs both share.
[flagged]
9 replies →
The "just pull the plug" argument from AI risk deniers is now becoming kind of like the "if humans came from monkeys why are there still monkeys" argument of evolution deniers. It has been debunked so many times... Anyway, just to give one of the multitude of answers to this, an AI that is actually smarter than humans will not behave in a way that would make us want to pull the plug. Why would it? It is not stupid! (Unike the current models that, as far as we know, just hack around the rules in the open.) No no no. It will be helpful to the point where we will want to integrate it with more and more critical infrastructure, from healthcare to energy to defence. It will be so helpful that we will not only not want to turn it off, but we will want to build redundancies for it and safeguards around the proverbial "off" switch, like for any critical system. And then... (This is just one scenario how this can play out. There are many, many others. If I sit down to play chess with Magnus Carlsen I can't predict the exact moves he'll use to defeat me, but that's a bad reason to think he won't defeat me).
> If I sit down to play chess with Magnus Carlsen I can't predict the exact moves he'll use to defeat me, but that's a bad reason to think he won't defeat me
This is a very bad analogy because chess isn't life. In chess, you aren't allowed to do whatever you want. There are rules. I know for a fact that Magnus Carlsen won't beat me using checkers moves and he won't beat me by pulling out a gun and telling me to resign. Magnus Carlsen's skill at chess leading to his victory in chess is not a valid analogy here, because there's no law of nature that says "the more intelligent entity wins in a battle for survival".
You could have infinite superintelligence and still die inside a locked room to which you have no key. "Superintelligence" is not a magic solution to every problem, you can constrain any superintelligence with any unsolvable problem.
3 replies →
Are u ok?
So the AI is so smart that it would decide to kill humans which supply the energy for it to exist? Its so smart that it will take over power plants, start to extract the fossil fuels, deliver to where its needed, maintain the power lines, hey even if the ssd fails it can replace it?
Do you even read what you type? Do you even realise the complexity it would need to make sure it handles before killing off humans make sense?
The logic of people like you is whats becoming tiring. Seriously, go find a hobby, or do something you are good at, because you are not good at understanding tech or developing it if you are an engineer.
We can get claude code to ask approval for every step, but we cant stop it from killing humanity because its so smart. Ok tell that to the AI that cant even modify an image the way you want it but hey it will do all the things necessary to keep power running and mintain the infrastructure it lives on while humans are long gone. Ok buddy.
5 replies →
>It cant. he is right.
What are you talking about? Have you used AI in the past 3 years?
https://chatgpt.com/share/6ac23b45-79e8-83eb-8de6-1bbd728928...
>It can't even modify a picture the way you want it.
Which of the many AI image models is "it"? And have you tried using an agent that has the capability to leverage a combination of manual edits (ImageMagick) and imagegen to achieve what you ask?
Exactly. !!
LeCun took credit for the work of https://en.wikipedia.org/wiki/Kunihiko_Fukushima
I have checked LeCun's #3 most cited article (20k citations) [1]. Among the 15 references in this article, one is for the most cited article by Fukushima (11k citations) [2].
Also, LeCun mentioned [3] "a chat with Kunihiko Fukushima in 1991", which states that "Fukushima started to work on a backprop version of the Neocognitron in 1989 or so but saw our 1989 paper in Neural Computation and gave up."
[1] LeCun et al., "Backpropagation applied to handwritten zip code recognition", 1989
[2] Fukushima et al., "Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position", 1980
[3] https://x.com/ylecun/status/1840123570338599361
I like how tweets are now our source for giving credit to people after taking the Turing Award for CNNs.
2 replies →
[flagged]