Comment by oxag3n
19 hours ago
[1] is not a valid proof LeCun was wrong, LLMs still can't do spacial reasoning when it can't be derived from the training data. He didn't argue that GPT 5000 won't be able to describe something with words.
19 hours ago
[1] is not a valid proof LeCun was wrong, LLMs still can't do spacial reasoning when it can't be derived from the training data. He didn't argue that GPT 5000 won't be able to describe something with words.
This really doesn't match my experience. I can ask an LLM to modify engineering plans using vague natural language prompts and it will find the right place in the plan from the description and then make appropriate modifications, which necessarily requires doing spacial reasoning.
Or it’s just taking common examples from training and applying those copied heuristics to your problem? Doesn’t mean it’s actually reasoning about the space and how to solve the problem. It’s the equivalent of a student writing an answer they saw somewhere else without understanding “why”.
Why does CoT significantly improve their performance? Most of what they are applying they learned in post-training by solving similar problems themselves. This isn't about regurgitating pre-trained knowledged.
Advancements in Math and coding are because RLVR at massive scale is so cheap.
Looking at the reasoning traces it sure seems like it's reasoning. It internally debates which of the possibly matching parts of the input are the one described by me in the prompt and picks the right one based on sound reasoning.
2 replies →
I think the point here is that LeCun was arguing that training on pure text would not grant spatial understanding. I believe most models are trained on spatial data as well, so you are both right.
Is that what he meant? He works on models with an explicitly spatial internal representation, whereas I was using a standard LLM that edited the provided plan by using a bajillion python calls to inspect small regions of the image at a time and then generate edits.
I've seen recent AIs make detailed and technically impressive 3D models. You might argue "they're not doing spatial reasoning, they're making measurements with code and doing math to configure relative positions". Fine, but at a certain point that becomes functionally indistinguishable from spatial reasoning.
I don't know - real-world tests leave me unconvinced: https://youtu.be/ENWVpqtOdRI?t=867
> when it can't be derived from the training data
This sounds like a goalpost on wheels. Can you define clearly where your stake in the ground is?
we have benchmarks proving it can do spatial reasoning.
...poorly? https://youtu.be/ENWVpqtOdRI?t=867
I’m not sure what you are trying to say
4 replies →
the statement was "can't do"
"doing poorly" is still doing
1 reply →
Aaah, the old benchmarks maxxing argument, having precise and clear definition of what "spatial reasoning" is, what, and most importantly WHY, the benchmarks of choice are would settle this debate, otherweise let's not delve into it.