Comment by brainwad
11 hours ago
This really doesn't match my experience. I can ask an LLM to modify engineering plans using vague natural language prompts and it will find the right place in the plan from the description and then make appropriate modifications, which necessarily requires doing spacial reasoning.
Or it’s just taking common examples from training and applying those copied heuristics to your problem? Doesn’t mean it’s actually reasoning about the space and how to solve the problem. It’s the equivalent of a student writing an answer they saw somewhere else without understanding “why”.
Why does CoT significantly improve their performance? Most of what they are applying they learned in post-training by solving similar problems themselves. This isn't about regurgitating pre-trained knowledged.
Advancements in Math and coding are because RLVR at massive scale is so cheap.
Looking at the reasoning traces it sure seems like it's reasoning. It internally debates which of the possibly matching parts of the input are the one described by me in the prompt and picks the right one based on sound reasoning.
I mean, it's a text predictor. When you say <BEGIN_REASONING>, you'll get reasoning-like output next, whether or not the model is capable of reasoning.
1 reply →
I think the point here is that LeCun was arguing that training on pure text would not grant spatial understanding. I believe most models are trained on spatial data as well, so you are both right.
Is that what he meant? He works on models with an explicitly spatial internal representation, whereas I was using a standard LLM that edited the provided plan by using a bajillion python calls to inspect small regions of the image at a time and then generate edits.