Comment by mym1990
9 hours ago
Or it’s just taking common examples from training and applying those copied heuristics to your problem? Doesn’t mean it’s actually reasoning about the space and how to solve the problem. It’s the equivalent of a student writing an answer they saw somewhere else without understanding “why”.
Why does CoT significantly improve their performance? Most of what they are applying they learned in post-training by solving similar problems themselves. This isn't about regurgitating pre-trained knowledged.
Advancements in Math and coding are because RLVR at massive scale is so cheap.
I have not seen the research on CoT and how it impacts spatial reasoning and outcomes, so I can’t comment on that part. I do see continuously that in almost every example of non-trivial image generation and 3d modeling, there are quirks that point to the fact that the model does not understand relationships within the space based on physics.
CoT seems to work really well for text based generation but once you’re past that and into physics models and detailed relationship mapping, it may work better but it’s not enough to make me believe it’s “reasoning” in a way that humans do.
Looking at the reasoning traces it sure seems like it's reasoning. It internally debates which of the possibly matching parts of the input are the one described by me in the prompt and picks the right one based on sound reasoning.
I mean, it's a text predictor. When you say <BEGIN_REASONING>, you'll get reasoning-like output next, whether or not the model is capable of reasoning.
It's not just text that appears at first blush to resemble reasoning, it's actual sound reasoning. And it can chain it for hours at a time without breaking down.