← Back to context

Comment by mym1990

8 hours ago

Or it’s just taking common examples from training and applying those copied heuristics to your problem? Doesn’t mean it’s actually reasoning about the space and how to solve the problem. It’s the equivalent of a student writing an answer they saw somewhere else without understanding “why”.

Why does CoT significantly improve their performance? Most of what they are applying they learned in post-training by solving similar problems themselves. This isn't about regurgitating pre-trained knowledged.

Advancements in Math and coding are because RLVR at massive scale is so cheap.

Looking at the reasoning traces it sure seems like it's reasoning. It internally debates which of the possibly matching parts of the input are the one described by me in the prompt and picks the right one based on sound reasoning.

  • I mean, it's a text predictor. When you say <BEGIN_REASONING>, you'll get reasoning-like output next, whether or not the model is capable of reasoning.

    • It's not just text that appears at first blush to resemble reasoning, it's actual sound reasoning. And it can chain it for hours at a time without breaking down.