Comment by dannyw

5 hours ago

Math is one of the most verifiable domains, esp thanks to LEAN, which also build coding skills.

The $$$ they're pouring isn't just for marketing. Think of these papers/results more as "useful side effects" from large-scale RL rollouts and post-training. Every token being generated contributes to post-training in some way.

There isn't a hard boundary between "training" or "inference", modern post-training is arguably inference-bound :)

> There isn't a hard boundary between "training" or "inference", modern post-training is arguably inference-bound :)

Ah this is an enlightening point. 8 hadn't thought about it this way, but you're right.

It was shown quite some time ago that training LLMs on programming tasks improves their logical reasoning skills also in other natural language domains. So I could see math also being a training gym for AI even if the final use case is not directly math-related. Having to solve math problems efficiently can build in skills that come handy in all kinds of more everyday tasks or science and engineering.

  • It’s also very useful signal that the reasoning trace is leading to solving open problems - you can be certain that you’re not landing somewhere inside the training data.