← Back to context

Comment by bonoboTP

6 hours ago

It was shown quite some time ago that training LLMs on programming tasks improves their logical reasoning skills also in other natural language domains. So I could see math also being a training gym for AI even if the final use case is not directly math-related. Having to solve math problems efficiently can build in skills that come handy in all kinds of more everyday tasks or science and engineering.

It’s also very useful signal that the reasoning trace is leading to solving open problems - you can be certain that you’re not landing somewhere inside the training data.

  • I think the entire thing here is that these solutions _are_ inherently interpolations of existing work in the field. That's the "super power" that LLMs have. To interpolate mass amounts of multi-dimensional data.

    • This depends on some handwavy use of the term "interpolation", not the mathematical definition. Mathematically interpolation usually means that the query point is in the convex hull of the data points, and that almost never happens in high-dimensional spaces.

      See: Learning in High Dimension Always Amounts to Extrapolation Randall Balestriero, Jerome Pesenti, Yann LeCun https://arxiv.org/abs/2110.09485

      I guess you mean by "interpolation" that it's some kind of nonlinear combination of the training data, but that's an almost vacuous statement. Any input-output relationship has to be so by definition.

      Or perhaps you mean that interpolation is when the test input comes from the same distribution as the training input (though this is not technically the meaning of "interpolation"). But this is also quite difficult to pin down.