← Back to context

Comment by FuckButtons

4 hours ago

It’s also very useful signal that the reasoning trace is leading to solving open problems - you can be certain that you’re not landing somewhere inside the training data.

I think the entire thing here is that these solutions _are_ inherently interpolations of existing work in the field. That's the "super power" that LLMs have. To interpolate mass amounts of multi-dimensional data.

  • This depends on some handwavy use of the term "interpolation", not the mathematical definition. Mathematically interpolation usually means that the query point is in the convex hull of the data points, and that almost never happens in high-dimensional spaces.

    See: Learning in High Dimension Always Amounts to Extrapolation Randall Balestriero, Jerome Pesenti, Yann LeCun https://arxiv.org/abs/2110.09485

    I guess you mean by "interpolation" that it's some kind of nonlinear combination of the training data, but that's an almost vacuous statement. Any input-output relationship has to be so by definition.

    Or perhaps you mean that interpolation is when the test input comes from the same distribution as the training input (though this is not technically the meaning of "interpolation"). But this is also quite difficult to pin down.