Comment by Kotlopou

1 day ago

I would like to know how much of the progress comes from effectively combining existing research programs plus massive persistence, and how much is AlphaZero-esque RLVR completely independent of training data. Since I cannot get anybody to care about this question (even though I think it's vital for guessing what the future trajectory will look like -- are we going to complete existing research programs or start new ones?), I live in ignorance and wait for the day when the answer becomes clear.

In looking at this over the past hour, I haven't seen clear evidence one way or the other. Some of the stuff is highly unexpected (like the multiplication algorithm), but counterexample-y, and about the rest the professional mathematicians online seem to have a consensus that it's not "breaking through fundamental obstacles". I suspect neither of us is competent to judge that.

Doesn't matter at this point i would say.

Alone the massive usage of us every day produces a massive amount of signals.

I build something and claude does something stupid? "hey thats not what i meant! Do this instead!" "Okay" <<< This is a signal.

The mathematician being unhappy about something from claude? Another signal.

This alone gives you enough progress i would argue. But additional its clear that certain tasks are worth to pay experts for for teaching one central AI once instead of every single human who needs to do the task.

IF RL is also working well, we are just faster f*ed than otherwise.