Comment by kingstnap

14 hours ago

It's implicitly trained against. There is like information leakage with researchers messing with the training parameters and checkpoints used.

It's not the direct feedback loop of RL but its not far.

It’s pretty far.

It’s the difference between “study law until you can pass any random bar exam” and “here are 200 legal questions and we’ll drill them, with me correcting and explaining when you get one wrong, until you can pass exactly these 200”.

Your right that tuning can aim for a benchmark, but it does not leak any information about the answers.