Comment by esafak 14 hours ago Not if you don't train against them. 1 comment esafak Reply kingstnap 13 hours ago It's implicitly trained against. There is like information leakage with researchers messing with the training parameters and checkpoints used.It's not the direct feedback loop of RL but its not far.
kingstnap 13 hours ago It's implicitly trained against. There is like information leakage with researchers messing with the training parameters and checkpoints used.It's not the direct feedback loop of RL but its not far.
It's implicitly trained against. There is like information leakage with researchers messing with the training parameters and checkpoints used.
It's not the direct feedback loop of RL but its not far.