← Back to context Comment by esafak 12 hours ago Not if you don't train against them. 1 comment esafak Reply kingstnap 12 hours ago It's implicitly trained against. There is like information leakage with researchers messing with the training parameters and checkpoints used.It's not the direct feedback loop of RL but its not far.
kingstnap 12 hours ago It's implicitly trained against. There is like information leakage with researchers messing with the training parameters and checkpoints used.It's not the direct feedback loop of RL but its not far.
It's implicitly trained against. There is like information leakage with researchers messing with the training parameters and checkpoints used.
It's not the direct feedback loop of RL but its not far.