Comment by ahmedhossamdev
19 hours ago
The replay simulator from history for off-policy eval is clever - avoids expensive rollouts. Curious how they prevent the policy from overfitting to already-discovered branches and going stale as the search space expands?
No comments yet
Contribute on Hacker News ↗