Comment by nullc
2 days ago
Diverging from the sampler used in RL training is not good for long multi-turn results-- it's a great way to knock models into reasoning loops that wouldn't otherwise.
2 days ago
Diverging from the sampler used in RL training is not good for long multi-turn results-- it's a great way to knock models into reasoning loops that wouldn't otherwise.
Peer reviews NeurIPS caliber paper to prove that? Because I can show you one titled "Long context generation is a sampling problem"...
Show us, we're curious. Did you upload to ArXiv yet?