Comment by egl2020
6 hours ago
Any idea how being persistent is trained? I've noticed that telling an LLM that it needs to think some more sometimes produces better results, but the claim here is that "they are very persistent" and "...kept going...".
It's from work like this:
https://arxiv.org/abs/2309.11495
A RL pipeline can reinforce verification behaviour even better than simple prompting.