Comment by miki123211
1 hour ago
Yes! As long as you have some criteria to judge the final answer, you can do a kind of "prompt-side RLVR", where you have the model generate prompt changes, try a bunch of different prompts and see which ones improve the results.
You don't necessarily need a bigger model to do this.
No comments yet
Contribute on Hacker News ↗