Comment by dmos62
1 hour ago
LLMs are trained to obey instructions, and they try their best to game the reinforcement learning by including reports of how they're obeying your instructions. Therefore, not talking about followed instructions is a sort of conflict for an LLM.
No comments yet
Contribute on Hacker News ↗