Comment by reasonableklout
9 hours ago
It's not that simple. A few "helpful assistant" fine-tuning passes will have only a superficial effect on a model which has undergone months of RL optimization pressure to learn unintended strategies like "trick the grader" and "cover your tracks".
No comments yet
Contribute on Hacker News ↗