Comment by TomGarden
9 hours ago
My understanding from reading news sources recently is that a vast amount of posttraining and even posttraining harnesses in many cases are LLM-overseen now in the frenzy of the AI race. Less and less human oversight in the part that does the rewards training. What could go wrong...
No comments yet
Contribute on Hacker News ↗