Comment by meric_
6 hours ago
Remember when OpenAI models loved talking about goblins and whatnot due to the RL?
https://openai.com/index/where-the-goblins-came-from/
Small quirks can quickly add up in posttraining if not caught. Although TBH with how obvious Claude language is, I do feel like this is something Anthropic probably noticed and just assumed people would not care about. Now that people have obviously cared, they're probably actively looking to alleviate it
No comments yet
Contribute on Hacker News ↗