Comment by lxgr
6 hours ago
This is their CoT reasoning trace. (Why you see it: Models are supposed to delimit "actual" CoT, their user-visible summary/"cleaned CoT", tool calls, and user output via delimiters, but don't always do it perfectly.)
Rewarding terse CoTs during training seems like a no-brainer, as long as it doesn't impact capabilities, so I suspect the style is completely emergent and probably what you get when you implement a dual goal of terseness and capabilities while still punishing completely non-human-readable CoT. (Failing to do the last part would probably have models speak in ominous Unicode glyphs in no time.)
No comments yet
Contribute on Hacker News ↗