Comment by nojs

9 hours ago

I suspect it’s a side effect of heavy RL that rewards solved problems but not writing clarity.

With human RL, sounding like you've solved a problem is even better. So many times Codex writes some enthusiastic paragraph, then I learn later that it never reran the tests, or had to add some insane hard-coded hack that renders the feature useless for the general case, etc.