Comment by mycall
7 days ago
That wrongness is the frontier labs trying to remove their benchmaxxing bias, so the models now have a concept of 'I don't know' and will rethink directions and goals better. There was lots of research last year on this topic, and it takes 6 to 12 months before it is implemented for general consumption.
2026 will see further improvements for you.
No comments yet
Contribute on Hacker News ↗