Comment by marsx-dev
1 hour ago
I wonder if this is partly because “pelican riding a bicycle” has become a kind of benchmark prompt by now. If so, could the models actually be getting better at the benchmark rather than getting better at following the prompt?
No comments yet
Contribute on Hacker News ↗