Comment by iainmerrick
4 hours ago
I had a similar thought -- rather than fixed benchmarks, you want dynamically-generated tests, specifically designed to exercise newly-exposed corner cases. So the way forward might be antagonistic benchmarks generated by another LLM.
No comments yet
Contribute on Hacker News ↗