Comment by lmeyerov
19 hours ago
It's been fun benchmarking AI investigations at botsbench.com . Part of it is checking for these kinds of issues - we recently started seeing contamination in our first generation challenge, and less obvious, agent sandbox escapes for other kinds of cheating. Fun times!
No comments yet
Contribute on Hacker News ↗