Comment by stefan_lec
3 days ago
AA isn’t testing different harnesses there, they’re testing different models on the same common harness:
“We run Terminal-Bench 2.1 with the Terminus 2 agent harness in an e2b sandbox and report pass@1 averaged over 3 repeats per task.”
This is a different agent harness than those Strands was testing with, and there’s no guarantee the pass criteria matches up the same either.
No comments yet
Contribute on Hacker News ↗