← Back to context

Comment by stefan_lec

3 days ago

AA isn’t testing different harnesses there, they’re testing different models on the same common harness:

“We run Terminal-Bench 2.1 with the Terminus 2 agent harness in an e2b sandbox and report pass@1 averaged over 3 repeats per task.”

This is a different agent harness than those Strands was testing with, and there’s no guarantee the pass criteria matches up the same either.