Comment by seizethecheese

4 days ago

Then I'm really confused. Terminal Bench 2.1 scores on Artificial Analysis are like 80-90%. https://artificialanalysis.ai/evaluations/terminalbench-2-1

AA isn’t testing different harnesses there, they’re testing different models on the same common harness:

“We run Terminal-Bench 2.1 with the Terminus 2 agent harness in an e2b sandbox and report pass@1 averaged over 3 repeats per task.”

This is a different agent harness than those Strands was testing with, and there’s no guarantee the pass criteria matches up the same either.