Comment by seizethecheese
4 days ago
> With Fable 5, Strands harness cost 77% less than Claude Code and scored higher on Terminal Bench 2.1.
Terminal Bench 2.1 is saturated. Many token saving techniques would save money and score basically the same running Fable 5 against Terminal Bench 2.1. (They claim a better score but don’t say how much better. I’d bet my favorite hat that it’s not statistically significant.)
This is at least the fourth time I’ve seen a project hit front page with a “save money with same score on saturated benchmark” claim.
The scores are in blog post's bar chart. For Terminal Bench 2.1, Strands harness (Fable 5) scored 69.7 while Claude Code (Fable 5) scored 61.8. This is on high effort.
I hear you tho about saturation. We're working on a follow-up deep dive post with more harnesses, so could look into Terminal Bench 4.0?
Then I'm really confused. Terminal Bench 2.1 scores on Artificial Analysis are like 80-90%. https://artificialanalysis.ai/evaluations/terminalbench-2-1
AA isn’t testing different harnesses there, they’re testing different models on the same common harness:
“We run Terminal-Bench 2.1 with the Terminus 2 agent harness in an e2b sandbox and report pass@1 averaged over 3 repeats per task.”
This is a different agent harness than those Strands was testing with, and there’s no guarantee the pass criteria matches up the same either.