Comment by xnorswap
5 hours ago
A really frustrating partial presentation, given an apparent lack of testing with a spread of efforts for each model.
Given that there's no reason to believe that Fable's xhigh is comparable to GPT-sol's xhigh, or Opus xhigh, for that matter, it would be far more useful to see the effort level where these tasks no longer achieved their goals.
These benchmarks are done with incorrectly and missing a lot of baselines. The article seems very vibe coded too.