Comment by spongebobstoes
3 days ago
this is not a good measure of current model capability. we need to test agents in harnesses, not models with a single prompt
test Codex, not Sol. test Claude code, not Opus
3 days ago
this is not a good measure of current model capability. we need to test agents in harnesses, not models with a single prompt
test Codex, not Sol. test Claude code, not Opus
There are other benchmarks for that