Comment by herrvogel-

4 hours ago

Is there a good place for harness comparison/scoring?

https://www.harness-bench.ai/leaderboard.html

This provides a method, but the data looks stale and perhaps a bit thin compared to say, Cursor, or even AntiGravity data.

  • never heard of nanobot. how reliable is this benchmark?

    • I can’t speak to the veracity of the benchmarks but it appears their methodology is sound. Nanobot has 47k stars, fwiw https://github.com/HKUDS/nanobot

      It has been more of an OpenClaw or Hermes alternative than a coding agent like OpenCode or Pi, so it’s likely to do well given less context bloat.

    • IMO it's not. It's benchmarking GPT 5.4 and Opus 4.6. It's also missing Claude Code... one of the most popular harnesses (the most?)

      1 reply →