Comment by frabia

4 months ago

Tangential to this: what are the most reliable benchmarks for LLM in coding these days?

1 comment

frabia

I found Terminal-Bench [0] to be the most relevant for me, even for tasks that go far outside the terminal. It's been very interesting to see tools climb up there, and it matches my own experimentation, that they generally get the most out of Sonnet (and even those that use a mix of models like Warp, typically default to Sonnet).

[0] https://www.tbench.ai/?ch=1