← Back to context

Comment by epolanski

6 hours ago

You touch a point I quickly skimmed in another comment.

Yes, the most valuable benchmarks and evaluations you can write are those that resemble your work.

The evaluations are extremely hard to write and test.

And yes, virtually all benchmarks are E2E one shots, they do not reflect multi turn processes or how most people interact with LLMs.

Which is why every Opus after 4.6 looks better on benchmarks, but is hard to work with interactively.