← Back to context

Comment by humbleferret

3 hours ago

Nice writeup! I imagine these results change as harnesses are updated, so you'd need to frequently rereview.

I'd love to see a tiny, reproducible benchmark repo that anyone can drop on their own hardware and then run against all harnesses at once to compare the per turn prefix token count, time to the first token, experienced tokens/sec (and prefill), cache reuse % and a pass rate on a deterministic set of small tasks. I think it could also be useful to have some way to share results and hardware for others to compare.