Comment by langs
5 hours ago
I don't get it. Why benchmark the latency instead of recall/precision? Optimizing for millisecond-level latency is meaningless in the context of LLM calls. Accuracy is the tool's greatest value, yet there is no testing for it?
No comments yet
Contribute on Hacker News ↗