Comment by jll29
4 days ago
I don't know who downvoted the parent or why, but it's a fair question IMHO.
The answer is there can be dramatic difference running a benchmark one time, because LLMs are not deterministic. A proper methodology would ask each question 20 times and calculate the mean correctness across experiments.
The reason is that the temperature parameter introduces random behavior.
No comments yet
Contribute on Hacker News ↗