← Back to context

Comment by epihelix

2 hours ago

That benchmark doesn't mean what you think it means. (See the test description that you quoted.)

A score of 51% means that out of the total answers the model failed to answer correctly (out of 6000 questions in the benchmark), 51% were factually incorrect rather than non-attempted or uncertain.

This doesn't mean that Astra hallucinated 3060/6000 answers in the benchmark! (The hallucination rate could be 51% in that scenario only if Astra failed to answer a single question correctly.)

If the model failed to give a correct answer to only 100 out of the 6000 questions, but gave a hallucinated answer to 51 of those rather than expressing uncertainty, that would also give a hallucination rate of 51%.

It's a useful metric, but not what you're looking for here. The "Score" or "Accuracy" benchmarks are more what you're after.

(The frontier models still generate hallucinations on this hard set of problems, but it's not as bad as you think.)