Comment by hnd9q09qk4

9 hours ago

Exact match graders are the real variable here, we had F1 or a judge model swing passage QA scores by ten points on identical answers.

A good point. I initially had a model for a judge, but it seemed to give very lenient scores. I'm open to learning about how best to benchmark the phenomenon, though.