Comment by AndrewOMartin
2 months ago
I think the "state of the art" of measuring the quality of outputs was to send the same task to multiple "agents" and only accept answers if over a certain amount agree. With some human review and reputation scoring sprinkled on top. It was a while since I was in this field though
This approach does work when there's a clear answer but what about tasks where the correct answer is multi-modal? Incentivizing agreement works only for tasks where there's clear answers.