Comment by XCSme

8 days ago

But 80% sounds far from good enough, that's 20% error rate, unusable in autonomous tasks. Why stop at 80%? If we aim for AGI, it should 100% any benchmark we give.

5 comments

XCSme

Davidzheng 7 days ago

I'm not sure the benchmark is high enough quality that >80% of problems are well-specified & have correct labels tbh. (But I guess this question has been studied for these benchmarks)

kenjackson 7 days ago

Are humans 100%?

XCSme 7 days ago
If they are knowledgeable enough and pay attention, yes. Also, if they are given enough time for the task.
But the idea of automation is to make a lot fewer mistakes than a human, not just to do things faster and worse.
- kenjackson 7 days ago
  
  Actually faster and worse is a very common characterization of a LOT of automation.
  
  1 reply →