Comment by mrieck
2 hours ago
This metric seems to be asking if years of research by specialists could be emulated by a few LLM api calls. Reminds me of this meme:
2 hours ago
This metric seems to be asking if years of research by specialists could be emulated by a few LLM api calls. Reminds me of this meme:
The difference to this meme is:
Often people who are critical of whether AI can lead to research breakthroughs are experts in the respective area, who nevertheless are afraid of their future career prospects in academia (getting a permanent position in academia is hard and it is deeply political who gets such a position).
This people are thus not scared by AI per se (it's basically their daily job to devise innovations that advance their field), but their fears are that
- because of the hype around AI the research into which they invested years, often decades, will be considered "unimportant",
- incompotent people in decision-making positions will think researchers can be replaced by AI.
It's still a bit weird though. For obvious reasons, The MO for benchmarks like this has been that the models remain pretty low until suddenly it's done. So even for a 'should i be worried yet' reality check, it's pretty terrible. You can't really keep track of what models are actually able to do. By the time you can replace years of research by specialists with a few api calls then...
> incompetent people in decision-making positions will think researchers can be replaced by AI.
This is a certainty, not a fear =[.
Love that. I myself first thought. How specialized is this and at what level of human complexity is this team working.