Comment by Certhas
6 hours ago
Good science, properly digested and presented takes time.
The idea that anything other than a breathless blog post about the latest model snapshot is useless is really poisonous to proper debate on AI issues
6 hours ago
Good science, properly digested and presented takes time.
The idea that anything other than a breathless blog post about the latest model snapshot is useless is really poisonous to proper debate on AI issues
Not sure how that vague truism applies to this paper.
Lots of papers have great results that don’t depend on the latest models.
However this case it’s problematic:
- They specifically make claims about the state of “current LLMs”. o3 is not representative of this.
- They ask are LLMs capable of X and arrive at a negative result.
If my claim were LLM’s can write coherent sentences, and my conclusion was positive. There would be no issue using old models because the result would be factual.
However, when you have a negative result that makes a claim about the current state of all LLMs and the ones you were using are not current, by definition it cause the whole conclusion into question.