Comment by WhitneyLand

2 hours ago

Not sure how that vague truism applies to this paper.

Lots of papers have great results that don’t depend on the latest models.

However in this case it’s problematic:

- They specifically make claims about the state of “current LLMs”. o3 is not representative of this.

- They ask are LLMs capable of X and arrive at a negative result.

If their claim was LLM’s can write coherent sentences, and their conclusion was positive, then there would be no issue using old models because the end result would be factual.

However, when you have a negative result that makes a claim about the current state of all LLMs and the ones you were using are not current, by definition it draws the whole conclusion into question.