Comment by autuni
7 hours ago
this is not entirely related to the tweet but to the topic in general, this prompted me to check their repo again and saw this:
> The vast majority of results were obtained with the same procedure using an unreleased internal OpenAI model. On average, each result used three hours of ChatGPT Pro thinking compute with that model. Over the course of the evaluation, the model was posed approximately 4,000 problems. Aggregating the output into result families and manuscripts and requiring an appropriate level of significance led to the catalog outlined above.
seeing the full list of problems would be the most interesting part of this whole situation. it could give some insights into what kind of attributes of problems cause issues / are easy to solve for LLMs. (edit: they posted results for ~700 of the 4000)
> seeing the full list of problems would be the most interesting part of this whole situation.
This is a "complaint" that Tao had (I think it was on his blog) is that if mathematicians could see the failures, it might provide insight of where/how the models struggle. Of course, it's unclear if these failures can be addressed with more chips/training/etc.
* "the most interesting part" => a tangentially interesting part