Comment by azan_
10 hours ago
I think treatment is not good benchmark - it requires lots of waiting and lots of regulatory work. The better benchmark - in my opinion- would be math discovery.
10 hours ago
I think treatment is not good benchmark - it requires lots of waiting and lots of regulatory work. The better benchmark - in my opinion- would be math discovery.
My point was that we won't care about benchmarks anymore because we would see an obvious and completely unprecedent increase in productivity (and I believe it will likely come from the same people who will develope such machine).
The reason most of the conversations are focused on benchmarks is because we are still in the age of weak AI.
That is an absolutely terrible benchmark. Inference over a bounded search space is not a good measure of what "intelligence" actually is. part of the reason they are using math and not something actually challenging like long distance interstate trucking is because it's so much simpler and easier than what make intelligence intelligent.
Just seems very weird to call getting Fields-medal-level results "inference over a bounded search space" and "not actually challenging".
Plagiarizing on a massive scale to generate works which appear to be Fields-medal-level results is not the same thing as inventing new conceptualizations in mathematics.
No matter how bodly they write the headlines, what has happened in mathematics using Large Language Models is very much "inference over a bounded search space" even if those bounds are immense.
For a comparison of true creation of novel conceptualization in mathematics is submit the works of Martin Hairer, one of which is Introduction to Regularity Structures, [https://arxiv.org/pdf/1401.3014] None of the so called, "novel math discoveries" by any LLM is as enlightening and expands the state of the art in math like any of his writings.
Claude Fable recently proved the existence of complex structures over S^6 (6-sphere).
If I had to guess, I think LLMs will be inventing highly original new mathematics within the next year. I think it will be approached as an optimisation problem, targeting how quickly LLMs can solve classes of maths problems as a function of the definitions they need to conjure up to do so.