Comment by jaykru
1 day ago
I wrote something [0] that might answer a bit a few weeks ago. Today's slopdrop certainly is challenging my stubbornness, but everything I've seen so far indicates that these (hugely impressive, world historic) capabilities won't extend past verifiable domains. Math yields especially impressive results because it so broad and deep that essentially no person can know of all of its parts; pretraining and deep search capacity is a huge advantage. Gowers has recently written about these capabilities and gestured [1] toward some human capabilities lacking from the current frontier models, though he isn't convinced they won't develop soon. If you assume we don't get a total mathematical superintelligence (which to me seems already sort of AGI-complete) and only amplify the capabilities we have today, it's not obvious to me that we get takeoff from recursive self-improvement, unless you happen to believe that 1) we can clearly specify what AGI or ASI is 2) all of the requisite ideas are out there and need only be combined and/or optimized.
[0] https://dank.systems/posts/2026-09-15-ai-bear.html
[1] https://gowers.wordpress.com/2026/08/12/what-sort-of-maths-a...
> we can clearly specify what AGI or ASI is
We'll have plenty of time for this, while living off UBI.
> everything I've seen so far indicates that these (hugely impressive, world historic) capabilities won't extend past verifiable domains
But they’re already extending into politics, military, journalism, art, and many other fields that aren’t verifiable in any meaningful sense of the word.
The magnitude of improvement in unverifiable domains is small, mostly down to models doing more careful research before answering and hallucinating less. They are more thoughtful, but I expect you'd have to drop 2 major versions of Opus before you'd start to see most people really clearly be able to differentiate them.
> The magnitude of improvement in unverifiable domains is small,
What makes you say that? What is an example of a domain where the improvement is small?
I can't think of any at all. Compare something as unverifiable as "Make good music". Models now are many times better than 3 years ago.
4 replies →
>> slopdrop
Really? Do better.
this is the term of art in the mathematics community. considering that the vast majority of the results don't come with a typechecking lean formalization, i don't think it's off base at all either.