Comment by ggreer
20 hours ago
Which LLM specifically? If it’s cloud based, you should be able to share the chat, right?
But seriously, I am still waiting for someone to wager that AI won’t be able to do a specific cognitive task in the next 5 years. This fact should be evidence enough that we have no idea how far AI capabilities will continue to advance.
I asked both gemini and chatgpt "do frontier models still have trouble counting letters?" and the first word of both responses was yes.
The fact that no one is taking you up on that bet I don't find to be particularly persuasive. I suspect there will be plenty of cognitive tasks LLMs struggle with in 5 years, maybe even 20. But I wouldn't hazard to guess which, I don't think anyone is capable of that level of foresight.
This is so strange I tried it on Gemini Flash: "Yes, but significantly less than before." When you read beyond the first word it explains where LLMs might fail and why.
Ok, share links to the conversations with both models. I asked ChatGPT Astra 6 medium effort and it said, "Much less than they used to." and provided stats showing how accurate they are.[1]
1. https://chatgpt.com/share/6ac5e4cc-02f0-83e8-8f05-99a7ea2bf9...
https://share.google/aimode/iKkrZtVYo4DSielWs
https://chatgpt.com/share/6ac63b6a-481c-83e9-a8fa-a13ce7402d...
I used whatever the default free model and thinking time was. If progress was really as fast and continually cheaper as some worry it is, wouldn't we expect free models by now to know (or even perform) what frontier models were capable of as much as 2 year ago?
This deep in the "comparing logs" tangent we risk missing the point. It's not what exactly frontier models are capable of at this particular point in time. But that there's entire categories of problems that seem easy to us which LLMs really struggle with. We've stumbled on several just a few replies into casual conversation. (Can they count? Can they know if they can count? Can they reproduce results? How quickly do new capabilities filter into free models? And that's just what's come up naturally, if we wanted to pick adversarial examples there's more to choose from.)
So while there's a number of difficult problems that are easy for LLMs (like bulk generating lean proofs), there are plenty of things where progress is not so impressive.
If LLMs can struggle so much with such easy problems, what hard problems have we yet to discover that they'll struggle with? The fact that no one knows, 5 years in advance, what those problems will be does not mean the chance of them is zero.
So far progress on the things LLMs are good at is fast and easy. It's like fire in a room full of oxygen. But once the low hanging fruit is gone, and the oxygen is out of the room. How fast will the fire burn through steel walls?
In my opinion it's a mistake to look at only rate of progress on one type of problem (whether it be what LLMs are good at OR what they're bad at) and assume progress on all tasks will progress at that rate indefinitely. Isn't there a saying about exponential curves, in nature, all being sigmoids eventually?
I guess we'll just have to see. I wish you good luck with your wagers.
2 replies →
Solve the Collatz conjecture in the next five years? If humans publish significant advances during that time, and A.I. copies it, then yes. Otherwise, I'd definitely bet money it won't happen. I'll give you 10,000 brownie points if I'm wrong.