Comment by kosh2
5 hours ago
> That's not going to reach AGI,
It has not been even 4 years since ChatGPT hit and LLMs + Transformers + Whatever they do has gotten us to solving millennium problems.
4 years ago, a program that could create photorealistic pictures, talk to you in any language of the world and solve the hardest math problems that we know, we would have called it AGI.
Now I don't know if what we have is AGI or not but I do not understand how you can see what has happened in the last 3 years and say "it will not get us there" no matter what "there" is.
> 4 years ago, a program that could create photorealistic pictures, talk to you in any language of the world and solve the hardest math problems that we know, we would have called it AGI.
I keep seeing this idea and I don't understand the reasoning behind it.
I think it could be a bit like saying if you showed someone 500 years ago a smartphone they would likely conclude at first it was magic. But once you had some time to let them use it and tell them how it all worked on a high level they would eventually obviously realise, no, it's not magic.
I guess just in the same way if you presented current LLM tech out of nowhere a few years ago to someone who'd never seen it, I concede they may be likely to imagine it was AGI in that first conversation, depending on their background.
But after using it for a bit and learning what an LLM is etc they'd land exactly where everyone is today - a great technology useful for some things, not AGI, not magic.
I know how they work pretty well and most days I still have moments where I am struck by how bizarre and magical these thinking machines are.
The fact that a monkey can pick up an artifact a size of a small mirror and do a video chat, see and talk to another monkey on tge other side of the planet is absolutely bizarre. There's no good reason why this should be allowed.
I'm finding it more bizarre, if compared to creating a virtual talking monkey from scratch, with a little bit of Python and a large calculator.
What would convince you that it is AGI?
I always here things like "oh it's useful but dumb on some things", but it's just vague.
What is the test? What is a question that it fails at compared to humans? And no, you can't just say "find me the cure for cancer", but I believe there is probably enough intelligence in the weights that there is likely a cure in there with enough compute and the right questions.
When it can replace a white collar remote worker without humans team realizing they work with the machine. And no, Agents are not like this whatever initial prompt and set of skills you give them. They still won't progress, learn and apply that knowledge like a human.
The fact that you need to ask "the right questions" is why it's not AGI. A general intelligence should be able to ask of its own volition the interesting questions required to advance its goals.
I can’t predict how a novel intelligence could prove to me that it is intelligent. A novel intelligence would have to work out how to do that for itself
i think we'll know agi when we see it, but we can't really predict what that will look like
> 4 years ago, a program that could [...] we would have called it AGI
If you had told someone in the 1800s that a machine could instantly multiply 100 digit numbers, that would have been considered dazzlingly intelligent. And yet we are not that dazzled by our calculators today (despite how useful they might be!).
Are you trying to explain how things once considered dazzling get normalized over time? Because otherwise this is a non-sequitur and has no bearing on the trivially verifiable, exponential explosion of capabilities we have seen in the last 4 years.
I keep saying this, until ChatGPT came out 4 years ago it was basically unimaginable that a single model could do any of, let alone all, the things they are doing today. Like, seriously, go take a look at the state of the art in NLP and NLU, the very first challenge in getting computers to even “understand” natural language, let alone other things like reasoning. Everything it does automatically was once a heavily experimental deep research field with long glorious careers for the researchers.
And now it’s all gone because the Bitter Lesson won again. If that’s not general enough to qualify for the G in AGI I don’t know what it is. And we’re sitting here going, “But it sometimes writes bad code though.”
Speak for yourself, I am dazzled by calculators!
In any case, I think this misses OP's point that LLM capabilities have rapidly made progress towards being more generally intelligent and capable, which is not true of most tech advances.
This is a motte & bailey moment. Parent comment stated something much sharper, that I responded to:
> I do not understand how you can see what has happened in the last 3 years and say "it will not get us there" no matter what "there" is.
--
Your statement is something much weaker, and I would still question what exactly "general" means when AI capabilities are commonly accepted to be so "jagged".
1 reply →
>I think this misses OP's point that LLM capabilities have rapidly made progress towards being more generally intelligent and capable, which is not true of most tech advances.
A PC of today can accomplish many more "general" tasks than one of 40 years ago. Much of the "why" is because of the huge infrastructure built up around them in the meantime. The abilities of LLMs to accomplish those same tasks through the PC is heavily piggybacking on that (both in the specific, with the existence of all the APIs and tools; and in the generic, using search engines to find specific sources and using that for instruction or troubleshooting).
In the world of "agents" much of the improvement appears to have been on a specific set of skills: impersonation of an 'I' that wants to accomplish a goal, and synthesizing existing information from documents with trial-and-error execution loops to move rapidly toward a solution much faster and with less boredom than a human would. The quality of the output when there is not a rapid-evaluation-and-validation harness lags considerably.
It's incredibly powerful automation but doesn't appear to be trending towards Matrix-style conscious AIs. The quality of an individual method written by the agent also is not particularly advanced compared to GPT-4 in early 2023, as far as I can tell—I was dabbling with trying to make such harnesses back then, where a major challenge was that the model itself was bad at staying on-track in a conversation, so instead much of that logic was moved to deterministic code, which was much more limited as it was super-tedious to enumerate all the necessary tool calls/etc to find its way out of corners. Staying on task is much better now, as is "read compiler error, fix try next thing" harness loop-handling. But the output remains—across Fable, Astra, whatever else I've tried—"iffy" in terms of the actual code structure on the first pass output. You can set it then on a different task to review and clean up the code, and it can do that well too, but it is a curious gap of generality where the "create" focus is much more limited than the "review" one (and conversely the "review" focus can make suggestions, but if it goes deep down the well of implementing them, loses that big-picture again).
If it kills us all, it will because someone decided to give the trial-and-error-loop-machine access to nukes or similar. The blame for that is on the "someone" not on some sort of "rogue" AI.
(I wonder if re-watching Terminator/Terminator 2 would support this sort of interpretation of it. Unlike in the Matrix, I don't think we get much sentient-AI POV/infodumping. Is it a plausible universe for "someone made ChatGPT control a fleet of soldier robots and gave it a bad harness with an insufficient sandbox"?)
1 reply →
Superhuman performance at chess probably would have blown people’s minds in the 1950’s. We’ve since learned that sometimes intelligence can be narrow and sometimes it can be spiky, even if you can have a decent conversation.
It’s hard to point to anything and say it’s impossible. AGI doesn’t break any laws of physics. But some things like driverless cars can still be a long slog to get to widespread deployment.
Waymo is the only self driving car that works well enough to even try widespread deployment with pricing that isn't going to be an operating cash flow disaster. But they're going to have to at least double their footprint to put a meaningful dent in the billion dollars a year Waymo spends on R&D.
Waymo isn't LLM powered. Everyone who thought LLMs would make Waymo Driver obsolete were wrong. In part because Waymo can't be spiky. In part because practical applications of AI require a lot of time and effort.
I am amazed at how usable and useful coding agents have become since they were a hot mess about a year ago. I would not be at all amazed if there are only one or two other use cases that have the same favorable evolutionary trajectory.
Those are just the same capabilities than before, but with a much bigger compute power and training data behind it.
AGI can't be reached by "training harder" as, the way I see it at least, it requires a qualitative leap, not just quantitative.
We are getting a machine that better navigates across the information in its training data, we are not getting a machine that can think out of that training process, even if it can fool a few people at that.
The entire field has repeatedly said that for many decades.
https://aeon.co/essays/how-close-are-we-to-creating-artifici...
https://xkcd.com/605/
Here's an actual log-scale trajectory with a few dozen real data points.
https://metr.org/time-horizons/
(May 2026, no longer applicable)
> 4 years ago, a program that could create photorealistic pictures, talk to you in any language of the world and solve the hardest math problems that we know, we would have called it AGI.
No. General means general.
[dead]