Comment by mvkel
12 hours ago
Take it from the mouth of the creator of ARC-AGI:
When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted"
That was 6 months ago, so the progress that Astra represents happened about 2x faster than I anticipated. I think the speed of progress will surprise a lot of people, and what the new models can do will challenge the views of AI that people developed by using prior generations of models.
I feel like AGI's definition got watered down, and these tests do not cover the original definition, what is your definition and thoughts on aligning with what all of us understood from the original claim?
I feel like this test is just helping someone like Sam Altman pretend like he implemented AGI as originally pitched for an IPO when in fact, he has not. Shameful.
> AGI is essentially the equivalent of a median human that could be hired as a remote co-worker... capable of performing any task that one would be satisfied with a remote colleague doing via a computer.
- Sam Altman on AGI
I can call a coworker right now and have a real time conversation with them without feeling like I’m talking to a frustrating machine. Most of them, anyway. But I guess that’s “moving the goalposts”.
I can call coworker right now and have conversation so frustrating that I wish I was talking to machine instead.
5 replies →
I feel like my odds are better with the AI than with random humans.
1 reply →
Reminds me of Turing test. https://en.wikipedia.org/wiki/Turing_test
I think they're claiming it's achieved by text models, not voice models, fwiw.
> I can call a coworker right now and have a real time conversation with them without feeling like I’m talking to a frustrating machine.
These machines were designed with the express purpose of being condescendingly sycophant to the point they hallucinate just to state "you are absolutely right".
No wonder some people even find these chatbots to be wife material.
1 reply →
I dunno, have you tried the voice chat in paid ChatGPT?
1 reply →
Here's another definition of AGI from Sam Altman:
https://www.nytimes.com/2023/11/20/podcasts/hard-fork-sam-al...
Sam Altman: Let’s say we make an A.I. that is really good, but it can’t go discover novel physics. Would you call that AGI?
Kevin Roose (New York Times): I probably would, yeah. Would you?
Sam Altman: Well, again, I don’t like the term, but I wouldn’t call that done with the mission.
So, "really good" is the boundary now, whatever it means.
2 replies →
If the new AGI benchmark is "be Einstein/Feynman" then we've hit AGI.
26 replies →
What is novel physics?
25 replies →
If it makes you feel any better the curmudgeon who drives the ARC-AGI tests feels the same way, that's why we're on 3 and I'm sure we'll see 4. Also, we can all avoid calling it "moving the goalposts" so no one feels talked down to.
It’s better to call a spade a spade.
People need to stop redefining and trying to capture the term AGI. None of this is AGI. Not even close. Can it do things an AGI could do? Yeah some of it, but the difference really matters. These things still regularly fail, and gaslight about answers to questions like "how many r's in strawberry" or "s's in espresso".
An AGI wouldn't struggle with that.
The last version to fail on those questions was GPT 4.5.
Meanwhile most humans fail to correctly answer how many f's are in the sentence, "Finished files are the result of years of scientific study combined with the experience of many years.".
9 replies →
A lot of movies gave the impression that making phone calls and balancing checkbooks would be the easiest tasks for consumer AI to solve, while math and science might require exceedingly advanced AI. Turns out to be the opposite: computers are great at math and bad at conversation.
1 reply →
> "how many r's in strawberry" or "s's in espresso".
And what percentage of your red retina receptors are firing?
The "number of letters" critique was broken before it was introduced the first time. The models were specifically designed with preprocessing to not be able to perceive their input as strings of letters. Blind people are not dumb. (True as a pun and in context.)
Sub-access sensory questions, or do-you-know-a-fact questions (which is what spelling becomes when you can't see the letters, and are not specifically trained to match all token encoded words to their letters) are not intelligence questions.
These are like saying someone isn’t human because they have a speech impediment or an auditory processing disorder.
AGI doesn’t mean infallible, it just means it can have a reasonable crack at things it hasn’t seen or done before.
1 reply →
I'm going to go out on a limb and guess that there are plenty of savants who can't tell you how many r's are in strawberry.
> gaslight about answers to questions like "how many r's in strawberry" or "s's in espresso".
this has been debunked too many times to bother rebutting. they struggle with those things because of the way they are.
it's completely irrelevant.
21 replies →
> I feel like AGI's definition got watered down
Typical result of venture capital and too many bag holders unfortunately.
I wonder if Altman's definition also includes taking on the same liability as a coworker would.
Probably not.
Would 100% need to be fully insured for liability, with a sizable war chest that OpenAI cannot even afford.
15 replies →
ARC-AGI was never "if this benchmark is saturated, we're at AGI". It was always about crafting adversarial tests that humans are good at, but current AIs are bad at. Point out the gap, get AI teams to attack them.
In practical terms? They usually get solved with a bigger badder LLM. "New ideas are needed?" Nah - ten times the params, ten times the test time compute.
ARC-AGI-3 was more of a failure in that regard than -1 or -2, because even on day 0, an off the shelf LLM with a harness could get 50%+. And messing with evals by forbidding "LLM with a harness" from scoring? Yeah no, that was just bad.
>When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted"
You're treating an off-hand comment by an ARC 3 researcher as some sort of a precise AI capability acceleration benchmark. Can we leave casual anecdotes (even from researchers) out of the discussions please?
François Chollet wrote in February that he expected ARC-3 to be saturated in "about one year".
"Frontier models today perform very poorly with a minimal harness. However if big labs start directly targeting the benchmark like they did for ARC-2, numbers will go up fast."
https://x.com/fchollet/status/2022054537293705260
Well it's not exactly saturated when OAI refused to use the harness explicitly provided by ARC-AGI. I'm not really familiar enough with the benchmark to declare whether it's a perfect measure for AGI, but I kind of doubt it is.
I think you're misunderstanding. Astra is at the top of the official ARC-AGI leaderboard, with an ARC-AGI approved harness. It's not a harness specialized for ARC-AGI. It just does the same thing the regular ChatGPT interface does: keeps conversation history across turns and compacts when it gets too long. Without the harness, it loses its entire context window every move. That's not how humans work and it's not how any real AI service works.
2x faster at what resolution? is 6mo vs 1 year really that different? Usually surprise comes in order of magnitude mismatches in expectations.
Yes, 6 months vs 1 year is huge for technology that has gained wider adoption only recently.
Adoption means nothing.
2x gains from a mature technology would be surprising.
2x gains from a new tech would still be called “low hanging fruit” in another setting.
I don’t read enough to know in what ways the training / other technical steps have really advanced.
You'll also need to compare the amount of compute used now and then, which seems exponential to me.
We are nowhere near AGI. They all talk the same, they can’t help but try to please and affirm us, and if you engage them for too long they become incoherent. They are facsimile machines. They are xeroxing language - but not even, because we can’t even duplicate our results. Too many people mistake the black box quality for magic.
Super useful, incredible tools, but not AGI. Try and roleplay a dialogue with one, make it whatever character and scenario, and see if it can sustain a coherent conversation for more than 30min with you AIM style (aol instant messenger, if that isn’t clear). Expert mode: never correct or adjust it mid conversation.
I’m not even talking about repetition and predictability. It’s nothing like talking to a person. And in a short amount of time it literally can’t form a coherent sentence.
Why is a >30-min context-length a requirement for AGI, or the naturalness of a human conversation?
Human brains have difficulty reasoning about exponential growth.
They keep saying that. I'd say it is more like human brains that don't remember high school math have trouble with it.
Reasoning is a strong statement here. But it is fair to say that it usually is not intuitive to us.
An example is if I gave you a huge sheet of thin paper (huge so that folding isn’t an issue) - how many times could you fold it in half until you couldn’t physically do it anymore? Could you do at least 10? Try this with random people and you’d be surprised how many say they could do 10 easily.
Or the chess board question. Works to rather get the financial equivalent of starting with a penny and then doubling it for every square on the board or a million dollars for each square? Again, if you ask people to pick one without giving them the time to work it out they will usually pick the million dollar per square.
1 reply →
If something at rest is accelerating at 9.8 m/s^2, how long in seconds will it take to reach 10% of c? Answer to the nearest order of magnitude - will it take approximately 1000, 10k, 100k, 1000k seconds?
I’m sure you know this is an exponential growth question but have no intuition of the answer.
3 replies →
Are there plans for ARC 4?
[dead]