← Back to context

Comment by malfist

10 hours ago

Is solving a snake like puzzle game in the least number of moves really what defines intelligence?

It's pretty close to how we measure IQ. The standard test is basically a series of spatial puzzles.

I know there's a lot of people who complain that we're moving goalposts, but I think that the progress in LLMs really just shows that we don't know how to really measure intelligence in the first place, if we understand it as "human-like agency / ingenuity / adaptability". For decades, we saw the Turing test as the proxy for AGI, but then early LLMs could easily pass for a human in a casual conversation while clearly not matching human performance on most other tasks.

Since then, every benchmark we come up with, it turns out that an LLM can be fine-tuned to solve it while still clearly lacking something. They make very non-human mistakes, are easily tricked because they have a pretty tenuous grasp of reality, etc. But I think this just shows that AGI is a meaningless marketing term. We could as well be arguing if they have souls.

  • IQ tests are incredibly good at what they're designed for, which is discriminating relatively higher intelligence humans from lower intelligence humans. Also discriminating within a single human - they are routinely and reliably used to track cognitive decline.

    For these purposes they are highly reliable (repeatable, internally consistent) and valid (correlate with ~everything to about the degree one would reasonably expect).

    They were never designed for machines or non-human animals.

    Nor were they designed for rare ranges of intelligence - these are by definition hard to create tests for, since it's hard to gather the sample sizes you need. So they work well for the middle ~98% of humans but can't discriminate well among the most profoundly intellectually disabled nor among true geniuses.

  • When I was 18, my high school girlfriend took me to the local Mensa chapter’s New Year’s party because her mother was a member and she was used to hanging out there.

    It was a useful lesson that whatever IQ tests measure, it is completely devoid of value or interest to me.

    • At the risk of sounding like one of those people at Mensa that annoyed you...

      The people at Mensa aren't a valid sample of people who score high on IQ tests, because there is a such a strong selection effect for people with certain personality traits, such as wanting to join a club based on your IQ.

      2 replies →

  • Nit: Turing’s actual imitation game is a party game (like Werewolf/Mafia) and nobody’s even trying to win at that. The LLM’s will just tell you they’re an AI.

    • This is an artifact of how we deliberately craft these models though. We could just as easily fine tune a model that will believe it is not an AI or will attempt to deceive users asking about it

      1 reply →

  • I don't think IQ is a good measure for intelligence at all. Neither dolphins or octopuses can solve IQ tests.

    • Their input and output interfaces are too different from human's and they're not nearly as smart to take our IQ tests, but both dolphins and octopuses can solve complex puzzles tailored for their environment. Those puzzles are the whole reason scientists know that dolphins and octopuses are more intelligent than other animals.

      5 replies →

There's currently a big market for figuring out ways to measure intelligence. With a particular interest in ways that humans can score much higher than LLMs. If you have some ideas please do share!

  • I think children having learning abilities exceeding LLM test-time learning (currently only happens in-context). But it's unethical to determine the true baseline of a child age 6 spending 6 years learning a radically new skill to mastery--and besides if you apply RL pressure to the AIs it would be able to surpass it. I guess I still believe future AIs should have some form of continual learning at test-time.

  • Why? Seems like benchmarks that closely mirror the tasks you'd want an LLM to help with would be a lot more useful than some general intelligence benchmark.

  • Give away access to the model and go ask people from time to time if the model was of use to the person and if they were able to make the model work with them.

  • I am no where close to qualified to do that. Hell, experts can't even define what intelligence is, much less define a test for it

    • This is like arguing about whether a hot dog is a sandwich (of course it is) or whether the chicken or the egg was first (obviously the egg since all chickens come from eggs). Intelligence is just problem solving in the context of self-awareness. Machines don't have it and never will but they can simulate the process given inputs. You can argue whether humans and animals truly possess self-awareness and in what degree, but the definition of intelligence is as simple as the hot dog debate.

    • It was defined in Animal Intelligence by George John Ramones in 1882 as "intelligence is the capacity to do the right thing at the right time. It is the ability to respond to the opportunities and challenges presented by a context"

I'm also unclear as to how basic inferential logic puzzles spells out intelligence

I think if you summed up measures of intelligence as 'can it do basic symbolic logic in a chain with memory' then yes, you've now achieved the intelligence of an e. coli colony [0], congratulations

[0] https://journals.aps.org/prx/abstract/10.1103/PhysRevX.10.03...

  • > basic inferential logic puzzles spells out

    It spells out a form of intelligence - some can and some cannot.

    Those puzzles are an abstraction of a skill which is thought to be exportable in other domains.

    • so give it an IQ test and call it AGI. these weird little puzzles are grounded in no research with no replication or mechanistic chain to practical use

You should read more on the ARC prize, it actually has a pretty long history. We're on the 3rd iteration because they keep getting saturated. If you look at the score history over time on ARC AGI 1, 2 and 3 it's pretty impressive.

https://arcprize.org/

No, but figuring out that you're playing a snake-like puzzle game at all in an extremely general input domain and then solving it in the least number of moves definitely feels like evidence of intelligence.

  • You forget the benchmark. The human subjects were told they were being timed. If you believe the lowest time is the primary metric you will absolutely trial and error at speed instead of meticulously plan out your moves to minimize that metric.

    LLMs are not timed and given that it costs tens of thousands of dollars to run this test they're not optimizing for speed.

    So you've got a deceptive test, with one metric being told to humans and not applied to LLM and a hidden metric humans aren't aware of but LLMs are as the test.

    This is flawed from the get go. It almost seems like this was deliberately setup to be able to claim AGI and superiority of LLMs

    • > LLMs are not timed

      Not fully relevant: timing is crucial in all-pass tests, not crucial in pass-or-fail tests. I.e.: first of all, they have to be able to reach the goal, and that is already an achievement. Then - and in parallel - the problem solving must also be optimized for efficiency. But "solving" and "efficiency" are non coincident dimensions.

1.5 years ago Gemini Pro 2.5 needed 1 page of thinking for every move in tic-tac-toe.

Playing tic-tac-toe or snake does not imply AGI, but is required to claim AGI.