Comment by bbor
5 days ago
First: this is really great technical writing, especially when you get into the rebuttals. Firm & clear without polemics -- props, and thanks for open-sourcing!
That said; I don't have the time, energy, or anywhere near the expertise to challenge you on the DL specifics, but I feel compelled to add another voice to the chorus of doubters nonetheless. Using other ARC examples at runtime (effectively, yes?) for "transduction" may not violate what some officer behind ARC said on Twitter --and is certainly a fantastic tool for certain problem spaces-- but it just seems like a glaring and unavoidable philosophical problem in this one. My issue isn't with using the eval set per-se (though that obviously sets off well-justified alarm bells), but rather building an AGI system whose performance relies on the arbitrary size and shape of this particular dataset.
There's a lot of ways to frame this, but given the transduction citations the most appropriate is probably the AI winter's infamous 'Frame Problem':
You say upfront that this works in the first place because ARC has "very few samples... in a high dimensional space"; to me, that seems like an extremely strong indicator that the datasets are not intended to capture anywhere near the full semantic space that we would consider relevant for AGI. If true, your approach would indeed be ""cheating"" by using an arbitrary & unavoidable feature of the dataset (that they didn't have the time or money to craft 100 million high quality cases by hand instead of 1000) to solve the frame problem upfront for you. This would explain why you don't even need a full LLM here -- that's the unsolvable problem that LLMs solve for us.
That is... even if the ARC train+eval sets contain the sum of human intuition between them, superficial differences IRL would render your model unable to identify which examples are relevant to which problems, and thus unable to transductively reason.
In plainer English: surely you'd agree that your model would do worse if we swapped it out with Opus behind the scenes than the next-highest-scoring ARC model would do in the same position, yes? For coding, research, dumb questions, SVG pelicans -- the lot?
If so, that seems like hard proof that this scores high on a benchmark at the cost of the benchmark itself. Like, if this transductive approach leads to ARC1 being claimed (which I thought it was ages ago but :shrug:), they'll either have to abandon the whole benchmark or ban this approach retroactively.
If not... well, I guess I encourage you to try it! It seems like you'd need 1000 truly stellar hand-picked examples to transductively cover that whole space, for one thing.
Thanks!
I think this is a great question. I have some thoughts on this but no hard evidence (neither does anyone else!)
Your argument relies on the AGI system being the model arch + weights. I think that the weights are irrelevant. The training algorithm is what is AGI: You choose/find a training set that covers a task, and then some form of deep learning with a big neural net.
For LLMs (which is a weak kind of AGI), this is freezing a model after NTP pretraining + RL postraining. This allows us to do well on a wide distribution of tasks.
But if we had a scoped task, I think its possible to just do the same thing with (1) a smaller dataset that covers this task + (2) not freezing the model (ie. test time training). This is easy for ARC because the data is small and has been curated well yes, but i don't see why can't this apply to more complex/ill-defined tasks (obv lots of things unsolved to make it work today)
> even if the ARC train+eval sets contain the sum of human intuition between them, superficial differences IRL would render your model unable to identify which examples are relevant to which problems, and thus unable to transductively reason.
In this hypothetical world, the dataset becomes incredibly large, and training on it makes it close to an LLM. You can then finetune transductively and we get the same thing (other approaches already do this iirc)
--
Note: BTW I'm not claiming that this is an AGI system, the benchmark creator also was clear that ARC is not sufficient for AGI (his goal was just to point out unsolved stuff, and the stuff ARC-2 pointed out was demonstrated very clearly by large reasoning models).
In the above work, I just wanted to show that AR transformers without pretraining can perform really well on this benchmark, which was incredibly non-obvious before.
The claims in this argument are separate and I haven't proved them yet
Kind of hijacking, would you say that LLM's have solved the frame problem?
To me, the frame problem is: Can you function in an open vs closed world, and to me the answer is yes, LLM's can definitely function in an open world where the rules are fuzzy, changing, undefined, etc. At the very least, much better than all GOFAI approaches by far.
The issue is now grounding - It can "function", but what would it take to "ground" them? A personality, maybe? Actual consequences? Making them interact only with constrained tools that are formally verified?
Right now it's a combination of harness engineering, and ml philosophers arguing about compression leading to the "objectively correct intelligence", whatever that means.
I think LLM's are "A[x]I" right now in the sense of "they have the capability to integrate with everything" - but obviously you can argue how much this actually reflects "A[x]I" (if you gave someone integration with everything, is that really your success or people handing you it)? But they are still missing some oomph factors that need to be clarified IMO. Maybe it's something as "mundane" as just having actual persistent memory, or maybe it's some deep philosophical thing like qualia. Who knows.
One of these is a much weaker claim than the other.
> yes, LLM's can definitely function in an open world where the rules are fuzzy, changing, undefined, etc.
> At the very least, much better than all GOFAI approaches by far
That's true. Again, I currently view LLMs as a function of integration - what they may lack in "intrinsic smarts", whatever that means, they can tool call and we build capacities (and they build capacities!) around them and to some degree can reason and be creative.
I do think the first claim has real merit even if it's not 100% on par with humans. Second claim is just true.
I'm glad you did! Of all the procrastination techniques I have mastered, engaging smart people on HN about artificial cognition is probably one of the more useful ;) Apologies in advance for the diatribe(s) -- I think about this stuff a lot.
First, a nit: I would describe your definition as a valid transformation (isomorphism?) of the original phrasings, which were about technical context and epistemological belief[1]. I mention this because A) (semi-)symbolically providing context to LLM calls is the challenge at the core of harnesses, routers, pipelines, 'orbs', and a long list of other marketing terms that must amount to an ∞-B\$/y industry by now, and B) it shows how arbitrary the phrasing was, at the end of the day. (I also prefer this to the wiki article btw, for the curious: https://plato.stanford.edu/entries/frame-problem/)
My actual answer here is a resounding "yes" and "no" at once, in the exact same way that the Turing test is both so obviously surmounted in 2023 (post-RLHF) to anyone applying 20c standards, while also somehow being so far away that we're not sure it'll ever be possible. The key is to 'dissolve the binary' for both, if you'll excuse the phil-ism: Turing's 1950 paper Computing Machinery & Intelligence was never intended to prescribe some yes/no evaluation procedure, and the people frustrated by the Frame Problem were not worried about a single yes/no "Frame Test", either.
Instead, Turing settled on behavioral comparison on an intuitive, human level as the best shared dimension to test, but only after calling Ed Zitron "absurd" and leaving room open for ESP & ghosts to end up proving souls (one of those is literally true, the other only figuratively).
By this metric, current LLM-backed agents are clearly able to behave like a reasonable-ish human over a long-ish timeframe -- that's just objectively an incredible achievement IMO, even from 2015 standards. The promised inversion is the retort that invites, namely: the '-ish' makes all the difference! An artificial mind that behaves in completely alien ways randomly is a much less useful tool even if those events are rare; ditto for an artifical mind that loses coherence across """mere""" days.
I'm cutting this as much as possible, but hopefully it's clear why all the above applies to the Frame Problem, too -- just replace 'behavioral' with 'epistemic'.
I think your use of 'ground' is nice and understandable, but is conflating too many things to work as a summary of the remaining work. All of the things you mentioned are absolutely being explored --both by scientists and by highschoolers collectively speedrunning 76 years of science live on HuggingFace to generate the best uncensored model for their polycule's DnD campaign-- but they hinge on distinct metrics.
For example, the last one deals with reliability, which is closely related to the "randomly alien" stuff I mentioned earlier.
"Actual consequences", OTOH, most directly relates to the camp(s) focused on "embodiment", which is basically the idea that truly human intuition is too spatial to reasonably emulate without the ability to experimentally interact with the world -- AKA the "AI needs robots" camp.
And finally, the "personality" bit... I personally think we have to rediscover the subfield of Affective Computing, but the closest lane so far is "Constitutional AI", an approach popularized by Anthropic that (wisely) just moves the whole problem over to the world of prose and optimizes it from there.
All three are important steps indeed, but I think deserve finer delineation than ~'does it connect the agent to the real world more/better/stronger/truer'.
Ha, totally agree on the distaste for intelligence as a single dimension. I will also say that 'harness engineering' is gonna end up being an outdated term for 'the rest of AI' over time, MMW. A more (in)famous voice beating this same drum is Gary Marcus (I know!) under the term 'neurosymbolic' (?).
Again, you have great intuitions here... I think it might help to consider how the human capacity for memory is simultaneously mundane and profound, at different levels of analysis. I think this situation is similar: we're not gonna need to invent Memory 2.0 (and can't, probably?), but there's a long list of human-specific heuristics, control planes, and other neural machines of some vague character that must exist, only a teeny tiny portion of which have been explored by "harness" engineering as of yet (for the best, probably...)
TL;DR: We're not through the Kuhnian paradigm shift just yet -- the new episteme has far from penetrated all the subfields of cognitive science, IMHO. Predicting the landing point feels a little pointless, for that both that reason and an even bigger one: if RSI ends up being realistic (which it very likely is for our 2026 human computers, to some significant extent), this is all just the anteshock anyway. As "the singularity" implies, that kind of exponential shift could really take us anywhere (or nowhere, forever).
P.S. Never done this before, but fuck it: I'm currently seeking exciting remote work ASAP -- if you found this interesting, please consider this my cover letter. Sorry mods if against the rules, but, y'know... one-time exceptions for the singularity?