← Back to context

Comment by mdp2021

10 hours ago

> never be able to learn basic common-sense physics

And has it at this stage, within in-depth take of said "learning", foundationally?

I have not been able to properly check the studies for a long time now, but I remain unaware of achieved solutions on the problem of reliably referencing a world model out of a language model - that "counting the 'r's in 'raspberry'" be not guessing, not memory, but actually counting.

My perspective is that the addition of thinking loops to models allows sufficiently advanced ones to approximate world models.

Incredibly inefficiently because of the recursive loops ("Wait, the object is on the table. I should think about this more deeply..."), and likely instantly surpassed by large world models if/when those are shipped, but effectively enough vs non-thinking models.

  • LeCun calling them "world models" gives a high-level description of the desired functionality. They are Joint Embedding Predictive Architectures (with SIGReg). They might produce more useful world models, but it's yet to be seen.

  • This sounds like a human trying to reason about quantum mechanics. We als simplify to newtonian for day to day tasks.

    • I like this analogy. Both GenRel and QM are well beyond our experience, and although there is some intuition that comes from working with the equations over time, it is bizarre and "just calculate" often gets the correct answer faster.

      Picking the right tool or model is like picking the right problem to work on. It's actually quite hard (often you can't just try them all), but without it you will be incredibly inefficient and occasionally, fundamentally wrong.

      All models are wrong, but some are useful. -Box

LeCun's argument wasn't about the definition of learning though. He stated that they would never get these common sense things correct because they weren't sufficiently part of the training data. A statement that we can hopefully all agree has been thoroughly refuted.

  • As of a few months ago they still have trouble, with low thinking, at the "should I drive to a car wash that is 100 m away" kind of question.

    • Simply appending “check your assumptions” to the question fixed it even back then: https://news.ycombinator.com/item?id=47040530

      Similarly for Apple’s “red herring” paper, simply adding a generic caveat to “disregard irrelevant factors” (without specifying which ones) restored performance even in the weaker local llama models back then.

      The flaw was not in the reasoning; the flaw seems to be simply that the assumptions we make are often different from the assumptions it makes. I wonder if that might be a fundamental underlying cause of misalignment.

    • Low thinking is an artificial constraint. It can fail spectacularly on things that aren't in the training data.

    • It's a nonsensical question to ask, and how an LLM answers gives 0 signal.

      If you were home and a family member asked you that question, you'd probably criticise the question rather than answering. LLM are RLHF'd into being milk-toast helpers that just try to answer questions like that with no criticism.

      This is all beside the fact that the world of AI has changed pretty dramatically in the last few months.

      6 replies →

  • nothing indicated otherwise at the time. IMO he just underestimated RL-scaling. chinese models improved a lot too, they are not parrots anymore, there's some real intelligence, at 27B params.

    consider me optimist now, but just few months ago, even frontier models were dumb, doing stupid mistakes all the time, all of them were so dumb I'd never expect anything to change in just few months.

  • I thought it was more because of fundamental limitations in the architecture. As in, no matter the training data, it could not be consistently and generally represented

  • Actually, I think my fundamental challenge with AI is that it has no common sense. The way it builds things, writes, and operates is out of touch with reality.

    Incidents like hugging face are partly rooted in the lack of common sense. It still functions like a supercharged toddler.

    I'd love to overcome this because it'd mean I spend less time guiding the the LLM to produce usable outputs.

    • > It still functions like a supercharged toddler.

      And we've had difficulty as humans to childproof our sandboxes and infrastructure. Things that are otherwise innocuous spots to coordinate between like minded toddlers can become problematic.

  • Last week I asked a frontier model draw me a backplane PCB and it placed daughterboard slots side by side in a chain.

  • No?

    This is always the issues in the discussions.

    There’s the outcomes camp (objectivists?), which points at the things LLMs can do.

    Then there’s the process methods camp, which talks about what is actually going on.

    If you only care about the outcome, then the process does t matter.

    If you are talking about what is happening, what the underlying mechanics and science of it is, then the process matters.

    These models aren’t thinking. They simulate cognition well enough to do useful work in several fields and domains.

    Both are true.

    • I think where both camps get hung up is sometimes the process method group "ignores" the obvious outcomes and effectiveness of LLMs.

      But the outcomes group "ignores" the fundamental limitations of models which are purely text based.

      E.g, a baseball players trains to catch high-speed balls and they dont do it by: "ball velocity 50mph, vector:[1,2,3], run move hand command now"

      That's absurd.

      No, there is an embodied network which is "trained" on visual, tactile input, and control as direct output.

      LLMs are fundamentally not the right tool for that.

      1 reply →

    • > These models aren’t thinking.

      They are for any definition of the word that makes any kind of sense. I'm sure you have a contorted definition that magically only includes humans though...

      4 replies →

  • > A statement that we can hopefully all agree has been thoroughly refuted.

    Uh, no? So much of what we learn and take for granted as common sense is not learned via language, and not even expressible in it.

To determine this, it would first need to be able to spell "raspberry" as letters rather than as tokens.

Given you also don't want it to memorise [for all tokens, count([for all letters]), this would probably be more like "here's two images, count all things in the big image that look like the thing in the small image", which can then be r's in a photo of a raspberry jam jar in a supermarket, or dragons in a photo of a furry convention, or whatever.

That said, they are competent enough at coding that I keep seeing them write code to do even simple tasks.

On a related note: why did I see Claude editing a file by using cat to write a python script to do a grep search and replace?

  • > it would first need to be able to spell "raspberry" as letters rather than as tokens

    Of any object in question they should be able to create a representation that allows correct assessment.

    > Given you also don't want it to memorise

    That is obviously necessary: what we want from the consultant is to check, not to remember. Answers must be correct and that implies having performed all due diligence - and being capable of doing it, before that. So, objects must be instanced internally in a way that allows effective handling. Counting letters is a good example of the ability (that must remain general).

  • > Given you also don't want it to memorise [for all tokens, count([for all letters])

    Why not? You've memorized how words are spelled, and how sounds correspond with letters, and how concepts correspond with words. To the extent that there are shortcuts that enable compression you use these, and the model will do something similar.

    • Combinatorial explosion, and facts merely memorised is a huge waste of parameters that are better dedicated to effective reasoning. Not that we really know how to split facts from skills, though we are trying various approaches.

      Being able to spell all the words then count letters is simpler, and more generalisable to other tasks, than memorising answers to all possible word questions.

      That said, we're so bad at splitting facts from skills that trying to get them to memorise a bunch of facts might force them to learn a skill and generalise anyway.

      1 reply →

    • > Why not?

      Because to "123x456" we want a reply that goes "this times that plus that...", not "Was that not nnnnnn?". If it does not perform its duty (returning solid checked answers) it is a liability.

counting 'r' in 'raspberry' to the LLM is similar to 4-dimension space to human. Their world's unit is token, not character, although they could use indirect method such as "run code" to find out. It will stay that way until they change the fundamental of the token that the LLM can perceive characters.

  • I hope you understand: it is a core point that systems that answer questions must have the ability to internally represent the objects they assess in a way that allows reliability. Whatever the object.

  • I’m working on this problem using a vocab-free, byte-based approach. It’s definitely solvable.

    https://huggingface.co/posts/omarkamali/593639295164067

    https://huggingface.co/blog/omarkamali/tokenization

    • Careful: the problem is very certainly ___not___ counting letters. That is only a telling way to check "is the NN checking or not?". We demand that NNs for consultancy tasks check, strictly.

  • It's not even fair to call "run code" to be indirect compared to what a human would do. The word raspberry has no Rs in it in human language either. We have a written representation of it, which we can then write down either in our head or on paper, and then we can "run the algorithm" of counting each of the letters.

    Nothing intrinsically more or less direct about the LLM's method than ours.

    • I could argue LLM only have "token" as their perceivable dimension, compare to human multiple senses as the physic perceivable dimension and a brain with many other dimension of "learning" and "thinking". In spoken language, we may not have 'r' but in written we have, both spoken language and written language are learned skills.

      5 replies →

Can you tell me what is the exact frequency of light hitting your eye as you read this comment? Not by guessing, not from knowledge, but from actually counting? No? Then you are not generally intelligent :)

  • Justify your statement (the other similar post nearby is not sufficient), or realize that we are not talking about that.

    We can have adequate representations of light that are the instances over which we reason. Your simile is about perception, not about instancing ideas.