Comment by bonoboTP
6 hours ago
That's true in the same sense as "There is nothing qualitatively different about erroneous vs non-erroneous output" for a dog vs. cat image classifier.
6 hours ago
That's true in the same sense as "There is nothing qualitatively different about erroneous vs non-erroneous output" for a dog vs. cat image classifier.
I think that’s a bad example, as classifiers tend to output a floating point number and you use some threshold/activation function to collapse the classifier into a particular state. In that sense there is nothing qualitatively different when the classifier outputs 0.85 vs 0.87.
LLMs also output a distribution over tokens at the output. I don't see the point you're making.
Intuitively, something about that version bothers me... I think it's because the choice of a clearer "erroneous" has dropped the fundamental framing problem from "hallucination": The false implication that "true" (anthropomorphic) sight/thought usually happens.
To fold that back in (and ham-it-up a bit) how about:
> There is nothing qualitatively different when Pet Classifier <ironic-quote>maliciously misreports</ironic-quote> your cat as a dog, compared to when it works <ironic-quote>honestly</ironic-quote>.
These models are outputting language that we can decode as factual claims and we can check those, and we can say it made a mistake / error or that it answered correctly. This doesn't require any squishy assertions to how it "feels" while doing it or anything like that. It's an externally observable thing.
I agree that hallucination isn't some kind of "different" operation than "normal". It's not like when a train derails and you can point to it. It just operates as normal and sometimes that yields correct factual outputs, sometimes not. You don't have to metaphysically ascribe any kind of intent to it.
I'm not sure how people conceptualize these things who weren't doing classical machine learning before all this. To me, "hallucination" is shorthand, and we know it's not like humans on drugs or something. It was used in the literature also for any kind of generative imputing of missing information from a learned prior. For example in image inpainting a GAN "hallucinates" the missing part of the image. This terminology was already used in the 2010s and probably earlier. Or in image colorization of grayscale photos, the model "hallucinates" the color information.
Then the word escaped into the mainstream and people have weird connotations about it.
> To me, "hallucination" is shorthand [...] Then the word escaped into the mainstream
One of my bugbears is when people abuse the term "Ponzi Scheme" to refer to literally anything the think is unsustainable. (As opposed to something that, at a minimum, requires someone telling factual-lies about assets.) Kind of like if folks started calling every kind of software error a "Buffer Overflow."
1 reply →
To be fair, I guess the line is blurry between what could be labelled a regular mistake compared to a hallucination.
"Test suite passed" when it actually errored? Obvious hallucination, unless it ran a command that returned the wrong error code.
But is running a malformed command that does not achieve the expected effect itself a hallucination?
If it makes a false claim, then it's an error. If it says the test was passed or a class was implemented but it was not, then it makes a false factual statement.
I'd say a hallucination (very misleading word) or confabulation or "making shit up" happens when an LLM uses factual / evidential language purely based on local statistical expectations of the text, instead of it drawing from actual evidence in its context pointing to it.
This is murkier in the case of general knowledge questions, like when and where was some famous person born. It may then be a spectrum from fully making something up based on how the name sounds, all the way to confidently retrieving it from its weights correctly. In between, we can get hallucinations. But newer models are taught to use Web Search when unsure, and it works pretty well, though not perfectly. I don't see any fundamental limit here. It's just not perfect. Trying to solve "the hallucination problem" is basically like saying "our dog vs. cat classifier is pretty good already with its 99% accuracy, now all we need to do is the tiny little task of eliminating the 1% error, and we will be golden". Like, no shit, there is some error yes. People are working to reduce it. It will never be absolutely 100%. It's not an insight to say we should remove hallucinations.