Comment by user43928
1 day ago
And why not?
For all that I saw over the last few hundred hours with AI on software engineering, hallucinations are no longer a problem at all.
Not once have I seen a task fail due to what would have been a "hallucination". If they still occur, they can apparently be detected and corrected automatically, or are subtle enough to escape notice with presumably no significant impact on the results.
Why would this not also be the case for mathematics?
I think OP is saying that hallucination or not is just semantics. There is nothing qualitatively different about hallucinated vs non-hallucinated output.
That's true in the same sense as "There is nothing qualitatively different about erroneous vs non-erroneous output" for a dog vs. cat image classifier.
I think that’s a bad example, as classifiers tend to output a floating point number and you use some threshold/activation function to collapse the classifier into a particular state. In that sense there is nothing qualitatively different when the classifier outputs 0.85 vs 0.87.
1 reply →
To be fair, I guess the line is blurry between what could be labelled a regular mistake compared to a hallucination.
"Test suite passed" when it actually errored? Obvious hallucination, unless it ran a command that returned the wrong error code.
But is running a malformed command that does not achieve the expected effect itself a hallucination?
1 reply →
Intuitively, something about that version bothers me... I think it's because the choice of a clearer "erroneous" has dropped the fundamental framing problem from "hallucination": The false implication that "true" (anthropomorphic) sight/thought usually happens.
To fold that back in (and ham-it-up a bit) how about:
> There is nothing qualitatively different when Pet Classifier <ironic-quote>maliciously misreports</ironic-quote> your cat as a dog, compared to when it works <ironic-quote>honestly</ironic-quote>.
4 replies →