Comment by user43928
7 hours ago
Hallucinations are no longer much of a practical problem in software engineering.
Two years ago, hallucinating that the code worked or that a task was accomplished was a common occurrence.
We have seen that now agent swarms across thousands of agents can coordinate to achieve a result.
Clearly hallucinations are no longer the problem they once were, since now we can get working results for long horizon tasks that require massive compute.
Consequently it would seem unwise to assume that current limitations will remain as they are and prevent LLMs from coming up with solutions that they can explain to humans.
It’s still a common occurrence.
It happens in more subtle ways, but it still happens often enough for me to notice. For example I have had hallucinated checksums show up in lock files as recently as yesterday using a SOTA model.
This is not surprising, since the whole basis of LLM training is to produce output that humans will accept _as a proxy for actual training goals_. In a sense, the training process of an LLM “wants” to produce output that is statistically plausible much more than it “wants” to produce correct output. It’s always going to be a struggle to drive that system towards other goals (and we see this bourne out in practice by the amount of effort that is required to be spent on RL).
I think there will be some threshold of correctness (something like 99.999% of the time) that if the model surpasses it, I can stop needing to check it, but I think we’re still at 99% or something which sounds good, but when you are producing a ton of output you hit that 1% frequently.
> Consequently it would seem unwise to assume that current limitations will remain as they are and prevent LLMs from coming up with solutions that they can explain to humans.
I 100% agree with this. In fact explaining things to humans is something LLMs are particularly well suited for.
The confidence with which you, anonymous user, keep commenting that "hallucination is not much of a practical problem in software engineering anymore" based solely on your own anecdotal evidence is really remarkable, in not a good way.
You're free to substantiate your comment by telling us about your apparently different experience.
I find hallucinations in my (mostly perfect) AI output every single day. If you're not finding them, you're just not looking hard enough. It's not surprising when everyone is screaming about how they don't read code these days.
This is just a fact. I'm sorry if it messes with your narrative.
https://arxiv.org/abs/2401.11817
1 reply →
The sum total of all human observations is still not proof of the lack of hallucinations as a problem (even if their observations were perfect, which they aren't considering the volume produced vs reviewed carefully). That's why you can use a counter example only to disprove and not prove anything.
And yeah I get hallucinations all the time still. Maybe it's because I'm working on harder/more niche problems (like a compiler with an unusual type system), but it happens quite a lot. I don't record all of them.
Although the most common one you can find is them misattributing the source of changes from themselves and also other agents (Fable, Opus 5.5, deepseek, whatever). They'll say "your changes" or "you changed" or "your ruling." I didn't decide anything and it's in their own chat log, and yet...