← Back to context

Comment by wilg

1 day ago

Making things up is only really a common issue on the non-thinking models which nobody should be using. The regular chatbots are Autogooglers and are very useful for research. This is just not a good argument anymore.

Edit: Guys, why are we downvoting this? Does no one use like ChatGPT or Claude and understand how it works? Do you all think its regularly hallucinating links still? Is everyone on HN using like free signed out accounts or something? What year is it?

We know they are, and the author of the article cites the proof. LLMs do hallucinate, there is no way to make them not do it, because of the way they work.

  • A “hallucination” is an authoritative counterfactual statement returned as a response. Why do you think it is impossible to engineer an LLM (by which I am including tool usage and RAG) that catches and prevents such statements?

I'm right there with you for a lot of stuff. I ask a question and can be very confident that ChatGPT is citing sources, then sometimes I go read the sources. The more critical the information I'm looking for is, the more careful I am about this.

The other day though I was seeing how well it could pull details of its own conversations with me. It often does this pretty well for broad strokes of things - it remembers, largely, what cameras I have and use when I ask photography questions. It's never made things up here, but it does forget details, such as whether I've bought something or am just considering it. However, when I asked it for a specific interaction I thought I remembered, it gladly went along with my false memory and provided an affirmative answer. It was the first time I'd been caught in a serious hallucination with a frontier model (Sol High on the web chat interface) in a long time.

>Do you all think its regularly hallucinating links still?

When did that stop? May 7th, 2026?

  • A tech blog is going to have more than it's fair share of enthusiasts using small self-hosted and similar models; it is entirely possible that that completely accounts for the behaviour described in the article.

    • Downvotes on a perfectly valid comment is like my opponent letting their clock tick down from 5 minutes rather than resigning when they've clearly lost: it adds a smile to my day.

      Thank you, may I have another?

Can't do much about the downvote parade, but I can second this. That said, when the models are not provided the right context, and cannot fetch it for themselves, things can be rocky still. A lot less so than even just a few months ago though.

  • "hallucinating less" is just shifting the goal posts into vagueness. Things can be rocky still = nothing fundamentally changed. Meanwhile, if I search my file system for "foo" a trillion gazillion times, it will not return "bar" once.

    • If a game on release is unplayable-tier buggy, then improves over time to the point where bugs are barely noticeable, is acknowledging that going to count as "goalpost moving" to you? It never became formally verified after all, and it's even running on physical hardware... Woe are the people lying to me (nobody), the issues have not been fundamentally ruled out!

      > Meanwhile, if I search my file system for "foo" a trillion gazillion times, it will not return "bar" once.

      Great! Nor will an agent any more likely, cause it just sends out a tool call and surfaces its output.

      It boggles the mind. One would think this is some highly secretive technology only a dozen people in the world have access to, the way one has to argue tooth and nail about trivially verifiable facts regarding it. You quite literally do not have to take either of our words or "vague" judgement for it.

      2 replies →