Comment by eru

7 hours ago

That seems like a weird standard.

I would be happy enough with: only produces what it can verify with sources.

If you eg try to remember a court case (ie produce the reference via LLM token generation only), it's easy enough to check with your data whether it really exists. Similar for following links and other references.

If your data or sources are wrong, obviously your report about them will be wrong. But I wouldn't call that a hallucination.

There isn’t a single human in this world and hasn’t ever been that meets your happy-enough standard. Make of it what you will.

  • Why is that at all relevant?

    Humans are known to hallucinate a lot. Ask 10 different witnesses at a crime scene what they saw and they'll all report different things.

    A good, non-hallucinating LLM would only report things for which it has evidence. It would consult the facts every single time.

    It's a pain in the butt for humans to fact-check everything but LLMs can quickly look up all kinds of stuff. That's what makes them useful.

    • Yes, and for the LLM you can do it in multiple passes.

      So you can bolt the fact-check / source-check pass onto whatever other system you have, without having to redesign the underlying system.

  • It's not a binary thing. You can get closer or further away from that standard.

    And humans also behave differently in different contexts. A conversation at the pub has more such hallucinations than a formal deposit in court. For the latter, a good lawyer will look at her shoes, when you ask him what colour her laces are.