Comment by COAGULOPATH
5 hours ago
This AI generated post (100% on Pangram) is pretty out of date.
>On SimpleQA, a benchmark of factual recall with no tools allowed, the current leader is Gemini 2.5 Pro at 53%, so the best recall money can buy still misses half the questions.
SimpleQA hasn't been updated in a long time. Gemini 2.5 Pro is a sixteen-month-old model, not "the best recall money can buy".
>The part I find most promising is what this does to hallucination. When a fact lives in weights, a wrong fact is unfindable and unfixable.
This seems confused. LLM hallucinations don't come from the weights containing "wrong facts", they are artifacts that appear at runtime.
>When the fact lives outside the model, a wrong answer has an address. The model cites a document, so you can open the document. If the document is wrong, you edit the document
You can make any modern LLM explain its reasoning and find sources for its claims. None of this has anything to do with facts needing to exist in weights or in harnesses.
The internet is full of wrong information and I cannot magically edit it to make it all correct, so this doesn't help me.
>if a model is factually wrong a claim with a source is checkable and a claim from weights isn't.
Why? If a model's weights claim that Bart Simpson became President in 2020, why does this fact suddenly become uncheckable?
I agree with everything you say except this:
> You can make any modern LLM explain its reasoning
You can make any modern LLM create a plausible, self-consistent explanation that looks like reasoning, but it's not "the reasoning it used to arrive at that answer".
Tangent: This is often true of humans as well.
We often make a decision based on a gut feeling, and then backfill a logical reason supporting our feeling, without even realizing we're doing it -- rationalization.
And we all know some people that rationalise poor choices and misbehavior, hide their mistakes, etc, to an unacceptable degree. Sometimes the individual knows they are rationalising but continues anyway, other times they seem incapable of seeing that.
When you ask people who are rationalising poor behaviour about the scenario, but it is someone else doing it, they may arrive at a better answer. Can we use multiple LLMs to achieve self criticism and critical thinking?
Tangent on the tangent: I think that's true in a minority of cases and in a majority of AI cases. Though in principle I think it should be possible for an LLM to have access to and faithfully represent its own reasoning.
This was beautifully shown by asking a model to explain how it added two numbers together (something like 45+21), and it told a plausible story, when in fact they showed it was some rotation on a helix living in some internal manifold.
Like asking a human "how did you catch that fast ball coming at you?"
It could be that the rotation in the helix manifold whatever is a low level representation of the logical steps (carry the 2, add the next column,...) it's describing. The point stands that the explanation it generates doesn't necessarily in all cases reflect what it "actually did" but your counterexample doesn't hold.
Seriously, anybody with a passing knowledge of LLMs knows thats not how they function. You can't encode logic in them because that's not how they work. It's a statistical model with useful emergent properties. It doesn't think, it doesn't reason, it isn't aware of facts or the rules of logic.
> It's a statistical model with useful emergent properties. It doesn't think, it doesn't reason, it isn't aware of facts or the rules of logic.
What makes you so sure your own brain doesn't work the same way?
> This AI generated post (100% on Pangram) is pretty out of date.
Quite ironic given the topic. It seems that the author’s model indeed contained too much knowledge about old Gemini releases, and did not do enough tool calling.
>>if a model is factually wrong a claim with a source is checkable and a claim from weights isn't.
>Why? If a model's weights claim that Bart Simpson became President in 2020, why does this fact suddenly become uncheckable?
Because in one case you have a source you can use to validate the fact, and in the other you don't. Though, as you explain earlier in your comment, the premise is misguided/hallucinated.
Yea the "When the fact lives outside the model, a wrong answer has an address" sentence seems aggressively AI written. Saw that and my senses went off.
Senses of what? LOL. The whole Internet is AI generated by now and we all contribute to that on daily basis. get used to it or dull your senses ...
>>You can make any modern LLM explain its reasoning and find sources for its claims.
>The internet is full of wrong information and I cannot magically edit it to make it all correct, so this doesn't help me.
My favorite RAG experience was asking Bart (or whatever they were calling Gemini back then) an answer to a question I knew.
It gave me the opposite of the truth (as was common with LLMs at the time).
But weirdly, it had cited sources for this "fact."
I checked the sources. Two of them, both AI SEO slop.
In this moment, andai was enlightened...
A true camper doesn't need to check Pangram, Jimbo. He goes by pure animal instinct!