Comment by Topfi
16 hours ago
> This is false. Perhaps you meant they have insufficient methods, or imperfect methods, but asserting none at all is facile and your links do not say that at all.
I did say "no reliable internalised way to assess the accuracy of their output", which is something very specific. If that Sonnet output is reasoning traces, there are many issues with using that as a source:
For one, back and forth reasoning does often correlate with less, not more accurate overall outputs in evaluations and one example could never seriously be extrapolated to be considered "reliable", i.e. happening consistently and dependably.
Secondly, self-correction is also not self-verification in regard to model output, revisions not necessarily mean internal accuracy assessment by themselves (again, over-revisioning has lead LLMs in my and even public evals like the one linked above to step away from accurate information written in their reasoning traces but discounted in the final output (if we must use anecdotal examples like your Sonnet quote)).
Then there is the fact that, unless that reasoning trace (if it indeed is one) was copied from a months old chat history, Anthropic has obfuscated their reasoning traces so this output is (if it isn't an ancient history you dug up) from the obfuscation model in between and not reflective of the actual models reasoning. So even for anecdotal evidence, this can likely not be used (unless again, you went for December 2025 history). If this was not reasoning traces (I am fairly confident it is having spend a long time reading Anthropic summarisations vs actual reasoning traces when the switch was on-going and you could get both for a limited time, what you shared reads very much like obfuscated rather than pre-obfuscation reasoning output), that still leaves this as anecdotal, one time, possibly erroneous (did this back and forth even prevent an error in the final output) and there are more issues still with just using that quote as evidence, this is simply unsuitable as a source in any situation.
Here are some papers I read lately, all published in 2026 and using the current crop of models which were what led me to make that specific statement. LLMs currently have no reliable internalised way to assess the accuracy of their output, at least as far as the literature is concerned:
> Even state-of-the-art models struggle to reliably discriminate between data uncertainty and model uncertainty.
Beyond “I Don’t Know”: Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty [0]
> LLMs cannot reliably revise their own errors without an external signal.
> The same models that confidently catch and repair errors in external content routinely fail to identify identical errors in their own reasoning traces.
The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models [1]
Simply, as of today, LLMs cannot reliably translate whatever internal signals they have into accurate self-assessment of their outputs. Even in papers that show limited, edge case capability, this often breaks with minimal prompt and/or task changes (happy to link those too, I just need to get to my Mac where I have the PDFs), so it is not reliable.
If you got a paper that shows that I am false, that there is reliable, internal self-verification of a models output (even a small research LLM not yet publicly released), I'd be happy to read it and retract my original statement.
No comments yet
Contribute on Hacker News ↗