Comment by pegasus
7 hours ago
These systems don't have a genuine capacity of introspection, beyond just mining the conversation trace. When you're asking them why they did this or that, they are basically guessing anew from the outside, and are just as likely to hallucinate as they are to hit the right answer. These are mechanical systems which brute-force chains of various (re)combinations of techniques acquired from the training data. The true authors are all those who have contributed those techniques in the past.
I think that's debatable. Anthropic's interpretability research suggest models do have self-introspection ability, at least in the "J-Space": https://transformer-circuits.pub/2026/workspace/ ; and this private working/'introspection' space is distinct and distinguishable from the tokens they output (CoT tokens are output too).