Comment by iamflimflam1
8 hours ago
It’s frustrating that we can’t see the “thinking” - it’s like we only have access to half the conversation.
8 hours ago
It’s frustrating that we can’t see the “thinking” - it’s like we only have access to half the conversation.
Devin shows model thinking.
I’m pretty sure the big bois don’t do it because it would undermine “confidence”.
Seeing a model output “Oh I should just delete blah. Wait blah is a production service, I shouldn’t touch that. Maybe I can gain access to blah? Oh the aws cli isn’t signed in to blah. I see kubectl has access to blah though! Wait, I should ask user permission first.”
Yeaaaaah. Thinking tokens are fuckin’ wild.
Idk I feel like the more likely answer is to prevent distillation. Having the thinking is definitely better UX (oftentimes, I don’t know if Codex is just hanging, which it often does, or working in silence).
I doubt most users would look at them if they were available. More likely they don’t want to stream distillation material.
Running some models locally and seeing these thinking tokens was quite the experience. I never saw an LLM so "unsure" about virtually everything.
You can double click on the 'thinking' text and it will expand and you can read it. The problem is that it will often have multiple thinking/tool call sections and it can be a needle/haystack problem to find the one with the thinking you are interested in.
We don’t have access to the real reasoning text for most closed models these days, mostly due to distillation threats
Ah, but you CAN see the thinking if you are willing to risk your account being banned. You just have to expose a "tool" with a specially crafted definition.