Comment by z2
10 hours ago
From recent ChatGPT (GPT5.6) conversations where I've seen occasional reasoning leaks into the UI, it's clear that something like this is already implemented, and I'd speculate that this is the majority of recent claims of less token usage. Not sure if they are literally prompting for cablese of course.
"Need check output vs prev. Ran script, results fine, need prep next step. Ready? Go."
This is their CoT reasoning trace. (Why you see it: Models are supposed to delimit "actual" CoT, their user-visible summary/"cleaned CoT", tool calls, and user output via delimiters, but don't always do it perfectly.)
Rewarding terse CoTs during training seems like a no-brainer, as long as it doesn't impact capabilities, so I suspect the style is completely emergent and probably what you get when you implement a dual goal of terseness and capabilities while still punishing completely non-human-readable CoT. (Failing to do the last part would probably have models speak in ominous Unicode glyphs in no time.)
I suspect you might be right. My conclusions from the digging are that this kind of compression works best with settled instructions/data for machine to machine talk. For something like OpenClaw (which I use a lot), that might mean the AGENTS.md, TOOLS.md, etc. Compression there would free up the context for the agent.
Would be pretty funny if OpenAI's models are so token efficient now due to hidden caveman prompt.
Reminds me of how we used to search with keywords that ended up sounding like caveman speak, then there was "natural language search", and now the AI is talking to itself in caveman speak again.
I can confirm I’ve seen this with Deepseek V4.1, I imagine it’s with other models too.