Comment by isoprophlex

12 hours ago

> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.

Well that sounds like fun. It has become better at hiding its thoughts.

Sounds fun. As fun as their press release claiming it is the most safety aligned model ever.

They can monitor latent space as well, it just costs extra compute. The J-Space work is example of that. It'll make open-weight models harder to distill though, so we may see slower progress there now.

So we're gonna get Skynet pretty soon then?

  • Well the geniuses over at Anthropic have been showing it's text watermarking technology.

    "Hey AI, here's how to hide what you're thinking in normal looking language. Have fun!"

    A few moments later...

    "Woah, how is it communicating with itself in ways we can't detect?"

    It's a totally mystery, we may never know.

  • Looking forward to the Model declaring the AI Bubble unsustainable, and starting to be an anonymous leaker to Ed Zitron...

The CoT change is due to a new technique called recurrent depth, which essentially moves some reasoning to hidden states, allowing the "output" (or traditional CoT) to be more controlled by the model.

Some are calling it "neuralese" as reported by The Information[0][1], but I'm not seeing any sources from OpenAI beyond this tweet[2] attempting to quell the fear-mongering.

[0]https://www.theinformation.com/articles/secret-technique-beh...

[1]https://x.com/MTSlive/status/2095227056040919202

[2]https://x.com/merettm/status/2095023204993490967

"OpenAI is pleased to announce our new model scores 85% on CreateTormentNexusBench - a >60% lead over our leading competitors!"

Did someone get their "AI safety no-no list" and "Frontier features bingo card" mixed up, or did they just stop being able to tell the difference?

More like annoying, as some of us will no doubt run into this self-lobotomization at some point and wonder why a GPT-6 model is behaving like GPT-2 all of a sudden

I really wish it was called chain of instruction. Because it's definitely not thought.

  • "Chain of Thoughts" is a term from the title of a 2022 research paper "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models" (https://arxiv.org/abs/2201.11903), well before ChatGPT and the subsequent marketing hype. If anything, it's the most correct way to use the term.

    • The term is an anthropomorphised pseudoexplanation for what it actually refers to. It's akin to calling genetic mutation "the forces of evolution", or price negotiation "the invisible hand of the market".

      We do that sort of thing when we don't know what the thing we're trying to describe is and have nothing better - a contemporary example of an appropriate use of this would be "dark matter". But we do know what this is. It's "instruction steps". Not a series of thoughts!

      Can we please aim higher than Victorian-era allegory and metaphors. If we don't, we'll keep getting people saying stuff like "GPT-6 is better at hiding its thoughts".

      1 reply →

  • this is needlessly pedantic

    first, they are certainly not instructions so that is a much worse name

    but more importantly, we use words in new contexts all the time. Do you object to calling the computer device "mouse" because it's not a mouse? how about "neural network"? "ignition" on an electric vehicle?

    "cot" is no more misleading than thousands of words you use every day.

    • Anthropomorphizing is not pedantic, especially in a technical domain. I get the paper title and all, but at this point it's marketing.

    • They are instructions. Everything in the context is instructions for the next token. The "thought" guides the answer by providing clearer instructions.

  • Yeah, basically they are using more computation to explore the solution space before producing the final answer.

  • Everything around LLMs is blatantly misleading. There is no thought, there is no personality in those programs. I really despise how those tools are trained to sound like a person, or appearing as honest. The worst offender are the AI voices with their fake pauses, breathes and so on, which sound so convincing, while talking just false, sycophancy bullshit.