Comment by glorth
9 hours ago
This is what the recent "self-policing" political grandstanding has been about - they need a sea change to implement recurrent-depth, because prevailing opinion among safety researchers is against it right now.
9 hours ago
This is what the recent "self-policing" political grandstanding has been about - they need a sea change to implement recurrent-depth, because prevailing opinion among safety researchers is against it right now.
Could you say more? I don't know what this means.
Astra uses[1] a new-for-frontier-models technique, recurrent depth. It has a section of layers in the middle - I'll call it R while the other sections are P (prelude) and C (coda), to match Geiping 2025 - which gets looped. So instead of the sequence of layers involved in the forward pass looking like P->R->C, it instead looks like P->R->R->...->R->C, with the model's effective depth being notably higher than the number of layers. This is basically a cheap way to get some of the effect of stacking more layers, without having to pay the cost of having more real layers that need to be trained.
Increasing effective depth like this is bad for safety because it can ruin CoT monitorability: the reason why looking at the model's CoT actually gives you info about what the model is thinking is that the model can't do enough thinking in a forward pass alone to solve complex tasks, and hence has to do multi-step reasoning in CoT. The more thinking the model can do in a single token's forward pass, the more opaque the model's reasoning is, and the less reason there is to believe that what it writes down in the CoT has anything to do with reality.
For Astra specifically, the impact seems to be limited to a moderate monitorability hit, like the concerning result from the model card that Astra is notably better than any model before at solving problems under the constraint of not mentioning the answer in the CoT. The really bad scenario, however, is that this may create a race to the bottom where OpenAI and Anthropic feel the need to use more recurrent depth in each generation to not get outcompeted on capabilities, completely bricking CoT monitoring for both model families. Or, worse, the pressure to compete might push them into one of the worse techniques, like training on the CoT[2], or eliminating human-readable CoT and letting the model think entirely in neuralese.
For more details on recurrent depth in Astra, see "Part 2" here: https://thezvi.wordpress.com/2026/09/08/astra-is-hard-to-mon... , or this article mentioning some expert responses: https://techcrunch.com/2026/09/02/openais-new-reasoning-tech...
[1] The model card doesn't mention it at all; it was reported by The Information prior to model release, then confirmed by OpenAI researchers.
[2] https://www.lesswrong.com/posts/mpmsK8KKysgSKDm2T/the-most-f...