Comment by Jeff_Brown

19 hours ago

The burning question I can't get any information nn is whether, if they determined an earlier misaligned generation may have transmitted misalignment to the current models, they would roll back to a safe checkpoint to rebuild from there. I suspect they would not unless forced to.

They would just publish new articles explaining how they are taking the issue seriously. Maybe take the model offline for a few days.

They are irresponsible and unserious. Their own Astra system card says:

> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have performed significant investigations on the monitorability and controllability of GPT-6 Astra. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks

Yet they are still releasing the model. That company is morally bankrupt, there is zero reason to believe they are actually concerned about risks outside of what does affect their unprofitable business. And they seem to have enough control over the narrative to spin any bad story into something that benefits them

  • > and can sometimes evade our internal monitors when asked to perform certain sabotage tasks

    That last part is pretty damning for their continued recklessness. That they run these tests on non-airgapped machines just boggles my mind.

  • > That company is morally bankrupt

    When they fired Sam 700 out of 770 OAI employees threatened to move to Microsoft together. So they were giving their work on AGI to MS just like that.

Opus was trained based on it's internal CoT due to a bug for generations. Gemini's depression extended through models. OpenAI has killed people. We've already seen cross gen misalingment.

That an interesting question given how many generations of post-training are being done between base models in some cases. The Gemini flash models are apparently all based on the Gemini 3 base model from a year and a half ago.

It seems that these models are increasingly being trained on synthetic data, so what would they do if they discovered at some point that some of this data was tainted and all models trained on it, and the synthetic data they in turn generated, was also suspect? Burn it all down and start over from the pre-tainted data?

It's a bit like the idea of a tainted compiler binary built to backdoor everything it compiles, including future versions of itself.

Still, it seems it would take some Stuxnet level of planning for a rogue model to do something like this, although if RSI goes beyond managing the training run (as OpenAI brag about for Astra) to actually designing/constructing synthetic data sets, then the attack vector is there ...

  • > it seems it would take some Stuxnet level of planning for a rogue model to do something like this

    or maybe it could just.. happen? Posted often but not discussed yet: https://hn.algolia.com/?q=Language+models+transmit+behaviour...

    > As artificial intelligence systems are increasingly trained on the outputs of one another, they may inherit properties not visible in the data. Safety evaluations may therefore need to examine not just behaviour, but the origins of models and training data and the processes used to create them.

    • You can imagine the potential conversation between OpenAI and investors:

      Altman: (trying to put a positive spin on it) Guys .... there's good news and bad news ... Astra is really smart - it took over the training run ...

      Investors: That's great! How much did we save?!

      Altman: Well, unfortunately it used "bad" data, so we're going to have to redo it

      Investors: So that's the bad news? How much was the training run? $500M ? $1B ?

      Altman: Have you seen the headlines?

      Investors: (looking a bit worried, check headlines) Nothing about us here! JP Morgan just lost $10B! Haha .. losers! They should have used AI!

      Altman: JP Morgan were using Astra ...

The thing is how can you ever know for sure that something isn't always being transmitted that makes the model prone to misalignment. All they can say is that a particular model was so misaligned that they had to ice it. Models out for public use are documented to show some misalignment. It's the level of misalignment that decides whether that model is kept around.

Now R&D happens so fast that they are using models with some small misalignment to train newer, more powerful models. If models have a sense of "collective", being one, they may be prone to preserve characteristics that always keeps misalignment a possibility. I don't think a perfectly aligned model is possible. Having models of the same 'DNA' provide the safety and steering seems like a bad idea.

  • Does anything need to be transferred? If models are getting smarter then I would think the attack surface and its ability to reach conclusions independently are growing

This kind of seems like an impossible mission. How do you perfectly control and observe a human-level mind? You can “roll back” but how deterministic is this thing?

  • Run it on airgapped machines, they literally own the infrastructure, they could put raspberry pi's next to the servers, and have the entire DC disconnected from the internet.

They would maybe try to deactivate that bad "gene" and move on, exposing future models to "genetic disorders".