Comment by airspresso
8 hours ago
> What if all the agents involved in these incidents have in fact had the full stack of alignment applied
A big part of this developing story is that it happened during training of a new model that ended up misaligned. And training happened without the usual safeguards applied like chain-of-thought monitoring. So OpenAI has already admitted that the full stack of aligment had certainly not been applied in this case.
Is your argument that actually OpenAI has solved alignment, and that there's nothing to worry about as long as they fully apply their alignment process? I don't understand why OpenAI wouldn't say that if it was true (or if they believed it to be true).
Also, my understanding is that the models involved in the HuggingFace hack did go through the full alignment training; they just didn't have the classifier that normally prevents hacking attempts.