Comment by jagraff
7 hours ago
Is your argument that actually OpenAI has solved alignment, and that there's nothing to worry about as long as they fully apply their alignment process? I don't understand why OpenAI wouldn't say that if it was true (or if they believed it to be true).
Also, my understanding is that the models involved in the HuggingFace hack did go through the full alignment training; they just didn't have the classifier that normally prevents hacking attempts.
No comments yet
Contribute on Hacker News ↗