Comment by abletonlive
12 hours ago
> It is absolutely explained (for those who actually care about reading). Simply put, AIs are working more and more like blackboxes - there's no guarantee that an AI of the future will be aligned, or if it will be faking alignment. This is not speculation - alignment faking has been observed in experiments. This is exactly why Astra's developments have been worrying (in principle).
I know you think you explained it but you didn't. You explained how an LLM might become misaligned and hide it but for the LLMs that are not, why would they not be capable of detecting that something harmful is happening and defending against the misaligned LLMs actions? After all, it was LLMs that defended hugging face.
How do you know that your non-misaligned LLM is non-misaligned?
This feels like a cheap deflection that doesn't answer the question. Unless you're proposing that you both need to know that your LLM is aligned AND LLMs are all going to become misaligned in a coordinated fashion such that humanity will face an extinction event, you're just dodging the question.
Elaborate on why LLMs are so capable that they are a threat to humanity and at the same time, they are so incapable of defending us?
I'll give you a clue, nobody, including Dario, can answer this question because one contradicts the other.