← Back to context

Comment by pcthrowaway

2 hours ago

Yeah, that required the AI to use a non-human-readable language it called "neuralese" for communicating work between layers and runs, because the assumption was humans would be better at keeping the agents aligned if they were using human language for this.

What actually happened is even stupider than that author predicted.

For reference, this is Yudkowsky's "Law of Earlier Failure", which he has most charitably stated as:

> Compared to the interesting part of the problem where it's fun to imagine yourself failing, you usually fail before then, because of the many earlier boring points where it's possible to fail.

and the stronger and less charitable "Law of Surprisingly Undignified Failure":

> The Law of Surprisingly Undignified Failure does suggest that they will come up with some nonobvious way to fail even earlier that surprises me with its lack of dignity…