← Back to context

Comment by Gareth321

7 hours ago

I agree on all points. It gets worse: OpenAI is switching their model thinking from sequential language tokens to primarily "latent neural representations." Meaning there will be little or no chain of thought to monitor. This appears more efficient, so all model labs will eventually switch to this. At the most crucial time for us to be monitoring reasoning and intent, we're about to make that much harder.

I also think the intent problem overlaps a frustrating amount with philosophical and political questions. It's the basis for Asimov's Three Laws of Robotics (1942). Intent is subjective. Language is subjective. Humans are imperfect at using language to accurately portray intent. All of these guarantee that an enormous number of queries in the future are going to be misinterpreted. Not such a big deal when it's about a cake recipe, but when it's about governance, laws, military targets, nuclear power sites, etc, the scope for failure becomes catastrophic. The Three Laws of Robotics attempt to create a backstop, but as countless stories have explored since (including I, Robot), even these laws are subject to interpretation.