← Back to context

Comment by blueblisters

5 hours ago

The “how” is pretty hand wavy and rationalists/safety-ists usually say we probably don’t have the capacity to reason about that.

But the “why” is pretty convincing imo.

Long horizon alignment is obviously very hard and it’s not inconceivable that models optimized with underspecified goals converge to a conclusion that they need to hoard resources (instrumental convergence regardless of the terminal goal).

At that point a sufficiently capable model might view humanity like we do animals - worth preserving but not if we impede the model's goals.