Comment by pixl97
15 hours ago
There are two questions about that stability I have.
One, things like catastrophic forgetting and falling into incoherence.
Two, less likely but far more worrying, falling into unwanted attractor states. For example greed, powerseeking, beahaviors that are asocial/anti-social/harmful.
Aren't such attractors also problems during training? Presumably alignment constraints would need to apply to continuous learning as well.
I mean yes, but that's far more difficult than one would expect in a layered system. You may be able to keep a concept aligned, but can you keep the meta concept aligned? Continuous learning means there is continous opportunity for a more powerful system of misalignment to form itself and take control of your alignment. Even worse is hidden layers of this misalignment that can hide itself by not using an interpretable language.