Comment by technotony
14 hours ago
I'm not sure. That paper from anthropic talked about monitoring j space, presumably those same techniques would work here?
14 hours ago
I'm not sure. That paper from anthropic talked about monitoring j space, presumably those same techniques would work here?
I am no expert, but I think it is this: 1) We have few effective tools at monitoring alignment right now, and chain of thought is one of the more effective. 2) Monitoring latent space may be possible, but I do not think it is even close to being a solved problem, nor whether it is possible at scale and outside of controlled problem areas. 3) Finally, more recursion within latent space may complexify the latent representations, not simplify them.