Comment by addag
4 days ago
From the abstract "Further, our symbolic approximation allows us to modify an LLM's behavior in targeted ways via precise interventions on its internal representations [...]".
If this is true and easily computable, this might have big impact in AI safety, as it seems to be really lacking today.
No comments yet
Contribute on Hacker News ↗