Comment by baby_souffle
14 hours ago
They do store the reasoning locally. It's encrypted, though.
Few weeks ago there was a new paper out where researchers took the encrypted reasoning tokens and injected it into a new session with a week or model in the same family that they could reliably jailbreak. They would then ask the model to repeat its reasoning and the results were pretty consistent.
They used the LLM as a decryption oracle of sorts.
No comments yet
Contribute on Hacker News ↗