Comment by ACCount37

9 hours ago

Continuous learning allows past behavior and past inputs to influence future inputs and future behavior. In humans.

Attention over KV cache allows past behavior and past inputs to influence future inputs and future behavior. In LLMs.

Until the cache runs out, that is. But even then, you could totally use any of 9000 methods of cache compression, truncation, dropping or streaming and get away with it.

The difference between continuous learning and in-context learning seems to be in capacity, not in principle. Both are doing a similar thing, but one has more length and depth to it.

2 comments

ACCount37

nomel 8 hours ago

Maybe, every night, you send the AI off to "sleep" where it uses those in cache "memories" to influence the long term weights [1].

[1] https://www.pnas.org/doi/10.1073/pnas.2220275120

ACCount37 7 hours ago

Context self-distillation does exist, but as is, it's used mostly in training rather than as a part of a continuous learning mechanism.