← Back to context

Comment by strbean

1 day ago

Would these alternatives to backprop make it more feasible to have constant live-training going on in a model? Giving it something akin to neuro-plasticity?

One aspect of how current training and continual learning are somewhat at odds is that the memory required to train a model is often times 2-3x the memory required to just run it (probably not as bad for PEFT, not sure).

DUST does have an advantage specifically along those lines because it doesn't have to save a ton of intermediate state other than each layer's input activations during a single forward pass.

There are many other issues that this algorithm does not address thoigh like catastrophic forgetting. it's still operating on a transformer which contains no inherent mechanism for selecting the relative value of a training step based on current knowledge, nor does it have segmentation of functionalities with specialized areas used for specific things that can be sequestered off and ignore new updates (we do not risk forgetting how to walk as we increase our French vocabulary)

Models suffer from "catastrophic forgetting" if you train them on new data.

People are working on this field, recent results suggest that continual learning can be possible by converting the input data to "LLMese"

Maybe the practical path is to separate fast-changing memory from slow-changing weights. Most things an agent learns during use probably don't need to become parameters immediately.