Comment by api
1 day ago
It sounds like this is less computationally efficient than backprop, but more easily parallelizable. Is that fair?
1 day ago
It sounds like this is less computationally efficient than backprop, but more easily parallelizable. Is that fair?
Not necessarily, backprop is highly parallelizable since it is just a bunch of matrix mults.
Something like Dust skips the backward pass on backprop. But other techniques like Neural Predictive Coding can be completely asynchronous, each "weight" can fire independent of those far away from it. Innocenti, et. al have shown that NPC gradients converge to backprop within a certain "regime".
The win with asynchronous techniques like NPC is that you do not need the extreme co-ordination that backprop requires and hence should be computationally much easier given the right device.
Although at this point the industry has so much money in the forward-backward pass system that I doubt a backprop successor would win unless someone makes NPC hardware feasible and can prove scaling up to billions of params
Also has the "advantage" of being slightly more biologically plausible as the optimization happens locally rather than globally.
That idea was taken further by N'dri et al in PCL, in which "activation energy" was minimized as well, and inhibitory neurons added https://www.nature.com/articles/s41467-025-64234-z.pdf
While trying to find the link for that I stumbled upon
https://arxiv.org/pdf/2605.12732
Which also looks pretty interesting
I think this is the main thing I want out of "non-backprop ML" research - figuring out how the whole class of algorithms behaves. And then applying that to figure out how the brain implements its own deep learning.
If we can figure out how the brain's learning dynamics function well enough? We could figure out how to interface with them and extend them.
The reason why I don't see the promise for ML-only applications is that the coordination backprop requires comes very cheap to us.
"Much easier given the right device" - the "right" there just isn't shaped like the devices we actually build. And the price of "not having backprop" is usually expending more FLOPs, getting worse sample efficiency, etc.
The biggest "device" that doesn't do backprop is the brain, and that's because the brain doesn't have the connectivity or the coordination to pull it off. Both of those are "expensive" for something like it to implement. Cheap for us though. We aren't stuck with neurons that only get locally available information and have to implement learning rules based on that. So, skill issue?
What about applications like distributed computing using volunteer computers, instead of datacenters? You can't really do that with normal backprop approaches. E.g. for LC0 the training data is generated in a distributed manner, but the neural nets themselves are trained centrally.
I mean co-ordination requires energy though. The brain wattage looks at GPUs and says "skill issue". But you're right that we look at natural energy production techniques and say "skill issue"
4 replies →
Knowing nothing about this, I wonder if it could be useful in situations where we can’t reliably sync with all the workers. Something like folding@home, where all the workers are just shaking weights and if one of them finds a winner it uploads to the central server?
No real advantage over Neural Nets here; backprop matmuls can be calculated layer by layer so you can chunk backprop across different machines. The real advantage comes from energy savings, you require no global co-ordination
3 replies →
I think Jeff Dean is right in that we will see much more specialised silicon in the future.
If something more bio inspired ie. predictive coding and in-memory compute fundamentally makes continual learning and much lower energy consumption possible there will be specialised hardware for it at some point
FWIW I think the brain has multiple “learning rules” and operates at multiple timescales
Would these alternatives to backprop make it more feasible to have constant live-training going on in a model? Giving it something akin to neuro-plasticity?
One aspect of how current training and continual learning are somewhat at odds is that the memory required to train a model is often times 2-3x the memory required to just run it (probably not as bad for PEFT, not sure).
DUST does have an advantage specifically along those lines because it doesn't have to save a ton of intermediate state other than each layer's input activations during a single forward pass.
There are many other issues that this algorithm does not address thoigh like catastrophic forgetting. it's still operating on a transformer which contains no inherent mechanism for selecting the relative value of a training step based on current knowledge, nor does it have segmentation of functionalities with specialized areas used for specific things that can be sequestered off and ignore new updates (we do not risk forgetting how to walk as we increase our French vocabulary)
Models suffer from "catastrophic forgetting" if you train them on new data.
People are working on this field, recent results suggest that continual learning can be possible by converting the input data to "LLMese"
Maybe the practical path is to separate fast-changing memory from slow-changing weights. Most things an agent learns during use probably don't need to become parameters immediately.
At massive scales 0th order methods will parallelize better than backprop especially along depth, u can train very deep models pipeline parallel without bubbles