Comment by ACCount39
1 day ago
The reason why I don't see the promise for ML-only applications is that the coordination backprop requires comes very cheap to us.
"Much easier given the right device" - the "right" there just isn't shaped like the devices we actually build. And the price of "not having backprop" is usually expending more FLOPs, getting worse sample efficiency, etc.
The biggest "device" that doesn't do backprop is the brain, and that's because the brain doesn't have the connectivity or the coordination to pull it off. Both of those are "expensive" for something like it to implement. Cheap for us though. We aren't stuck with neurons that only get locally available information and have to implement learning rules based on that. So, skill issue?
What about applications like distributed computing using volunteer computers, instead of datacenters? You can't really do that with normal backprop approaches. E.g. for LC0 the training data is generated in a distributed manner, but the neural nets themselves are trained centrally.
I mean co-ordination requires energy though. The brain wattage looks at GPUs and says "skill issue". But you're right that we look at natural energy production techniques and say "skill issue"
I think both can be true - ie. We can harness energy from the sun at much bigger scale than nature has so far, yet we can still learn a lot from how nature has spent massive parallel search over millions of years on coming up with incredible nano engineering we still have no clue how to do
Even today's LLMs suddenly get power-competitive when you compare by "power per task". Sure, a GPU can draw 1000W under load. But it also works very fast, and doesn't have to spend any time on things like "sleep".
The trick about comparing the two is that different things are expensive to different substrates.
Coordination is cheap for GPGPU and expensive for brain. When you have a fixed number of reusable general purpose computational units, coordinating execution is more natural than not coordinating execution, and the power cost is nil. When your computational units are independent, purpose specific, and fully embedded into the data path, coordinating them can get less natural and, frankly, optional. When wiring is expensive, coordination can become expensive in turn.
Another thing in the same "cheap for GPGPU but expensive for brain" regime is bandwidth. Look no further than optic nerve to see just how hard it is for nerves to push any appreciable amount of data. Another thing is connectivity. For GPGPU, global connectivity is natural - but the brain has to pay in physical wires for all the connectivity it has, and, see "bandwidth": it doesn't have any good wires. Yet another thing is weight reuse: a big part of why humans get "handedness" is that the brain can't just reuse the motion control circuitry for one hand for another nearly identical hand.
And the final thing I can name off the top of my head is memory - but specifically, memory capable of fast R/W. The capacity of human "working memory" is a disgrace, and not because there was no use for more. Humans rapidly lose visual fidelity of representations for objects they aren't directly looking at, and not at all because "being able to check how things looked 2 seconds ago" is useless. Those capabilities were just too expensive for the substrate to afford them easily.
It's why brains, broadly, favor dataflow-like and SSM-like dynamics, with largely fixed asynchronous dataflows and recurrence over updated local information - instead of something that would require a lot of global connectivity and transformer-like many-to-many attention ops. SSM is not necessarily the "best" tool for the job in ML land, for most jobs - but when you struggle to fit "attention" into your connectivity/bandwidth budget, and your memory is extremely expensive but hard-coupled to processing, SSM starts looking very appealing.
Now, something that might be expensive for GPGPU but cheap for brain, for once? Online learning. Maybe it's substrate dependent, or maybe it's going to get cheap in GPU land too once we figure out the trick. But so far? No one figured out how to make it cheap, stable and usable. You'd be lucky to get "pick one".
I mean the reason why models are using less energy is because they are getting smarter per token and also engineering algorithms/chips that make inference cheaper.
If we could have success with spiking neural networks in silico they would take even less energy, because they don't require global co-ordination. Co-ordination is information and "information = energy by the second law of thermodynamics" is my crank proof
Also the brain has way more parameters than LLMs and also has different neurotransmitters, loops, branching etc so they probably have WAY more capacity than LLMs.
But coding output/W LLMs have us beat
1 reply →