Comment by Marha01
4 days ago
> Yeah cause every car needs 8xH200 pulling 10kW to run a VLM at realtime speeds.
When the models stop improving, we will get model-specific ASICs that are much more power-efficient.
4 days ago
> Yeah cause every car needs 8xH200 pulling 10kW to run a VLM at realtime speeds.
When the models stop improving, we will get model-specific ASICs that are much more power-efficient.
> When the models stop improving
Soo, never? Granted Cerebras is a thing, if the process can be commoditized.
At the moment the area of edge inference at speed seems pretty bleak though.