Comment by Marha01

4 days ago

> Yeah cause every car needs 8xH200 pulling 10kW to run a VLM at realtime speeds.

When the models stop improving, we will get model-specific ASICs that are much more power-efficient.

> When the models stop improving

Soo, never? Granted Cerebras is a thing, if the process can be commoditized.

At the moment the area of edge inference at speed seems pretty bleak though.