Comment by TeMPOraL
4 hours ago
I look at it this way. A year ago is a long time in AI terms, but not that long. Those models were already decent for the tasks we're using SOTA models for today.
So imagine taking a year-old SOTA model and running it at 100 tokens per second on an edge device. That's enough to feed a screen's worth of content through it and power decent multilingual message suggestions on IM.
Imagine running it at 1000 tps. That's enough to reparse that screen mid-keystroke, and give you semantic autocomplete in text. Or fully general "the phone has a good idea of what you're attempting to do" context at all times.
There's many, many new classes of features that will open up if decent enough models can be run on edge devices at 100+ "intelligence per second".
No comments yet
Contribute on Hacker News ↗