← Back to context

Comment by jjcm

5 hours ago

Most responses here are along the lines of "model capabilites move too fast to build hardware for".

I think the fact that there are plenty of 1yr+ old models on openrouter serving hundreds of billions of tokens a month shows that there's plenty of use case for models that are "good enough. Cerebras' entire business is serving older models at high speed. I would happily use an opus 4.7 at 15k tokens per second. The intelligence per second of an ASIC still makes sense even with rapidly evolving models.

"Intelligence per second" is a striking phrase! I'll be turning it around my head at a moderate IPS until I hopefully make something of it.

But you're right in sense: moderate intelligence at superhuman rates (and presuming moderate energy usage) is very compelling compared to an intelligence that takes 1000 years to return "42"

  • I look at it this way. A year ago is a long time in AI terms, but not that long. Those models were already decent for the tasks we're using SOTA models for today.

    So imagine taking a year-old SOTA model and running it at 100 tokens per second on an edge device. That's enough to feed a screen's worth of content through it and power decent multilingual message suggestions on IM.

    Imagine running it at 1000 tps. That's enough to reparse that screen mid-keystroke, and give you semantic autocomplete in text. Or fully general "the phone has a good idea of what you're attempting to do" context at all times.

    There's many, many new classes of features that will open up if decent enough models can be run on edge devices at 100+ "intelligence per second".

Totally. But it's worth noting that this is a pretty new thing! I wouldn't have bet on that a year ago, but now I would.

  • It is still a bet. Is the current model still going to be good enough next year? We have no idea what will change. If the change is minor improvements than current fable on a cheap is cheaper and better than next years sonnet on GPU. However there are plenty of things people want that maybe they will deliver and suddenly I wouldn't touch today's fable when I can run next years sonnet instead.

    • Yep, that's why I also described it as a bet :)

      But I think it's a good bet. I think that in two years, if I can get opus/sonnet 5.5 or the gpt-6 models for much cheaper and faster than whatever the "frontier" is at that point, that this will probably be a great trade for most of my work. I certainly don't know that for sure, that's why it's a bet, but it's what I think right now.

      I wouldn't quite say that about any of the open weight models at this point. But I'm hopeful that will change in the next generation or two of those models.