← Back to context

Comment by zdragnar

7 hours ago

Model SOTA moves faster than chips can be designed or produced. You'd need to commit to a particular model for years to get payoff while still burning buckets of money producing new SOTA models to keep up with the competition.

It's why everyone and their dog runs these things on GPUs. When a new model supercedes the previous one, so long as you've got the memory for it your chips aren't obsolete.

I'm looking forward to someone picking a model to be "good enough" (say, qwen 4.0 or something) and selling them as peripheral hardware

At this point, LLM's are "good enough" for all kinds of tasks. Instead of making them more capable, now the efforts are making them smaller and cheaper.

All aboard! We're racing to the bottom now.

  • IMO this is the dream scenario! Cheaper and faster at the current level of capability gives us incredibly useful tools without the worst of the risks people fear. (Though there are certainly already great risks at the current level of capability as well.)

  > Model SOTA moves faster than chips can be designed or produced.

From what I remember working in that area the hardest part is getting masks for a design. Masks were developed in the span of half an year. Masks also reusable, they can be mixed and matched and this is why fabless companies work with fabs to produce specialized masks for them, it saves time for consumer to have masks for some macroblocks prebuilt.

Here's my analysis of how to etch relatively big LM into silicon: https://news.ycombinator.com/item?id=47109252

Given some amount of work with the fab before main pipeline set (I think a year long process), one can then spew LM-on-a-chip in six months or less and much more than 2 per year, because there can be several LMs in pipeline.

I don’t disagree with most of what you’re saying, except for one point: I must have gotten a dumb dog, I’m a little jealous…

  • Mine currently just helps me haul firewood, but I'm going to get him started on linear algebra next week. We'll see how it goes from there.

They have trillions..

Lots of people would have happily taken GPT-4o as good enough for a lot of use cases a year ago and not lived to regret it.

  • They have trillions worth of obligations to their partners in terms of compute purchase agreements and equity. Burning models to chips isn't really something they can blow money on just for funsies, they need to be able to justify it. The above comments are hypotheses as to why they haven't yet.

I know FPGAs are more expensive than GPUs, but are they fast enough to justify the extra cost?

  • FPGAs are FPGAs by virtue of putting on the chips vast, vast arrays of wiring that can be controlled by software. Any given utilization of the FPGA will leave large fractions of the chip resources unused. If you've got a highly stereotypical use case FPGAs will have a "highly stereotypical" set of components being unused, where it would be better to use that space instead to do real work. A lot of people only see the "pro" side of the FPGA proposition without realizing they come with some very substantial "cons" that are intrinsic to the way they work.

    • I guess that's why they work in particular niche spaces like a synthesizer where you have a max of 8 voices and every voice goes through the same pipeline (osc / filter / env / amp) and everything is necessarily running all the time. In that sense I suppose they're very good for modelling any kind of analog circuitry?

      Even then, while there are some amazing FPGA-based synths available, companies like Korg just put their code on a raspberry pi and call it a day. The same is true for emulators (SNES Mini etc. are also just raspberry pis under the hood iirc)

      4 replies →

  • It isn't just a matter of speed, it's also a matter of model quality. If they take 6 months to burn Fable to chips, and it takes 2 years to break even between design, custom fab, energy savings, etc, are those chips even worth running when the new models that are running on GPUs at that point are producing 10x better quality results?

    Sure, your 2.5 year old models are running faster, but you can't drop prices on them without pushing the break even point further out.

    If the cost difference isn't incredibly significant, will people even want to pay for the 2.5 year old model, or will they get more value for their money paying more to get better results from the newer model?

    There's a lot of open ended questions that I don't have the insiders knowledge for to suggest whether or not such a capital outlay would be a worthy investment.

    My guess is that state of the art stuff will stay on GPUs and models burned into chips will be for "good enough" applications that people are still teasing out. Probably highly specialized models in automated sensor units and such.

  • They're not magical go faster juice. I don't know of a microarch where they're faster than modern GPUs at ML training or inference.

  • They are more like a way to proving the architecture of the accelerator before committing 100's of millions into a custom ASIC with TSMC

  • 1. No

    2. They don't have enough capacity either

    The current largest FPGA, the AMD Versal Premium VP1902 has 18.5 million logic cells. That's not even enough for the smallest whisper.cpp model (75M).

    You'd have to order hundreds of thousands of them (or millions) to serve even a single copy of a frontier model, and at that scale inference quickly becomes starved by the speed of light.

    • Well, you'd use BRAM to store model weights, not fabric. But still, you only get a couple hundred MB for probably close to US $100k per chip.

      It's likely that the major FPGA vendors will soon announce parts specifically architected to support LLMs and similar models. But the current generation isn't suitable for that at all.

      2 replies →