Comment by zozbot234

3 hours ago

> Companies who run locally, are perfectly able to buy a few H200/B200 and get a setup that run a model that almost rivals Opus 5.0 in their office.

I agree with your broader point about Flash being about speed not total model size, but I think we should also point out that H200/B200's are seriously overkill for the "run a model in your office" scenario. That sort of hardware is optimized (in a roofline analysis sense) for running hundreds of concurrent sessions on a 24/7 basis. You're severely overpaying for your VRAM in basically any typical local-inference scenario, you should most likely be buying gear based on LPDDR and Flash memory instead which will slash your cost by orders of magnitude.

> I think we should also point out that H200/B200's are seriously overkill for

I simply mention what came to mind ;)

A quad 6000 with 96GB, can run this model at NVFP4. That is 60.000 Euro for the GPUs and lets be generous with another 20.000 for the rest of the system. The price of a single developer for a year.