← Back to context

Comment by bitexploder

7 hours ago

Flash next is /really/ close. It is at parity with 4.7 as far as I can tell and basically where Opus 4.8 was. It is a genuinely good model. And I run it at home on $1500 of GPU at 125 t/s :)

Seconding this. Flash-next (and let's not forget, it's a PREVIEW of the 4 architecture - with the "real" 4 rumored coming later this month) is the first model I can run on reasonable hardware (2 thoroughly obsolete P100s off ebay at ~$100 each plus the RAM I could scavenge from other PCs at home) at a reasonable speed (22, with GPUs in layer-split due to llama-cpp's limitation on qwen4-exp arch, and no MTP. Strata could double these numbers).

It's.... the real thing, for the first time. If you cut me off cloud models today, I would get plenty of utility out of this thing.

(Others may have had the same feeling from GLM5.3 or Deepseek 4.1 flash but I never had a chance of running those.)

Running an nvidia card at full load, would cost me ~100 euro of electricty each month (europe). Of course one wouldn't have usage caps.

Where the internet was a subscription 15 euro subscription to encyclopaedic knowledge, an genAI subscription is renting a researcher/programmer for 100 euro.

  • > (europe)

    Given the massive difference in electricity price between different european countries, adding "Europe" doesn't bring much context.

    • Even with a country. Here in Norway we have multiple price zones, and at times there can be 100x difference between them, often 10x. All due to lack of transmission capacity between northern and southern zones.