← Back to context

Comment by nowittyusername

7 hours ago

512 option isnt worth it imo, you get severe slowdowns when weights are that large. 256 is the sweet spot, you can run large open weight models at decent speeds for full private inference.

a) we don't actually know what the prices will look like yet, b) what about same weights + huge context? or, same weights that you'd run on 128gb/256gb, but multiple models running for different tasks?

  • I think it's safe to start the conversation as about bad as the jump from 256 to 512 on the M3, which was a little more than double base to 256. If it's surprisingly different at launch then it can be a party, but there is no sense getting your hopes up for that at the moment.

    Longer context also slows token prediction proportional to the context size. If it wasn't regularly referenced then there would be no need to keep it in RAM.

    Usually the pitch for more memory is "I can run a massive model/context and get my answer in a while instead of next weekend from disk".

> 512 option isnt worth it imo, you get severe slowdowns when weights are that large.

I think most people are getting 512 for running Chrome with a bunch of tabs open. /s