Comment by bob1029

6 hours ago

I think Luna might be just small enough to provide some kind of stepwise improvement in how it is hosted.

Going from 81GB of weights to 79GB of weights can mean a 50% reduction in GPU capacity required.

If you can fit a model in just one GPU (or rack) as opposed to across an entire datacenter, the latency gains can be substantial too. If you can reduce token latency by half, that would double the amount of customers you could support.