Comment by a11r
3 hours ago
Yes, you can probably get the whole thing up and running in about an hour the first time. If you pause and restart, it takes about 15 minutes to load the models from disk into GPU memory, so budget for cold startup time.
3 hours ago
Yes, you can probably get the whole thing up and running in about an hour the first time. If you pause and restart, it takes about 15 minutes to load the models from disk into GPU memory, so budget for cold startup time.
How does it take fifteen minutes to read <100 GB into GPU memory? Shouldn't that be limited by SSD speed with everything slower than a minute being a terrible ssd?
It also depends on the runtime, vllm is unbelievably slow at model loading compared to llama.cpp
A lot of cloud platforms have terrible slow network storage. They also might need to compile the GPU kernels fresh as they might not have a persistent CUDA cache.