Comment by piterrro

18 hours ago

Could this vllm port be faster to install? Im starting gpu machine multiple times a day and it takes 5 minutes to set vllm up. If Inise this port that time is minimized?

This is on my list to evaluate, I absolutely do not want to download 9gb of supply chain risk into prod every time we upgrade, when I can compile 70mb of binary. We run vLLM in a container with hardware passthrough for gitops, having the entire environment in a single container would drastically improve things and move local LLM into a pattern that more closely follows our other CI/CD systems, rather than this hulking behemoth snowflake deployment.