Comment by boredatoms 3 hours ago It also depends on the runtime, vllm is unbelievably slow at model loading compared to llama.cpp 0 comments boredatoms Reply No comments yet Contribute on Hacker News ↗
No comments yet
Contribute on Hacker News ↗