Comment by gpugreg
8 hours ago
llama.cpp is already extremely fast for single-token responses (<5 ms). I can't see Jev being faster when taking network latency into account, except maybe for multimodal inputs.
8 hours ago
llama.cpp is already extremely fast for single-token responses (<5 ms). I can't see Jev being faster when taking network latency into account, except maybe for multimodal inputs.
Sounds like you are running a tiny toy model if you can get generations in under 5 ms? Typical response times from LLMs for typical "jev-like" queries from OpenAI and Anthropic are 2s-10s. Not milliseconds. Seconds. Same queries from Jev are like 0.2s. and the cost is 1000x.