← Back to context

Comment by Der_Einzige

5 hours ago

Switching to mlx-vlm is basically harmful since it (like vllm and sglang) have such garbage support for modern samplers. To be clear, I am one of the authors on the min_p paper, and if min_p is the best you have (when llamacpp supports the far superior top-n-sigma), than I have no reason to switch even if you are somehow faster.

(https://arxiv.org/abs/2411.07641)

And if you do care to support modern samplers, you can start with the following:

1. https://arxiv.org/abs/2509.23234

2. https://arxiv.org/abs/2509.02510

3. https://arxiv.org/abs/2604.11012

Thanks for bringing this up. You're 100% right. But most people, even technical ones are oblivious to how much of a difference modern samplers and higher quality quantization algorithms make for on-device LLM inference and are stuck with good old top-p, top-k samplers and RTN quantization.

TBF mlx-vlm does support min-p sampling, but none of the other modern samplers that you list. Ollama and LM Studio are even worse with only top-p and top-k samplers.