Comment by nblgbg
2 days ago
Is there any advantage to using the model from Unsloth compared with https://huggingface.co/Qwen/Qwen3.8-27B-FP8 ?
2 days ago
Is there any advantage to using the model from Unsloth compared with https://huggingface.co/Qwen/Qwen3.8-27B-FP8 ?
Depends on what software/hardware you'll run it. GGUFs from Unsloth can run on pretty much every single potato; full weights need beefy gpus
We also made NVFP4 ones if that helps! https://huggingface.co/unsloth/Qwen3.8-27B-NVFP4
This is the version we'll be testing on our rtx 6000 today! Thank you
Why not just run FP8 on vLLM with that much vRAM? It's plenty fast.
2 replies →
Unsloth one is gguf for llama.cpp (and some other on-device engines).
So advantage is not having to produce your own quantisation / gguf from .safetensors you've linked.
Run the unsloth if you are using llama.cpp (GGUF)
Run the one you linked if you are running vllm (safetensors)
Unsloth usually also fixes the models when they bork something, which always happens. For Gemma for example the tool calling wasn't working for the longest time.
That wasn't our problem right? Gemma officially updated tool calling which we adopted
if you have the VRAM, use offical release. quantized model lose focus after long context and can do damages or thinking loop