← Back to context Comment by stymaar 3 hours ago Flash[1]: 309B total / 15B activated parametersPro [2]:, 1.02T total / 42B activated parameters[1]: https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL[2]: https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL 9 comments stymaar Reply verdverm 3 hours ago There's also a Qwen 3.5 9B distillhttps://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B gandreani 3 hours ago Those this mean they've fine-tuned this Qwen 3.5 9B on output from the V2.6 model? mydreamof 3 hours ago It is a 9B agentic model developed by Xiaomi MiMo through supervised fine-tuning of Qwen3.5-9B on MiMo-generated data simonedepertis 3 hours ago [flagged] verdverm 3 hours ago curious why the HF pill (on the right) always has inaccurate values bopbop9876 37 minutes ago I believe it's because this model is natively fp8 (for the most part), and that display struggles native quants. stymaar 3 hours ago I noticed the same, and I wonder as well. verdverm 2 hours ago I suspect they are calculating something in the weights or config, I see it pretty consistently with quants segmondy 2 hours ago more like 500B in FP8
verdverm 3 hours ago There's also a Qwen 3.5 9B distillhttps://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B gandreani 3 hours ago Those this mean they've fine-tuned this Qwen 3.5 9B on output from the V2.6 model? mydreamof 3 hours ago It is a 9B agentic model developed by Xiaomi MiMo through supervised fine-tuning of Qwen3.5-9B on MiMo-generated data simonedepertis 3 hours ago [flagged]
gandreani 3 hours ago Those this mean they've fine-tuned this Qwen 3.5 9B on output from the V2.6 model? mydreamof 3 hours ago It is a 9B agentic model developed by Xiaomi MiMo through supervised fine-tuning of Qwen3.5-9B on MiMo-generated data simonedepertis 3 hours ago [flagged]
mydreamof 3 hours ago It is a 9B agentic model developed by Xiaomi MiMo through supervised fine-tuning of Qwen3.5-9B on MiMo-generated data
verdverm 3 hours ago curious why the HF pill (on the right) always has inaccurate values bopbop9876 37 minutes ago I believe it's because this model is natively fp8 (for the most part), and that display struggles native quants. stymaar 3 hours ago I noticed the same, and I wonder as well. verdverm 2 hours ago I suspect they are calculating something in the weights or config, I see it pretty consistently with quants
bopbop9876 37 minutes ago I believe it's because this model is natively fp8 (for the most part), and that display struggles native quants.
stymaar 3 hours ago I noticed the same, and I wonder as well. verdverm 2 hours ago I suspect they are calculating something in the weights or config, I see it pretty consistently with quants
verdverm 2 hours ago I suspect they are calculating something in the weights or config, I see it pretty consistently with quants
There's also a Qwen 3.5 9B distill
https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
Those this mean they've fine-tuned this Qwen 3.5 9B on output from the V2.6 model?
It is a 9B agentic model developed by Xiaomi MiMo through supervised fine-tuning of Qwen3.5-9B on MiMo-generated data
[flagged]
curious why the HF pill (on the right) always has inaccurate values
I believe it's because this model is natively fp8 (for the most part), and that display struggles native quants.
I noticed the same, and I wonder as well.
I suspect they are calculating something in the weights or config, I see it pretty consistently with quants
more like 500B in FP8