Comment by zkmon

3 hours ago

I don't get it. It's file size is about 6 times larger than 27B model for the same quant, but the performance improvement is hardly 10% across all benchmarks, according the metrics on it's hf page. Why should one devote so much more hardware for so little benefit?

6B activated weights per token vs 27B. Something like DGX Spark is way better suited for Flash Next.