I am surprised at the lack of open-weights models in the >35B, but <200B range. I keep thinking about devices like the NVIDIA Spark and AMD Ryzen Halo, which have their 128GB of combined memory, but there are so few models made for that range. Nearly all the open weights distillations are for larger customer bases with <24GB VRAM.
American exceptionalism states that America is special and unique so everyone else must be a copycat. American ai labs don't need this kind of optimization and fable will outright refuse to do it.
I am surprised at the lack of open-weights models in the >35B, but <200B range. I keep thinking about devices like the NVIDIA Spark and AMD Ryzen Halo, which have their 128GB of combined memory, but there are so few models made for that range. Nearly all the open weights distillations are for larger customer bases with <24GB VRAM.
Qwen Flash Next 3.8 … even at 3 bit quant it is very solid.
The market is too small.
American exceptionalism states that America is special and unique so everyone else must be a copycat. American ai labs don't need this kind of optimization and fable will outright refuse to do it.