Comment by simonw
12 hours ago
I ran some pelicans at the four different reasoning levels (none, low, medium, xhigh - apparently high and xhigh are aliases of each other) on a DGX Spark using Unsloth's unsloth/Qwen3.8-Flash-Next-GGUF (UD-IQ1_S):
https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
Surprised I didn't get one I liked as much as the Qwen 3.8 27B one https://simonwillison.net/2026/Aug/16/qwen-38-27b/#the-defau... , maybe because of quantization.
Why did you use 1-bit quantization vs 3-bit quantization?
It looks like the 3-bit requires 90 GB[1] which, I imagine, would fit within the DGX Spark's 128GB of unified memory.
[1] https://unsloth.ai/docs/models/qwen3.8-next
If I read correctly, that's based on a 1-bit quantization, and can we really expect that to produce any useful output at all?
You should find an excuse to offer 3D printed extruded pelicans from various models as awards for something. I have no idea for what, but the idea captivates and I'd love to win one somehow. They'd be collector's items in a few decades
If Simon would pitch for example PCBWay that and I am pretty sure they will sponsor it (assuming their logo stays). They can do laser engraved versions also ;)
The spark can easily run UD-Q4_K_XL on this model... using IQ1_S doesn't make much sense.
Doing this on a 1-bit quant is unfair
Tried again with a different quant, UD-Q2_K_XL:
https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
... and once more with UD-IQ4_XS
https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
the xhigh version looks amazingly good for a 2 bit quant.
The Pelican Brief