← Back to context

Comment by jacquesm

3 hours ago

DS4 is an odd model. I have it working on way too many GPUs and yet for many tasks Qwen 3.8 will do much better. It also tends to loop, which is super annoying.

GLM5.3 runs on similar hardware and is much better so if you're going to burn cycles and brain power on this maybe look at GLM5.3 as a comparison as well?

Other than that, when you're done with that card...

Agreed that glm-5.3-flash is a strong model. On a single 6000 pro you can (barely) fit a q2 quant in vram. I haven't used it extensively, but my initial feeling is that the q2 is quite a bit worse than q4 for this model. To run it comfortably at q4 with reasonable context length you really need 2 6000s.

When the qwen 4 series is released I am hopeful there'll be a strong model with the same architecutre. as qwen3.8-flash-next.