← Back to context

Comment by AntiRush

1 hour ago

Agreed that glm-5.3-flash is a strong model. On a single 6000 pro you can (barely) fit a q2 quant in vram. I haven't used it extensively, but my initial feeling is that the q2 is quite a bit worse than q4 for this model. To run it comfortably at q4 with reasonable context length you really need 2 6000s.

When the qwen 4 series is released I am hopeful there'll be a strong model with the same architecutre. as qwen3.8-flash-next.