Comment by vardalab

5 hours ago

Well, if you are serious about it and you have Strix Halo, there are better ways of getting more context and capability and speed. Lookup halogen for Strix The most cost-effective local option right now, I think, is dual R9700. You can run a 27B dense Qwen at FP8 around with a full context and 2-3 concurrent sessions of 260K context. If you go down to an MXFP4, you get 4 to 5 concurrent sessions. And speed is on par with anything you'll get from hosted providers. You're getting between 60-80 for FP8 and 150+ tokens per second speed for MXFP4 and pp is 4K+. Lookup vllm radiance There is also a lot of progress in running Qwen 3.8 next flash with dual R9700. Obviously one gets less context and speed is a little bit less, but it's still very acceptable. Better than what you're getting with Llama on Strix Halo, that's for sure.

These are small dense models,meaning Qwen, have gotten capable and fast. And there's been a lot of progress in the area. So sticking with Llama you are not taking advantage of the hardware you have. And yeah, for Strix Halo, you should just look into halogen and you shouldn't be just using 96 gigs for the VRAM. You should give it most of the VRAM to the inference and connect to it from your laptop or something. People ar egetting 1000+ pp with qwen 3.8 Next Flash