Comment by jjcm
2 days ago
Image->html test for this.
Original images: https://image.non.io/neonRamenDesigns.webp
Qwen 3.8 build: https://html.non.io/neonRamenQwen3.8-27b
Overall I'm very impressed with how well this did. It's a big improvement over 3.6, and it feels on-par with some much, much larger models. I think this one is on-par with Gemini 3.7 Flash.
One thing to note - the build for this on my RTX 6000 pro blackwell took a long time. Easily one of the longest builds I've done. It took around 2 hours to build the site. Obviously we'll have some quants for this soon that will accelerate things, but I was still surprised with how long it took.
Comparison builds from this week:
https://html.non.io/neonRamenGemini3.7
https://html.non.io/neonRamenGLM5.3 (note: non-multimodal)
what token/s?
27 t/s. I suspect there will be significant speed ups in the coming weeks.
You should be getting way more than that on a 6000 pro even today. I'm getting 40tok/s on a pair of 3060s. You can ask a SOTA model to optimize your setup for you.
Getting 30 t/s on a Mac M5 Max laptop. You should be getting close to 200 if you can use the 5090 acceleration tools I see posted here. There are forks for the 4090 and the 3090. Maybe there's a fork for the 6000?
I'm seeing the exact same number on my Blackwell box. MTP put it at almost 70 which is pretty decent
Any idea why it’s so slow? the entire model should fit in the vram of one card.
1 reply →
harness setup? how much vram ? how are you handling a 2hr long build ? multiple sessions? fan out sessions (subagents)?