← Back to context

Comment by geye1234

9 hours ago

I find 27B more accurate -- maybe because I'm running at FP8 instead of NVFP4? Flash Next starts making spelling mistakes when I get to 150K context or so. Also it sometimes ignores .md file instructions. Not sure if others have found that.

Spelling mistakes?

What inference engine are you using for flash next?

  • I'm running Pennyroyal's Docker image (on Podman) which uses sglang. I have a single RTX 6000 Blackwell and 128GB RAM. I turned off disk caching. I'm running with a ~500K context, but have been limiting it to 256K in the client (pi).

    It always detects its spelling mistakes, btw, but it worried me. It may turn 'rm -rf ' into 'rm -rf /' one day.

    Almost certainly the problem is my config, not the image.

  • Yep, can confirm that is NOT normal. Are you using Nvidia’s NVFP4 quant? There are other NVFP4s floating around but they are not as good. The quality of the calibration data really matters.

    Qwen Flash Next is just excellent, all the way to the very end of the native 262k context. (I haven’t tried YaRN scaling to 1M, so I don’t know about that.)