← Back to context

Comment by Iolaum

18 hours ago

I am really curious why after Qwen-3.8-flash and deepseek-v4.1-flash newer open weight models (or proprietary but they don't tell) don't use n-grams. It looks (to me) like they are a cheep way to add more knowledge to the model.

What do I mean by cheap? You can rely on the SSD to retrieve the relevant tokens as no computation is needed meaning you can leverage storage (or cpu ram if you don't have unified memory) to serve part of the model which (to my understanding) is much cheaper to get than GPU RAM.

Anyone know what am I missing? Or is it that the pace of iteration for labs slow enough that they can't actually leverage it yet?

DeepSeek Engram paper published: 12th January

Qwen3.8 Flash Next release date: 26th August

DeepSeek V4.1 Flash release date: 10th September

Current date: 6th October

I think they'll become more popular in the coming months. Also Gemma 4 PLE (April) is similar to DeepSeek Engram in a lot of ways, just with 1-grams.

On the proprietary model point: I'm personally curious about whether heavy n-gram offload is one reason Anthropic keep driving down their token vocabulary size (the other reason being eliminating the LM head gradient bottleneck).