DeepSeek has introduced an updated version of its Flash model, called Flash V4.1, which features significant architectural improvements. Despite having 763 billion parameters—more than 2.5 times the size of its previous version—the model uses less memory. This is made possible by changes in how the model manages key-value (KV) caches, which help track the model's state during interactions. These modifications have reduced the memory needed for KV caches by up to 75%, allowing the model to support four to eight times more users with the same memory resources. A major advancement in Flash V4.1 is the integration of a new type of parameter called N-grams. Out of the 763 billion total parameters, 196 billion are N-grams, forming what DeepSeek refers to as a "conditional memory module." These N-grams function as a kind of lookup table, providing quick access to relevant information without needing to be fully loaded into the GPU memory each time a token is generated. This allows them to be stored in system memory or even on fast storage, maintaining performance without compromising efficiency. While the model theoretically requires 763 GB of GPU memory for its weights in FP8, the use of N-grams reduces this need to about 567 GB. However, in real-world applications, the actual memory usage will be higher due to the KV caches, which depend on the length of user input and the number of concurrent users. DeepSeek is not the only company exploring this approach. Google has already implemented a similar technique, called Per-Layer Embedding (PLE), in its Gemma models, allowing certain weights to be stored on local storage. Similarly, Alibaba has introduced Qwen 3.8-Flash-Next, a 180 billion parameter model that includes 51 billion N-gram parameters, inspired by DeepSeek’s earlier research. Alibaba says this architecture will form the foundation for its next generation of Qwen models. The use of N-grams in large language models may become more common as companies seek to manage the high computational and energy costs of running massive AI systems. This approach helps reduce the need for an exponential increase in servers, which are both expensive to build and energy-intensive to operate. By leveraging N-grams and similar techniques, AI developers aim to create more efficient and scalable models for the future.