DeepSeek published the weights of DeepSeek V4.1 Flash on September 10, 2026 under the MIT licence, and the files come to 510.3 GB. DeepSeek’s announcement calls it a 552B-parameter mixture-of-experts model that reads images and text. Hugging Face counts 763,205,315,794 parameters in the repository, and the model card says 8 billion are active for each input token and 16 billion for each output token. We have not run the model. Our test machine is an Apple M4 Pro with 48 GB of memory, and the download is more than ten times that size.
Takeaway points
- Hugging Face counts 763,205,315,794 parameters in the repository, while DeepSeek’s announcement calls it a 552B-parameter model.
- The model card says 8 billion parameters are active for each input token and 16 billion for each output token.
- We have not run DeepSeek V4.1 Flash: its 510.3 GB of weight files is more than ten times the 48 GB of memory on our test machine.
Every score and design figure below is DeepSeek’s claim, taken from its announcement, model card, inference notes and technical report. The file sizes and parameter counts come from the Hugging Face listing of the repository, which we saved on October 5, 2026.
How big is DeepSeek V4.1 Flash, 552B or 763.2B parameters?
The answer depends on which document you read, because DeepSeek’s three documents and the Hugging Face data do not agree.
The announcement calls the model a “552B-parameter MoE.” The model card and the technical report say 552 billion is the backbone, and add 196 billion parameters for Engram, a lookup memory that the report says is split evenly across two modules at layers 1 and 14. Those two figures sum to 748 billion. Hugging Face counts 763,205,315,794 parameters in the weight files, about 15 billion more than that sum. None of the three DeepSeek documents itemises the difference. Our reading is that the 552B in the announcement names only the backbone, but DeepSeek does not say so.
| Source | Parameters |
|---|---|
| Announcement | 552B |
| Model card and technical report | 552B backbone plus 196B Engram, 748B in all |
| Hugging Face count of the weight files | 763,205,315,794 |
Hugging Face records 557,171,343,360 of those parameters as 8-bit integers, 204,015,223,296 as 8-bit floating point, 1,976,441,856 as 16-bit and 42,307,282 as 32-bit. The config file names the quantisation method fp8, with the expert weights in fp4.
The model card says each layer holds 1 shared expert and 384 routed experts, and that 6 routed experts are active for each token. The repository’s config file gives the same three numbers. The card describes a 40-layer network split into a 20-layer encoder and a 20-layer decoder, and says the split lets the model use 8 billion parameters while it reads a prompt and 16 billion while it writes a reply. The announcement puts it as “just 8B active parameters for input, 16B for output.”
The config names the architecture DeepseekV41ForCausalLM, with a hidden size of 5,120, 64 attention heads, a vocabulary of 129,280 tokens and 1,048,576 maximum positions. The sliding attention window is 128 tokens. Images pass through a vision encoder that DeepSeek says it trained from scratch, and the config gives it 32 layers and a ceiling of 1,024 image tokens.
The card’s title is “Pushing the Limits of KV Cache Compression”. It says the cache is 890 bytes per token, roughly 1/4 of the figure for DeepSeek-V4-Flash, and that the persistent cache, which the report says is kept on disk or in host memory, falls to roughly 1/8. The card also says reasoning effort is set with a whole number from 1 to 100.
Can DeepSeek V4.1 Flash run on a Mac with 48 GB of memory?
No, not from the official files. The repository holds 48 safetensors files that total 510.3 GB, and the two largest are 101.5 GB each. Each of those two files is more than twice the memory of our test machine, and the full set is more than ten times it. The repository holds no GGUF or MLX build, the formats that llama.cpp, LM Studio and Apple’s MLX read. If you want to see how a GGUF file is loaded once you have one, we wrote how to import a local GGUF file into Ollama with a Modelfile, using a model that does fit.
DeepSeek’s own inference notes describe the supplied code as “A readable reference implementation rather than a production serving engine.” The notes say the weights must first be converted into one checkpoint file per parallel rank, and the example sets that number to 8 and starts 8 processes. The requirements file asks for PyTorch 2.10.0 or later and version 0.1.8 of the tilelang package. The model card states no minimum hardware.
The announcement invites operators planning “a large-scale deployment with 2,000 GPUs + a storage cluster” to get in touch. It adds: “We’ll work closely with the open-source community on V4.1-Flash inference support and explore more deployment options.”
Other accounts have published conversions. We searched Hugging Face for the name DeepSeek-V4.1 on October 5, 2026 and saved the result. It returned 189 repositories, of which 188 sit outside DeepSeek’s own account. In those 188, 23 carry GGUF in the name or tags, 20 carry MLX and 11 carry NVFP4, a 4-bit format. These counts come from that one name search on that one day, and a repository that leaves the name out would not appear. We have not tested any of the conversions, and DeepSeek’s card does not mention them.
For a sense of what does run on our machine, we measured Gemma 4 31B QAT on an M4 Pro with 48 GB. We have no speed or memory figure for DeepSeek V4.1 Flash.
Is there a chat template for DeepSeek V4.1 Flash?
Yes, although the model card still says otherwise. The card says: “This release does not include a Jinja-format chat template.” The repository holds a 15,348-byte file named chat_template.jinja, added in a commit dated October 1 with the title “Adding chat template (#68)”. The last commit that changed the card is dated September 10, and its text still carried the older sentence when we fetched it on October 5.
Two settings differ between documents as well. The card recommends a temperature of 1.0, and the interactive example in the inference notes passes 0.6. The chat template sets reasoning effort to “high” by default and maps that word to 75, while the card says every instruct score on it was produced at 100.
How does DeepSeek V4.1 Flash score against Opus-5.0, GPT-5.6 Sol and Kimi K3?
DeepSeek’s card puts it ahead of Opus-5.0 on two software tests and well behind on four harder ones. Every score is DeepSeek’s own evaluation, run at the maximum reasoning effort, and we have not reproduced any of them.
| Test | DeepSeek V4.1 Flash | Opus-5.0 | Others on the card |
|---|---|---|---|
| DeepSWE v1.1 | 74.2 | 74.0 | GPT-5.6 Sol 73.0, K3 67.5, DeepSeek-V4-Flash 54.4 |
| Terminal-Bench 2.1 | 90.6 | 89.1 | |
| Terminal-Bench 3.0 | 30.0 | 43.3 | |
| Terminal-Bench 4.0 | 31.2 | 51.8 | GPT-5.6 Sol 39.9 |
| HLE | 36.8 | 56.3 | |
| ProgramBench | 20.3 | 37.0 |
The announcement says tests by “multiple parties” put V4.1 Flash ahead of V4-Pro on performance, cost and speed, and it names none of the parties. In its table of base models the card says “Scores within 0.3 of each other are considered equivalent.” It states no such rule for the table that compares V4.1 Flash with other companies’ models, where the DeepSWE margin over Opus-5.0 is 0.2. The technical report says a gap with the largest models remains on science tasks such as Terminal-Bench 4.0.
The DeepSWE score depends on the software that drives the model. A second table on the card reports 74.2 with the mini-SWE agent, 69.8 with Claude Code, 65.6 with Codex and 65.5 with OpenCode. The announcement names OpenCode as an official partner.
The report’s conclusion measures the model against two systems that are absent from the card’s table. It names Fable-5 and GPT-6 Astra as top-tier models and says “a performance gap remains on the most challenging tasks.” The card gives no scores for either.
The cost of the top setting is in the report too. Raising reasoning effort from 25 to 100 lifted the DeepSWE score from 66.0 to 74.2 and used roughly 2.5 times as many output tokens, DeepSeek says.
What changed in the DeepSeek API with V4.1 Flash?
The model went live on DeepSeek’s paid service on the day of the release, under the name deepseek-flash. The announcement says V4-Flash and V4-Flash-Vision-Exp are retired. From 04:00 UTC on September 14, it says, requests for deepseek-v4-pro are routed to V4.1 Flash, an arrangement that “will continue until V4.1-Pro launches.” Off-peak rates are 50% of peak rates, according to the announcement. It gives no prices in the text we saved.
DeepSeek says it trained V4.1 Flash from scratch on 45 trillion tokens of text and images. The card says training began at a sequence length of 64K and moved to 1 million tokens once 34 trillion tokens had been processed. The report gives the batch size as 100.6 million tokens. The vision encoder was trained separately, the report says, first on about 47 billion image and text pairs and then on 236 billion tokens alongside a 4-billion-parameter language model that was later discarded. Of the later tuning stage, the report says “our post-training introduces no algorithmic innovation” and that the changes are in the training data.
The card’s base-model table lists 284 billion backbone parameters for DeepSeek-V4-Flash and 1.6 trillion for DeepSeek-V4-Pro. Hugging Face records the repositories for both as created on April 22. The V4.1 Flash repository was created on September 10 and last changed on October 1. It carries an MIT licence file and is not gated.
Have we run DeepSeek V4.1 Flash, and what has DeepSeek not published?
We have not run it. The official weights do not fit on our test machine, there is no GGUF or MLX build from DeepSeek, and we have not tested the community conversions. This piece therefore carries no speed or memory measurement of our own. Every figure about the model is the maker’s claim, and every file size and count is from the Hugging Face listing we saved on October 5, 2026.
What we did do was read the card, the announcement, the config, the inference notes, the chat template and the technical report against one another, and against the repository listing. That is where the disagreements above came from: the 552B figure against the 748B sum and the 763.2B count, the card’s chat template sentence against the file in the repository, and the temperature and reasoning effort settings.
DeepSeek has not said when V4.1-Pro will arrive. The card and the report give scores for a DeepSeek-V4.1-Flash-Base model, and the company’s Hugging Face account listed no repository under that name on October 5. Hugging Face records Base repositories for both V4 models as created on April 22.
The card gives no minimum hardware and no official build for llama.cpp or MLX. The report gives no energy figure for training. It also says errors in the model’s sparse attention and in its cache reconstruction “may still cause capability degradation in untested boundary cases.” If we get a build that fits in 48 GB, we will measure it the way we measured Gemma 4, and how we measure tokens per second in LM Studio describes the method.
For scale, Aikido’s Altar-1, another open-weight release we have reported on, comes to 327.99 GB of files.
Sources
- DeepSeek announcement, DeepSeek-V4.1-Flash release, September 10, 2026
- DeepSeek-V4.1-Flash model card, DeepSeek
Cite this page
Hill. "DeepSeek V4.1 Flash weights are out: 552B in the announcement, 763.2B in the files". Onticpost, published 5 October 2026. https://onticpost.com/deepseek-v4-1-flash-open-weights/
Updated 5 October 2026. This address stays the same when the piece is updated.