On September 26, 2026, IST Austria’s DASLab published a compressed GGUF build of Qwen3.8 Flash Next that keeps 256 of the model’s 512 routed experts in each layer, according to its model card. Hugging Face lists the build’s two files at 58.41 GB, against 360.00 GB for the original weights from Qwen. We have not run any of these builds. The smallest is 10.41 GB larger than the 48 GB of memory in our test machine, an Apple M4 Pro, so this piece reports what the documents say and where they disagree.
Takeaway points
- DASLab’s Coder build keeps 256 of the 512 routed experts in each layer, and Hugging Face lists its two files at 58.41 GB.
- The Coder model card reports 75.60 on SWE-bench Verified, against 82.80 for the uncompressed model.
- We have not run it: at 58.41 GB the smallest build is still larger than the 48 GB of memory on our test machine.
How big are the Qwen3.8 Flash Next GGUF builds from DASLab?
The five builds run from 58.41 GB to 83.62 GB, according to the file sizes Hugging Face lists. The new repository, ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF, was created on September 26 at 09:42 UTC, and its two model files were uploaded that evening, the commit history shows. On the same day DASLab added a fourth build, IQ3_S, to its earlier repository. Hugging Face records that repository as created on September 7, and its commit history shows the first three builds arrived on September 15.
Every build comes as two files. The first holds the transformer weights. The second holds a per-layer n-gram embedding table and is 28,800,138,432 bytes in all five builds. The cards say the table has 51.2 billion parameters and is stored at a fixed 4.5 bits per weight.
| Build | First file | Both files | Card’s figure |
|---|---|---|---|
| Coder (IQ1_M) | 29.61 GB | 58.41 GB | 58.4 GB |
| Q2_0 | 37.62 GB | 66.42 GB | 66.4 GB |
| IQ2_XS | 39.23 GB | 68.03 GB | 68.0 GB |
| IQ3_XXS | 47.04 GB | 75.84 GB | 75.8 GB |
| IQ3_S | 54.82 GB | 83.62 GB | 83.6 GB |
The sizes come from the Hugging Face API and the card figures are DASLab’s. They agree after rounding. Each repository also holds a vision projector of 907,543,008 bytes, or 0.91 GB, for image input. With it, the GGUF files in the Coder repository come to 59.32 GB.
The base repository, Qwen/Qwen3.8-Flash-Next, holds 131 safetensors files of 360,000,192,888 bytes in total, and Hugging Face counts 179,999,981,459 parameters in them. DASLab’s cards give the base model 176.9 billion parameters and 354 GB at 16 bits. The metadata Hugging Face reads from the unpruned GGUF files records 176,943,899,520 parameters, which is 3,056,081,939 fewer than its count for the base repository. The cards do not explain the gap. The same metadata gives the Coder build 116,514,464,640 parameters, 60,429,434,880 fewer than the unpruned builds.
How do GSQ and RCO compress Qwen3.8 Flash Next?
The Coder build deletes half of the routed experts and stores the rest at 3.5 bits per weight, the card says. Qwen3.8 Flash Next is a mixture-of-experts model, a design in which each token passes through a few small sub-networks instead of the whole model. The base repository’s config file lists 48 layers, 512 experts and 10 experts per token. Qwen’s card adds one shared expert to the 10 routed ones.
The card calls the experts deleted, not quantised. Its headline figure of 1.89 bits is an average over the original parameter count. “No individual weight is stored at 1.89 bits,” the card says. The files and their directory carry the label IQ1_M all the same, and the commit that added them describes the build as “50% expert-pruned, 1.89 bpw effective.” The label says one thing, the card says another. Our reading is that a reader who sees IQ1_M and expects a 1-bit file will be misled: the card puts the stored weights at 3.5 bits.
RCO, short for Riemannian Constrained Optimization, picked which experts to remove. The card says the search minimised the KL divergence between the pruned and unpruned model on calibration data, and did not rank experts by a heuristic importance score. The GGUF format stores a single expert count for the whole model, the card says, so every layer had to keep the same number, and DASLab set one budget per layer. The RCO paper on arXiv was first submitted on May 1, 2026 and revised on September 27, the day after the release. Its abstract says the method “enforces the expected budget exactly at every iterate.”
GSQ, short for Gumbel-Softmax Quantization, does the quantising. Its paper on arXiv was submitted on April 20, 2026. The abstract describes a post-training method that “jointly learns the per-coordinate grid assignments and the per-group scales,” and reports tests on the Llama-3.1-8B and 70B Instruct models. In the unpruned builds, the card says, RCO assigns a quantisation type to each of 352 tensors: 304 dense tensors and 48 fused expert matrices, one per layer.
The calibration data decided what survived. A first search used code and agentic data only, the Coder card says, and “image capability degraded substantially while coding scores were unaffected.” The released build was searched again with vision in the mix.
Can a 48 GB Mac run Qwen3.8 Flash Next?
Not by file size alone: the smallest complete build is 58.41 GB, which is 10.41 GB more than our test machine’s 48 GB of memory, and the smallest unpruned build, Q2_0, is 66.42 GB. We have not tested whether any of them loads, so we have no speed or memory figure of our own.
DASLab’s cards argue that the download size is not the memory requirement. Only the first file needs to be held in memory, they say, because the n-gram table is read one row per token and can stay on disk. The Coder card puts the part that must be in memory at 29.6 GB and calls that “within the capacity of a single 32 GB accelerator.” The card for the unpruned builds names two llama.cpp options that keep the table memory-mapped on disk. The Coder card’s own example command does not include them.
By the cards’ figures, the first file is smaller than 48 GB in four of the five builds: 29.61 GB for Coder, 37.62 GB for Q2_0, 39.23 GB for IQ2_XS and 47.04 GB for IQ3_XXS. IQ3_S, at 54.82 GB, is not. The card for the unpruned builds also says: “Add headroom for the KV cache and the vision projector (0.91 GB) if used.” On a 48 GB machine the IQ3_XXS first file would leave under 1 GB before the cache is counted. That is our arithmetic, not a measurement.
The card says the files are standard GGUF and run unmodified in llama.cpp, Ollama and LM Studio. The GGUF metadata names the architecture qwen4exp and a context length of 262,144 tokens. Neither card names the llama.cpp version required. For readers who want to try a local GGUF file in Ollama, we wrote how to import a local GGUF file into Ollama with a Modelfile, using Gemma 4 31B. How much memory a long context adds on this machine is in our piece on Ollama’s default context length for Gemma 4 31B on a 48 GB Mac. Both concern a different model, so treat them as method, not as a prediction for this one.
Are DASLab’s scores and licence labels reliable?
Every score below is DASLab’s claim, and the licence labels on the repositories conflict with the base model’s licence. The Coder card reports 75.60 on SWE-bench Verified against 82.80 for the 16-bit base model, or 91.3% retained, and 86.28 on LiveCodeBench v6 against 87.43, or 98.7%. It says the numbers were measured at “xhigh reasoning effort.” The card reports no score outside coding. It says the pruning was directed at code, agentic tool use, vision and spatial reasoning, and adds: “Degradation outside this set is an accepted cost of the method.” For general use it recommends the unpruned builds.
For those, the card gives a task average of 89.07 for Q2_0, 89.16 for IQ2_XS, 92.57 for IQ3_XXS and 93.26 for IQ3_S, against 93.12 for the base model. The card says recoveries slightly above 100% “reflect benchmark variance, not a model that is better than the one it was quantized from.”
The base model’s own scores differ by document, and the same test gives three different numbers.
| Document | Base model on LiveCodeBench v6 |
|---|---|
| Qwen’s model card | 91.9 |
| DASLab GGUF cards | 87.43 |
| DASLab card for a format for Nvidia GPUs, created October 2 | 57.03, at 131,072 new tokens |
The cards describe different test settings. That card also gives the base 79.8 on SWE-bench Verified, against 82.80 on the GGUF card. We did not run the benchmark, so we cannot say which figure fits which setting.
The licence labels conflict too. Both GGUF cards declare Apache 2.0 in their metadata, and Hugging Face tags the repositories accordingly. The text of each card says the weights “inherit the license of the base model.” That licence is the Qwen Community License 1.0, according to the licence file in Qwen’s repository. It requires products with more than 100,000,000 monthly active users or US$ 20,000,000 in monthly revenue to display the model name. It also says a licensee that runs a “Model as a Service or AI Work Assistant business” must “obtain a separate license from Qwen” before commercial use. DASLab’s October 2 repository declares the Qwen licence, not Apache 2.0. Our reading is that the Qwen licence is the one that governs, since the cards say the weights inherit it. Anyone building a product on these files should read the licence file in the base repository.
Qwen created the base repository on August 24 and uploaded the weights on August 26, Hugging Face data shows. Its card calls the model an “experimental preview of the architecture that will underpin Qwen4” and lists 125 billion parameters with 6 billion activated, plus 51 billion in the n-gram embedding.
Have we run Qwen3.8 Flash Next, and what has DASLab not published?
No. We have not downloaded or run any of the five builds, and the figures in this piece are the documents’ figures, not ours. The only numbers that are ours are two pieces of arithmetic: the 10.41 GB gap between the Coder build and 48 GB of memory, and the sums of the two files in each build. Our own speed work on this machine so far is on another model, in our test of Gemma 4 31B QAT on an M4 Pro with 48 GB.
DASLab’s cards do not name the calibration data that decided which experts were kept. The speed table on the card for the unpruned builds does not say what hardware it was measured on. The Coder card does not measure how much quality was lost outside coding. Neither card reconciles the Apache 2.0 tag with the Qwen licence, or the 176.9 billion parameter figure with Hugging Face’s count of 179,999,981,459. The Hugging Face organisation page spells the group’s name as the Distributed Algorithms and Systems Lab, and the cards as the Deep Algorithms and Systems Lab. The Coder card says: “This is an experimental release and feedback is welcome.”
For scale, IFM’s K2 Horizon family, another open-weight release we have reported on, has a largest model of 758.34 GB of weight files.
Sources
Cite this page
Hill. "DASLab publishes a 58.41 GB Qwen3.8 Flash Next GGUF build that keeps half the experts". Onticpost, published 5 October 2026. https://onticpost.com/qwen3-8-flash-next-gsq-rco-gguf/
Updated 5 October 2026. This address stays the same when the piece is updated.