Qwen3.8 Flash Next shrunk from 354 GB to 29.6 GB
Qwen3.8 Flash NextквантизацияMoE
How the model was compressed without blind quantization
What stands out here is not GGUF itself, but the combination of two levers: reducing the number of MoE experts and applying non-uniform quantization to the remaining weights. In ISTA-DASLab's Hugging Face model card, Qwen3.8 Flash Next is described as a 512-expert MoE model compressed with GSQ and RCO for use in llama.cpp-compatible environments.
The published build reportedly removes half of the experts and lowers weight precision to 3.5 bpw. The original BF16 base occupied 354 GB, while the processed working set is 29.6 GB. This is more than a cosmetic file-size reduction: it moves the model into a different class of accessible hardware.
GSQ performs post-training scalar quantization. The method learns coordinate-to-grid assignments and scales for groups of weights using Gumbel-Softmax relaxation. The point is to avoid forcing every tensor into one coarse scheme when different model components clearly have different sensitivity.
RCO addresses the next level of the problem: it selects one of K quantization types for each tensor while meeting an exact final-size constraint. No separate penalty coefficient for the budget needs tuning. From an engineering perspective, this is stronger than simply trying GGUF presets, because size becomes an optimization constraint rather than an accidental settings outcome.
The reported quality retention is impressive: 91.3% of the original SWE-bench Verified result and 98.7% on LiveCodeBench v6. However, the official model-card materials substantiate GSQ, RCO, and the 512-expert architecture in detail, but do not make every compression and benchmark number independently verifiable. I therefore treat them as release results rather than reproduced measurements.
What this changes for local coding models
The practical shift is straightforward: a large MoE model can be brought closer to local inference not through one aggressive quantization pass, but through joint optimization of its structure and weight representation. That cuts memory requirements while preserving much of the claimed coding-task quality.
Still, expert removal is the first part I would validate. An average benchmark score can conceal regressions in rare languages, unusual repositories, and tasks whose routing relied on specialized experts. Real throughput, memory use during long sessions, and quality stability beyond the two cited tests are equally important.
The overall approach looks less like a trick and more like a serious budget-constrained compression strategy. The key open question is whether quality on rare MoE routes survives as well as it does on the familiar distribution of coding benchmarks.