Qwen3.8-Flash-Next: 50–60 tokens/s on two GPUs
Qwen3.8-Flash-NextStrataлокальный инференс
What the home-server run actually showed
The striking part is the claimed throughput: in a post by user 122159126, Qwen3.8-Flash-Next running on the Strata engine produced 50–60 tokens per second. The hardware is unusual: an RTX PRO 2000 Blackwell with 16 GB and an RTX A2000 with 12 GB in one home server. This is a real user report, but not an independently reproduced benchmark.
According to Qwen's official materials, this is a 125-billion-parameter MoE model with 6 billion parameters activated for each token. That sparse design helps explain why a model of this scale can even be discussed for a local machine. Strata community materials describe the engine as a local inference solution that distributes work across GPUs, system memory, and the CPU.
Qwen3.8-Flash-Next also has a demanding context specification: its native context is 262,144 tokens, with YaRN extending it to a claimed 1,000,000 tokens. The context length used in this run was not disclosed, so the 50–60 tokens/s result should not be extrapolated to long conversations. Performance on a short prompt can differ substantially from performance with a filled context window.
The main gap in the report is not the headline number but the missing details around it. The available description does not state the quantization settings, system RAM capacity, model-splitting scheme, or PCIe utilization. It is also unclear whether the measurement covers generation alone or includes prompt processing.
Why the mixed-GPU pair matters more than the model itself
My conclusion is straightforward: running a large MoE model locally on heterogeneous GPUs no longer looks like a trick; it looks technically plausible. One accelerator is considerably newer than the other, yet Strata apparently used the pair without throughput collapsing to only a few tokens per second.
The biggest beneficiaries are local experimenters who value data control and freedom from a remote API. Still, one result does not become a universal benchmark: performance will depend on quantization, context length, system memory, and inter-device overhead.
As of September 30, 2026, there is no broad set of independent Strata benchmarks for this exact GPU combination. I therefore treat 50–60 tokens/s as a strong signal, not as a promise that will automatically repeat on the next desk.