10M Tokens per Second: Record or Metric Sleight of Hand?
инференсLLMбенчмарки
What is behind 10 million tokens per second
I would not call 10 million tokens per second a model's speed until it is clear what was measured and how. The story began with a post by Rysana on X, referenced by User 144406. The discussion understandably raised questions, and the cost of that speed was even compared to a gold bar per million tokens.
The core issue is not the large number itself, but its denominator. It could be aggregate cluster throughput under batching, prefill speed, processing of cached states, or the decode rate for a single request. These modes look completely different to a user, even though marketing labels all of them as tokens per second.
Long-context claims create another revealing source of confusion. Lightbits and Inferra state that they support a 10-million-token context on L40S hardware, but the cited result concerns time to first token: 26 seconds versus 8.3 hours. That is a major acceleration of one stage, not the generation of 10 million new tokens every second.
For perspective, NVIDIA's benchmark page lists 1,183,327 tokens per second for DeepSeek R1 on a Vera Rubin NVL72 system. It is an enormous figure, but it applies to a specialized multi-processor system and remains well below the 10 million under discussion.
Why this record says little about real-world use so far
Without the hardware configuration, batch size, and a clear separation between prefill and decode, the record cannot be mapped onto a typical API request. As of September 30, 2026, the available context does not show a single model producing 10 million output tokens per second for one user.
Three engineering measures matter here: time to first token, sustained generation speed, and processing cost. Aggregate throughput can be an excellent measure of cluster utilization while saying nothing about the latency of an individual session. The larger the batch, the more impressive the headline number and the weaker its connection to interactive use.
My conclusion is simple: this is not necessarily a fabricated result, but it is clearly an ambiguous claim so far. In the inference race, the winner is not the one with the largest number, but the one who labels the chart axes most honestly.