A Simplex Cut Per-Token Memory Traffic by 44%
ассоциативная памятьсимплексыоптимизация памяти
What address stretching and the simplex changed
What stands out here is not geometry for its own sake, but a concrete result: in the described experiment, a 4D simplex requires five memory-cell accesses instead of the nine used by a 2D cube. Per-token traffic falls from roughly 111 KB to 61 KB, a reduction of 44%.
At the time of the discussion, the primary source for these figures was a message from the experiment’s author in a technical Telegram community. This is not a published benchmark or a completed research paper, so the results should be treated as a report on a current prototype rather than a universal property of simplex memory.
The first optimization is even simpler: read and write addresses are multiplied by a single factor, expanding the available address space. For 1,024 pairs, the original metric was in the 0.69–0.73 range; with stretching at γ = 2, it rose to 0.998–0.999. For 4,095 pairs, it went from 0.04–0.05 to 0.94–0.95 at γ = 4.
According to the author, this form of stretching does not increase computational work. The cost is paid in memory because more cells are needed. In other words, it is not a free optimization for the whole system, but a trade of physical capacity for a more favorable address layout.
Replacing the cube with a simplex targets a different cost center. As dimensionality rises, a cubic scheme rapidly increases the number of cells that must be read, while the tested simplex variant provides a more compact access neighborhood. The author also reports notably better extrapolation without stretching and a slightly higher ceiling with it.
Where the gain ends
The result looks meaningful from an engineering perspective, but its scope is still narrow: the gain applies to a specific memory design and addressing method. The 44% figure cannot be generalized to every neural memory system or vector-search workload.
The first things to evaluate are not just traffic, but total memory budget, access latency, and behavior under noisy queries. The author already notes that the simplex scheme handles noise in addresses, including paraphrases, less well and uses slightly more memory. This is where an elegant reduction in reads can meet the actual query distribution.
The next test, a 6D simplex, was still being computed when the message was posted. The theoretical expectation was optimistic, but the source data contains no result, so attributing an additional gain to it would be premature.
The most interesting part is not a record number but the shape of the trade-off: less data movement in exchange for more storage and potentially lower resilience to noise. The approach will be decided not by geometry on paper, but by whether this balance survives heterogeneous addresses and paraphrased queries.