C2C transfers meaning between LLMs through the KV cache
Cache-to-CacheKV-кешмультиагентные системы
What Cache-to-Cache actually transfers
What stands out most is that C2C transfers meaning between LLMs without an intermediate text reply. In the Cache-to-Cache paper on arXiv, the authors propose taking the source model’s KV cache, mapping it into the recipient model’s representation space with a neural network, and merging it with the recipient’s own cache. The recipient can then continue decoding as if the relevant context were already embedded in its internal state.
First, both models run prefill over shared context and build their own KV caches. A trainable fuser then projects selected source states layer by layer into the corresponding recipient layers; residual fusion adds the transferred semantics, while a learned gate determines which layers should be changed at all. That detail matters: blindly injecting information into every layer would look more like a way to corrupt representations than a communication protocol.
Oracle experiments showed that KV-cache semantics can be enriched without increasing cache size. At publication time, the authors reported average accuracy improvements of 6.4 to 14.2 percentage points over standalone models, and 3.1 to 5.4 points over text-based transfer. So the idea is supported not only by appealing visualizations of hidden states, but by task results.
The work is dated October 2025, so it is more useful to read it as an emerging direction than as breaking news. An official implementation repository makes C2C more concrete than a purely conceptual proposal, but it does not make it a universal protocol.
What changes for multi-agent systems
The practical benefit is straightforward: models no longer need to package working state into text and have another model parse that package again. This can reduce token usage and text-handoff latency while retaining semantics that an ordinary message could lose.
Prebuilt model combinations with access to internal caches stand to benefit the most. Universality is the trade-off: C2C depends on a trained projector, layer alignment, and compatible interfaces. An adapter for an arbitrary pair of models does not appear automatically.
My main question is not accuracy but controllability. Text exchanges are easy to log and inspect, whereas a fused KV cache remains an opaque internal state. The key question now is not whether a cache can carry meaning, but how many model pairs will need to be trained to understand one another from scratch.