3 min read

C2C transfers meaning between LLMs through the KV cache

Cache-to-CacheKV-кешмультиагентные системы

Cache-to-Cache transfers meaning between LLMs without generating an intermediate text message. It projects one model’s KV cache into another model’s internal space and merges it with the recipient cache. Reported results show accuracy gains, while practical use still depends on trained adapters for specific model pairs.

What Cache-to-Cache actually transfers

What stands out most is that C2C transfers meaning between LLMs without an intermediate text reply. In the Cache-to-Cache paper on arXiv, the authors propose taking the source model’s KV cache, mapping it into the recipient model’s representation space with a neural network, and merging it with the recipient’s own cache. The recipient can then continue decoding as if the relevant context were already embedded in its internal state.

First, both models run prefill over shared context and build their own KV caches. A trainable fuser then projects selected source states layer by layer into the corresponding recipient layers; residual fusion adds the transferred semantics, while a learned gate determines which layers should be changed at all. That detail matters: blindly injecting information into every layer would look more like a way to corrupt representations than a communication protocol.

Oracle experiments showed that KV-cache semantics can be enriched without increasing cache size. At publication time, the authors reported average accuracy improvements of 6.4 to 14.2 percentage points over standalone models, and 3.1 to 5.4 points over text-based transfer. So the idea is supported not only by appealing visualizations of hidden states, but by task results.

The work is dated October 2025, so it is more useful to read it as an emerging direction than as breaking news. An official implementation repository makes C2C more concrete than a purely conceptual proposal, but it does not make it a universal protocol.

What changes for multi-agent systems

The practical benefit is straightforward: models no longer need to package working state into text and have another model parse that package again. This can reduce token usage and text-handoff latency while retaining semantics that an ordinary message could lose.

Prebuilt model combinations with access to internal caches stand to benefit the most. Universality is the trade-off: C2C depends on a trained projector, layer alignment, and compatible interfaces. An adapter for an arbitrary pair of models does not appear automatically.

My main question is not accuracy but controllability. Text exchanges are easy to log and inspect, whereas a fused KV cache remains an opaque internal state. The key question now is not whether a cache can carry meaning, but how many model pairs will need to be trained to understand one another from scratch.

We previously examined how Claude Opus 4.6 configurations and context costs shape LLM system architecture. C2C extends that discussion by proposing direct transfer of already computed KV-cache state between models.