×
Community Blog When AI Transitions from 'Conversation' to 'Collaboration,' Who Will Manage the Working Memory of Millions of Tokens?

When AI Transitions from 'Conversation' to 'Collaboration,' Who Will Manage the Working Memory of Millions of Tokens?

This article introduces Mooncake, a KVCache infrastructure designed to eliminate memory and latency bottlenecks in multi-agent LLM inference.

Editor's note: When large language models (LLMs) evolve from simple single-turn conversations to tool calling, and then to multiple agents collaborating to complete complex tasks, an issue ignored by most people is quietly becoming the core contradiction of performance bottlenecks. Every LLM invocation repeatedly computes the context that has already been computed. Recently, at the Agent Application and Architecture Engineering Conference held in Shanghai, a developer from the OpenAnolis community gave a technical share titled "Multi-agent Oriented KVCache Optimization." From the underlying perspective of Agentic inference, he systematically deconstructed the memory and latency challenges faced in the multi-agent era, and deeply demonstrated the disaggregated architecture design and large-scale production practice of Mooncake as a KVCache data foundation.

1

The Multi-Agent Era: "Working Memory" is Exploding

The agent systems we see today are no longer simple "Q&A" conversations. In a typical multi-agent workflow, the context that needs to be carried in each LLM invocation includes: system prompts, tool definitions and function signatures, multi-turn conversation history, and the intermediate status transfer between upstream and downstream agents. These together constitute the "working memory" of the agent.

An intuitive figure: when 10 agents collaborate for 20 turns of interaction, the working memory can easily reach millions of tokens.

The problem is that current inference engines treat each invocation as an independent request. The complete context is repeatedly prefilled, and the same prefix is recomputed over and over again. The direct consequences are: aggravated GPU memory fragmentation and a sharp increase in Time to First Token (TTFT) latency. In certain scenarios, when there is no KVCache persistence, TTFT degrades by up to 136 times. The dynamic loading of tools by MCP causes the cache hit ratio to plummet from 85% to 0%. Even the field order differences in serialized JSON can slow down TTFT by 65%.

A highly insightful analogy was proposed in this speech: KVCache is to agentic systems what the database Buffer Pool is to a traditional database. 30 years ago, databases became independent from general-purpose operating systems and began to manage memory themselves. Today, agentic inference is becoming independent from general-purpose LLM Serving and moving towards the path of autonomously managing KVCache.

Agentic Inference vs Traditional Inference: a Fundamental Paradigm Shift

Traditional inference workloads are dominated by single-turn or a small number of multi-turn interactions. The input-to-output ratio is approximately 10:1. The workflow presents a linear topology, and KVCache can be released after the request ends. In contrast, the feature of agentic inference is dozens of turns of LLM-tool recursive invocations. The input-to-output ratio can reach 100:1 or more. The workflow presents a DAG or even a dynamic graph structure. KVCache requires persistence across tool calling, and needs to be shared among multiple agents.

This means that cache policies designed for traditional inference almost completely fail in agentic scenarios. We need a completely new layered optimization technology stack from the infrastructure layer to the application protocol layer.

Mooncake: a KVCache Data Foundation Born for Disaggregated Architecture

Mooncake is the infrastructure core of this technology stack. As a disaggregated architecture project for large language model (LLM) inference, Mooncake provides key capabilities such as Transfer Engine (full-link zero-copy, multiple network interface controller (NIC) pooling, supporting up to 8 × 400 Gbps aggregation bandwidth), KVCache Store (transparent tiered cache, covering VRAM/DRAM/SSD/Remote), and Elastic EP. Currently, Mooncake has received over 4,000 stars on GitHub and integrated with 12+ ecosystem projects, covering three major scenarios: inference, middleware, and reinforcement learning training.

2

In terms of disaggregated architecture, Mooncake supports various disaggregation patterns: Prefill-Decode disaggregation (PD Disaggregation), elastic PD disaggregation (EPD), reinforcement learning disaggregation (RL Disaggregation), and Attention-Free disaggregation (AF Disaggregation). The common target of these patterns is to independently extend compute resources at different stages to maximize resource utilization efficiency.

In the in-depth integration with the SGLang inference engine, Mooncake KVCache Store acts as an L3 cache to implement cross-machine KV sharing and persistence. It supports "infinite long context" (Infinite Context), and overlaps the execution of GPU compute and KV data transmission through pipeline masking and zero-copy RDMA transmission, effectively hiding I/O latency. Actual measurement data shows that after introducing the L3 GPU cache, the input token throughput increased from 6,576 tokens/s to 15,022 tokens/s, and the request throughput increased from 0.61 req/s to 1.39 req/s, both achieving an over 2-fold performance jump.

3

Frontier Exploration: Agent-Specific Cache Optimization Policies and Mooncake Foundation

If Mooncake built the "highway network" for KVCache management, then the various emerging agent-specific optimization policies at the upper layer are the "intelligent scheduling systems" running on this highway.

In recent years, the academic community and the open source community have generated a batch of cutting-edge KVCache optimization research around agent scenarios. They approach from different angles such as scheduling policies, cache reuse, and workflow awareness, attempting to solve the memory bottleneck in multi-agent collaboration. The presentation emphasized that for these policies to be truly implemented in a production environment, a high-performance, low-latency underlying KV data foundation is indispensable. Mooncake is exactly the infrastructure built for this. Whether it is space-time joint scheduling, cross-agent cache reuse, or global workflow optimization, the cross-node KV transmission, tiered cache persistence, and shared memory pool they require can be supported out-of-the-box by Mooncake.

Tokencake: Space-Time Joint Scheduling

When the agent waits for the tool to return, the GPU is in an idle state, while multiple agents compete for limited KVCache capacity. Tokencake performs optimization from both time and space dimensions. In terms of time, KVCache is offloaded to the CPU through event-driven mechanisms and is transmitted back to the GPU in advance by combining prediction models. In terms of space, fine-grained memory management is achieved through DAG critical path analysis and dynamic partitioning of shared pools and reserved pools. The measured data shows that the latency is reduced by 47%, and the memory utilization is increased by 16.9%. The efficient operation of this scheduling policy highly depends on the rapid transfer capability of underlying KV data—the zero-copy RDMA transmission provided by Mooncake Transfer Engine (60 ms level latency, which is much lower than the 9000 ms recomputation overhead) is exactly the physical foundation for Tokencake to achieve real-time offloading and transmission back.

KVCOMM: Lossless Reuse of KVCache Across Agents

The core assumption of the standard prefix cache is the "same prefix". However, in multi-agent scenarios, the prefix of each agent is different, and the cache becomes completely invalid. KVCOMM (ICML'25) stores the cache through the "anchor pool" (Anchor Pool), and achieves position-independent KVCache reuse by utilizing rotary position embedding (RoPE) de-rotation/re-rotation and offset correction technologies. In a 5-agent scenario, the time to first token (TTFT) is reduced from 430 ms to 55 ms, achieving an acceleration of 7.82 times. The reuse rate reaches 70% to 87.6%, and the output quality is completely lossless. The "anchor pool" here essentially requires a cross-node shared, high-bandwidth, and low-latency KV storage tier. This is exactly the core capability of Mooncake KVCache Store. The tiered cache architecture (VRAM→DRAM→SSD→Remote) combined with the global shared pool turns "compute once, reuse globally" from a concept into reality.

Helium: Workflow-Aware Global Optimization

The core idea of Helium is to optimize the agent workflow as the "query plan" of a database. By building a task relationship tree (TRT) and implementing proactive caching and cache-aware scheduling policies, Helium achieves an end-to-end acceleration of 1.56 times in an end-to-end workflow containing 19 agents and 88 large language model (LLM) operators. The ablation experiment reveals a key conclusion: "seeing the global structure" is more than 6 times more important than "optimizing single-point cache". Removing the workflow pruning results in a performance drop of 23.35%, while removing the KVCache optimization only results in a drop of 3.55%. The global scheduling policy of Helium needs to efficiently distribute and reuse KV data among the nodes of the workflow DAG. The cross-machine transmission capability and topology-aware routing of Mooncake provide a solid underlying channel for this workflow-aware cache policy.

Production Validation: A Revolution in Tail Latency

In the integration with the multi-session inference architecture of OpenClaw, Mooncake demonstrates its value in the production environment. The testing adopts the configuration of the Qwen3-14B model and 4 rounds of interactions for each of the 2 independent sessions, and the result is impressive. The P95 latency of Turn 1 is reduced from 5295 ms to 339 ms, achieving an optimization of 15.6 times. The P95 latency of Turn 2+ is reduced from 4909 ms to 770 ms, achieving an optimization of 6.4 times.

4

A core insight was particularly emphasized in the speech: Mooncake does not change the speed of the fastest requests, but rather changes the allowed slowness of the slowest requests. The system transforms the unstable performance of "usually fast but occasionally very slow" into a "consistently smooth" predictable performance. For enterprise-level services, this is often more valuable than simply improving the median latency.

From the Protocol Layer to the Physical Layer: an End-To-End Agent Memory Solution

Finally, a complete technical panorama is outlined: the upper layer is the multi-agent collaboration protocol (A2A/MCP), the middle layer is the memory-aware scheduling layer (llm-d prefix indexes, Helium workflow DAG, Tokencake space-time scheduling), and the bottom layer is Mooncake as the physical memory infrastructure. The Transfer Engine implements cross-node zero-copy KV transmission, the KVCache Store provides a tiered cache, and the global shared pool ensures "one-time compute, global reuse". Further down is the support of hardware acceleration technologies such as CXL memory pooling, RDMA, and NVMe-of.

Agent memory should not stay at the vector database layer. It needs to sink to the physical inference infrastructure to truly release the potential of the multi-agent system. This is the direction that the Mooncake community is promoting, and it is also the inevitable path for agentic inference to move towards production.

Currently, Mooncake has completed integration with mainstream inference engines and middleware such as SGLang, vLLM, LMDeploy, LMCache, and Dynamo, covering the three major scenarios of inference, middleware, and reinforcement learning (RL) training. As an important open source project of the OpenAnolis community in the LLM inference infrastructure realm, Mooncake is continuously pushing the performance border of agentic inference. More developers are welcome to join the co-construction.

5

The Mooncake project GitHub has obtained 4000+ Stars. You can star and contribute:
https://github.com/kvcache-ai/Mooncake

0 0 0
Share on

OpenAnolis

121 posts | 6 followers

You may also like

Comments