A solution architect's guide to the storage products announced at Apsara Conference 2026 and how they solve real AI workload challenges.
AI workloads are not generic. Training a trillion-parameter model, serving long-context inference, building a RAG pipeline, and running thousands of agents — each has distinct storage requirements for throughput, latency, concurrency, and data structure. Generic cloud storage forces compromises at every layer.
At Apsara Conference 2026, Alibaba Cloud announced a set of storage products that eliminate those compromises by specializing for each AI workload stage. This guide walks through each product — what's new, the solutions it enables, and the benefits for your architecture.
KVCacheStore is a dedicated storage tier for KV cache offload during LLM inference — positioned between GPU memory and remote shared storage. It is the first public cloud product to address this gap.
KVCacheStore is an entirely new product — the first public cloud G3.5-tier KV cache storage product. It creates a storage tier between GPU memory and remote shared storage that did not exist before. Previously, when KV cache exceeded GPU memory, the only options were token recomputation (expensive in compute) or remote shared storage (too slow for real-time inference).
Longer context windows without GPU pressure. When KV cache exceeds GPU memory, KVCacheStore provides an intermediate tier with latency low enough for real-time inference — no need to truncate context or recompute tokens.
Lower inference cost. Combined with Tair KVCM (KVCache Management), which orchestrates multi-level caching across GPU memory, KVCacheStore, and remote storage, the system achieves a 99% effective cache hit rate and 50% lower per-token cost. Fewer recomputations means less GPU compute spent on redundant work.
Better user experience. The 54% reduction in first-token latency for cache misses directly improves response time for end users — particularly in multi-turn conversations where session context grows over time.
Higher batch throughput. Offloading KV cache to G3.5 frees GPU memory for larger batch sizes, improving inference throughput per GPU.
Specify Tair KVCM as the cache orchestration layer with KVCacheStore as the G3.5 backend. For inference workloads with long context windows (>128K tokens), multi-turn agent sessions, or high-concurrency real-time serving, KVCacheStore eliminates the binary choice between "fit in GPU memory" and "accept recomputation cost."
OSS Tables combines object storage, vector storage, and table storage in a single bucket — eliminating the need for separate services to manage different data modalities.
OSS previously supported only object storage. OSS Tables introduces native vector indexing and table storage capabilities directly within the OSS bucket — three data modalities unified into a single namespace.
Simplified RAG architecture. The standard RAG pipeline requires object storage for documents, a separate vector database for embeddings, and a table store for metadata — connected by ETL pipelines. OSS Tables collapses this into one bucket. No vector database to provision, scale, or manage. No ETL to maintain. No data duplication.
Guaranteed data consistency. Because objects, vectors, and tables live in the same bucket, there is no sync lag between raw data and its embeddings. When a document is updated, the vector representation is updated in the same namespace.
Lower TCO. Remove the vector database line item from your cost model. Remove the ETL compute cost. Remove the duplicated storage cost. For RAG workloads where vectors are derived from documents stored in OSS, the savings are immediate.
Faster queries. Vector similarity search and metadata filtering happen within the same storage layer — no cross-service round trips.
For RAG and multimodal AI applications, use OSS Tables as the single data foundation. Store raw documents as objects, generate and store embeddings as vectors, and maintain structured metadata as tables — all within one bucket. For advanced vector workloads requiring complex hybrid search or fine-grained index control, a dedicated vector database may still be appropriate — evaluate case by case.
CPFS (Cloud Parallel File System) delivers a purpose-built parallel file system for large-scale AI training clusters — from hundreds to millions of GPUs.
The new-generation CPFS announced at Apsara 2026 delivers major upgrades across throughput, scale, and cost efficiency:
Lower training cost. CPFS reduces AI storage costs by 69%. For large training clusters, storage is a significant portion of TCO — this reduction shifts the economics of whether to use managed parallel storage versus self-managed alternatives.
Faster model iteration. Model startup time — the time to load checkpoints and begin a training run — drops by 50%. For teams running frequent checkpoint-restart cycles (fault recovery, hyperparameter tuning, elastic scaling), this directly increases productive GPU hours.
Higher GPU utilization. Peak compute utilization increases by 30%, meaning GPUs spend less time waiting on storage I/O and more time computing. This is the metric that matters — every percentage point of GPU utilization recovered is real money saved.
Simplified architecture. The 100 PiB single-filesystem scale (5x increase) means most training clusters operate within a single namespace — no data sharding, no multi-filesystem placement logic, no complexity.
For training clusters of any size, CPFS serves as the single hot data layer — checkpoint storage, training data staging, and intermediate results all in one filesystem. The multi-tier staging pattern (separate storage for preprocessing, training, and checkpoints) is no longer necessary.
AgenticFS is a file storage service under the NAS product family, designed specifically for Agent Sandbox environments running on Container Compute Service (ACS) and Function Compute (FC).
AgenticFS is a brand-new product — the first file storage built specifically for agent sandbox workloads. Agent workloads previously used standard NAS volumes designed for human-driven access patterns.
Native agent file storage. Agent workloads need file access — but traditional NAS is designed for human-driven access patterns with long-lived sessions. AgenticFS is built for the agent pattern: short-lived compute instances, machine-driven access, high-fan-out concurrent reads.
Multi-tenant security. AccessPoints with RAM policies ensure each agent only sees files it is authorized to access. Multiple agent workloads can share a single AgenticFS instance without cross-tenant data exposure — critical for multi-customer or multi-team deployments.
Operational simplicity. Volumes are created, mounted, and managed through the same control plane used to build and run agents. No separate storage configuration, no manual volume provisioning, no storage team involvement needed.
For agent sandbox architectures, use AgenticFS for shared datasets, model files, and persistent agent state. Use EBS for instance-local high-IOPS scratch space. AccessPoints provide the isolation layer for multi-tenant agent deployments.
EBS (Elastic Block Storage) received capability upgrades targeting two high-value patterns in AI infrastructure: data preprocessing pipelines and agent high-concurrency scenarios.
EBS gains improved throughput and IOPS characteristics specifically tuned for I/O-intensive data preprocessing workloads. It also adds support for massive concurrent volume mounts, snapshots, and read/write operations — a pattern introduced by agent sandbox workloads.
Faster data preparation. Data preprocessing — tokenization, augmentation, cleaning — is often the longest phase of AI projects. Enhanced EBS throughput directly reduces this cycle time, getting models into training faster.
Agent deployment at scale. Agent workloads involve hundreds or thousands of short-lived sandbox instances, each mounting volumes, reading data, and exiting. EBS concurrency improvements mean agent deployments scale without hitting storage-level bottlenecks.
Predictable pricing. ESSD PL0-PL3 tiers bundle IOPS and throughput into the storage price — no separate per-IOPS or per-MiBps charges. This makes cost estimation more predictable compared to unbundled pricing models where IOPS and throughput are billed separately.
Use EBS for instance-local scratch and high-IOPS preprocessing workloads. Pair with AgenticFS for shared datasets that agents need to read — EBS handles local high-speed I/O, AgenticFS handles shared file access.
With the full portfolio available, each stage of the AI pipeline has a purpose-built storage solution:
| AI Stage | Storage Solution | What's New | Benefit |
|---|---|---|---|
| Model inference | KVCacheStore | New G3.5 tier — first public cloud KV cache storage product | 99% cache hit rate, 50% lower per-token cost, 54% lower first-token latency |
| Data management & retrieval | OSS Tables | Native vector + table storage in OSS bucket | Unified object + vector + table eliminates separate services, ETL, and data duplication |
| Model training | CPFS | 100 TB/s throughput, 100 PiB scale, -69% cost, -50% startup, +30% utilization | Lower training cost, faster iteration, higher GPU utilization |
| Agent execution | AgenticFS | New product — purpose-built file storage for agent sandboxes | Identity-scoped file access, multi-tenant security, AgentCore integration |
| Data preprocessing | EBS (enhanced) | Agent concurrency support, preprocessing-optimized throughput | High-throughput block I/O reduces data prep cycle time |
Raw Data → EBS (preprocessing) → OSS Tables (objects + vectors + tables)
↓
CPFS (training checkpoint I/O)
↓
KVCacheStore + Tair KVCM (inference cache)
↓
AgenticFS (agent file access)
Each stage flows into the next. Data moves from preprocessing through management, training, inference, and agent execution — with storage purpose-built for the I/O pattern at each step.
The portfolio is no longer a set of generic storage services that you adapt to AI workloads. Each AI stage now has a storage solution designed for its specific I/O pattern, latency requirement, and scale characteristic. For solution architects designing AI infrastructure on Alibaba Cloud, the default starting point has shifted — these purpose-built products should be the foundation of new designs.
Product documentation: KVCacheStore | OSS | CPFS | AgenticFS | EBS
OSS Tables: The Missing Piece in the AI Data Lakehouse Infrastructure
13 posts | 2 followers
FollowAlibaba Cloud Community - September 25, 2026
Alibaba Cloud Community - September 23, 2026
ApsaraDB - May 21, 2026
Alibaba Cloud Community - September 23, 2026
ApsaraDB - July 22, 2026
Alibaba Cloud Community - September 22, 2026
13 posts | 2 followers
Follow
Storage Capacity Unit
Plan and optimize your storage budget with flexible storage services
Learn More
Simple Log Service
An all-in-one service for log-type data
Learn More
Hybrid Cloud Storage
A cost-effective, efficient and easy-to-manage hybrid cloud storage solution.
Learn More
Hybrid Cloud Distributed Storage
Provides scalable, distributed, and high-performance block storage and object storage services in a software-defined manner.
Learn MoreMore Posts by Justin See