×
Community Blog Alibaba Cloud's AI Storage Portfolio: Five Purpose-Built Solutions from Apsara 2026

Alibaba Cloud's AI Storage Portfolio: Five Purpose-Built Solutions from Apsara 2026

Alibaba Cloud announced five purpose-built storage products at Apsara 2026 — KVCacheStore, OSS Tables, CPFS, AgenticFS, and enhanced EBS — each target.

Alibaba Cloud Storage Solutions for AI Workloads — Apsara 2026 Portfolio Update

A solution architect's guide to the storage products announced at Apsara Conference 2026 and how they solve real AI workload challenges.


Introduction

AI workloads are not generic. Training a trillion-parameter model, serving long-context inference, building a RAG pipeline, and running thousands of agents — each has distinct storage requirements for throughput, latency, concurrency, and data structure. Generic cloud storage forces compromises at every layer.

At Apsara Conference 2026, Alibaba Cloud announced a set of storage products that eliminate those compromises by specializing for each AI workload stage. This guide walks through each product — what's new, the solutions it enables, and the benefits for your architecture.


KVCacheStore: Dedicated KV Cache Storage for LLM Inference

Solution

KVCacheStore is a dedicated storage tier for KV cache offload during LLM inference — positioned between GPU memory and remote shared storage. It is the first public cloud product to address this gap.
KVCacheStore

What's New

KVCacheStore is an entirely new product — the first public cloud G3.5-tier KV cache storage product. It creates a storage tier between GPU memory and remote shared storage that did not exist before. Previously, when KV cache exceeded GPU memory, the only options were token recomputation (expensive in compute) or remote shared storage (too slow for real-time inference).

  1. New tier: G3.5 — purpose-built for KV cache offload during LLM inference
  2. Cache scale: Hundred-billion-level KV cache entries
  3. Cache coverage window: 900% expansion
  4. First-token latency (cache miss): -54%
  5. Per-token cost (with Tair KVCM orchestration): -50%
  6. Cache hit rate (with Tair KVCM orchestration): 99% effective

Key Capabilities

  1. G3.5 storage tier — between GPU memory (G3) and remote shared storage (G1/G2)
  2. Support for hundred-billion-level KV cache entries
  3. 900% expansion of effective cache coverage window
  4. 54% reduction in first-token latency for cache misses
  5. Designed for long-context, multi-turn, and agent inference workloads

Benefits for Your Architecture

Longer context windows without GPU pressure. When KV cache exceeds GPU memory, KVCacheStore provides an intermediate tier with latency low enough for real-time inference — no need to truncate context or recompute tokens.

Lower inference cost. Combined with Tair KVCM (KVCache Management), which orchestrates multi-level caching across GPU memory, KVCacheStore, and remote storage, the system achieves a 99% effective cache hit rate and 50% lower per-token cost. Fewer recomputations means less GPU compute spent on redundant work.

Better user experience. The 54% reduction in first-token latency for cache misses directly improves response time for end users — particularly in multi-turn conversations where session context grows over time.

Higher batch throughput. Offloading KV cache to G3.5 frees GPU memory for larger batch sizes, improving inference throughput per GPU.

Design Pattern

Specify Tair KVCM as the cache orchestration layer with KVCacheStore as the G3.5 backend. For inference workloads with long context windows (>128K tokens), multi-turn agent sessions, or high-concurrency real-time serving, KVCacheStore eliminates the binary choice between "fit in GPU memory" and "accept recomputation cost."


OSS Tables: Unified Object, Vector, and Table Storage

Solution

OSS Tables combines object storage, vector storage, and table storage in a single bucket — eliminating the need for separate services to manage different data modalities.
OSS

What's New

OSS previously supported only object storage. OSS Tables introduces native vector indexing and table storage capabilities directly within the OSS bucket — three data modalities unified into a single namespace.

  1. New capability: Native vector indexing and similarity search in OSS bucket
  2. New capability: Structured table storage with high-throughput point queries in OSS bucket
  3. Services eliminated: Separate vector database and table store no longer needed
  4. ETL pipelines eliminated: No data sync between object storage and vector/table stores
  5. Data duplication eliminated: Single copy across all three modalities
  6. API: Unified access across object, vector, and table from one bucket

Key Capabilities

  1. Object storage — raw files (documents, images, audio, video)
  2. Vector storage — native vector indexing and similarity search
  3. Table storage — structured data with high-throughput point queries and scans
  4. Unified API access across all three modalities
  5. Single bucket, single namespace, single billing line

Benefits for Your Architecture

Simplified RAG architecture. The standard RAG pipeline requires object storage for documents, a separate vector database for embeddings, and a table store for metadata — connected by ETL pipelines. OSS Tables collapses this into one bucket. No vector database to provision, scale, or manage. No ETL to maintain. No data duplication.

Guaranteed data consistency. Because objects, vectors, and tables live in the same bucket, there is no sync lag between raw data and its embeddings. When a document is updated, the vector representation is updated in the same namespace.

Lower TCO. Remove the vector database line item from your cost model. Remove the ETL compute cost. Remove the duplicated storage cost. For RAG workloads where vectors are derived from documents stored in OSS, the savings are immediate.

Faster queries. Vector similarity search and metadata filtering happen within the same storage layer — no cross-service round trips.

Design Pattern

For RAG and multimodal AI applications, use OSS Tables as the single data foundation. Store raw documents as objects, generate and store embeddings as vectors, and maintain structured metadata as tables — all within one bucket. For advanced vector workloads requiring complex hybrid search or fine-grained index control, a dedicated vector database may still be appropriate — evaluate case by case.


CPFS: High-Performance Parallel Storage for AI Training

Solution

CPFS (Cloud Parallel File System) delivers a purpose-built parallel file system for large-scale AI training clusters — from hundreds to millions of GPUs.
CPFS

What's New

The new-generation CPFS announced at Apsara 2026 delivers major upgrades across throughput, scale, and cost efficiency:

  1. Aggregate throughput: 100 TB/s level per filesystem
  2. Single filesystem capacity: 100 PiB — 5x increase
  3. IOPS: Hundreds of millions
  4. Model startup time: 50% reduction
  5. Peak compute utilization: +30%
  6. AI storage cost: -69%

Key Capabilities

  1. Aggregate throughput at 100 TB/s level per filesystem
  2. Hundreds of millions of IOPS
  3. Single filesystem capacity up to 100 PiB (5x increase)
  4. Built on end-to-end self-developed technology stack
  5. 50% reduction in model startup time
  6. 30% increase in peak compute utilization
  7. 69% reduction in AI storage costs

Benefits for Your Architecture

Lower training cost. CPFS reduces AI storage costs by 69%. For large training clusters, storage is a significant portion of TCO — this reduction shifts the economics of whether to use managed parallel storage versus self-managed alternatives.

Faster model iteration. Model startup time — the time to load checkpoints and begin a training run — drops by 50%. For teams running frequent checkpoint-restart cycles (fault recovery, hyperparameter tuning, elastic scaling), this directly increases productive GPU hours.

Higher GPU utilization. Peak compute utilization increases by 30%, meaning GPUs spend less time waiting on storage I/O and more time computing. This is the metric that matters — every percentage point of GPU utilization recovered is real money saved.

Simplified architecture. The 100 PiB single-filesystem scale (5x increase) means most training clusters operate within a single namespace — no data sharding, no multi-filesystem placement logic, no complexity.

Design Pattern

For training clusters of any size, CPFS serves as the single hot data layer — checkpoint storage, training data staging, and intermediate results all in one filesystem. The multi-tier staging pattern (separate storage for preprocessing, training, and checkpoints) is no longer necessary.


AgenticFS: File Storage for Agent Sandbox Environments

Solution

AgenticFS is a file storage service under the NAS product family, designed specifically for Agent Sandbox environments running on Container Compute Service (ACS) and Function Compute (FC).
AgenticFS

What's New

AgenticFS is a brand-new product — the first file storage built specifically for agent sandbox workloads. Agent workloads previously used standard NAS volumes designed for human-driven access patterns.

  1. New product: Purpose-built file storage for Agent Sandbox on ACS (Container Compute Service) and FC (Function Compute)
  2. New capability: AccessPoints with per-agent RAM policy enforcement for identity-scoped file access
  3. New capability: Volume lifecycle managed through the Agent Sandbox control plane — no separate storage provisioning
  4. New capability: Native AgentCore integration for end-to-end agent lifecycle management
  5. Concurrency: Designed for hundreds to thousands of short-lived agent instances with high-fan-out reads
  6. Compute platforms: ACS, FC

Key Capabilities

  1. Mountable POSIX file volumes for agent sandbox instances
  2. AccessPoints with RAM policy enforcement for identity-scoped file access
  3. Volume lifecycle managed through the Agent Sandbox control plane
  4. Integration with AgentCore for end-to-end agent lifecycle management

Benefits for Your Architecture

Native agent file storage. Agent workloads need file access — but traditional NAS is designed for human-driven access patterns with long-lived sessions. AgenticFS is built for the agent pattern: short-lived compute instances, machine-driven access, high-fan-out concurrent reads.

Multi-tenant security. AccessPoints with RAM policies ensure each agent only sees files it is authorized to access. Multiple agent workloads can share a single AgenticFS instance without cross-tenant data exposure — critical for multi-customer or multi-team deployments.

Operational simplicity. Volumes are created, mounted, and managed through the same control plane used to build and run agents. No separate storage configuration, no manual volume provisioning, no storage team involvement needed.

Design Pattern

For agent sandbox architectures, use AgenticFS for shared datasets, model files, and persistent agent state. Use EBS for instance-local high-IOPS scratch space. AccessPoints provide the isolation layer for multi-tenant agent deployments.


EBS: Enhanced Block Storage for AI Data Pipelines and Agent Workloads

Solution

EBS (Elastic Block Storage) received capability upgrades targeting two high-value patterns in AI infrastructure: data preprocessing pipelines and agent high-concurrency scenarios.
EBS

What's New

EBS gains improved throughput and IOPS characteristics specifically tuned for I/O-intensive data preprocessing workloads. It also adds support for massive concurrent volume mounts, snapshots, and read/write operations — a pattern introduced by agent sandbox workloads.

  1. New capability: Preprocessing-optimized throughput and IOPS for tokenization, augmentation, and data cleaning pipelines
  2. New capability: Support for hundreds to thousands of short-lived agent sandbox instances with high-fan-out ephemeral volume mounts
  3. New capability: Concurrent snapshot and read/write operations from agent sandbox environments
  4. Pricing: ESSD PL0-PL3 tiers with bundled IOPS + throughput — no separate per-IOPS or per-MiBps charges

Key Capabilities

  1. Improved throughput and IOPS for I/O-intensive data preprocessing workloads
  2. Support for massive concurrent volume mounts, snapshots, and read/write operations from agent sandbox environments
  3. ESSD performance tiers (PL0-PL3) with bundled IOPS and throughput in the storage price

Benefits for Your Architecture

Faster data preparation. Data preprocessing — tokenization, augmentation, cleaning — is often the longest phase of AI projects. Enhanced EBS throughput directly reduces this cycle time, getting models into training faster.

Agent deployment at scale. Agent workloads involve hundreds or thousands of short-lived sandbox instances, each mounting volumes, reading data, and exiting. EBS concurrency improvements mean agent deployments scale without hitting storage-level bottlenecks.

Predictable pricing. ESSD PL0-PL3 tiers bundle IOPS and throughput into the storage price — no separate per-IOPS or per-MiBps charges. This makes cost estimation more predictable compared to unbundled pricing models where IOPS and throughput are billed separately.

Design Pattern

Use EBS for instance-local scratch and high-IOPS preprocessing workloads. Pair with AgenticFS for shared datasets that agents need to read — EBS handles local high-speed I/O, AgenticFS handles shared file access.


Complete AI Pipeline Architecture

With the full portfolio available, each stage of the AI pipeline has a purpose-built storage solution:

AI Stage Storage Solution What's New Benefit
Model inference KVCacheStore New G3.5 tier — first public cloud KV cache storage product 99% cache hit rate, 50% lower per-token cost, 54% lower first-token latency
Data management & retrieval OSS Tables Native vector + table storage in OSS bucket Unified object + vector + table eliminates separate services, ETL, and data duplication
Model training CPFS 100 TB/s throughput, 100 PiB scale, -69% cost, -50% startup, +30% utilization Lower training cost, faster iteration, higher GPU utilization
Agent execution AgenticFS New product — purpose-built file storage for agent sandboxes Identity-scoped file access, multi-tenant security, AgentCore integration
Data preprocessing EBS (enhanced) Agent concurrency support, preprocessing-optimized throughput High-throughput block I/O reduces data prep cycle time

Pipeline Data Flow

Raw Data → EBS (preprocessing) → OSS Tables (objects + vectors + tables)
                                        ↓
                                   CPFS (training checkpoint I/O)
                                        ↓
                              KVCacheStore + Tair KVCM (inference cache)
                                        ↓
                                  AgenticFS (agent file access)

Each stage flows into the next. Data moves from preprocessing through management, training, inference, and agent execution — with storage purpose-built for the I/O pattern at each step.


Solution Benefits Summary

Cost Reduction

  1. Inference cost: 50% lower per-token cost with KVCacheStore + Tair KVCM
  2. Data management: Eliminate vector database and ETL costs with OSS Tables
  3. Training storage: 69% reduction with new-gen CPFS
  4. Block storage: Bundled IOPS and throughput pricing with ESSD PL tiers

Performance

  1. Inference latency: 54% lower first-token latency for cache misses
  2. Cache coverage: 900% larger effective cache window
  3. Training throughput: 100 TB/s aggregate, hundreds of millions of IOPS
  4. Training startup: 50% faster checkpoint loading
  5. GPU utilization: 30% higher peak compute utilization

Architecture Simplification

  1. KV cache: Dedicated G3.5 tier eliminates recomputation or overprovisioning
  2. RAG pipelines: Collapse object storage + vector database + table store into one OSS bucket
  3. Training storage: Single CPFS filesystem replaces multi-tier staging designs
  4. Agent file storage: Native solution replaces NAS workarounds

Security

  1. Agent isolation: AccessPoints + RAM policies for multi-tenant agent file access
  2. Identity-scoped access: Each agent sees only authorized files

Conclusion

The portfolio is no longer a set of generic storage services that you adapt to AI workloads. Each AI stage now has a storage solution designed for its specific I/O pattern, latency requirement, and scale characteristic. For solution architects designing AI infrastructure on Alibaba Cloud, the default starting point has shifted — these purpose-built products should be the foundation of new designs.


Product documentation: KVCacheStore | OSS | CPFS | AgenticFS | EBS

0 2 1
Share on

Justin See

13 posts | 2 followers

You may also like

Comments