Configure Fluid data caching to balance performance, stability, and read/write consistency.
Each section includes recommended configurations and YAML examples for common workload patterns.
When data caching helps
Caching benefits workloads that repeatedly read the same data. Determine whether your workload fits this pattern before configuring a cache system.
| Workload pattern | Caching benefit | Example |
|---|---|---|
| Repeated reads of the same data | High | AI training iterating over a dataset across multiple epochs |
| Concurrent reads of shared files | High | Inference service loading a shared model into GPU memory at startup |
| Shared data across tasks | High | SparkSQL jobs processing order data referenced by multiple analytics queries |
| One-time reads | None | ETL data cleansing or one-time data migration |
| Write-only workloads | None | No read operations on the same data path |
| Very low-frequency access | None | Cache expires or is evicted before reuse |
If caching applies to your workload, configure the policies below.
Performance optimization
Configure infrastructure (ECS instance types and cache media), tune cache parameters, and align Pod scheduling with cache placement.
Select ECS instance types for the cache system
A distributed cache aggregates storage and bandwidth across nodes. Estimate the upper limits for your cache cluster:
-
Available cache capacity = Cache capacity per Worker Pod x Number of Worker Pod replicas
-
Available cache bandwidth = Number of Worker Pod replicas x min{Maximum available bandwidth of the ECS node where the Worker Pod runs, I/O throughput of the cache medium used by the Worker Pod}
-
Theoretical maximum application bandwidth = min{Available bandwidth of the ECS node where the application Pod runs, Available cache bandwidth}
Actual bandwidth depends on the application node's available bandwidth and access pattern (sequential vs. random). When multiple application Pods access data concurrently, cache bandwidth is shared.
Estimation example
Add two ecs.g7nex.8xlarge instances to your ACK cluster as cache Worker Pods, each with 100 GiB memory on a separate node. The application Pod runs on one ecs.gn7i-c8g1.2xlarge instance (8 vCPUs, 30 GiB memory, 16 Gbps bandwidth).
| Metric | Calculation | Result |
|---|---|---|
| Available cache capacity | 100 GiB x 2 | 200 GiB |
| Available cache bandwidth | 2 x min{40 Gbps, Memory access I/O throughput} | 80 Gbps |
| Max application bandwidth (on cache hit) | min{80 Gbps, 16 Gbps} | 16 Gbps |
The bottleneck is the application node's 16 Gbps bandwidth, not the cache cluster. To increase throughput, use a higher-bandwidth instance type or distribute the workload across multiple application Pods.
Recommended ECS instance types
Choose high-bandwidth instance types for cache nodes. Use memory for maximum throughput, or local SSDs for larger capacity at lower cost.
| ECS instance family | ECS instance type | Configuration |
|---|---|---|
| g7nex, network-enhanced general-purpose | ecs.g7nex.8xlarge | 32 vCPUs, 128 GiB memory, 40 Gbps bandwidth |
| ecs.g7nex.16xlarge | 64 vCPUs, 256 GiB memory, 80 Gbps bandwidth | |
| ecs.g7nex.32xlarge | 128 vCPUs, 512 GiB memory, 160 Gbps bandwidth | |
| i4g, local SSD storage | ecs.i4g.16xlarge | 64 vCPUs, 256 GiB memory, 2 x 1920 GB local SSD, 32 Gbps bandwidth |
| ecs.i4g.32xlarge | 128 vCPUs, 512 GiB memory, 4 x 1920 GB local SSD, 64 Gbps bandwidth | |
| g7ne, network-enhanced general-purpose | ecs.g7ne.8xlarge | 32 vCPUs, 128 GiB memory, 25 Gbps bandwidth |
| ecs.g7ne.12xlarge | 48 vCPUs, 192 GiB memory, 40 Gbps bandwidth | |
| ecs.g7ne.24xlarge | 96 vCPUs, 384 GiB memory, 80 Gbps bandwidth | |
| g8i, general-purpose | ecs.g8i.24xlarge | 96 vCPUs, 384 GiB memory, 50 Gbps bandwidth |
| ecs.g8i.16xlarge | 64 vCPUs, 256 GiB memory, 32 Gbps bandwidth |
See Instance families for specifications.
Select cache media
The cache medium determines the I/O throughput ceiling for each node. Even if an ECS instance has high network bandwidth, a slow cache medium becomes the bottleneck.
An enterprise SSD (ESSD) often cannot meet the performance needs of data-intensive workloads. For example, a single PL2 disk has a maximum throughput of 750 MB/s. If you use one PL2 disk as the cache medium on an instance with 40 Gbps bandwidth, the effective cache bandwidth is limited to 750 MB/s, wasting the instance's network capacity.
Available cache media:
| Cache medium | mediumtype |
volumeType |
path |
Best for |
|---|---|---|---|---|
| Memory | MEM | emptyDir | /dev/shm | Maximum throughput; dataset fits in memory |
| System disk (local SSD) | SSD | emptyDir | /var/lib/fluid/cache | Larger capacity at lower cost; cache lifecycle tied to Worker Pod |
| Mounted SSD data disk | SSD | hostPath | /mnt/disk1 | Dedicated disk for cache storage |
| Multiple SSD data disks | SSD | hostPath | /mnt/disk1,/mnt/disk2 | Highest local capacity; capacity evenly distributed across disks |
Configure the cache medium and capacity in spec.tieredstore on the Runtime resource:
Memory as cache medium
spec:
tieredstore:
levels:
- mediumtype: MEM
volumeType: emptyDir
path: /dev/shm
quota: 30Gi # Cache capacity per Worker Pod replica.
high: "0.99"
low: "0.95"
Local storage as cache medium
Choose based on your disk configuration:
System disk storage:
spec:
tieredstore:
levels:
- mediumtype: SSD
volumeType: emptyDir # emptyDir ties the cache lifecycle to the Worker Pod, preventing residual data.
path: /var/lib/fluid/cache
quota: 100Gi # Cache capacity per Worker Pod replica.
high: "0.99"
low: "0.95"
Mounted SSD data disk:
spec:
tieredstore:
levels:
- mediumtype: SSD
volumeType: hostPath
path: /mnt/disk1 # Mount path of the local SSD on the host.
quota: 100Gi # Cache capacity per Worker Pod replica.
high: "0.99"
low: "0.95"
Multiple SSD data disks:
spec:
tieredstore:
levels:
- mediumtype: SSD
volumeType: hostPath
path: /mnt/disk1,/mnt/disk2 # Mount paths of the data disks on the host.
quota: 100Gi # Cache capacity per Worker Pod replica. Capacity is evenly distributed: /mnt/disk1 and /mnt/disk2 each get 50 GiB.
high: "0.99"
low: "0.95"
Thehighandlowwatermarks control cache eviction. When cache usage exceeds thehighthreshold, the cache system evicts data until usage drops to thelowthreshold.
Configure scheduling affinity between the cache and applications
Cross-zone network latency between application Pods and cache Worker Pods degrades data access performance. Deploy both in the same zone.
Specifically:
-
Deploy cache Worker Pods in the same zone whenever possible.
-
Deploy application Pods in the same zone as cache Worker Pods whenever possible.
Concentrating all Pods in a single zone reduces disaster recovery capability. Balance performance against availability based on your service-level agreement (SLA).
Set the cache zone by configuring spec.nodeAffinity on the Dataset resource:
apiVersion: data.fluid.io/v1alpha1
kind: Dataset
metadata:
name: demo-dataset
spec:
...
nodeAffinity:
required:
nodeSelectorTerms:
- matchExpressions:
- key: topology.kubernetes.io/zone
operator: In
values:
- <ZONE_ID> # The zone to deploy cache Worker Pods in, for example, cn-beijing-i.
This places cache Worker Pods on nodes in the specified zone. Fluid can also auto-inject affinity rules so application Pods schedule alongside cache Worker Pods. See Optimize scheduling based on data cache affinity.
Tune prefetch parameters for large file sequential reads
Workloads that perform full sequential reads on large files can benefit from aggressive prefetch settings. Common examples include:
-
AI training on datasets in TFRecord or Tar format
-
Inference service startup loading model parameter files
-
Distributed analytics reading Parquet files
Aggressive prefetch increases read amplification. For random reads or small files, these settings waste bandwidth on unused data. Use defaults for mixed or random-read workloads.
Prefetch configuration varies by runtime.
JindoRuntime
Configure Filesystem in Userspace (FUSE) prefetch in spec.fuse.properties:
kind: JindoRuntime
metadata:
...
spec:
fuse:
properties:
fs.oss.download.thread.concurrency: "200"
fs.oss.read.buffer.size: "8388608"
fs.oss.read.readahead.max.buffer.count: "200"
fs.oss.read.sequence.ambiguity.range: "2147483647" # About 2 GB.
| Parameter | Description |
|---|---|
fs.oss.download.thread.concurrency |
Number of concurrent prefetch threads. |
fs.oss.read.buffer.size |
Size of a single prefetch buffer. |
fs.oss.read.readahead.max.buffer.count |
Maximum number of buffers for a single-stream prefetch. |
fs.oss.read.sequence.ambiguity.range |
Range used to determine whether a read pattern is sequential. |
JuiceFSRuntime
Configure prefetch in spec.fuse.options and spec.worker.options:
kind: JuiceFSRuntime
metadata:
...
spec:
fuse:
options:
buffer-size: "2048"
cache-size: "0"
max-uploads: "150"
worker:
options:
buffer-size: "2048"
max-downloads: "200"
max-uploads: "150"
master:
resources:
requests:
memory: 2Gi
limits:
memory: 8Gi
...
| Parameter | Description |
|---|---|
buffer-size |
Read/write buffer size. |
max-downloads |
Prefetch download concurrency. |
max-uploads |
Upload concurrency. |
fuse cache-size |
Local cache capacity for FUSE. |
Setting cache-size: "0" disables the FUSE local cache, so it uses node memory as Linux Page Cache instead. On a cache miss, it reads from the JuiceFSRuntime Worker distributed cache. See the JuiceFS documentation for detailed tuning.
Stability optimization
The following policies address common stability risks: metadata overload, data loss on Pod restart, FUSE OOM crashes, and mount target failures.
Avoid mounting directories with too many files
The cache system maintains metadata for every file in the mounted directory. Mounting a directory with too many files (such as the root of a large-scale storage system) consumes excessive memory and CPU.
Mount specific subdirectories that match your data scope. Create separate Datasets for different data collections.
apiVersion: data.fluid.io/v1alpha1
kind: Dataset
metadata:
name: demo-dataset
spec:
...
mounts:
- mountPoint: oss://<BUCKET>/<PATH1>/<SUBPATH>/
name: sub-bucket
Multiple cache systems add operational complexity. Choose the right granularity based on your needs.
| Scenario | Approach |
|---|---|
| Small dataset with few files and strong relevance | One Dataset and one cache system |
| Large dataset with many files | Split into multiple Datasets by data directory. Application Pods can mount multiple Datasets. |
| Multiple users or data isolation requirements | Create short-lived Datasets per user or job. Consider a Kubernetes Operator for dynamic management. |
Persist metadata with durable storage
JindoRuntime and similar cache systems use a Master-Worker architecture. The Master Pod stores file metadata and cache status for the mounted backend storage. Without persistent metadata, a restart requires a full state rebuild.
Use a persistent volume claim (PVC) backed by an ESSD to store metadata durably:
apiVersion: data.fluid.io/v1alpha1
kind: JindoRuntime
metadata:
name: sd-dataset
spec:
...
volumes:
- name: meta-vol
persistentVolumeClaim:
claimName: demo-jindo-master-meta
master:
resources:
requests:
memory: 4Gi
limits:
memory: 8Gi
volumeMounts:
- name: meta-vol
mountPath: /root/jindofs-meta
properties:
namespace.meta-dir: "/root/jindofs-meta"
Here, demo-jindo-master-meta is a pre-created PVC backed by an ESSD. Metadata survives Pod restarts and migrates with the Pod. See Use JindoRuntime to persist the state of the Master Pod.
Allocate FUSE Pod resources
The FUSE Pod mounts a FUSE file system on the node and bind-mounts it into the application Pod, exposing a POSIX interface for reading remote data as local files.
To prevent OOM-related mount failures, avoid setting a FUSE Pod memory limit, or set it close to the node's allocatable memory:
spec:
fuse:
resources:
requests:
memory: 8Gi
# limits:
# memory: <ECS_ALLOCATABLE_MEMORY>
If the FUSE Pod is OOM-killed, the application Pod's mount target breaks. Restart the application to re-establish the mount, unless FUSE self-healing is enabled (see below).
Enable FUSE self-healing
By default, a FUSE crash leaves the mount target inaccessible. The application container must restart to re-trigger the bind mount.
Fluid's FUSE self-healing mechanism automatically recovers mount target access after a FUSE restart without requiring an application container restart.
During the FUSE restart window, the mount target is inaccessible. The application must handle I/O errors and include retry logic.
FUSE self-healing has limitations. Enable it only where restarting is disruptive, such as interactive Jupyter Notebook or VS Code environments.
Cache read/write consistency
Caching introduces consistency challenges — strong consistency often degrades performance. Choose a policy that matches your workload's read/write pattern.
The following table summarizes the five scenarios. See each subsection for configuration details.
| Scenario | Backend data changes | Access mode | Configuration approach |
|---|---|---|---|
| Read-only, static data | None | ReadOnlyMany (default) | Default Dataset configuration |
| Read-only, periodic changes | Periodic (scheduled) | ReadOnlyMany | DataLoad with Cron policy |
| Read-only, event-driven changes | On-demand | ReadOnlyMany | Disable metadata cache; set FUSE metadata timeouts |
| Read/write, separate directories | Via separate write path | ReadOnlyMany + ReadWriteMany | Two Datasets with different access modes |
| Read/write, same directory | In-place | ReadWriteMany | Single Dataset; POSIX-compatible backend |
Scenario 1: Read-only data with no backend changes
When to use: A single AI training job reads a fixed dataset across multiple epochs. Cache is cleared after training.
Configuration: Fluid defaults to read-only mode. Use the default Dataset configuration or explicitly set it:
apiVersion: data.fluid.io/v1alpha1
kind: Dataset
metadata:
name: demo-dataset
spec:
...
# accessModes: ["ReadOnlyMany"] ReadOnlyMany is the default value.
Setting accessModes: ["ReadOnlyMany"] prevents accidental writes to the dataset during training.
Scenario 2: Read-only data with periodic backend changes
When to use: Business data is collected daily in backend storage. Nightly analysis jobs process the day's data. Results are written directly to backend storage, bypassing the cache.
Configuration: Use a DataLoad resource with a Cron policy to synchronize changes from backend storage at regular intervals:
apiVersion: data.fluid.io/v1alpha1
kind: Dataset
metadata:
name: demo-dataset
spec:
...
# accessModes: ["ReadOnlyMany"] ReadOnlyMany is the default value.
---
apiVersion: data.fluid.io/v1alpha1
kind: DataLoad
metadata:
name: demo-dataset-warmup
spec:
...
policy: Cron
schedule: "0 0 * * *" # Prefetch data at 00:00 every day.
loadMetadata: true # Sync metadata changes from backend storage during prefetch.
target:
- path: /path/to/warmup # Path in backend storage to prefetch.
loadMetadata: true also refreshes metadata during each prefetch, keeping the cache consistent with backend storage.
Scenario 3: Read-only data with event-driven backend changes
When to use: Users upload custom models to backend storage and select them for inference. The cache must reflect new uploads without waiting for a scheduled sync.
Configuration: Disable the server-side metadata cache and set short FUSE metadata timeouts to force backend storage checks on each access.
Using JindoRuntime:
apiVersion: data.fluid.io/v1alpha1
kind: Dataset
metadata:
name: demo-dataset
spec:
mounts:
- mountPoint: <MOUNTPOINT>
name: data
path: /
options:
metaPolicy: ALWAYS # Disable server-side metadata cache.
---
apiVersion: data.fluid.io/v1alpha1
kind: JindoRuntime
metadata:
name: demo-dataset
spec:
fuse:
args:
- -oauto_cache
# Metadata timeout in seconds. Setting to 0 provides strong consistency but may greatly reduce read efficiency.
- -oattr_timeout=30
- -oentry_timeout=30
- -onegative_timeout=30
- -ometrics_port=0
Setting metaPolicy: ALWAYS disables the server-side metadata cache, forcing each metadata access to query backend storage directly.
Settingattr_timeout,entry_timeout, andnegative_timeoutto0provides strong consistency but can significantly reduce read performance. A value of30(seconds) balances freshness and performance for most event-driven scenarios.
Scenario 4: Read and write in separate directories
When to use: Large-scale distributed training reads data from directory A and writes checkpoints to directory B each epoch. Caching writes improves checkpoint efficiency.
Configuration: Create two Datasets. Set one to read-only for the training data and the other to read-write for checkpoints:
apiVersion: data.fluid.io/v1alpha1
kind: Dataset
metadata:
name: train-samples
spec:
...
# accessModes: ["ReadOnlyMany"] ReadOnlyMany is the default value.
---
apiVersion: data.fluid.io/v1alpha1
kind: Dataset
metadata:
name: model-ckpt
spec:
...
accessModes: ["ReadWriteMany"]
Mount both Datasets to the application Pod:
apiVersion: v1
kind: Pod
metadata:
...
spec:
containers:
...
volumeMounts:
- name: train-samples-vol
mountPath: /data/A
- name: model-ckpt-vol
mountPath: /data/B
volumes:
- name: train-samples-vol
persistentVolumeClaim:
claimName: train-samples
- name: model-ckpt-vol
persistentVolumeClaim:
claimName: model-ckpt
This isolates the training dataset from checkpoints. The read-only train-samples Dataset prevents write-related consistency issues, while the read-write model-ckpt Dataset enables cached writes.
Scenario 5: Read and write in the same directory
When to use: Interactive development, such as an online Jupyter Notebook or VS Code workspace, where files are frequently created, modified, and deleted in a shared directory.
Configuration: Set the Dataset access mode to read-write. Use a POSIX-compatible storage backend:
apiVersion: data.fluid.io/v1alpha1
kind: Dataset
metadata:
name: myworkspace
spec:
...
accessModes: ["ReadWriteMany"]
ReadWriteMany enables concurrent read/write access to the same directory by multiple users or processes.