The storage-operator add-on automates cross-zone disk migration and multi-zone spreading for StatefulSets. If an error occurs during migration, the add-on restores the application in the original zone through precheck and rollback to ensure service availability.
Use cases
| Scenario | Description |
|---|---|
| Zone planning changes | Move workloads to a different zone due to infrastructure or capacity updates. |
| Multi-zone spreading | Distribute replicas and their disks across multiple zones to improve availability. |
| Resource constraints | Insufficient capacity in the current zone for continued operation or scale-out. |
NAS and OSS support cross-zone and multi-mount. Disks are zone-bound — they cannot float across zones or reuse existing persistent volume claims (PVCs) and persistent volumes (PVs). Create new disks in the target zone from snapshots.
Key constraints
Review these constraints before migration:
-
Business interruption required: To ensure data consistency, migration scales the StatefulSet to 0 replicas, then restores all at once after disk migration — not a rolling update. Plan for downtime. Duration depends on replica count, container startup time, and disk capacity.
-
ESSD disks required: All storage used by the StatefulSet must be ESSD disks. Migration uses instant access snapshots, which only support ESSD disks.
-
Target zone requirements: The target zone must support ESSD disks, and the cluster must have nodes in that zone available for scheduling.
If your application uses non-ESSD disks, do one of the following before migrating:
-
Manually create a snapshot for a single disk volume and rebuild the disk across zones.
How it works
Cross-zone migration creates snapshots of source disks and uses instant access to minimize creation time. See Snapshot billing.
storage-operator runs these steps:
-
Precheck: Verifies the application is running and identifies disks to migrate. Stops if precheck fails.
-
Scale to zero: Scales the StatefulSet to 0 replicas, pausing the application.
-
Create snapshots: Creates instant access snapshots for all mounted disks. Snapshots are zone-agnostic.
-
Provision new disks: After confirming snapshots are available, creates new disks in the target zone with the same data.
-
Rebuild PVCs and PVs: Rebuilds PVCs with the same names and their corresponding PVs, bound to the new disks.
-
Restore replicas: Restores the original replica count. Replicas bind to rebuilt PVCs and mount the new disks.
-
(Optional) Delete original resources: After confirming application health, delete the original PVs and disks. See Block storage billing.
Each step after precheck has a rollback strategy. Confirm the StatefulSet runs correctly after migration before deleting original disks — this ensures the application can remount original disks if rollback is needed.
Prerequisites
Ensure the following:
-
A cluster running Kubernetes 1.20 or later with the Container Storage Interface (CSI) driver installed
-
storage-operator v1.26.2-1de13b6-aliyun or later installed
-
csi-plugin and csi-provisioner installed, with csi-provisioner using the non-managed version
If the managed version is installed, switch to non-managed. Then restart the storage controller:
kubectl delete pod -n kube-system <storage-controller-pod-name> -
(ACK dedicated clusters only) Worker and master RAM roles have
ModifyDiskSpecpermission on the ECS API. See Create a custom policy. View the required RAM policy:ACK managed clusters do not require
ModifyDiskSpecpermission.{ "Version": "1", "Statement": [ { "Effect": "Allow", "Action": [ "ecs:CreateSnapshot", "ecs:DescribeSnapshot", "ecs:DeleteSnapshot", "ecs:ModifyDiskSpec", "ecs:DescribeTaskAttribute" ], "Resource": "*" } ] }</details>
Migrate a StatefulSet across zones
Step 2: Create a migration task
Create a ContainerStorageOperator resource:
cat <<EOF | kubectl apply -f -
apiVersion: storage.alibabacloud.com/v1beta1
kind: ContainerStorageOperator
metadata:
name: default
spec:
operationType: APPMIGRATE
operationParams:
stsName: web
stsNamespace: default
stsType: kube
targetZone: cn-beijing-h,cn-beijing-j
checkWaitingMinutes: "1"
healthDurationMinutes: "1"
snapshotRetentionDays: "2"
retainSourcePV: "true"
EOF
Parameters:
| Parameter | Required | Default | Description |
|---|---|---|---|
operationType |
Required | — | Set to APPMIGRATE for stateful application migration. |
stsName |
Required | — | Name of the StatefulSet to migrate. Only one StatefulSet per task. Multiple tasks run sequentially in deployment order. |
stsNamespace |
Required | — | Namespace of the StatefulSet. |
targetZone |
Required | — | Comma-separated target zones, e.g., cn-beijing-h,cn-beijing-j. Disks already in a listed zone are skipped. Multiple zones distribute remaining disks in list order. |
stsType |
Optional | kube |
StatefulSet type. Valid values: kube (native) and kruise (OpenKruise Advanced StatefulSet). |
checkWaitingMinutes |
Optional | "1" |
Polling interval (minutes) for replica availability checks after migration. Increase for large StatefulSets or slow startups to avoid premature rollback. |
healthDurationMinutes |
Optional | "0" |
Wait time (minutes) after replicas reach the expected count before a secondary health check. Set to "0" to skip. |
snapshotRetentionDays |
Optional | "1" |
Retention period for instant access snapshots. Valid values: "1" (one day) and "-1" (permanent). |
retainSourcePV |
Optional | "false" |
Whether to retain the original disk and PV after migration. "false" deletes both. "true" retains them — disk stays in the ECS console, PV enters Released state. |
Examples
The following examples use an ACK Pro cluster with nodes in three zones:
-
Zone B: cn-shanghai.192.168.5.245
-
Zone G: cn-shanghai.192.168.2.214
-
Zone M: cn-shanghai.192.168.3.236, cn-shanghai.192.168.3.237
Step 1: Create a StatefulSet with ESSD disks
Create a test StatefulSet with ESSD disks. Skip this step if you already have a StatefulSet to migrate.
-
Deploy the StatefulSet. View the YAML for the Nginx StatefulSet:
cat << EOF | kubectl apply -f - apiVersion: apps/v1 kind: StatefulSet metadata: name: web spec: selector: matchLabels: app: nginx serviceName: "nginx" replicas: 2 template: metadata: labels: app: nginx spec: containers: - name: nginx image: anolis-registry.cn-zhangjiakou.cr.aliyuncs.com/openanolis/nginx:1.14.1-8.6 ports: - containerPort: 80 name: web volumeMounts: - name: www mountPath: /usr/share/nginx/html volumeClaimTemplates: - metadata: name: www labels: app: nginx spec: accessModes: [ "ReadWriteOnce" ] storageClassName: "alicloud-disk-essd" resources: requests: storage: 20Gi EOF</details>
-
Verify that both pods are running:
kubectl get pod -o wide -l app=nginxThe output shows both pods scheduled to zone M (actual placement depends on the scheduler):
NAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE READINESS GATES web-0 1/1 Running 0 2m 192.168.3.243 cn-shanghai.192.168.3.237 <none> <none> web-1 1/1 Running 0 2m 192.168.3.246 cn-shanghai.192.168.3.236 <none> <none>
Step 2: Create a migration task
Example 1: Cross-zone migration
Migrate all pods to a single target zone (zone B in this example).
Confirm the target zone has sufficient node resources and supports ESSD disks.
-
Create the migration task:
cat <<EOF | kubectl apply -f - apiVersion: storage.alibabacloud.com/v1beta1 kind: ContainerStorageOperator metadata: name: migrate-to-b spec: operationType: APPMIGRATE operationParams: stsName: web stsNamespace: default stsType: kube targetZone: cn-shanghai-b # Target zone for migration. healthDurationMinutes: "1" # Wait 1 minute after migration to confirm the application is running properly. snapshotRetentionDays: "-1" # Retain snapshots permanently until manually deleted. retainSourcePV: "true" # Retain the original disks and PVs. EOF -
Check the migration status:
If the status is
FAILED, see FAQ for troubleshooting.kubectl describe cso migrate-to-b | grep StatusA
SUCCESSstatus confirms the migration completed:Status: Status: SUCCESS -
Verify pod placement after migration:
kubectl get pod -o wide -l app=nginxBoth pods are now on the
cn-shanghai.192.168.5.245node in zone B:NAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE READINESS GATES web-0 1/1 Running 0 2m36s 192.168.5.250 cn-shanghai.192.168.5.245 <none> <none> web-1 1/1 Running 0 2m14s 192.168.5.2 cn-shanghai.192.168.5.245 <none> <none> -
Confirm the results in the ECS console:
-
Snapshots page: 2 new snapshots created with permanent retention.
-
Block Storage page: 2 new disks in zone B; the 2 original disks in zone M are retained (because
retainSourcePVis"true").
-
Example 2: Multi-zone spreading
Distribute pods across two zones (zones B and G) to improve availability.
-
Create the migration task:
cat <<EOF | kubectl apply -f - apiVersion: storage.alibabacloud.com/v1beta1 kind: ContainerStorageOperator metadata: name: migrate spec: operationType: APPMIGRATE operationParams: stsName: web stsNamespace: default stsType: kube targetZone: cn-shanghai-b,cn-shanghai-g # Target zones. Multiple zones trigger automatic spreading. healthDurationMinutes: "1" # Wait 1 minute after migration to confirm the application is running properly. snapshotRetentionDays: "-1" # Retain snapshots permanently until manually deleted. retainSourcePV: "true" # Retain the original disks and PVs. EOF -
Check the migration status:
If the status is
FAILED, see FAQ for troubleshooting.kubectl describe cso migrate | grep StatusA
SUCCESSstatus confirms the migration completed:Status: Status: SUCCESS -
Verify pod placement after migration:
kubectl get pod -o wide -l app=nginxThe pods are spread across zone B (
cn-shanghai.192.168.5.245) and zone G (cn-shanghai.192.168.2.214):NAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE READINESS GATES web-0 1/1 Running 0 4m59s 192.168.2.215 cn-shanghai.192.168.2.214 <none> <none> web-1 1/1 Running 0 4m38s 192.168.5.250 cn-shanghai.192.168.5.245 <none> <none> -
Confirm the results in the ECS console:
-
Snapshots page: 2 new snapshots created with permanent retention.
-
Block Storage page: 2 new disks across zones B and G; the 2 original disks in zone M are retained.
-
FAQ
If a migration task returns FAILED, get the error message:
kubectl describe cso <ContainerStorageOperator-name> | grep Message -A 1
Example output:
Message:
Consume: failed to get target pvc, err: no pvc mounted in statefulset or no pvc need to migrated web
The component could not find a PVC to migrate. Common causes:
-
The StatefulSet has no mounted storage.
-
All disks are already in the target zone — no migration is needed.
-
The component could not retrieve PVC information.
Resolve the issue based on the error message, then reapply the migration task.