Choose a DR strategy by RTO, RPO, and budget, then deploy it with ACK One across zones or regions.
Key concepts
Recovery Time Objective (RTO) is the maximum acceptable downtime after a service interruption.
Recovery Point Objective (RPO) is the maximum acceptable data loss, measured as time since the last recovery point.
Smaller RTO and RPO require more resources and operational complexity. Set targets based on business criticality and budget.
Choose a DR strategy
Select a strategy based on business criticality, data loss tolerance, and budget.
| Strategy | How it works | Cost | Best for |
|---|---|---|---|
| Backup-Restore | Applications and data are backed up on a schedule. On failure, restore from backup in another location. | Low | Non-critical workloads, last-resort protection |
| Active-Standby | Primary location handles most traffic. Secondary location runs fewer instances. Test traffic is periodically sent to check system effectiveness. On failure, perform database primary-standby switchover, scale up the secondary, and switch traffic. | Medium | Workloads with moderate availability requirements |
| Active-Active | Both locations run equal instances and handle traffic simultaneously. On failure, route all traffic to the healthy location. | High | Business-critical workloads requiring near-zero downtime |
DR scope
Across availability zones (multi-AZ)
A region contains multiple availability zones (AZs) with separate power and network infrastructure. Cross-AZ DR handles localized failures and suits stateful applications such as databases, caches, and message queues due to low inter-AZ latency.
See Regions and zones.
Across regions (multi-region)
Cross-region DR handles disasters that affect all AZs in a region, but higher latency makes implementation more complex and costly.
Design principle
Before designing a DR solution, verify that your stateful applications (databases, caches, message processors) and their dependent cloud services support the intended DR scope.
Backup-Restore solutions
Backup-Restore has the lowest cost but the highest RTO and RPO. The actual duration depends on the data volume and application complexity. Combine full and incremental backups (supported by the ACK One backup center) to reduce both.
Backup-Restore also serves as the last line of protection. Maintain regular schedules and verify backup integrity.
Backup-Restore also supports application migration across clusters:
-
Migrate workloads from aging clusters to newer versions instead of upgrading in place.
-
Reorganize account permissions or restructure organizations. See cross-platform cluster management with ACK One Fleet and cross-region application migration.
Solution 1: Cross-AZ and cross-region backup and recovery on Alibaba Cloud
The ACK One backup center backs up both stateless and stateful applications in ACK clusters. For stateful applications, storage data is backed up alongside the YAML configuration.
The backup center integrates these Alibaba Cloud services for automated backup of YAML data, PVs backed by cloud disks, and PVs backed by file systems:
Backup data can be restored to ACK clusters in any region and AZ. Alibaba Cloud databases such as ApsaraDB RDS for MySQL also support backup and restoration and data migration between instances.
Solution 2: Data backup and restoration for hybrid clouds
Connect on-premises or third-party Kubernetes clusters to ACK using registered clusters in ACK One. Then use the ACK One backup center to back up both stateless and stateful workloads, including storage data alongside the YAML configuration.
Application data (Deployments, StatefulSets) and storage data (PVs, PVCs) can be restored to ACK clusters in any region and AZ.
Cross-cluster access during migration with multi-cluster Services
During batch migration, applications across clusters may need to communicate. Use ACK One multi-cluster Services (MCS) to enable cross-cluster access.
MCS injects a Kubernetes Service (including its endpoints) from one cluster into another. For example, injecting Application 2's Service from Cluster 1 into Cluster 2 lets Application 1 reach it across clusters.
To register on-premises or third-party clusters, connect via a leased line, use the ACK One registered cluster feature, and enable MCS.
Active-Active DR solutions for a single region (multi-AZ)
Multi-AZ Active-Active DR outperforms Active-Standby in three ways:
-
Lower cost with higher utilization: Resources in both AZs serve live traffic, reducing idle capacity.
-
Higher service quality and stronger fault tolerance: More replicas improve response speed and peak traffic handling. Failover avoids service interruptions, and updates can proceed without downtime.
-
Cross-zone scaling: If one zone runs short on resources, scale application replicas in other zones with available capacity.
Solution: Multi-cluster gateway based on ACK One
This solution deploys applications across two ACK clusters in separate AZs and uses the ACK One multi-cluster gateway for Layer 7 traffic routing and health-based failover.
Normal operation:
After AZ1 failure — automatic failover to Cluster 2 in AZ2:
How it works:
-
Deploy applications to two ACK clusters using GitOps, with Git repositories as the single source of truth.
-
Define Kubernetes Ingress rules using the ACK One multi-cluster gateway, which is itself deployed across AZs for HA.
-
When Cluster 1 or its applications become unavailable, the multi-cluster gateway automatically reroutes traffic to Cluster 2 without manual intervention.
-
As traffic grows in Cluster 2, the Horizontal Pod Autoscaler (HPA) scales application replicas, triggering the autoscaler to add cluster nodes.
-
For cross-AZ DR of ApsaraDB RDS, see Build a high availability architecture.
This solution uses Layer 7 HTTP forwarding with health checks, reducing traffic loss during switchover compared to DNS-based distribution. It also consolidates the L4 load balancer and L7 Ingress gateway into a single multi-cluster gateway, reducing system complexity and maintenance costs.
Advantages over DNS-based traffic distribution:
| Dimension | ACK One multi-cluster gateway | DNS-based distribution |
|---|---|---|
| Failover speed | Milliseconds to seconds | Minutes (blocked by DNS TTL caching) |
| Routing capabilities | Advanced Layer 7 routing, session persistence, QUIC 0-RTT | Limited; no cross-cluster session persistence |
| Management | Single control plane (Fleet) manages all Ingress configurations | Separate DNS records per cluster |
| Cluster migration | Transparent — traffic shifts to the healthy cluster and back automatically | IP address changes disrupt clients until TTL expires |
DNS TTL workarounds (such as reducing TTL values) increase request volume and costs without eliminating caching delays.
Architecture of the DNS-based alternative (for reference):
Cloud + data center DR for a single region
This solution extends multi-AZ DR to a hybrid cloud setup, combining an ACK cluster with an on-premises Kubernetes cluster.
Setup:
-
Establish a leased line between the VPC and the data center for management and data channels.
-
Connect the on-premises cluster to ACK One using the registered cluster feature to manage both clusters from a single control plane with Alibaba Cloud observability and security.
-
Deploy applications to both clusters using ACK One GitOps.
Architecture (single-region on-premises and cloud Active-Active):
Multi-region DR solutions
Deploy business systems independently in multiple regions when your user base is geographically distributed or a single-region outage is unacceptable.
Solution 1: Multi-cluster gateway based on ACK One (recommended)
Use this solution when:
-
You need cross-region HA and the primary region has constrained resources (for example, GPU scarcity due to AI workload demand).
-
Your applications have moderate latency sensitivity but require advanced multi-cluster traffic management.
How it works:
-
An Application Load Balancer (ALB) multi-cluster gateway in Region 1 handles Layer 7 cross-region traffic routing, including QUIC 0-RTT and header-based routing. Region 2 runs a single-cluster ALB in cold standby.
-
Global Traffic Manager (GTM) provides DNS resolution and load distribution, monitors ALB health in both regions, and triggers DR automatically.
-
Failure handling has two paths:
-
If Region 1 or its ALB instance fails, GTM switches DNS resolution to the Region 2 ALB.
-
If only a cluster or service in Region 1 fails, or if Region 2 goes down, the ALB multi-cluster gateway reroutes traffic to the healthy cluster without a GTM DNS switch.
-
-
Connect clusters across regions using Cloud Enterprise Network (CEN) or VPC peering. Cross-region traffic travels over leased lines for reliability.
-
Use Global Distributed Cache for Tair for multi-region cache HA.
-
For database HA, see Architecture for multi-zone deployment.
Key advantages:
-
Advanced traffic management: Content-based routing and flexible health checks beyond what traditional GTM provides.
-
Centralized management: One Fleet control plane manages Ingress configurations and services across all clusters.
-
Faster failover: Seamless failover in seconds for cluster or service failures, avoiding DNS propagation delays.
Solution 2: DNS-based traffic distribution with a single ACK One Fleet instance
Use this solution to route users to the nearest region when DNS-level failover latency is acceptable.
How it works:
-
Protect against volumetric and web application attacks — including SQL injection, cross-site scripting (XSS), and command injection — using Anti-DDoS Proxy and Web Application Firewall (WAF). See GTM works with WAF, GA, and SLB and Protect a website service by using Anti-DDoS Pro or Anti-DDoS Premium and WAF.
-
Route user requests to the nearest region using GTM.
-
Deploy applications to both ACK clusters using GitOps.
-
Use Global Distributed Cache for Tair for multi-region cache HA.
-
For database HA, see Architecture for multi-zone deployment.
-
Apply multi-AZ DR within each region.
Solution 3: DNS-based traffic distribution with multiple ACK One Fleet instances
This solution follows the same structure as Solution 2 but uses multiple ACK One Fleet instances instead of one. Use this when organizational boundaries or compliance requirements mandate separate management planes per region.
The traffic protection, routing, application deployment, cache HA, database HA, and per-region multi-AZ DR configuration are identical to Solution 2.
Cross-region unit-based Active-Active solution
This pattern shards application traffic and data across geographic units. Each unit has complete service capability for its data shard, achieving failure isolation, security isolation, and horizontal scaling.
Architecture:
How it works:
-
Business is split into subunits (holding sharded data) and a central unit (holding user data).
-
Traffic sharding rules determine which unit handles each request based on user identity or geography.
-
Units interact for cross-shard operations.
This architecture requires custom traffic distribution, data splitting, and cross-unit interaction logic. It is the most complex DR pattern and is typically reserved for large-scale platforms.
FAQ
Why use the ACK One multi-cluster gateway instead of DNS-based distribution?
DNS failover depends on TTL expiry, which takes minutes. The ACK One multi-cluster gateway uses Layer 7 health checks to reroute traffic in milliseconds to seconds and supports features such as QUIC 0-RTT and session persistence that DNS cannot provide.
What happens during failover if my Cluster 2 is not pre-scaled?
If Cluster 2 lacks capacity at failover, HPA scales replicas and the autoscaler provisions nodes, but this takes time. Pre-scale Cluster 2 or set appropriate HPA minimum replica counts to avoid delays.
How do I handle stateful applications (databases, caches) in a cross-AZ setup?
Verify that your database and cache services support cross-AZ replication. For ApsaraDB RDS, see Build a high availability architecture. For cache HA, use Global Distributed Cache for Tair.
Can I use these DR solutions for on-premises clusters?
Yes. Use ACK One registered clusters to connect on-premises Kubernetes clusters to ACK One. Once registered, apply the same backup, GitOps, and multi-cluster gateway configurations.
Which solution should I choose for my use case?
| Scenario | Recommended solution |
|---|---|
| Non-critical workloads, cost is the primary constraint | Backup-Restore |
| Single-region HA with moderate availability requirements | Active-Standby |
| Single-region HA with near-zero downtime requirement | Multi-AZ Active-Active (ACK One multi-cluster gateway) |
| Hybrid cloud (on-premises + cloud) | Cloud + data center DR |
| Multi-region HA with advanced traffic management | Multi-region Solution 1 (ACK One multi-cluster gateway) |
| Multi-region HA with geo-based routing, DNS latency acceptable | Multi-region Solution 2 or 3 (DNS-based) |
| Massive scale, large user groups | Cross-region unit-based Active-Active |