When you use zone-level cloud services as route next hops, a zone failure traditionally requires manual traffic switchover to a standby instance. A route target group supports active-standby configuration with two next-hop instances. The system continuously monitors instance health and automatically performs zone-level disaster recovery switchover when an instance becomes unhealthy, reducing the recovery time objective (RTO).
How it works
A route target group sits between a VPC route entry and its next-hop instances. When you set the next hop of a route entry to a route target group, all traffic goes to the active instance. The group's health monitoring continuously probes both instances and triggers failover automatically.
Active-standby weight model
The two member instances have fixed weights that determine which one handles traffic:
| Instance | Weight | Role |
|---|---|---|
| Active | 100 | Handles all traffic when healthy |
| Standby | 0 | Takes over only when the active instance fails |
Both instances must be of the same type and deployed in different zones to enable cross-zone disaster recovery.
Health monitoring
The system continuously sends health probes to both instances:
-
Failure detection: Three consecutive failed health probes mark an instance as
Unhealthyand trigger automatic failover. -
Recovery verification: Three consecutive successful health probes mark the instance as
Healthyagain.
Automatic and manual switching
| Scenario | Trigger | Switchover time |
|---|---|---|
| Failover | Automatic (health check failure) | < 30s |
| Failback | Manual | < 10s |
After the active instance recovers, the system does not automatically return traffic to it. This is intentional — automatic failback can cause repeated switching during network jitter or brief service disruptions. Fail back manually during a maintenance window or off-peak hours.
Scope and limits
| Item | Description |
|---|---|
| Supported member instance types | Gateway Load Balancer endpoints (GWLBe) and NAT gateways. All member instances within the same route target group must be of the same type. |
| Instance deployment requirements | One active + one standby instance. Both must belong to the same VPC and be deployed in different zones to enable cross-zone disaster recovery. When the target instance type is NAT Gateway, you must deploy single-zone NAT gateways in different zones within the same VPC. |
| Route target groups per VPC | 10 |
| Supported route entry type | IPv4 only |
| Manual switchover prerequisite | The inactive instance must be Healthy |
A route target group supports the same routing operations as its member instances. For example, if a system route in a gateway route table can use a GWLB endpoint as the next hop, it can also use a route target group that contains GWLB endpoints as members.
Configure automatic failover
Set the next hop of a route entry to a route target group. Traffic is forwarded based on the currently active instance. When an instance becomes unhealthy, the system automatically performs zone-level disaster recovery switchover.
Console
-
Go to the VPC - Route Target Groups page and click Create Target Group.
Field Setting VPC Select the VPC where your target instances reside. Mode Currently only Active/Standby Mode is supported. Member Type Select GWLB Endpoint or NAT Gateway. When the member type is NAT Gateway, only single-zone NAT gateways are supported.
Member Settings Select an active and a standby instance deployed in different zones. After creation, you can modify inactive instances from the target group details page by clicking Modify Member. Active instances cannot be modified. Instance Name, Resource Group and Tags Optional fields for categorizing and managing instances. -
Go to the VPC - Route Tables page and click the target route table ID.
-
Configure a route entry to point to the route target group. You can either create a new route entry or modify an existing one:
-
Create a new route entry: On the Custom Route tab, click Add Route Entry. Enter the Destination CIDR Block for traffic that needs to be forwarded through the active and standby instances. Set Next Hop Type to Route Target Group and select the corresponding route target group.
-
Modify an existing route entry: On the Custom Route tab, find the target route entry and click Edit. Change the Next Hop Type to Route Target Group, then select the corresponding route target group.
-
API
-
Call CreateRouteTargetGroup to create the route target group.
-
Call CreateRouteEntry to create a route entry. Set
NextHopTypetoRouteTargetGroup.
Perform a manual switchover
-
Applicable scenarios:
-
Manually trigger traffic switchover between active and standby instances for disaster recovery drills or planned maintenance.
-
To avoid frequent automatic switching caused by network jitter on the active instance (service disruption), route target groups do not automatically fail back to the active instance after recovery. After the active instance recovers, you can manually fail back during off-peak hours.
-
-
Restriction: Switching is not allowed when the inactive instance is unhealthy.
Console
-
Go to the VPC - Route Target Groups page and click the route target group ID.
-
In the upper-right corner, click Switch Member.
API
Call SwitchActiveRouteTarget to perform the switchover.
Billing
The route target group feature itself is free of charge.
Target instances and their backend services are billed according to the pricing rules of each product:
-
Gateway Load Balancer endpoints (GWLBe): For billing details, see PrivateLink billing and GWLB billing.
-
NAT Gateway: For billing details, see single-zone NAT Gateway billing.
Apply in production
-
Connection interruption risk: For stateful applications (such as long-lived firewall connections or NAT sessions), an active-standby switchover interrupts existing connections, and services need to re-establish connections. Consider implementing reconnection mechanisms and thoroughly evaluate the momentary impact of a switchover on your services.
-
Disaster recovery drill: Before introducing production traffic, perform a manual switchover during off-peak hours or a maintenance window to conduct a disaster recovery drill. Verify the availability of the standby path, security group policies, and backend services to ensure the backup path works as expected during an actual failure.