You can configure alerts to detect leased line faults in a timely manner and minimize business impact.
Background
When a leased line fails, the business side typically immediately detects packet loss or unreachability. Whether alert configurations are complete and timely determines whether you can detect exceptions in a timely manner and start the emergency response process.
Express Connect relies on CloudMonitor to provide alert capabilities and supports the following two types of alerts:
Alert rules: Alerts based on metric thresholds. Suitable for routine threshold monitoring of quantitative metrics such as up/down status and traffic volume. You must manually define thresholds.
System events: Subscriptions based on events. System events are faults or exceptions identified by Alibaba Cloud through comprehensive analysis of multiple backend metrics, such as BGP faults, BFD DOWN, and VBR traffic drops. No need to define thresholds, out of the box.
Key alert configurations
We recommend that you configure the following key alerts.
1. Physical port faults
Configure the following settings for the target physical port Add an alert rule:
Product: Express Connect - Physical Connections.
Metric: PhysicalConnectionStatus
Period: 1 consecutive period (1 period = 1 minute).
Threshold: Status < 1, indicates a physical port fault.
2. VBR traffic exceptions
Configure the following settings Subscribe to system events:
Products: Network Intelligence Service.
Event name: risk-ec-inTrafficDroppedToZero or risk-ec-outTrafficDroppedToZero.
3. ECR traffic exceptions
Configure the following settings for the target ECR Add an alert rule. For example, if the historical outbound traffic from the ECR to a specific TR is stably above 100 Mbps 24/7, you can consider a monitored value of < 1 Mbps as a traffic drop:
Product: Express Connect Router.
Metric: TR Instance -> RateOutFromECRToTR.
Period: 1 consecutive period (1 period = 1 minute).
Threshold: Monitored value < 1 Mbps (This is only an example. Set a reasonable alert threshold based on the actual traffic bandwidth).
Dimension: Select the target region and target TR.
4. BGP and BFD exceptions
Configure the following settings Subscribe to system events:
Product: Network Intelligence Service.
Event name:
risk-ec-bgpRouterFail: If the BGP connection state changes from Connected to another state, traffic transmission is affected.
Notification of BFD in Down: A BFD state exception affects traffic transmission.
5. Traffic probing packet loss
Configure the following settings for the target probing task Add an alert rule:
Product: New Sitemonitor.
Metric: Packet loss rate
Period: 1 consecutive period (1 period = 1 minute).
Threshold: Average = 100%.
6. Health check exceptions (in scenarios without ECR)
Note: Health check applies only to scenarios without ECR, including VBR directly connected to a VPC (VBR upper connection, not enabled by default) or VBR directly connected to a TR.
Configure the following settings for the target probing task Add an alert rule:
Product: Express Connect - VBR.
Metric: VbrHealthyCheckLossPercent.
Period: 1 consecutive period (1 period = 1 minute).
Threshold: Monitored value = 100%.
Procedure
Add an alert rule
Go to the alert rule configuration page. Depending on the target resource:
Physical port or Virtual Border Router (VBR):
In the Monitoring column of target physical port/target VBR, click the corresponding monitoring icon
to open the Monitoring dialog.In the upper-right corner of the Monitoring dialog, click Set Alarm to go to the CloudMonitor console. The Create Alert Rule-Configure Rule Description dialog automatically appears.
Express Connect Router (ECR):
In the CloudMonitor console, in the target ECR column, click Monitoring Charts on the Actions.
Then, in the upper-right corner of the chart page, click Create Alert Rule. The Create Alert Rule-Configure Rule Description dialog automatically appears.
Traffic probing:
In the CloudMonitor console, on the Alert Service - Alert Rules page, click Create Alert Rule.
In the Create Alert Rule dialog: set Product to Site Monitor; set Resource Range to Instances, and in Associated Resources, click Add Instance to select the target traffic probing task; in Rule Description, click Add Rule -> Simple Metric to open the Configure Rule Description dialog.
Health check:
In the CloudMonitor console, on the Alert Service - Alert Rules page, click Create Alert Rule.
In the Create Alert Rule dialog: set Product to Express Connect - VBR; set Resource Range to Instances, and in Associated Resources, click Add Instance to select the target VBR; in Rule Description, click Add Rule -> Simple Metric to open the Configure Rule Description dialog.
In the Configure Rule Description dialog:
Alert Rule: Enter a meaningful name, such as xxx exception.
Metric Type: Keep the default Simple Metric.
Metric: Select the target Monitoring indicators based on Key alert configurations, and set Period and Threshold in the Critical section of Threshold and Alert Level.
After the configuration is complete, click OK.
In the Create Alert Rule dialog:
Mute Period: After an alert notification is sent, the same alert is not repeatedly sent within the specified period even if it is continuously triggered. The default period is 24 hours, which means that each alert is sent at most once every 24 hours.
Alert Contact Group: Select the target group from the drop-down list. If the drop-down list is empty, Create alert contacts and alert contact groups first.
After the configuration is complete, click Confirm.
Subscribe to system events
Go to the CloudMonitor Event Subscription page and click Create Subscription Policy.
On the Create Subscription Policy page:
Basic information: Enter a meaningful Name, such as xxx exception.
Alert Subscription: Set Subscription Type to System Events, select the target Product and Event name based on Key alert configurations for Subscription Scope, and keep the other settings as default.
Combined Noise Reduction: Keep the default setting.
Notification: Select the target Notification Configuration from the drop-down list. If the list is empty, Create a notification policy first.
Push and Integration: Leave this field empty.
Click Submit at the bottom of the page.
Verify alerts
After you complete the configurations in this topic, we recommend that you perform an alert drill on a regular basis to verify that the contact group, notification channels, and on-call response chain remain valid. This prevents alert configurations from being forgotten after setup and ensures that alerts take effect.
The following example uses a physical leased line fault drill to verify alert effectiveness:
A fault drill puts the drilled resource into an artificial fault state by shutting it down. Make sure that you have configured redundancy for the drill resource. Otherwise, business may be interrupted.
Go to the Failure Drill page and click Create Task.
On the Create Task page,
Region: Select the region where the target physical port is located.
Drill Resource: Select Express Connect Circuit, and then select the target physical port on the lower-left side and move it to the right Selected instances panel by clicking the arrow.
Drill Mode: Select Start Now.
Drill Duration: Select an acceptable duration. In this topic, 5 minutes is used as an example.
After you confirm the settings, click OK.
After The failure drill task is created. is displayed, click View Details to confirm the basic information of the drill task.
The drill status first changes to Starting, and then enters the In Drill state.
After the status changes to In Drill, Alert Contacts receives an alert. The following shows an example of an email alert:
ECR traffic exception
Time: 2026-05-22 10:26:49
Alert level: Critical
Instance name: [instanceId=ecr-fih7************, nodeRegion=eu-central-1, nodeId=tr-gw8m************]
Instance details: instanceId = ecr-fih************;
nodeRegion = eu-central-1;
userId = 11081************;
nodeId = tr-gw8************;
Metric: Outbound rate from ECR to TR
Alert condition: Triggered 1 consecutive time. Current value < 1 Mibit/s
Current value: 0 bit/s
Data details: __ts__=177************,
__count__=1,
instanceId=ecr-fih7sr************,
nodeRegion=eu-central-1,
Value=0,
nodeId=tr-gw8m************,
userId=1108************,
timestamp=1779************,
Duration: 1 minute,
Alert rule: [Outbound rate from ECR to TR < 1Mbps](Outbound rate from ECR to TR < 1Mbps).
BGP status alert
[CloudMonitor] ************ triggered 4 alerts. Highest level: CRITICAL.
User: Dear ************(110811************)
Start time: 2026-05-21 23:03:31 CST
Last trigger time: 2026-05-21 23:14:43 CST
Alert suppression: Direct notification
Policy rule: ************
Related alert events:
CRITICAL 2026-05-21 23:14:43 CST
Service name: Network Intelligence Service
Instance name: vbr-gw************
Affected resources: acs:EC:eu-central-1:11081************:instance/vbr-gw822************
Region: Germany 1 (Frankfurt) (eu-central-1)
Event name: BGP connection fault. Description: Supported cloud services and their system events
Event content: impact : {"EC":["vbr-gw8************"]}
description : Express Connect: vbr-gw822e************ BGP connection fault caused by a physical network connectivity failure or BGP configuration exception, resulting in route loss. BGP details: bgpGroupId:bgpg-gw8ns************, bgpPeerId:bgp-gw8hw************. We recommend that you contact your account manager.
event_time : 2026-05-21T15:14:00Z
Traffic probing
Dear user xxx, hello:
Site monitoring instance: taskName=alert, address=172.16.0.1. A packet loss alert was triggered at 2026-05-22 10:26:58. The average value is 100% and the duration is 0 minutes.
Rule details: Alert rule xxx. The 1-minute statistical value of the packet loss rate met the expression average == 100% for 1 consecutive time.
Health check
[CloudMonitor] Dear xxxx, xxx-Express Connect-Physical Connection triggered an alert
Time: 2026-05-22 10:29:08
Alert level: Critical
Instance name: [a-idc]
Instance details: instanceId = vbr-gw8v************;
userId = 1108************;
Metric: VBR health check packet loss rate
Alert condition: Triggered 3 consecutive times. Current value == 100%
Current value: 100%
Data details: __ts__=1779416910000,
__count__=1,
instanceId=vbr-gw8v************,
Value=100,
userId=11081************,
timestamp=1779416820000,
Duration: 4 minutes,
Alert rule: [6. Health check exceptions (in scenarios without ECR)](6. Health check exceptions (in scenarios without ECR))
If the alert rule alert is not triggered, first confirm that the contact information of the alert contact is valid, and then check the metric changes:
In the target alert rule column, click Modify on the Actions.
In the Modify Alert Rule pane, click the edit icon
on the rightmost side of Rule Description to open the Configure Rule Description pane.In the Configure Rule Description pane, view the metric trends in Chart Preview to check whether the metrics match the alert Period and Threshold.
Related documents
CloudMonitor - Alert rules: Create and manage alert rules.
CloudMonitor - System events: View and manage system events.
NIS system events: Multiple system events related to Express Connect belong to the NIS product. You can view the detailed rule definitions of Express Connect-related system events in this document.