All Products
Search
Document Center

Express Connect:Detect leased line faults in a timely manner through monitoring, alerts, and events

Last Updated:Jul 08, 2026

You can configure alerts to detect leased line faults in a timely manner and minimize business impact.

Background

When a leased line fails, the business side typically immediately detects packet loss or unreachability. Whether alert configurations are complete and timely determines whether you can detect exceptions in a timely manner and start the emergency response process.

Express Connect relies on CloudMonitor to provide alert capabilities and supports the following two types of alerts:

  • Alert rules: Alerts based on metric thresholds. Suitable for routine threshold monitoring of quantitative metrics such as up/down status and traffic volume. You must manually define thresholds.

  • System events: Subscriptions based on events. System events are faults or exceptions identified by Alibaba Cloud through comprehensive analysis of multiple backend metrics, such as BGP faults, BFD DOWN, and VBR traffic drops. No need to define thresholds, out of the box.

Key alert configurations

We recommend that you configure the following key alerts.

1. Physical port faults

Configure the following settings for the target physical port Add an alert rule:

  • Product: Express Connect - Physical Connections.

  • Metric: PhysicalConnectionStatus

    • Period: 1 consecutive period (1 period = 1 minute).

    • Threshold: Status < 1, indicates a physical port fault.

2. VBR traffic exceptions

Configure the following settings Subscribe to system events:

  • Products: Network Intelligence Service.

  • Event name: risk-ec-inTrafficDroppedToZero or risk-ec-outTrafficDroppedToZero.

    Click to view the risk-ec-inTrafficDroppedToZero rule definition

    Monitor the per-minute inbound rate from the IDC to the VPC for the VBR instance. An alert is triggered if all of the following conditions are met:

    Condition 1: For 3 consecutive minutes, the rate drops by ≥ 99% compared with the average rate of the preceding 7 minutes in each minute.

    Condition 2: For 3 consecutive minutes, the absolute drop in rate compared with the average rate of the preceding 7 minutes in each minute is ≥ 1 Mbps.

    Condition 3: For 3 consecutive minutes, the absolute drop in rate compared with the average rate of the preceding 15, 30, and 60 minutes in each minute is ≥ 0.5 Mbps.

    Condition 4 (smart baseline alert): By learning the periodic patterns of the inbound rate of the VBR instance, the system predicts the stable range of the inbound rate for the next cycle. If the rate breaks the lower bound of the predicted range by ≥ 99% for 2 minutes within a 3-minute window at the start of the cycle, an exception drop is detected.

    The outbound rule is defined in the same way. For more information, see the NIS documentation: Event center.

3. ECR traffic exceptions

Configure the following settings for the target ECR Add an alert rule. For example, if the historical outbound traffic from the ECR to a specific TR is stably above 100 Mbps 24/7, you can consider a monitored value of < 1 Mbps as a traffic drop:

  • Product: Express Connect Router.

  • Metric: TR Instance -> RateOutFromECRToTR.

    • Period: 1 consecutive period (1 period = 1 minute).

    • Threshold: Monitored value < 1 Mbps (This is only an example. Set a reasonable alert threshold based on the actual traffic bandwidth).

  • Dimension: Select the target region and target TR.

4. BGP and BFD exceptions

Configure the following settings Subscribe to system events:

  • Product: Network Intelligence Service.

  • Event name:

    • risk-ec-bgpRouterFail: If the BGP connection state changes from Connected to another state, traffic transmission is affected.

    • Notification of BFD in Down: A BFD state exception affects traffic transmission.

5. Traffic probing packet loss

Configure the following settings for the target probing task Add an alert rule:

  • Product: New Sitemonitor.

  • Metric: Packet loss rate

    • Period: 1 consecutive period (1 period = 1 minute).

    • Threshold: Average = 100%.

6. Health check exceptions (in scenarios without ECR)

Note: Health check applies only to scenarios without ECR, including VBR directly connected to a VPC (VBR upper connection, not enabled by default) or VBR directly connected to a TR.

Configure the following settings for the target probing task Add an alert rule:

  • Product: Express Connect - VBR.

  • Metric: VbrHealthyCheckLossPercent.

    • Period: 1 consecutive period (1 period = 1 minute).

    • Threshold: Monitored value = 100%.

Procedure

Add an alert rule

  1. Go to the alert rule configuration page. Depending on the target resource:

    • Physical port or Virtual Border Router (VBR):

      1. In the Monitoring column of target physical port/target VBR, click the corresponding monitoring icon image to open the Monitoring dialog.

      2. In the upper-right corner of the Monitoring dialog, click Set Alarm to go to the CloudMonitor console. The Create Alert Rule-Configure Rule Description dialog automatically appears.

    • Express Connect Router (ECR):

      1. In the CloudMonitor console, in the target ECR column, click Monitoring Charts on the Actions.

      2. Then, in the upper-right corner of the chart page, click Create Alert Rule. The Create Alert Rule-Configure Rule Description dialog automatically appears.

    • Traffic probing:

      1. In the CloudMonitor console, on the Alert Service - Alert Rules page, click Create Alert Rule.

      2. In the Create Alert Rule dialog: set Product to Site Monitor; set Resource Range to Instances, and in Associated Resources, click Add Instance to select the target traffic probing task; in Rule Description, click Add Rule -> Simple Metric to open the Configure Rule Description dialog.

    • Health check:

      1. In the CloudMonitor console, on the Alert Service - Alert Rules page, click Create Alert Rule.

      2. In the Create Alert Rule dialog: set Product to Express Connect - VBR; set Resource Range to Instances, and in Associated Resources, click Add Instance to select the target VBR; in Rule Description, click Add Rule -> Simple Metric to open the Configure Rule Description dialog.

  2. In the Configure Rule Description dialog:

    • Alert Rule: Enter a meaningful name, such as xxx exception.

    • Metric Type: Keep the default Simple Metric.

    • Metric: Select the target Monitoring indicators based on Key alert configurations, and set Period and Threshold in the Critical section of Threshold and Alert Level.

      After the configuration is complete, click OK.

  3. In the Create Alert Rule dialog:

    • Mute Period: After an alert notification is sent, the same alert is not repeatedly sent within the specified period even if it is continuously triggered. The default period is 24 hours, which means that each alert is sent at most once every 24 hours.

    • Alert Contact Group: Select the target group from the drop-down list. If the drop-down list is empty, Create alert contacts and alert contact groups first.

    After the configuration is complete, click Confirm.

Subscribe to system events

  1. Go to the CloudMonitor Event Subscription page and click Create Subscription Policy.

  2. On the Create Subscription Policy page:

    1. Basic information: Enter a meaningful Name, such as xxx exception.

    2. Alert Subscription: Set Subscription Type to System Events, select the target Product and Event name based on Key alert configurations for Subscription Scope, and keep the other settings as default.

    3. Combined Noise Reduction: Keep the default setting.

    4. Notification: Select the target Notification Configuration from the drop-down list. If the list is empty, Create a notification policy first.

    5. Push and Integration: Leave this field empty.

    Click Submit at the bottom of the page.

Verify alerts

After you complete the configurations in this topic, we recommend that you perform an alert drill on a regular basis to verify that the contact group, notification channels, and on-call response chain remain valid. This prevents alert configurations from being forgotten after setup and ensures that alerts take effect.

The following example uses a physical leased line fault drill to verify alert effectiveness:

Warning

A fault drill puts the drilled resource into an artificial fault state by shutting it down. Make sure that you have configured redundancy for the drill resource. Otherwise, business may be interrupted.

  1. Go to the Failure Drill page and click Create Task.

  2. On the Create Task page,

    • Region: Select the region where the target physical port is located.

    • Drill Resource: Select Express Connect Circuit, and then select the target physical port on the lower-left side and move it to the right Selected instances panel by clicking the arrow.

    • Drill Mode: Select Start Now.

    • Drill Duration: Select an acceptable duration. In this topic, 5 minutes is used as an example.

    • After you confirm the settings, click OK.

  3. After The failure drill task is created. is displayed, click View Details to confirm the basic information of the drill task.

    • The drill status first changes to Starting, and then enters the In Drill state.

    • After the status changes to In Drill, Alert Contacts receives an alert. The following shows an example of an email alert:

      ECR traffic exception

      Time: 2026-05-22 10:26:49

      Alert level: Critical

      Instance name: [instanceId=ecr-fih7************, nodeRegion=eu-central-1, nodeId=tr-gw8m************]

      Instance details: instanceId = ecr-fih************;

      nodeRegion = eu-central-1;

      userId = 11081************;

      nodeId = tr-gw8************;

      Metric: Outbound rate from ECR to TR

      Alert condition: Triggered 1 consecutive time. Current value < 1 Mibit/s

      Current value: 0 bit/s

      Data details: __ts__=177************,

      __count__=1,

      instanceId=ecr-fih7sr************,

      nodeRegion=eu-central-1,

      Value=0,

      nodeId=tr-gw8m************,

      userId=1108************,

      timestamp=1779************,

      Duration: 1 minute,

      Alert rule: [Outbound rate from ECR to TR < 1Mbps](Outbound rate from ECR to TR < 1Mbps).

      BGP status alert

      [CloudMonitor] ************ triggered 4 alerts. Highest level: CRITICAL.

      User: Dear ************(110811************)

      Start time: 2026-05-21 23:03:31 CST

      Last trigger time: 2026-05-21 23:14:43 CST

      Alert suppression: Direct notification

      Policy rule: ************

      Related alert events:

      CRITICAL 2026-05-21 23:14:43 CST

      Service name: Network Intelligence Service

      Instance name: vbr-gw************

      Affected resources: acs:EC:eu-central-1:11081************:instance/vbr-gw822************

      Region: Germany 1 (Frankfurt) (eu-central-1)

      Event name: BGP connection fault. Description: Supported cloud services and their system events

      Event content: impact : {"EC":["vbr-gw8************"]}

      description : Express Connect: vbr-gw822e************ BGP connection fault caused by a physical network connectivity failure or BGP configuration exception, resulting in route loss. BGP details: bgpGroupId:bgpg-gw8ns************, bgpPeerId:bgp-gw8hw************. We recommend that you contact your account manager.

      event_time : 2026-05-21T15:14:00Z

      Traffic probing

      Dear user xxx, hello:

      Site monitoring instance: taskName=alert, address=172.16.0.1. A packet loss alert was triggered at 2026-05-22 10:26:58. The average value is 100% and the duration is 0 minutes.

      Rule details: Alert rule xxx. The 1-minute statistical value of the packet loss rate met the expression average == 100% for 1 consecutive time.

      Health check

      [CloudMonitor] Dear xxxx, xxx-Express Connect-Physical Connection triggered an alert

      Time: 2026-05-22 10:29:08

      Alert level: Critical

      Instance name: [a-idc]

      Instance details: instanceId = vbr-gw8v************;

      userId = 1108************;

      Metric: VBR health check packet loss rate

      Alert condition: Triggered 3 consecutive times. Current value == 100%

      Current value: 100%

      Data details: __ts__=1779416910000,

      __count__=1,

      instanceId=vbr-gw8v************,

      Value=100,

      userId=11081************,

      timestamp=1779416820000,

      Duration: 4 minutes,

      Alert rule: [6. Health check exceptions (in scenarios without ECR)](6. Health check exceptions (in scenarios without ECR))

If the alert rule alert is not triggered, first confirm that the contact information of the alert contact is valid, and then check the metric changes:

  1. In the target alert rule column, click Modify on the Actions.

  2. In the Modify Alert Rule pane, click the edit icon image on the rightmost side of Rule Description to open the Configure Rule Description pane.

  3. In the Configure Rule Description pane, view the metric trends in Chart Preview to check whether the metrics match the alert Period and Threshold.

Related documents