All Products
Search
Document Center

Cloud Network Well-architected Design Guidelines:Intelligent network O&M design

Last Updated:Jun 02, 2026

Overview

Summary

As enterprises adopt cloud computing, cloud network O&M becomes critical to platform efficiency, data security, and service availability.

Compared with traditional IT architectures, cloud environments are more complex and abstract. Manual configuration no longer scales as parameters multiply, so automation tools must assist decision-making. An intelligent O&M system identifies and resolves potential issues to ensure service continuity and stability.

O&M aims to locate, fix, and prevent failures while optimizing network architecture and performance. Alibaba Cloud provides the following solution:

  1. Alerting: Deploy CloudMonitor to monitor system status in real time and trigger alerts on anomalies, minimizing service interruptions.

  2. Inspection: Run periodic full-dimensional inspections to identify and fix potential risks before they cause major incidents.

  3. Observation: Use AIOps to continuously monitor network environments. Track key metrics to discover trends, plan ahead, and improve network stability and performance.

Keywords

  • Network Intelligence Service (NIS): NIS provides AIOps tools for managing cloud networks across the full lifecycle, including traffic analysis, inspections, performance monitoring, diagnostics, path analysis, and topology creation. NIS helps you optimize your network architecture, improve O&M efficiency, and reduce network operations costs.

  • CloudMonitor: CloudMonitor is a service that monitors resources and Internet applications.

  • Virtual Private Cloud (VPC): A VPC is a logically isolated private network on Alibaba Cloud. VPCs are logically isolated from each other at Layer 2. You can create and manage cloud resources in your VPC, such as ECS, SLB, and ApsaraDB RDS instances.

  • Elastic IP Address (EIP): An EIP is a public IP address that you can purchase and hold as an independent resource.

  • NAT Gateway: NAT gateways translate network addresses.

  • Application Load Balancer (ALB): ALB is a Layer 7 load balancer optimized for HTTP, HTTPS, and QUIC traffic. ALB supports complex routing, scales elastically, and is deeply integrated with other cloud-native services, serving as a cloud-native Ingress gateway.

  • Network Load Balancer (NLB): NLB is a Layer 4 load balancer that auto-scales and supports up to 100 million concurrent connections.

  • Classic Load Balancer (CLB): CLB distributes inbound traffic across backend servers based on forwarding rules to improve application performance and availability.

  • Cloud Enterprise Network (CEN): CEN is a high-availability network that runs on the Alibaba Cloud global private network. CEN uses transit routers to connect VPCs across regions and with data centers, enabling flexible, reliable, and enterprise-class cloud networks.

  • VPN Gateway: VPN Gateway provides secure and reliable connections between enterprise data centers, office networks, and Internet clients and Alibaba Cloud through encrypted private tunnels.

  • Express Connect circuit: Express Connect circuits are physical cables or fibers that connect data centers, deployed and maintained by ISPs. They are classified into dedicated and shared types.

  • Express Connect: What is Express Connect is a networking service that establishes high-speed, secure private connections between data centers and Alibaba Cloud.

  • Virtual border routers (VBRs): VBRs are virtualized abstractions of Express Connect circuits in the SDN architecture. A VBR sits between customer-premises equipment (CPE) and a VPC to exchange data between the VPC and data center.

Design principles

Consider the following principles:

Alert-driven O&M response mechanism

  1. Event subscription mechanism: Configure alert rules to notify you of system anomalies, performance issues, or security risks as early as possible.

  2. Emergency response to high-severity alerts: Define response plans with assigned owners for high-severity alerts.

  3. Periodic audit in the event center: Periodically audit historical events to identify error trends and potential risks before they cause service interruptions.

Troubleshooting mechanism for high-severity risks

Perform periodic network inspections to identify and fix potential risks. Build a network O&M system to monitor status and respond quickly to risks that may compromise performance and security.

Observation-oriented network optimization mechanism

  • Keep traffic analysis enabled to continuously monitor throughput, packet loss, latency, and user distribution. Use these metrics to optimize your service architecture.

  • Use topology generators to track network status in real time and optimize the network structure.

  • Use network insight providers to monitor Internet status and detect issues for better Internet management.

Key design

Use alerts to detect and locate errors

Configure alert rules

Alert rules for system events

System events: System events cover failure and O&M events across cloud services. Subscribing sends alert notifications to you or a third-party system when events trigger. Configure the subscription scope: services, event types, event names, severity levels, application groups, event content, and event resources.

Enable all CloudMonitor modules related to system events. This ensures you receive business-critical alerts promptly, improving system stability and security.

For more information about the system events supported by CloudMonitor, see Supported cloud services and their system events.

Network system events fall into the following categories:

  1. Bandwidth and performance limits

  • Over-limit events: The upper limit on private bandwidth, Internet bandwidth, ALB, CLB, or NLB bandwidth, or number of connections on ALB, CLB, or NLB is reached.

  • Packet loss events: Packets are dropped due to bandwidth exhaustion on ALB, CLB, VPCs, or NAT gateways.

  • Over-limit QPS and request events: The HTTP 503 error is triggered when the upper limit on the ALB QPS is reached.

  1. Connect management and session control

  • Over-limit sessions and dropped connections: New connections are dropped because the number of sessions on the ALB or CLB instance has reached the upper limit or the number of new connections on the NLB instance suddenly increases.

  • Connection failures: The number of connection failures on the CLB or NLB instance suddenly increases.

  1. Routes and network stability

  • Over-limit routes: The number of CEN routes or dynamically allocated BGP routes reaches the upper limit.

  • Network jitters: CEN or VPC network jitters.

  • Connection errors: Errors on Express Connect circuits or BGP connections.

  1. VPN and IPsec events

  • Over-limit bandwidth and connections: The upper limit on VPN bandwidth and IPsec negotiations is reached.

  • Health checks: A VPN gateway or an IPsec connection passes or fails health checks.

  1. Endpoint and connection management

  • Operations on endpoints: Accept, reject, add, or delete an endpoint.

  1. Certificate issues

  • Certificate and security issues: The certificate of an SLB or VPN gateway expires.

Business alerts

Threshold-triggered events: Events trigger when conditions in a threshold-triggered alert rule are met. Subscribe to these events to configure fine-grained notifications, including alert merging, denoising, and custom notification methods. Configure the subscription scope: services, metrics, severity levels, and application groups.

Configure fine-grained alert rules and thresholds for business-critical metrics in CloudMonitor. The system then performs trend analysis and anomaly inspection to identify potential errors and risks, ensuring service availability.

For more information about the monitoring metrics supported by CloudMonitor, see Appendix 1: Metrics for Alibaba Cloud products.

Subscribe to alert notifications

Alert notifications are classified into Critical, Warning, Notification (Info), and Resolved based on the alert severity.

Configure a notification method for each alert level. For Critical alerts that directly impact your business, use telephone calls as the primary notification method and respond immediately. For low-impact alerts, review and manage them daily during a specific time window.

For more information about alert templates, see Manage notification templates.

Alerts triggered by system events

In the CloudMonitor console, choose Event Center > Event Subscription and create a subscription policy.

Business alerts
  1. Create alert rules

    Create alert rules to monitor cloud resource usage. When metrics meet specified conditions, CloudMonitor automatically sends alert notifications.

    You can create alert rules based on CloudMonitor metrics or custom business metrics. In the CloudMonitor console, choose Alerts > Alert Rules and click Create Alert Rule.

  2. Subscribe to threshold-triggered events

    Use event subscription to configure custom alert notifications, including alert merging, denoising, contact group upgrades, custom notification methods, and JSON-based alert delivery to destination channels.

    In the CloudMonitor console, choose Event Center > Event Subscription and click Create Subscription Policy.

Manage alerts

Alerts triggered by system events

CloudMonitor displays detected events on the Event Center > Notification History page. O&M engineers can manage and fix issues based on the event details.

Critical alerts require immediate response. For low-severity alerts, check the event center daily to ensure system stability.

Business alerts

View business alerts on the Event Center > Notification History page in the CloudMonitor console.

Configure Function Compute or automation scripts to automatically fix issues based on your business requirements. Alternatively, manage alerts on the Notification History page on a regular basis to optimize resource utilization.

Use inspections to identify and eliminate potential risks

Configure inspections for different types of risks

  1. Stability risks

    In HA architecture design, improper primary/secondary configuration can cause switchover failures. Improper resource deployment may also expand the blast radius, affecting more servers and reducing overall stability.

    Run inspections to optimize resource deployment and verify switchover configurations, improving disaster recovery and eliminating potential risks.

  2. Security risks

    Coarse-grained ACLs may fail to block unauthorized access. Security groups may expose unnecessary ports, violating the principle of least privilege (PoLP) and increasing attack risk.

    Run inspections to verify that ACLs and security groups allow only authorized access to necessary destinations.

  3. Performance risks

    Network latency may be increased by performance bottlenecks or bypasses. Packet loss may occur if network traffic frequently exceeds the maximum bandwidth.

    Use inspections to monitor latency and scale out resources as needed, ensuring QoS as data transfer volumes grow.

  4. Resource waste

    Low resource unitization results in resource waste. If you select an improper billing method, spending on resources may unexpectedly increase, which reduces the cost-benefit ratio.

    Run inspections to optimize resource deployment and increase utilization. Select a billing method based on cost-benefit analysis to control your budget.

For more information, see Network inspection.

Run inspections

Run inspections weekly to monitor your network status and identify potential issues that reduce resource utilization. Continuous monitoring helps maintain a stable network architecture, reduce costs, and ensure service continuity.

To view weekly inspection reports, in the NIS console, click Network Inspection, click View historical reports in the Newest Inspection Report column, and then click Re-start.

  1. Assess the overall network status based on the pass rate: O&M engineers can quickly determine overall network performance and identify potential issues based on the trend of inspection scores.

  2. Handle risks by risk level: Inspection items are sorted by priority from highest to lowest risk. Follow the suggestions in the inspection reports to fix high-risk issues first and then optimize your network environment.

Handle potential risks

Examples:

  1. Control costs

    • EIPs: Run inspections to detect and release idle EIPs to prevent resource waste.

    • CEN: Allocate inter-region bandwidth resources based on actual traffic volumes to prevent resource waste.

  2. Improve stability

    • Over-limit risks: Bandwidth exhaustions or insufficient resources specifications.

    • Single points of failure (SPOFs) in a zone: If you deploy an ALB instance, an NLB instance, or a transit router in a single zone, instability issues may arise.

    • SPOFs on connections: If you use only one Express Connect circuit, one GA acceleration one, or one VPN tunnel, connectivity issues may arise.

    • Service unavailability: Service errors may occur.

Perform global network optimization based on observability

Use observation tools

  1. Generate topologies — Virtualize the entire network

    Network topologies visualize the connections and relationships between network resources. They help you understand your Alibaba Cloud network architecture, verify configurations, troubleshoot issues, and centralize O&M.

    Topology

    Displayed information

    VPC

    Resources, including ECS instances, vSwitches, and routers

    Routes, including network elements inside and outside VPCs and route tables

    CEN

    Transit routers worldwide, VPCs connected to transit routers, and transit routers connected to each other

    SLB

    SLB zones, virtual IP addresses (VIPs), EIPs, and security groups

  2. Traffic analysis — Sort network traffic from multiple dimensions

    The traffic analysis feature monitors real-time and historical network traffic, generating visualized time series charts in the NIS console. Use the traffic data and collected metrics for troubleshooting.

    • Internet traffic analysis: Analyze traffic by resource type associated with public IP addresses, including CLB, ECS, Internet NAT gateways, EIPs, and EIPs in Internet Shared Bandwidth instances.

    • Hybrid cloud traffic analysis: Analyze inbound and outbound traffic through VBRs connected to transit routers in hybrid clouds.

    • Inter-region traffic analysis: Analyze inbound and outbound traffic through transit routers across regions. Data is displayed as 1-tuple, 2-tuple, and 5-tuple.

    • Intra-region traffic analysis: Analyze inbound and outbound traffic through transit routers connected to VPCs within the same region.

    • Internet NAT gateway traffic analysis: Analyze Internet NAT gateway traffic and view visualized time series charts on the Overview page in the NIS console.

  3. Internet quality — Impacts caused by Internet quality degradation

    • Detect Internet quality degradation based on the round-trip time (RTT) and retransmission rate.

    • Detect Internet quality degradation events, including the time range, ISP, area, and traffic volume.

    • Detect the public IP addresses affected by Internet quality degradation.

On-demand observation

  1. Network topology

    In the NIS console, select a network instance in the Network Topology module and click Generate Topology. Topology drilldowns provide information from different network layers, helping you analyze resource allocation and manage network O&M.

    1. VPC topologies: Display resource and route topologies within VPCs. View basic information about network instances, analyze them, and check reachability.

    2. CEN topologies: Display intra-region and inter-region connections between transit routers on a CEN instance. View global cloud resource connections and network instance details.

    3. SLB topologies: Display connections between listeners and backend server groups of an SLB instance. Analyze network instances to verify that traffic is routed as expected.

  2. Traffic analysis

    Use NIS traffic analysis to monitor real-time and historical network traffic. Analyze traffic by source IP, source and destination IPs, or the full 5-tuple (source IP, source port, destination IP, destination port, protocol). Sort network traffic to identify top N instances.

    Enable the following features separately before use: Internet traffic analysis, hybrid cloud traffic analysis, inter-region traffic analysis, and intra-region traffic analysis.

    • Enable Internet traffic analysis for specific regions or public IP addresses. Selecting a region enables this feature for all public IP addresses in that region.

    • Enable hybrid cloud traffic analysis for specific VBR connections on transit routers.

    • Enable inter-region traffic analysis for specific inter-region connections on transit routers.

    • Enable intra-region traffic analysis for specific VPC connections on transit routers.

  3. Insight provider

    Use insight providers to get real-time Internet quality assessments, detect quality degradation, and receive impact analysis for Internet quality events.

    When you create an insight provider, configure monitored objects. Ten minutes after creation, the insight provider starts collecting resource traffic and pushing metrics. Click the insight provider name to view quality scores, degradation events, and affected public IP addresses. This information helps you make timely Internet adjustments to prevent business loss.

Analysis and optimization

  1. Optimization based on network topology observation

    1. A network topology shows the entire architecture, providing architecture summaries, path analysis, and resource allocation status.

    2. Network topologies help you identify potential issues through the following checks:

      • Redundancy check: ensures redundancy mechanisms exist to prevent SPOFs.

      • Configuration check: verifies configurations follow best practices and helps correct improper settings.

      • Security check: identifies potential risks, such as unnecessarily exposed ports and services.

    3. Manage low-utilization or idle resources:

      • Resource recycling: Release IP addresses and ports that are no longer used.

      • Configuration optimization: Optimize resource allocation and disable unused services.

  2. Traffic and business optimization based on traffic analysis

    1. Internet optimization

      Internet traffic analysis identifies the geographic distribution of users. Deploy services in popular areas to reduce latency and improve user experience.

      Internet traffic analysis continuously monitors Internet status using key metrics such as bandwidth utilization, source/destination IPs, ports, protocols, and RTT. This data helps you identify peak hours and plan capacity to maintain high availability during high workloads.

    2. Internal network optimization

      To optimize internal network traffic, detect the top N sources generating the highest traffic volume and perform drilldown analysis to identify and fix anomalies. This helps you prioritize key business traffic and reduce degradation from non-critical traffic. Inspect the TCP retransmission rate regularly to assess packet loss that may compromise business continuity. Make adjustments based on observation results to improve network quality.

  3. Internet issue identification based on insight providers

    Insight providers collect client location and ISP data, use intelligent baseline algorithms to detect performance or availability degradation, and provide event details for troubleshooting (including traffic analysis and Internet probes). View RTT and traffic data on the Internet traffic source map to monitor Internet status in real time and make timely adjustments.

Best practices

The following best practices are based on the preceding design principles:

Check alerts and fix issues

Check alerts daily. Ensure that high-severity alerts are pushed to your mobile phone in real time.

image

Run inspections to eliminate potential risks

Run inspections weekly.

image

Observe and optimize

Choose a suitable analysis tool.

image

Scenarios

Network O&M alerting

  • Notifications of risks and anomalies: When resource availability or performance issues occur, Alibaba Cloud pushes events to the NIS or CloudMonitor event center. Examples include performance degradation from excessive resource usage, Internet packet loss causing business unavailability, and subscription expiration. Handle these events promptly to prevent business interruptions.

  • Automatic O&M: Alibaba Cloud defines event status in the NIS event center. New events and status changes are reported to CloudMonitor, enabling you to build an event-driven automated O&M system.

Network O&M inspection

When you deploy or maintain networks, configurations may not follow best practices if you are unfamiliar with the cloud services. As your network grows, managing and inspecting numerous instances requires significant effort. The network inspection feature diagnoses your network architecture and deployed resources, and provides optimization suggestions.

Network O&M observation

  • Network topology analysis: Network topologies provide comprehensive information about network architectures to help you identify and optimize the deployment and communication between network nodes. They visualize connections and relationships between network resources to help you understand your Alibaba Cloud network, verify configurations, troubleshoot issues, and centralize O&M.

  • Network traffic monitoring and management: Monitor network traffic from a single console. Use traffic analysis to monitor real-time and historical network traffic.

  • Internet quality assessment: Run periodic or continuous tests to assess Internet quality based on key metrics such as latency, packet loss rate, and jitter. Use the results to identify performance issues and improve user experience.