Monitor Network Usage and Performance
Monitoring network usage and performance involves systematically observing how network resources are used and how key performance metrics behave. This helps ensure business continuity and service quality in real time. Build a multi-dimensional monitoring system that covers critical metrics such as bandwidth utilization, traffic patterns, latency distribution, and packet loss rate. Combine active network inspection with real-time alerting to detect potential risks early. Continuously track resource consumption and performance baselines to quantify your architecture’s elasticity—such as its ability to handle burst traffic—and optimize overall system performance.
Priority
Middle
Not Recommended
Focus only on network bandwidth and ignore latency. Overemphasize bandwidth utilization while ignoring key metrics such as TCP retransmission rate and network latency. This leads to slow application response times.
Mismatch monitoring metrics with business needs. Design a monitoring system that does not reflect real business scenarios. It fails to identify network issues that affect user experience.
Expected Outcome
Build a layered monitoring, alerting, and inspection system that covers the Network Layer (traffic, connections), Transport Layer (latency, packet loss), and Application Layer (service availability). This system detects performance risks, identifies issues proactively, and supports prompt response and resolution.
Implementation Guide
Monitoring network usage and performance is part of building a cloud network operations and maintenance system. Follow these steps:
Establish performance baselines. Use Alibaba Cloud NIS, Cloud Monitor, and ARMS to collect key network-layer metrics, such as network latency, packet loss rate, TCP retransmission rate, bandwidth utilization, Connections, and new connection rate. Set different thresholds for different business scenarios to ensure monitoring data reflects real user experience.
Use the Network Dashboard to see the big picture and gain insights across your entire environment. The Network Dashboard is more than a visualization dashboard. It is the central hub for cloud network operations—combining monitoring, analysis, decision-making, and collaboration. Design your dashboard using these guidelines:
Support specific roles with data to solve specific problems.
Show network architecture in layers to avoid information overload.
Focus on key business metrics. Place the most important metrics on the dashboard. Track other metrics only when investigating related issues.
Use alerts to detect and locate issues. Set up monitoring and alerting for performance metrics across different products to help you spot performance gaps quickly.
Subscribe to events. Subscribe to events that affect your business and set up alerts to detect system anomalies, performance issues, or security threats as soon as they occur.
Respond immediately to critical alerts. Create a strict emergency response plan. For alerts marked as critical, define clear procedures and assign a dedicated person to coordinate resolution until the issue is fully resolved.
Review the Event Center regularly. Schedule periodic reviews of the Event Center history. Analyze this data to identify emerging trends or chronic issues, and then take preventive action to avoid service interruption.
Run routine inspections to find and fix hidden risks. Perform regular inspections to identify performance risks—such as scaling out network resources when usage reaches a threshold. NIS provides some network inspection capabilities: Network Inspection.
Use tools to analyze root causes and resolve issues. For example, use NIS for instance diagnosis, path analysis, and network traffic analysis to identify performance risk points and resolve performance issues: Detect and Troubleshoot Anomalies.