Appendix: Supported diagnostic scenarios

Updated at:
Copy as MD

This document covers the diagnostic scenarios supported by Cloud Service Diagnosis. For more details, refer to other help documents in this section.To suggest new diagnostic features or provide feedback, join our DingTalk group (ID: 86570007290).

Diagnostic scenarios

See the supported diagnostic scenarios below. We are continually adding more.

Compute

  1. ECS comprehensive diagnosis

    1. Diagnostic product: ECS

    2. Diagnostic target: Running ECS instance

    3. Description: If you suspect an issue with your ECS instance but are unsure of the cause, use the ECS comprehensive diagnosis diagnostic tool to troubleshoot issues related to compute resources, operating system configurations, instance configurations, network services, security controls, and billing. Follow the provided suggestions to resolve the issue and restore your service.

    4. Diagnostic entry: ECS comprehensive diagnosis

  2. ECS Remote Connection Failure

    1. Diagnostic product: ECS

    2. Diagnostic target: Running ECS instance

    3. Description: If you cannot remotely connect to an ECS instance, use the ECS Remote Connection Failure tool to quickly diagnose the cause. The diagnosis covers instance configuration management, network services, operating system settings, compute services, and billing. Follow the provided suggestions to resolve the issue and restore your service.

    4. Diagnostic entry: ECS Remote Connection Failure

  3. ECS Instance Security Risks

    1. Diagnostic product: ECS

    2. Diagnostic target: ECS instance

    3. Description: If you suspect your ECS instance has been attacked or compromised, use the ECS Instance Security Risks tool to quickly check the ECS instance for security risks. If risks are found, you can follow the suggestions to resolve them and secure your instance.

    4. Diagnostic entry: ECS Instance Security Risks

  4. ECS instance high load

    1. Diagnostic product: ECS

    2. Diagnostic target: ECS instance

    3. Description: If you notice high CPU, disk, or memory utilization, or slow system response on your ECS instance, use the ECS instance high load tool to quickly diagnose the cause of the high load. The tool diagnoses the load on the instance's CPU, memory, disk IOPS or BPS, and bandwidth. Follow the provided suggestions to resolve the issue and restore your service.

    4. Diagnostic entry: ECS instance high load

  5. ECS Instance Security Control

    1. Diagnostic product: ECS

    2. Diagnostic target: ECS instance

    3. Description: If you find that your ECS instance is locked or suspect it has been blocked, use the ECS Instance Security Control tool to quickly identify the cause and impact of any security control events on the instance. Follow the provided suggestions to resolve the issue and restore your service.

    4. Diagnostic entry: ECS Instance Security Control

  6. ECS Instance Downtime

    1. Diagnostic product: ECS

    2. Diagnostic target: ECS instance

    3. Description: If your ECS instance experiences a system crash, blue screen, freeze, automatic restart, or downtime, use the ECS Instance Downtime tool to quickly diagnose the causes of these issues. Follow the provided suggestions to resolve the issue and restore your service.

    4. Diagnostic entry: ECS Instance Downtime

  7. ECS Network Performance Degradation

    1. Diagnostic product: ECS

    2. Diagnostic target: ECS instance

    3. Description: If you experience slow network speed, frequent packet loss, or abnormal network sessions on your ECS instance, use the ECS Network Performance Degradation tool to quickly diagnose the cause of these issues. Follow the provided suggestions to resolve the issue and restore your service.

    4. Diagnostic entry: ECS Network Performance Degradation

  8. Insufficient ECS Resource Quota

    1. Diagnostic product: ECS

    2. Diagnostic target: region

    3. Description: If you cannot create a security group or an image, or save data to a cloud disk in a specific region, use the Insufficient ECS Resource Quota tool to quickly check if your resource quota in that region is sufficient. If the quota is insufficient, follow the suggestions to increase it and restore your service.

    4. Diagnostic entry: Insufficient ECS Resource Quota

  9. ECS Cost and Security Behavior Audit

    1. Diagnostic product: ECS

    2. Diagnostic target: ECS instance

    3. Description: If you notice that a security group has been modified, an instance has been stopped unexpectedly, costs have increased unexpectedly, or the number of instances has changed for no apparent reason, use the ECS Cost and Security Behavior Audit tool to quickly check for unexpected changes related to instances, security groups, and costs. If anomalies are found, follow the provided suggestions to resolve the issue and restore your service.

    4. Diagnostic entry: ECS Cost and Security Behavior Audit

  10. ECS GPU Device Health Check

    1. Diagnostic product: ECS

    2. Diagnostic target: Running Linux ECS instance

    3. Description: If the GPU of your ECS instance is not working correctly, or you cannot find or connect to it, use the ECS GPU Device Health Check tool to quickly check the health of the GPU device. If an issue is found, follow the provided suggestions to resolve the issue and restore your service.

    4. Diagnostic entry: ECS GPU Device Health Check

  11. ECS Cloud Disk Resize Failure

    1. Diagnostic product: ECS

    2. Diagnostic target: Running Linux ECS instance

    3. Description: If a cloud disk resize has not taken effect on your ECS instance, use the ECS Cloud Disk Resize Failure tool to quickly check the cloud disk status. If an issue is found, follow the provided suggestions to resolve the issue and restore your service.

    4. Diagnostic entry: ECS Cloud Disk Resize Failure

  12. ECS instance startup failure

    1. Diagnostic product: ECS

    2. Diagnostic target: Stopped Linux ECS instance

    3. Description: If your ECS instance fails to start, you cannot access the operating system, or the instance cannot be stopped or shut down, use the ECS instance startup failure tool to quickly diagnose the startup or shutdown failure. Follow the provided suggestions to resolve the issue and restore your service. This diagnostic tool may attach a temporary repair disk, which can modify your instance's system disk. Before you run the diagnosis, create a snapshot of the system disk to prevent data loss. After the diagnosis, if no further repairs are needed, detach the repair disk by using the "Detach Repair Disk" feature in the report.

    4. Diagnostic entry: ECS instance startup failure

  13. ECS SSH Connection Failure

    1. Diagnostic product: ECS

    2. Diagnostic target: Running ECS instance

    3. Description: If you cannot connect to your ECS instance over SSH, use the ECS SSH Connection Failure tool to quickly diagnose the cause. Follow the provided suggestions to resolve the issue and restore your service.

    4. Diagnostic entry: ECS SSH Connection Failure

  14. ECS Workbench public remote connection failure

    1. Diagnostic product: ECS

    2. Diagnostic target: Running ECS instance

    3. Description: If you cannot connect to your ECS instance through Workbench over the public network, use the ECS Workbench public remote connection failure tool to quickly diagnose the cause. Follow the provided suggestions to resolve the issue and restore your service.

    4. Diagnostic entry: ECS Workbench public remote connection failure

  15. ECS Workbench internal remote connection failure

    1. Diagnostic product: ECS

    2. Diagnostic target: Running ECS instance

    3. Description: If you cannot connect to your ECS instance through Workbench over the internal network, use the ECS Workbench internal remote connection failure tool to quickly diagnose the cause. Follow the provided suggestions to resolve the issue and restore your service.

    4. Diagnostic entry: ECS Workbench internal remote connection failure

  16. ECS Ping Failure

    1. Diagnostic product: ECS

    2. Diagnostic target: Running ECS instance

    3. Description: If you cannot ping your ECS instance, use the ECS Ping Failure tool to quickly diagnose the cause. Follow the provided suggestions to resolve the issue and restore your service.

    4. Diagnostic entry: ECS Ping Failure

  17. Simple Application Server Remote Connection Failure

    1. Diagnostic product: Simple Application Server

    2. Diagnostic target: Running Simple Application Server instance

    3. Description: If you cannot remotely connect to a Simple Application Server instance, use the Simple Application Server Remote Connection Failure tool to quickly diagnose the possible causes. The diagnosis covers instance configuration management, network services, operating system settings, compute services, and storage services. Follow the provided suggestions to resolve the issue and restore your service.

    4. Diagnostic entry: Simple Application Server Remote Connection Failure

  18. Linux OS comprehensive diagnosis

    1. Diagnostic product: ECS

    2. Diagnostic target: Running ECS instance

    3. Description: This tool performs comprehensive OS health checks and kernel-level diagnostics. It analyzes the correlation between application behavior and cloud system events to help you identify problematic processes and find solutions.

    4. Diagnostic entry: Linux OS comprehensive diagnosis

Container

  1. ACK Pod Anomaly

    1. Diagnostic Product: Container Service for Kubernetes

    2. Diagnostic Object: ACK pod

    3. Description: If you suspect a pod in Container Service for Kubernetes (ACK) is failing to start or restarting frequently, use the ACK Pod Anomaly diagnostic tool to troubleshoot pod anomalies. Use its recommendations to resolve issues, restore services, and improve operational efficiency.

    4. Diagnostic Entry: ACK Pod Anomaly

  2. ACK Node Anomaly

    1. Diagnostic Product: Container Service for Kubernetes

    2. Diagnostic Object: ACK node

    3. Description: If you suspect issues with a node in Container Service for Kubernetes (ACK), use the ACK Node Anomaly diagnostic tool to troubleshoot node anomalies. Use its recommendations to resolve issues, restore services, and improve operational efficiency.

    4. Diagnostic Entry: ACK Node Anomaly

  3. ACK Ingress Anomaly

    1. Diagnostic Product: Container Service for Kubernetes

    2. Diagnostic Object: ACK ingress

    3. Description: If you suspect issues with an ingress in Container Service for Kubernetes (ACK), use the ACK Ingress Anomaly diagnostic tool to troubleshoot issues with ingress resources, configuration, and connectivity. Use its recommendations to resolve issues, restore services, and improve operational efficiency.

    4. Diagnostic Entry: ACK Ingress Anomaly

  4. ACK Service Anomaly

    1. Diagnostic Product: Container Service for Kubernetes

    2. Diagnostic Object: ACK service

    3. Description: If you suspect issues with a service in Container Service for Kubernetes (ACK), use the ACK Service Anomaly diagnostic tool to troubleshoot issues with service configuration, quotas, and abnormal events. Use its recommendations to resolve issues, restore services, and improve operational efficiency.

    4. Diagnostic Entry: ACK Service Anomaly

  5. ACK Comprehensive Diagnosis

    1. Diagnostic Product: Container Service for Kubernetes

    2. Diagnostic Object: ACK cluster

    3. Description: If you suspect issues with Container Service for Kubernetes (ACK), use the ACK Comprehensive Diagnosis diagnostic tool to troubleshoot issues with cluster security, performance, stability, cost, and service limits. Use its recommendations to resolve issues, restore services, and improve operational efficiency.

    4. Diagnostic Entry: ACK Comprehensive Diagnosis

Storage

  1. Cloud Disk Diagnosis

    1. Product: block storage

    2. Target: cloud disk instance

    3. Description: If you suspect a cloud disk has issues, such as slow I/O or poor performance, use the Cloud Disk Diagnosis tool to quickly troubleshoot the problem. Use the provided suggestions to resolve the issue, promptly restore your services, and improve operational efficiency.

    4. Entry point: Cloud Disk Diagnosis

Network and CDN

  1. Classic Load Balancer

    1. Diagnostic Product: Load Balancer

    2. Diagnostic Target: CLB instance

    3. Description: If you suspect a Classic Load Balancer (CLB) instance has issues such as packet loss, connectivity problems, an abnormal health status, overdue payments, or security policy problems, use the Classic Load Balancer to investigate its health, configuration, security, capacity, and costs. If any problems are identified, use the provided remediation suggestions to resolve them, restore services, and improve O&M efficiency.

    4. Diagnostic Entry: Classic Load Balancer

  2. Application Load Balancer

    1. Diagnostic Product: Load Balancer

    2. Diagnostic Target: ALB instance

    3. Description: If you suspect an Application Load Balancer (ALB) instance has issues such as packet loss, connectivity problems, an abnormal health status, overdue payments, or security policy problems, use the Application Load Balancer to investigate its health, configuration, security, capacity, and costs. If any problems are identified, use the provided remediation suggestions to resolve them, restore services, and improve O&M efficiency.

    4. Diagnostic Entry: Application Load Balancer

  3. Network Load Balancer

    1. Diagnostic Product: Load Balancer

    2. Diagnostic Target: NLB instance

    3. Description: If you suspect a Network Load Balancer (NLB) instance has issues such as packet loss, connectivity problems, an abnormal health status, overdue payments, or security policy problems, use the Network Load Balancer to investigate its health, configuration, security, capacity, and costs. If any problems are identified, use the provided remediation suggestions to resolve them, restore services, and improve O&M efficiency.

    4. Diagnostic Entry: Network Load Balancer

  4. NAT Gateway

    1. Diagnostic Product: NAT Gateway

    2. Diagnostic Target: NAT Gateway instance

    3. Description: If you suspect a NAT Gateway instance has issues such as packet loss, slow speeds, connectivity problems, insufficient quota, overdue payments, or security policy problems, use the NAT Gateway to investigate its health, configuration, security, capacity, and costs. If any problems are identified, use the provided remediation suggestions to resolve them, restore services, and improve O&M efficiency.

    4. Diagnostic Entry: NAT Gateway

  5. Elastic IP Address

    1. Diagnostic Product: Elastic IP Address (EIP)

    2. Diagnostic Target: EIP instance

    3. Description: If you suspect an Elastic IP Address (EIP) instance has issues such as packet loss, connectivity problems, insufficient bandwidth, overdue payments, or security policy problems, use the Elastic IP Address to investigate its health, configuration, security, capacity, and costs. If any problems are identified, use the provided remediation suggestions to resolve them, restore services, and improve O&M efficiency.

    4. Diagnostic Entry: Elastic IP Address

  6. Global Accelerator

    1. Diagnostic Product: Global Accelerator (GA)

    2. Diagnostic Target: GA instance

    3. Description: If you suspect a Global Accelerator (GA) instance has issues such as packet loss, connectivity problems, insufficient bandwidth, overdue payments, or security policy problems, use the Global Accelerator to investigate its health, configuration, security, capacity, and costs. If any problems are identified, use the provided remediation suggestions to resolve them, restore services, and improve O&M efficiency.

    4. Diagnostic Entry: Global Accelerator

  7. VPN Gateway

    1. Diagnostic Product: VPN Gateway

    2. Diagnostic Target: VPN Gateway instance

    3. Description: If you suspect a VPN Gateway instance has issues such as packet loss, slow speeds, connectivity problems, insufficient bandwidth, overdue payments, or security policy problems, use the VPN Gateway to investigate its health, configuration, security, capacity, and costs. If any problems are identified, use the provided remediation suggestions to resolve them, restore services, and improve O&M efficiency.

    4. Diagnostic Entry: VPN Gateway

  8. Virtual Border Router

    1. Diagnostic Product: Express Connect

    2. Diagnostic Target: Virtual Border Router (VBR) instance

    3. Description: If you suspect a Virtual Border Router (VBR) instance has issues such as packet loss, connectivity failures, an abnormal health status, overdue payments, or security policy problems, use the Virtual Border Router to investigate its health, configuration, security, capacity, and costs. If any problems are identified, use the provided remediation suggestions to resolve them, restore services, and improve O&M efficiency.

    4. Diagnostic Entry: Virtual Border Router

  9. Transit Router

    1. Diagnostic Product: Transit Router (TR)

    2. Diagnostic Target: TR instance

    3. Description: If you suspect a Transit Router (TR) instance has issues such as packet loss, connectivity failures, slow speeds, overdue payments, or security policy problems, use the Transit Router to investigate its health, configuration, security, capacity, and costs. If any problems are identified, use the provided remediation suggestions to resolve them, restore services, and improve O&M efficiency.

    4. Diagnostic Entry: Transit Router

  10. PrivateLink Endpoint

    1. Diagnostic Product: PrivateLink

    2. Diagnostic Target: PrivateLink endpoint instance

    3. Description: If you suspect a PrivateLink endpoint instance has issues such as packet loss, connectivity failures, slow speeds, or overdue payments, use the PrivateLink Endpoint to investigate its health, configuration, security, capacity, and costs. If any problems are identified, use the provided remediation suggestions to resolve them, restore services, and improve O&M efficiency.

    4. Diagnostic Entry: PrivateLink Endpoint

  11. PrivateLink Endpoint Service

    1. Diagnostic Product: PrivateLink

    2. Diagnostic Target: PrivateLink endpoint service instance

    3. Description: If you suspect a PrivateLink endpoint service instance has issues such as configuration problems or overdue payments, use the PrivateLink Endpoint Service to investigate its health, configuration, and costs. If any problems are identified, use the provided remediation suggestions to resolve them, restore services, and improve O&M efficiency.

    4. Diagnostic Entry: PrivateLink Endpoint Service

MaxCompute

  1. MaxCompute job diagnosis

    1. Product: MaxCompute, a cloud-native big data computing service

    2. Target: MaxCompute SQL/SQLRT jobs

    3. Description: If you suspect performance issues in a MaxCompute job, use the MaxCompute job diagnosis tool to quickly troubleshoot common problems such as data skew, resource contention, and data inflation. It also detects anomalies such as insufficient permissions and runtime errors. Follow the provided troubleshooting suggestions to resolve these issues, restore services, and improve O&M efficiency.

    4. Access: MaxCompute job diagnosis

Database

  1. RDS whitelist check

    1. Diagnostic product: ApsaraDB for RDS

    2. Diagnostic target: RDS instance

    3. Description: If you cannot connect to an RDS instance, use the RDS whitelist check tool to quickly check multiple private or public IPs against the RDS instance whitelist. You can also add IPs to the whitelist with a single click. For more solutions to RDS connection issues, see Troubleshoot instance connection issues.

    4. Diagnostic entry: RDS whitelist check

  2. RDS resource anomaly

    1. Diagnostic product: ApsaraDB for RDS

    2. Diagnostic target: RDS for MySQL instance

    3. Description: If you notice or suspect issues such as high space utilization, CPU utilization, memory utilization, or IOPS utilization on an RDS for MySQL instance, use the RDS resource anomaly tool to quickly diagnose the instance. If the tool finds an anomaly, follow its recommendations to resolve the issue, restore services, and improve operational efficiency.

    4. Diagnostic entry: RDS resource anomaly

  3. PolarDB resource anomaly

    1. Diagnostic product: ApsaraDB for PolarDB

    2. Diagnostic target: PolarDB for MySQL instance

    3. Description: If you notice or suspect issues such as high space utilization, CPU utilization, memory utilization, or IOPS utilization on a PolarDB for MySQL instance, use the PolarDB resource anomaly tool to quickly diagnose the instance. If the tool finds an anomaly, follow its recommendations to resolve the issue, restore services, and improve operational efficiency.

    4. Diagnostic entry: PolarDB resource anomaly

Other

  1. Website Unavailable

    1. Products: Multiple products

    2. Diagnostic object: Website domain name

    3. Description: If a website hosted on Alibaba Cloud is unavailable or has access issues, use the Website Unavailable diagnostic tool to quickly troubleshoot causes such as connectivity failures, access anomalies, or services being blocked. By entering the website's domain name, the tool automatically analyzes its dependent cloud resources (limited to resources under your account), such as Server Load Balancer (SLB), Elastic IP Address (EIP), and ECS, to identify the root cause. If the tool detects an issue, follow the provided recommendations to resolve the problem and quickly restore your service.

    4. Diagnosis entry: Website Unavailable

  2. Network Path Connectivity

    1. Products: Multiple products

    2. Diagnostic objects: ECS, public IP address, vSwitch, load balancing, and more

    3. Description: If you suspect connectivity issues between network nodes such as ECS instances, public IP addresses, vSwitches, or load balancing resources, use the Network Path Connectivity tool to quickly diagnose the network path between a source and a destination. Select a source and destination, and the tool automatically analyzes the end-to-end network path and displays information for each node along the way. If the tool finds an issue, follow the recommendations to resolve it and restore your service. For example, if a route configuration is missing a next hop, you may need to add it; if the instance cannot connect to the internet, you may need to associate an Elastic IP Address; or if traffic is being dropped by a security group rule, you may need to adjust the rule. If no issues are found, you can also initiate a reverse path diagnosis with a single click to verify round-trip connectivity.

    4. Diagnosis entry: Network Path Connectivity

  3. RAM permission error

    1. Products: Multiple products

    2. Diagnostic object: RAM error message

    3. Description: If an API call returns a Resource Access Management (RAM) permission error, use the RAM permission error tool to find the cause by entering the request ID. The tool provides authorization recommendations based on system policies or custom policies to help you grant the necessary permissions and quickly restore your service.

    4. Diagnosis entry: RAM permission error

  4. API/SDK error

    1. Products: Multiple products

    2. Diagnostic object: Error message or request ID

    3. Description: If you receive an error when calling an Alibaba Cloud API or using an SDK, use the API/SDK error tool to find the cause by entering the error message or request ID. You can then use the provided error details and recommended actions to resolve the issue and restore your service.

    4. Diagnosis entry: API/SDK error

One-click diagnosis

Note

The One-click Diagnosis feature is now available for early access. It performs a comprehensive diagnosis of your cloud resources and helps you resolve issues in a single operation, eliminating manual troubleshooting. Join our DingTalk group (ID: 86570007290) to receive an invitation link.

One-click diagnosis

With a scenario-specific diagnosis, you first assess the problem and then use specific tools to resolve issues individually. In contrast, one-click diagnosis scans all your cloud resources in a single run. It checks each resource for issues, prioritizes them by severity, and provides recommendations so you can address them all from one place. One-click diagnosis is like a full-body health check that can identify both obvious and hidden problems, while a scenario-specific diagnosis is like a specialist consultation offering a more in-depth analysis. After a one-click diagnosis, you can run a scenario-specific diagnosis on any issue found. This "health check + specialist" workflow makes troubleshooting and problem resolution more efficient.

Diagnosis coverage

Category

Product name

Diagnostic object

Limitations

Compute (2)

ECS

instance

Only running instances are supported.

Simple Application Server

instance

Only running instances are supported.

Container (1)

Container Service for Kubernetes (ACK)

cluster

Only running clusters are supported.

Network and CDN (8)

Server Load Balancer

CLB/ALB/NLB instances

NAT Gateway

instance

Elastic IP Address

instance

Global Accelerator

instance

VPN Gateway

instance

Express Connect

Virtual Border Router instance

Cloud Enterprise Network

instance

PrivateLink

PrivateLink endpoint instances and endpoint service instances

Database (2)

ApsaraDB for RDS

instance

Only running MySQL instances are supported.

ApsaraDB for PolarDB

cluster

Only running MySQL-compatible clusters are supported.

One-click diagnosis

Note

This feature is currently available to beta users only. To get an invitation link, join our DingTalk discussion group (ID: 86570007290).

One-click diagnosis

Entry 1: Log on to the console. You can start a diagnosis from the sidebar on the console home (if the sidebar is collapsed, click image in the lower-right corner to expand it).

image

Entry 2: Log on to the console. You can start a diagnosis from the operations and monitoring card on the console home.

image

Entry 3: Log on to the console. You can start a diagnosis by navigating to console home - operations and monitoring and clicking Create Diagnosis.

image

The system displays the products and resources in your account that you can diagnose with a single click. By default, the resources on the first page for each product are selected. You can select which instances to diagnose. A single diagnostic task can include up to 50 resources per product. To diagnose more resources, start a new diagnostic task after the current one is complete.

image

Click Start Diagnosis to run a one-click diagnosis. You can monitor the progress of the overall diagnosis and for each individual resource. The entire process typically takes a few minutes.

image

View the diagnostic results once the diagnostics are complete.

image

If the diagnosis finds an issue, the affected resources are listed at the top. Click the arrow to view the issue details and suggested fixes. Follow the suggested fixes to resolve the issue. If the issue persists, submit a ticket.

Click "Helpful" or "Not Helpful" to provide feedback on the diagnostic results. We review all feedback to continuously improve this feature.

Note

Diagnostic result labels:

Normal: Indicates no significant issues. You can rule out this item during troubleshooting.

Information: Indicates a deviation from best practice. This does not currently affect your service, and addressing it is optional.

Warning: Indicates a significant issue that may affect your service. You should resolve it as soon as possible.

Critical: Indicates a major issue that will likely impact your service. This requires immediate action.

Failed: Indicates the diagnosis failed due to an unexpected error. You can re-diagnose.