Features
High Availability Service (HAS) focuses on IT risk prevention and control. Its main features include risk management, routine inspections, fault diagnosis, emergency plans, and fault drills.
Risk management
Risk management is the core of HAS. It is the central platform for collecting and handling risk events.
Risk events
Risk event collection: Collects risk and alert information from monitoring, inspections, and diagnostics.
Risk event handling: You can handle risk events directly from the risk event list. When handling a risk event, the risk management feature recommends executable emergency plans and lets you trigger them directly.
Risk scenarios
A risk scenario is a module for centrally handling specific risk events. A risk scenario contains the information needed to handle a risk event, such as a diagnostic decision tree, an emergency plan, and business impact information. After the emergency scenario feature is upgraded, risk scenarios must be linked with emergency responses, which requires adding more properties.
Routine inspections
Routine inspections are the most frequently used feature of HAS. You can use this feature to automatically check system stability and availability on a regular basis. The inspection results are pushed in real time to a specified DingTalk group. This helps O&M engineers identify application risks as soon as possible. The feature also generates inspection reports that O&M engineers can archive. Inspection plugins support multiple types, such as python, Shell, automated testing images, and page health checks. You can customize inspection plugins based on your application's needs. HAS also provides out-of-the-box inspection rules that are based on the long-term experience of Ant Group and other users.
Fault diagnosis
The core capability of fault diagnosis is to codify and display the troubleshooting procedures and expertise of O&M engineers.
O&M engineers can graphically orchestrate the fault diagnosis process and design a troubleshooting sequence using a decision tree. When a risk event occurs, this standardized process is automatically executed through the fault decision tree, and the diagnosis result is provided directly. The fault diagnosis platform greatly shortens troubleshooting time and eliminates discrepancies caused by different levels of experience and skill among O&M engineers. This enables quick fault identification.
Emergency plans
Emergency plans allow you to orchestrate atomic O&M operations, such as application restarts, removing applications from traffic, database switchovers, and physical server restarts.
O&M engineers can select and orchestrate the required atomic operations based on procedures for handling common fault scenarios to create an executable emergency plan. When a risk event occurs, the Event Center recommends executable emergency plans. O&M engineers can then quickly select and automatically execute a plan. This standardized process enables rapid fault recovery.
Fault drills
Fault drills provide a fault injection capability. You can use the drill platform to proactively trigger faults and test the high availability of your application.
The fault drill platform can trigger common faults, such as high CPU and memory utilization, network packet loss, container breakdowns, and physical machine breakdowns. You can create detailed drill and recovery plans for these faults. This lets you systematically measure and observe your application's high availability.
