Configure node self-healing notifications
Configure node self-healing message notifications so you receive alerts when abnormal nodes are detected in Lingjun intelligent computing resources. After you receive a notification, stop or close the workloads on the affected node as soon as possible so the node self-healing process can complete.
How node self-healing notifications work
When the system detects an abnormal node, it automatically switches to a standby node to keep your workloads available. Notifications are sent for the following events:
-
Node scheduling blocked: The system identifies an abnormal node and temporarily blocks scheduling on it.
-
Node self-healing blocked: Running jobs on the abnormal node are blocking self-healing. Take the following actions:
-
DSW instances: Save the environment and close the instance manually, or configure automatic restart in Scheduling Settings.
-
DLC jobs: Stop the jobs manually.
-
-
Node self-healing started: The system has started the node self-healing process.
-
Node self-healing completed: The system has completed the node self-healing process.
The following table summarizes the events and the actions you might need to take.
|
Event |
What it means |
Action required |
|
Node scheduling blocked |
The system identifies an abnormal node and temporarily blocks scheduling on it. |
None |
|
Node self-healing blocked |
Running jobs on the abnormal node are blocking self-healing. |
For DSW instances, save the environment and close the instance, or configure automatic restart. For DLC jobs, stop the jobs. |
|
Node self-healing started |
The system has started the node self-healing process. |
None |
|
Node self-healing completed |
The system has completed the node self-healing process. |
None |
Limitations
This feature is available only for Lingjun intelligent computing resources.
Set up message notifications
China site
You can receive node self-healing notifications through in-console notifications, email , SMS, or chatbots. To make sure you receive these notifications in time, confirm that the following settings are enabled in the Alibaba Cloud Message Center.
SMS, email, and in-console notifications
-
Log in to the PAI console.
-
Click
in the upper-right corner to open Message Center. -
In the left-side navigation pane, choose Message receiving management > Basic receiving management.
-
In the Message type column, find Product O&M notifications, and select In-console notification, Email, or SMS.
-
The default notification recipient is the Alibaba Cloud primary account. You can also click Modify in the Actions column to configure more contacts.
-
When an abnormal node is detected, the system notifies you of the affected node name, resource quota, and jobs running on the node after the configuration is complete.
Chatbot notifications
The Alibaba Cloud Message Center chatbot receiving platform currently supports DingTalk, WeCom, Lark, and Slack. For more information, see Alibaba Cloud Message Center.
-
Log in to the PAI console.
-
Click
in the upper-right corner to open Message Center. -
In the left-side navigation pane, choose Message receiving management > Bot receiving management.
-
Add a chatbot. Skip this step if you have already added one.
-
In the upper-right corner of the Bot receiving management page, click Bot management.
-
On the Bot management page, follow the platform links at the top to get the chatbot webhook information, and then click Add chatbot to complete the configuration.
-
Click Test in the Actions column to test chatbot connectivity. A dialog showing Test successful indicates that the connection is normal.
-
-
On the Bot receiving management page, find Product O&M notifications in the Message type column, and click Modify in the Actions column.
-
Switch to the Message receiving chatbot tab, select the added chatbot, and click Save.
Custom notification rules
You can configure different chatbots and notification rules based on your business requirements. For example, if one team cares only about PAI messages and another team needs all Alibaba Cloud product O&M notifications, configure two chatbots:
-
Chatbot test1: receives all messages. No notification rule is required.
-
Chatbot test2: receives only messages that contain the PAI keyword. To configure this rule:
-
In the left-side navigation pane of the Message Center page, choose Message receiving management > Bot receiving management.
-
Find Product O&M notifications in the Message type column and click Modify in the Actions column.
-
Switch to the Message receiving chatbot tab. In the Notification rule column of the target recipient, click Edit. On the Configure custom notification rule page, set an allowlist keyword, such as PAI. After this configuration, you receive only PAI-related notifications.
You can also select multiple recipients and click Batch edit rules to configure rules in batches.
-
After the configuration is complete, click OK, and then click Save on the Modify message receiving configuration page. After the configuration succeeds, Enabled allowlist is displayed in the Notification rule column.
Respond to a blocked self-healing notification
After you receive a node self-healing blocked notification, follow these steps to close the DSW instances and stop the DLC jobs on the abnormal node so that the node replacement can proceed.
|
Resource on the abnormal node |
Required action |
Where to find the steps |
|
DSW instance |
Save the environment and close the instance, or enable automatic restart. |
Migrate DSW instances |
|
DLC job |
Stop the job and resubmit it. |
Stop DLC jobs |
Migrate DSW instances
Manual migration
For DSW instances on an abnormal node, if you have the instance open in the browser, a dialog prompts you to save the environment and close the instance as soon as possible to ensure node self-healing.
Automatic instance migration
Use Automatic Instance Migration (from Abnormal Node) to migrate DSW instances automatically when a node becomes abnormal.
-
Log in to the PAI console.
-
In the left-side navigation pane, click Workspaces, find the workspace in the list, and click its name to open it.
-
On the Workspace Details page, choose Workspace configuration Scheduling Settings.
-
In the DSW section, turn on the Enable Automatic Instance Migration from Abnormal Node switch.
After this feature is enabled, when a node in the underlying Lingjun infrastructure becomes abnormal, the system automatically closes and restarts the instance to support node self-healing and keep resources available. The restart process saves the environment image, but running processes cannot be recovered.
For DSW instances on an abnormal node, if you have the instance open in the browser, a dialog prompts you to save the environment and close the instance. The dialog also displays the remaining time before automatic restart so that node self-healing can proceed.
Stop DLC jobs
-
Click the details link in the in-console notification, email , or SMS to go to the resource quota page.
-
Based on the node information provided, view the job list under the node. On the Nodes tab of the resource quota page, click Details in the Job count column of the target node.
-
Click the DLC job name to go to the job details page. Then click More > Stop to stop the DLC job.
-
Click Clone. Your job reuses the original configuration and is scheduled onto a healthy node. For details, see Clone a job.