HCCL provides multiple fault detection functions, including the link setup fault detection time configuration, cluster heartbeat monitoring switch, and process suspension detection switch. After these detection functions are enabled, fault information can be quickly located and displayed when service exceptions occur, facilitating troubleshooting and handling.
The following options are supported:
When the link setup times out, HCCL starts locating the root node where the link setup fails and propagates the information about the root node. The entire process takes the time specified by the connection_fault_detection_time parameter plus 10s required for propagating the root node information.
The value of connection_fault_detection_time can be 0 or in the range of [20, 7200]. The unit is second and the default value is 20.
If this parameter is set to 0, the link setup fault detection function is disabled. That is, when a connection fails to be set up, there is no extra waiting time and the link setup process exits immediately.
This parameter can be set to on (indicating to enable the heartbeat monitoring function) or off (indicating to disable the heartbeat monitoring function). The default value is on.
Note: After the cluster heartbeat monitoring function is disabled, communication operation timeouts cannot be detected, the cluster fault propagation capability will be lost, and root node fault information will not be recorded in runtime logs.
The value can be on (indicating to enable the process suspension detection function) or off (indicating to disable the process suspension detection function). The default value is on.
In scenarios that are sensitive to communication performance, you can use this parameter to disable the process suspension detection function. However, after the process suspension detection function is disabled, service suspension faults are not proactively detected and reported.
The value can be on (indicating to enable the operator delivery inconsistency detection function) or off (indicating to disable the operator delivery inconsistency detection function). The default value is off.
You can use this parameter to enable the operator delivery inconsistency detection function, but the performance will deteriorate to some extent. Note that by default, after this function is disabled, the system does not proactively detect and record operator delivery inconsistency issues.
Note: This function does not support the HcclBatchSendRecv operator and graph mode scenarios. After this function is enabled, data cache is generated, occupying the host memory.
Notes:
The detection capability provided by this environment variable is only for cluster fault localization. The detected events may not correspond to the root causes of service failures in complex scenarios. Confirm the faulty root node by combining the generation time of detected events and specific error information on the monitored node.
export HCCL_DFS_CONFIG="connection_fault_detection_time:30,cluster_heartbeat:on,stuck_detection:on,inconsistent_check:off,task_monitor_interval:0"
None