HCCL_DFS_CONFIG

Description

HCCL provides multiple fault detection functions, including the link setup fault detection time configuration, cluster heartbeat monitoring switch, and process suspension detection switch. After these detection functions are enabled, fault information can be quickly located and displayed when service exceptions occur, facilitating troubleshooting and handling.

The following options are supported:

The detection capability provided by this environment variable is only for cluster fault localization. The detected events may not correspond to the root causes of service failures in complex scenarios. Confirm the faulty root node by combining the generation time of detected events and specific error information on the monitored node.

Example

export HCCL_DFS_CONFIG="connection_fault_detection_time:30,cluster_heartbeat:on,stuck_detection:on,inconsistent_check:off,task_monitor_interval:0"

Constraints

None

Applicability

Atlas A3 training product/Atlas A3 inference product

Atlas A2 training product/Atlas A2 inference product (For Atlas A2 training product/Atlas A2 inference product, only the Atlas 800T A2 training server, Atlas 900 A2 PoD cluster basic unit, and Atlas 200T A2 Box16 heterogeneous subrack are supported.)