Configures the expansion mode of communication operators.
The following lists supported configurations and corresponding scenario descriptions for each product model. Products not listed do not support this environment variable. If an unsupported environment variable is set, the default value is used.
This configuration only supports the following operators: Broadcast, Reduce, AllReduce, Scatter, ReduceScatter, AllGather, AlltoAll, AlltoAllV, and AlltoAllVC.
Notes:
In graph mode (Ascend IR) or graph capture (aclgraph) scenarios, if the communication algorithm uses the AI_CPU mode, the number of concurrent graphs on a single device cannot exceed 6. Otherwise, communication congestion may occur due to full occupation of AI CPU cores.
If the number of Vector Cores allocated during service compilation fails to meet the algorithm orchestration requirements, HCCL will throw an error and prompt the minimum number of Vector Cores required.
In MS mode, the on-chip CCU MS serves as a transfer buffer for communication with multiple remote ends to reduce memory read/write bandwidth consumption. The MS features small capacity but high access speed.
The system will automatically fall back to the AI_CPU mode when CCU resources are insufficient.
In this mode, the CCU acts as a scheduler to schedule UB WQE tasks to the UB engine. The on-chip CCU MS is not used; data is transmitted directly between HBM of two ranks.
For the AllReduce, ReduceScatter, and Reduce operators in single-server communication, once the data volume exceeds a certain threshold, the system will automatically switch to the AI_CPU mode to avoid performance degradation. This threshold is not fixed and varies with factors such as the operator execution mode and network scale.
The system will automatically fall back to the AI_CPU mode when CCU resources are insufficient.
Full communication operators are supported within a supernode and between supernodes. For the Reduce, ReduceScatter, ReduceScatterV, and AllReduce operators, the data type can only be int8, int16, int32, float16, float32, or bfp16, and the reduce operation type can only be sum, max, or min. For details about the data types supported by other communication operators, see the corresponding collective communication APIs.
Notes:
AI CPU cache allows HCCL to reuse the execution results generated during the first run of a communication operator when the same operator is executed again, reducing operator expansion overhead. Enabling this feature incurs additional device memory overhead. Therefore, in service scenarios with frequently changing communication data sizes, disable this feature to reduce device memory consumption.
If the number of Vector Cores allocated during service compilation fails to meet the algorithm orchestration requirements, HCCL will throw an error and prompt the minimum number of Vector Cores required.
Notes:
When the algorithm orchestration expansion location is set to AIV, if the HCCL_DETERMINISTIC environment variable is set to true or strict to enable deterministic computation, deterministic computation takes higher priority, and the AIV expansion may fail to take effect in certain scenarios.
This option supports only the AllGather, AlltoAll, AlltoAllV, and AlltoAllVC operators.
Notes:
In graph mode (Ascend IR) or graph capture (aclgraph) scenarios, if the communication algorithm uses the AI_CPU mode, the number of concurrent graphs on a single device cannot exceed 6. Otherwise, communication congestion may occur due to full occupation of AI CPU cores.
If the number of Vector Cores allocated during service compilation fails to meet the algorithm orchestration requirements, HCCL will throw an error and prompt the minimum number of Vector Cores required.
Notes:
export HCCL_OP_EXPANSION_MODE="HOST"
1 2 3 4 5 | [ERROR] KERNEL(5044,sklogd):2024-07-29-10:33:22.646.254 [klogd.c:247][257382.266115] [ascend] [ERROR] [devmm] [devmm_page_fault_d2h_query_flag 810] <kworker/u16:2:14887,14887> Host page fault send message fail.(hostpid=2131021; devid=0; vfid=0; ret=-22; va=0x12c700300000; hostpid=2131021; devid=0; vfid=0) [ERROR] KERNEL(5044,sklogd):2024-07-29-10:33:22.646.284 [klogd.c:247][257382.266124] [ascend] [ERROR] [devmm] [devmm_svm_device_fault 468] <kworker/u16:2:14887,14887> Vm fault failed. (hostpid=2131021; devid=0; vfid=0; ret=64; fault_addr=0x12c700300000; start=0x12c700300000) [ERROR] KERNEL(5044,sklogd):2024-07-29-10:33:22.659.429 [klogd.c:247][257382.282181] [ascend] [ERROR] [tsdrv] [ipc_fault_msg_para_check 309] <swapper/3:0> Invalid node id. (devid=0; node_type=100; node_id=40; node_num=25) ................ [ERROR] KERNEL(5044,sklogd):2024-07-29-10:33:24.874.211 [klogd.c:247][257384.473533] [ascend] [ERROR] [tsdrv] [tsdrv_hb_cq_callback 332] <kworker/0:0:20353> receive ts exception msg, call excep_code=0xb4060006, time=1722249204.850014098s, devid=0 tsid=0 |