In cluster networks featuring hierarchical topologies with inter-layer bandwidth convergence, collective communication confronts two primary technical challenges. First, traditional single-level collective algorithms suffer from severe performance degradation induced by cross-region bandwidth convergence. Second, conventional hierarchical algorithms become inapplicable due to the asymmetric distribution of compute units—specifically, device-count asymmetry—across different regions. For example, in a cluster, a communicator may span two SuperPoDs with different numbers of devices (for example, one SuperPoD has 64 devices and the other has 128 devices). This poses great challenges to the performance of the collective communication algorithm.
In the preceding figure, AHC regroups NPUs within a communicator and partition the data on each NPU in a topology-aware manner. It fully uses the high-speed network bandwidth for intra-group communication, while enabling an asymmetric combination across groups leveraging logical devices with identical IDs. The implementation involves the following three steps:
The algorithm of the intra-group and inter-group ReduceScatter, AllGather, and AllReduce operations can be any known algorithm, such as NB, NHR, and Ring. Currently, the AHC algorithm selects a stitching algorithm with better performance based on the scenario and policy.
When the NB algorithm is used for both intra-group and inter-group operations, the time required by the AllReduce operator is as follows.
|
Operation |
Time Required |
|---|---|
|
AllReduce |
In the formula, m indicates the minimum number of groups, m+d indicates the maximum number of groups, G indicates the total number of groups, and C indicates the convergence ratio of the inter-group bandwidth to the intra-group bandwidth. |