Description: Quantizes token data (optional). When there is TP domain communication, AllToAllV communication in the EP domain is performed first, and then AllGatherV TP domain communication is performed. When there is no such communication, AllToAllV communication in the EP domain is performed.
Compared with the
[object Object]API, this API has the following changes:Dynamic scale-in support: The operator can run properly without recompilation after faulty ranks are removed from the created communication domain. Enable this feature by passing the
[object Object]parameter.Special expert scenarios are supported:
[object Object]: Enabled by setting the[object Object]parameter to a value greater than 0.[object Object]: Enabled by setting the[object Object]parameter to a value greater than 0 and setting a valid value for the[object Object]parameter.[object Object]: Enabled by setting the[object Object]parameter to a value greater than 0 and setting valid values for the[object Object],[object Object],[object Object], and[object Object]parameters.
For details, see the following parameter description. For details about the
[object Object],[object Object],[object Object], and[object Object]parameters, see the[object Object]file.
Formula:
- [object Object]Atlas A2 training products/Atlas A2 inference products[object Object]: This API must be used together with
[object Object]. - [object Object]Atlas A3 training products/Atlas A3 inference products[object Object]: This API must be used together with
[object Object]or[object Object].
- [object Object]Atlas A2 training products/Atlas A2 inference products[object Object]: This API must be used together with
[object Object]
Each operator has calls. First, [object Object] is called to obtain the workspace size required for computation and the executor that contains the operator computation process. Then, [object Object] is called to perform computation.
Parameters
[object Object][object Object][object Object]- The value of
[object Object]can be[object Object],[object Object],[object Object], or[object Object]. It is recommended to use[object Object]with driver version 25.0.RC1.1 or later. When set to[object Object]or[object Object], the communication algorithm is selected based on HCCL environment variables (not recommended).[object Object]indicates that tokens are directly transmitted through RDMA.[object Object]indicates a two-stage communication process: intra-server communication followed by inter-server communication, which reduces cross-server data transmission. [object Object]must be passed as a null pointer when[object Object]is[object Object], or when HCCL_INTRA_PCIE_ENABLE=1 and HCCL_INTRA_ROCE_ENABLE=0[object Object]depends on the[object Object]value: For[object Object], it requires a 1D tensor with shape (Bs, ), where[object Object]must precede[object Object](for example, {true, false, true} is invalid); for[object Object], it is currently not supported and a null pointer should be passed.- The value of
[object Object]must be a 2D tensor with the shape of (Bs, K). - The value of
[object Object]depends on the[object Object]value: For[object Object], it supports 16, 32, 64, 128, 192, and 256; for[object Object], it supports 16, 32, and 64. - The value of
[object Object]must be in the range (0, 512] and satisfy moeExpertNum / (epWorldSize - sharedExpertRankNum) ≤ 24. [object Object]is not supported in the current version. Pass an empty string.- The current version does not support
[object Object],[object Object],[object Object],[object Object], and[object Object]. Pass 0 for these parameters. - The shape of
[object Object]is (moeExpertNum + 2globalBsK × serverNum,). (The first[object Object]elements indicate the number of received tokens, and the remaining elements indicate the[object Object]information before communication.) - Currently, TP domain communication is not supported.
[object Object]must be a 1D tensor with shape (A ).[object Object]supports 0 (non-quantization) and 2 (dynamic quantization).[object Object]is not supported in the current version. Pass a null pointer.- When
[object Object]is[object Object], the value of[object Object]must be in the range [0, MAX_INT32), where MAX_INT32 = 2^31 - 1. Valid zero expert IDs must be in the range [[object Object],[object Object]). - When
[object Object]is[object Object], the value of[object Object]must be in the range [0, MAX_INT32), where MAX_INT32 = 2^31 - 1. Valid copy expert IDs must be in the range [[object Object],[object Object]). [object Object]is not supported in the current version. Pass 0.[object Object]
[object Object]supports three input modes: "", "fullmesh_v1", and "fullmesh_v2".- "" (default): The
[object Object]template is disabled. - "fullmesh_v1": The
[object Object]template is disabled. - "fullmesh_v2": The
[object Object]template is enabled. The[object Object]template takes effect only when[object Object]is set to 1 and cannot be enabled in scenarios where[object Object]values of each rank are inconsistent, dynamic scale-in is performed,[object Object]is input, or special experts are used.
- "" (default): The
[object Object]must be a 1D tensor with shape (Bs, ) or a 2D tensor with shape (Bs, K). If it is a 1D tensor,[object Object]must be placed before[object Object]. If it is a 2D tensor and the K values corresponding to tokens are all[object Object], the tokens do not participate in communication.[object Object]is not supported in the current version. Pass a null pointer.- The value of
[object Object]must be in the range [2, 768]. - The value of
[object Object]must be in the range (0, 1024]. [object Object]must be a string of length [0, 128) and cannot be the same as[object Object]. This parameter can be left empty only when there is no TP domain communication.- The value of
[object Object]must be in the range [0, 2]. 0 and 1 indicate no TP domain communication. 2 is required when TP domain communication is used. - The value of
[object Object]must be in the range [0, 1].[object Object]of each rank in the same TP domain must be unique. If TP domain communication is not used, pass 0. - The value of
[object Object]must be 0, indicating that shared expert ranks are placed in front of MoE expert ranks. - The value of
[object Object]must be in the range [0, 4]. - The value of
[object Object]must be in the range [0, epWorldSize). If the value is 0,[object Object]is 0 or 1. If the value is not 0,[object Object]is 0. - The shape of
[object Object]is (epWorldSize × max(tpWorldSize, 1) × localExpertNum, ). - When there is TP domain communication,
[object Object]is a 1D tensor with shape (tpWorldSize, ). [object Object]is not supported in the current version.[object Object]supports 0 (non-quantization) and 2 (dynamic quantization).- You can pass valid data or a null pointer for
[object Object]. If a null pointer is passed, the dynamic scale-in feature is disabled. If valid data is passed, it must be a 1D tensor with shape[object Object]. The first four numbers in the tensor indicate: whether scale-in is performed, the actual number of ranks after scale-in, the number of ranks used by shared experts after scale-in, and the number of MoE experts after scale-in. The remaining 2 ×[object Object]indicates two rank mapping tables. After scale-in, some ranks on the current device may be removed from the EP communication domain due to failures. The mapping for the first table is Table1[epRankId]=localEpRankId or Table1[epRankId]=-1. Here,[object Object]denotes the rank index in the new EP communication domain, and[object Object]indicates that the rank with the corresponding[object Object]has been removed from the communication domain. The mapping for the second table is Table2[localEpRankId] = epRankId . - The value of
[object Object]must be in the range [0, MAX_INT32), where MAX_INT32 = 2^31 - 1. Valid zero expert IDs must be in the range[object Object]. When[object Object]is set to "fullmesh_v2", this parameter is not supported in the current version. In this case, set this parameter to 0. - The value of
[object Object]must be in the range [0, MAX_INT32), where MAX_INT32 = 2^31 - 1. Valid expert IDs must be in the range[object Object]. When[object Object]is set to "fullmesh_v2", this parameter is not supported in the current version. In this case, set this parameter to 0. - The value of
[object Object]must be in the range [0, MAX_INT32), where MAX_INT32 = 2^31 - 1. Valid expert IDs must be in the range[object Object]. When[object Object]is set to "fullmesh_v2", this parameter is not supported in the current version. In this case, set this parameter to 0.
- The value of
Returns:
[object Object]: status code. For details, see .The first-phase API implements input parameter verification. The following errors may be thrown.
[object Object]
Deterministic computing:
[object Object]defaults to a deterministic implementation.
API constraints:
[object Object]and CombineV3 operators must be used together. The[object Object],[object Object],[object Object], and[object Object]outputs of[object Object]must be directly passed to the corresponding parameters of[object Object]. The service logic cannot depend on the specific values of these tensors.
Parameter consistency constraints:
- The values of
[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object],[object Object], and[object Object]must be consistent across all ranks, and be the same as the value of CombineV3. - The deployment information after dynamic scale-in is transferred to the operator through the
[object Object]parameter. Other parameters do not need to be modified. The scale-in parameters take effect only when[object Object]is set to 1. After dynamic scale-in, the number of MoE experts deployed on the current rank must be the same as that before scale-in. Configurations where no MoE expert ranks remain after scale-in are not supported.
- The values of
Product constraints:
- [object Object]Atlas A3 training products/Atlas A3 inference products[object Object]: In this scenario, a single rank contains dual dies. Therefore, the "rank" in the parameter description indicates a single die.
- The dynamic scale-in feature cannot be enabled in the tensor parallelism scenario.
Shape variable constraints:
[object Object]undefined
Environment variables constraints:
HCCL_BUFFSIZE:
Before calling this API, check whether the value of the
[object Object]environment variable is proper. This environment variable indicates the buffer size occupied by a single communication domain, in MB. If this environment variable is not set, the default value 200 MB is used.- [object Object]Atlas A2 training products/Atlas A2 inference products[object Object]:
- If
[object Object]is set to[object Object]or[object Object], select the[object Object]or[object Object]formula based on the[object Object]and[object Object]environment variables. - If
[object Object]is set to[object Object], the value requirement is (≥ 2 × (Bs × epWorldSize × min(localExpertNum, K) × H × sizeof(uint16) + 2 MB)). - If
[object Object]is set to[object Object], the value must be[object Object], and[object Object]is not required.
- If
- [object Object]Atlas A3 training products/Atlas A3 inference products[object Object]:
- In an EP communication domain, when
[object Object]is[object Object], a null character string, or a null pointer, the value must satisfy ≥ 2 × (localExpertNum × maxBs × epWorldSize × Align512(Align32(2 × H) + 64) + (K + sharedExpertNum) × maxBs × Align512(2 × H)). - In an EP communication domain, when
[object Object]is[object Object], the value must satisfy ≥ 2 × (localExpertNum × maxBs × epWorldSize × 480Align512(Align32(2 × H) + 64) + (K + sharedExpertNum) × maxBs × Align512(2 × H)). - In a TP communication domain, the value must satisfy ≥ (A × Align512(Align32(h × 2) + 44) + A × Align512(h × 2)) × 2.
- 480Align512(x) = ((x + 480 - 1) / 480) × 512, Align512(x) = ((x + 512 - 1) / 512) × 512 and Align32(x) = ((x + 32 - 1) / 32) × 32.
- In an EP communication domain, when
- [object Object]Atlas A2 training products/Atlas A2 inference products[object Object]:
HCCL_INTRA_PCIE_ENABLE and HCCL_INTRA_ROCE_ENABLE:
- [object Object]Atlas A2 training products/Atlas A2 inference products[object Object]: This environment variable is not recommended. You are advised to set
[object Object]to[object Object].
- [object Object]Atlas A2 training products/Atlas A2 inference products[object Object]: This environment variable is not recommended. You are advised to set
Constraints on the use of the communication domains:
[object Object]and[object Object]in a model support only the same EP communication domain, and no other operators are allowed in the communication domain.[object Object]and[object Object]in a model support only the same TP communication domain or both do not support a TP communication domain. If a TP communication domain is supported, no other operators are allowed in the communication domain.- [object Object]Atlas A3 training products/Atlas A3 inference products[object Object]: Nodes in a communication domain must be in the same SuperPoD. Cross-SuperPoD nodes are not supported.
Networking constraints:
- [object Object]Atlas A2 training products/Atlas A2 inference products[object Object]: In multi-server scenarios, only switch-based networking is supported, and direct point-to-point networking between two servers is not supported.
Other constraints:
- In the formulas, / denotes integer division.
- [object Object]moeExpertNum + zeroExpertNum + copyExpertNum + constExpertNum < MAX_INT32[object Object]
[object Object]Atlas A2 training products/Atlas A2 inference products[object Object]: Similar to the following example for [object Object]Atlas A3 training products/Atlas A3 inference products[object Object]. For the new scenario parameters of V3 compared with V2, set the parameter values based on the preceding parameter description.
[object Object]Atlas A3 training products/Atlas A3 inference products[object Object]: The sample code is as follows (for reference only). Call the [object Object] and [object Object] APIs.
Preparing files:
Create a
[object Object]directory. Follow the instructions to create[object Object]and[object Object]files in the[object Object]directory, and modify them according to the code.Install the CANN package and compile and run
[object Object].
Compilation script:
[object Object]Compilation and execution:
[object Object]The sample code is as follows:
[object Object]