[object Object]

[object Object][object Object]undefined
[object Object]
  • Description: Quantizes token data (optional). When there is TP domain communication, AllToAllV communication in the EP domain is performed first, and then AllGatherV TP domain communication is performed. When there is no such communication, AllToAllV communication in the EP domain is performed.

    Compared with the [object Object] API, this API has the following changes:

    Added the capability to collect the communication duration, that is, record the communication time of each rank. This function is enabled by passing the [object Object] parameter. It is recommended that this function be used together with . The communication duration per rank for each operator call is accumulated in this tensor. Clear it as needed before use.

  • Formula:

    agOut=AllGatherV(X)expandXOut=AllToAllV(agOut)agOut = AllGatherV(X)\\ expandXOut = AllToAllV(agOut)\\
    • [object Object]Atlas A2 training products/Atlas A2 inference products[object Object]: This API must be used together with [object Object].
    • [object Object]Atlas A3 training products/Atlas A3 inference products[object Object]: This API must be used together with [object Object] or [object Object].

    Note: [object Object] and [object Object] operators are collectively referred to as CombineV4 series operators in subsequent documents.

[object Object]

Each operator has calls. First, [object Object] is called to obtain the workspace size required for computation and the executor that contains the operator computation process. Then, [object Object] is called to perform computation.

[object Object]
[object Object]
[object Object]
  • Parameters

    [object Object][object Object][object Object]
    • The value of [object Object] can be [object Object], [object Object], [object Object], or [object Object]. It is recommended to use [object Object] with driver version 25.0.RC1.1 or later. When set to [object Object] or [object Object], the communication algorithm is selected based on HCCL environment variables (not recommended). [object Object] indicates that tokens are directly transmitted through RDMA. [object Object] indicates a two-stage communication process: intra-server communication followed by inter-server communication, which reduces cross-server data transmission.
    • [object Object] must be passed as a null pointer when [object Object] is [object Object], or when HCCL_INTRA_PCIE_ENABLE=1 and HCCL_INTRA_ROCE_ENABLE=0
    • [object Object] depends on the [object Object] value: For [object Object], it requires a 1D tensor with shape (Bs, ), where [object Object] must precede [object Object] (for example, {true, false, true} is invalid); for [object Object], it is currently not supported and a null pointer should be passed.
    • The value of [object Object] must be a 2D tensor with the shape of (Bs, K).
    • The value of [object Object] depends on the [object Object] value: For [object Object], it supports 16, 32, 64, 128, 192, and 256; for [object Object], it supports 16, 32, and 64.
    • The value of [object Object] must be in the range (0, 512] and satisfy moeExpertNum / (epWorldSize - sharedExpertRankNum) ≤ 24.
    • [object Object] is not supported in the current version. Pass an empty string.
    • The current version does not support [object Object], [object Object], [object Object], [object Object], and [object Object]. Pass 0 for these parameters.
    • The shape of [object Object] is (moeExpertNum + 2globalBsK × serverNum,). (The first [object Object] elements indicate the number of received tokens, and the remaining elements indicate the [object Object] information before communication.)
    • Currently, TP domain communication is not supported.
    • [object Object] must be a 1D tensor with shape (A ).
    • [object Object] supports 0 (non-quantization) and 2 (dynamic quantization).
    • [object Object] is not supported in the current version. Pass a null pointer.
    • When commAlg is [object Object], the value of [object Object] must be in the range [0, MAX_INT32), where MAX_INT32 = 2^31 - 1. Valid zero expert IDs must be in the range [[object Object], [object Object]).
    • When commAlg is [object Object], the value of [object Object] must be in the range [0, MAX_INT32), where MAX_INT32 = 2^31 - 1. Valid copy expert IDs must be in the range [[object Object], [object Object]).
    • [object Object] is not supported in the current version. Pass 0.
    • You can pass valid data or a null pointer for [object Object]. If you pass a null pointer, the function of recording the communication duration is disabled. If you pass valid data, it must be a 1D tensor with shape (ep_world_size,), and its data type and data format must be int64 and ND, respectively.
    [object Object][object Object][object Object]
    • [object Object] is not supported in the current version. Pass a null pointer.
    • [object Object] must be a 1D tensor with shape (Bs, ) or a 2D tensor with shape (Bs, K). If it is a 1D tensor, [object Object] must be placed before [object Object]. If it is a 2D tensor and the K values corresponding to tokens are all [object Object], the tokens do not participate in communication.
    • [object Object] is not supported in the current version. Pass a null pointer.
    • The value of [object Object] must be in the range [2, 768].
    • The value of [object Object] must be in the range (0, 1024].
    • [object Object] must be a string of length [0, 128) and cannot be the same as [object Object]. This parameter can be left empty only when there is no TP domain communication.
    • The value of [object Object] must be in the range [0, 2]. 0 and 1 indicate no TP domain communication. 2 is required when TP domain communication is used.
    • The value of [object Object] must be in the range [0, 1]. [object Object] of each rank in the same TP domain must be unique. If TP domain communication is not used, pass 0.
    • The value of [object Object] must be 0, indicating that shared expert ranks are placed in front of MoE expert ranks.
    • The value of [object Object] must be in the range [0, 4].
    • The value of [object Object] must be in the range [0, epWorldSize). If the value is 0, [object Object] is 0 or 1. If the value is not 0, [object Object] is 0.
    • The shape of [object Object] is (epWorldSize × max(tpWorldSize, 1) × localExpertNum, ).
    • When there is TP domain communication, [object Object] is a 1D tensor with shape (tpWorldSize, ).
    • [object Object] is not supported in the current version.
    • [object Object] supports 0 (non-quantization) and 2 (dynamic quantization).
    • You can pass valid data or a null pointer for [object Object]. If a null pointer is passed, the dynamic scale-in feature is disabled. If valid data is passed, it must be a 1D tensor with shape [object Object]. The first four numbers in the tensor indicate: whether scale-in is performed, the actual number of ranks after scale-in, the number of ranks used by shared experts after scale-in, and the number of MoE experts after scale-in. The remaining 2 × [object Object] indicates two rank mapping tables. After scale-in, some ranks on the current device may be removed from the EP communication domain due to failures. The mapping for the first table is Table1[epRankId]=localEpRankId or Table1[epRankId]=-1. Here, [object Object] denotes the rank index in the new EP communication domain, and [object Object] indicates that the rank with the corresponding [object Object] has been removed from the communication domain. The mapping for the second table is Table2[localEpRankId] = epRankId .
    • The value of [object Object] must be in the range [0, MAX_INT32), where MAX_INT32 = 2^31 - 1. Valid zero expert IDs must be in the range [object Object][moeExpertNum, moeExpertNum + zeroExpertNum)[object Object].
    • The value of [object Object] must be in the range [0, MAX_INT32), where MAX_INT32 = 2^31 - 1. Valid expert IDs must be in the range [object Object][moeExpertNum + zeroExpertNum, moeExpertNum + zeroExpertNum + copyExpertNum)[object Object].
    • The value of [object Object] must be in the range [0, MAX_INT32), where MAX_INT32 = 2^31 - 1. Valid expert IDs must be in the range [object Object].
    • [object Object] is a reserved parameter, which is not supported in the current version. Pass a null pointer.
    [object Object]
  • Returns:

    [object Object]: status code. For details, see .

    The first-phase API implements input parameter verification. The following errors may be thrown.

    [object Object]
[object Object]
  • Parameters

    [object Object]
  • Returns

    [object Object] status code. For details, see .

[object Object]
  • Deterministic computing:

    • [object Object] defaults to a deterministic implementation.
  • API constraints:

    • [object Object] and CombineV4 operators must be used together. The [object Object], [object Object], [object Object], and [object Object] outputs of [object Object] must be directly passed to the corresponding parameters of [object Object]. The service logic cannot depend on the specific values of these tensors.
  • Parameter consistency constraints:

    • The values of [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], and [object Object] must be consistent across all ranks, and be the same as the value of CombineV4.
    • The deployment information after dynamic scale-in is transferred to the operator through the [object Object] parameter. Other parameters do not need to be modified. The scale-in parameters take effect only when [object Object] is set to 1. After dynamic scale-in, the number of MoE experts deployed on the current rank must be the same as that before scale-in. Configurations where no MoE expert ranks remain after scale-in are not supported.
  • Product constraints:

    • [object Object]Atlas A3 training products/Atlas A3 inference products[object Object]: In this scenario, a single rank contains dual dies. Therefore, the "rank" in the parameter description indicates a single die.
    • The dynamic scale-in feature cannot be enabled in the tensor parallelism scenario.
  • Shape variable constraints:

    [object Object]undefined
  • Environment variables constraints:

    • HCCL_BUFFSIZE: Before calling this API, check whether the value of the [object Object] environment variable is proper. This environment variable indicates the buffer size occupied by a single communication domain, in MB. If this environment variable is not set, the default value 200 MB is used.

      • [object Object]Atlas A2 training products/Atlas A2 inference products[object Object]:
        • If [object Object] is set to [object Object] or [object Object], select the [object Object] or [object Object] formula based on the [object Object] and [object Object] environment variables.
        • If [object Object] is set to [object Object], the value must satisfy [object Object].
        • If commAlg is set to [object Object], the value must be [object Object], and [object Object] is not required.
      • [object Object]Atlas A3 training products/Atlas A3 inference products[object Object]:
        • When [object Object] is [object Object], a null character string, or a null pointer, the value must satisfy ≥ 2 × (localExpertNum × maxBs × epWorldSize × Align512(Align32(2 × H) + 64) + (K + sharedExpertNum) × maxBs × Align512(2 × H)).
        • When [object Object] is [object Object], the value must satisfy ≥ 2 × (localExpertNum × maxBs × epWorldSize × 480Align512(Align32(2 × H) + 64) + (K + sharedExpertNum) × maxBs × Align512(2 × H)).
        • [object Object], [object Object] and [object Object].
    • HCCL_INTRA_PCIE_ENABLE and HCCL_INTRA_ROCE_ENABLE:

      • [object Object]Atlas A2 training products/Atlas A2 inference products[object Object]: This environment variable is not recommended. You are advised to set [object Object] to [object Object].
  • Constraints on the use of the communication domains:

    • [object Object] and [object Object] in a model support only the same EP communication domain, and no other operators are allowed in the communication domain.
    • [object Object] and [object Object] in a model support only the same TP communication domain or both do not support a TP communication domain. If a TP communication domain is supported, no other operators are allowed in the communication domain.
    • [object Object]Atlas A3 training products/Atlas A3 inference products[object Object]: Nodes in a communication domain must be in the same SuperPoD. Cross-SuperPoD nodes are not supported.
  • Networking constraints:

    • [object Object]Atlas A2 training products/Atlas A2 inference products[object Object]: In multi-server scenarios, only switch-based networking is supported, and direct point-to-point networking between two servers is not supported.
  • Other constraints:

    • In the formulas, / denotes integer division.
    • [object Object]moeExpertNum + zeroExpertNum + copyExpertNum + constExpertNum < MAX_INT32[object Object]
[object Object]

[object Object]Atlas A2 training products/Atlas A2 inference products[object Object]: Similar to the following example for [object Object]Atlas A3 training products/Atlas A3 inference products[object Object]. For the new scenario parameters of V4 compared with V3, set the parameter values based on the preceding parameter description.

[object Object]Atlas A3 training products/Atlas A3 inference products[object Object]: The sample code is as follows (for reference only). Call the [object Object] and [object Object] APIs.

  • Preparing files:

    1. Create a [object Object] directory. Follow the instructions to create [object Object] and [object Object] files in the [object Object] directory, and modify them according to the code.

    2. Install the CANN package and compile and run [object Object].

  • Compilation script:

    [object Object]
  • Compilation and execution:

    [object Object]
  • The sample code is as follows:

    [object Object]