[object Object][object Object][object Object]undefined
[object Object]
  • Description: Implements grouped matrix multiplication, supporting non-uniform matrix dimension sizes across multiple groups. The basic function is matrix multiplication, for example, yi[mi,ni]=xi[mi,ki]×weighti[ki,ni],i=1...gy_i[m_i,n_i]=x_i[m_i,k_i] \times weight_i[k_i,n_i], i=1...g, where gg indicates the number of groups and mim_i, kik_i, and nin_i define the shapes for each group. Both inputs and outputs are of the aclTensorList type, with the following functions:

    • K-axis grouping: kik_i varies across groups, while mim_i and nin_i remain the same for each group. In this case, xix_i and weightiweight_i can be concatenated along the K-axis.
    • M-axis grouping: kik_i remains the same for each group. In this case, weightiweight_i and yiy_i can be concatenated along the N-axis.

    Compared with , this API provides the following new features:

    • [object Object]Atlas A2 training products/Atlas A2 inference products[object Object]:

      • Supports weight transposition in non-quantization scenarios. Transposition refers to the case where the shape is [M, K], the stride is [1, M], and the data layout is [K, M].
      • Supports M-axis and K-axis grouping, represented by [object Object].
      • Supports FLOAT32 input for [object Object] and [object Object] when [object Object], [object Object], and [object Object] are all single-tensor in non-quantization scenarios.
      • Supports weight transposition and single-tensor weights in quantization and fake-quantization scenarios.
      • For the features supported by , this API does not support scenarios where [object Object] is single-tensor while [object Object] and [object Object] are multi-tensor. Notes:
    • "Single-tensor" means that tensors of all groups in a tensor list are concatenated into one tensor along the axis specified [object Object].

    • Tensor transpose: If the tensor shape is [M, K], the stride is [1, M], and the data layout is [K, M], then the tensor is a non-contiguous tensor.

  • Formula:

    • Non-quantization scenario:
    yi=xi×weighti+biasiy_i=x_i\times weight_i + bias_i
    • Quantization scenario:
    yi=(xi×weighti+biasi)scalei+offsetiy_i=(x_i\times weight_i + bias_i) * scale_i + offset_i
    • Dequantization scenario:
    yi=(xi×weighti+biasi)scaleiy_i=(x_i\times weight_i + bias_i) * scale_i
    • Fake-quantization scenario:
    yi=xi×(weighti+antiquant_offseti)antiquant_scalei+biasiy_i=x_i\times (weight_i + antiquant\_offset_i) * antiquant\_scale_i + bias_i
[object Object]

Each operator has calls. First, [object Object] is called to obtain the input parameters and compute the required workspace size based on the process. Then, [object Object] is called to perform computation.

  • [object Object]
  • [object Object]
[object Object]
  • Parameters

    • x (aclTensorList*, computation input): required parameter, aclTensorList on the device, xx in the formula. The can be ND, and the maximum length is 128.
      • [object Object]Atlas A2 training products/Atlas A2 inference products[object Object]: The data type can be FLOAT16, BFLOAT16, INT8, or FLOAT32.
      • [object Object]Atlas inference products[object Object]: The data type can be FLOAT16.
    • weight (aclTensorList*, computation input): required parameter, aclTensorList on the device, weightweight in the formula. The can be ND, and the maximum length is 128.
      • [object Object]Atlas A2 training products/Atlas A2 inference products[object Object]: The data type can be FLOAT16, BFLOAT16, INT8, or FLOAT32.
      • [object Object]Atlas inference products[object Object]: The data type can be FLOAT16.
    • biasOptional (aclTensorList*, computation input): optional parameter, aclTensorList on the device, biasbias in the formula. The can be ND, and the length is the same as that of [object Object].
      • [object Object]Atlas A2 training products/Atlas A2 inference products[object Object]: The data type can be FLOAT16, FLOAT32, or INT32.
      • [object Object]Atlas inference products[object Object]: The data type can be FLOAT16.
    • scaleOptional (aclTensorList*, computation input): optional parameter, aclTensorList on the device, indicating the scale factor for quantization parameters. The data type can be UINT64, the can be ND, and the length is the same as that of [object Object].
      • [object Object]Atlas inference products[object Object]: This parameter is not supported currently and needs to be passed as a null pointer.
    • offsetOptional (aclTensorList*, computation input): optional parameter, aclTensorList on the device, indicating the offset for quantization parameters. The data type can be FLOAT32, the can be ND, and the length is the same as that of [object Object].
      • [object Object]Atlas inference products[object Object]: This parameter is not supported currently and needs to be passed as a null pointer.
    • antiquantScaleOptional (aclTensorList*, computation input): optional parameter, aclTensorList on the device, indicating the scale factor for fake-quantization parameters. The can be ND, and the length is the same as that of [object Object].
      • [object Object]Atlas A2 training products/Atlas A2 inference products[object Object]: The data type can be FLOAT16 or BFLOAT16.
      • [object Object]Atlas inference products[object Object]: The data type can be FLOAT16. This parameter is not supported currently and needs to be passed as a null pointer.
    • antiquantOffsetOptional (aclTensorList*, computation input): optional parameter, aclTensorList on the device, indicating the offset for fake-quantization parameters. The can be ND, and the length is the same as that of [object Object].
      • [object Object]Atlas A2 training products/Atlas A2 inference products[object Object]: The data type can be FLOAT16 or BFLOAT16.
      • [object Object]Atlas inference products[object Object]: The data type can be FLOAT16. This parameter is not supported currently and needs to be passed as a null pointer.
    • groupListOptional (aclTensor*, computation input): optional parameter, aclTensor type on the device, indicating the Matmul size distribution along the grouping axis for inputs and outputs. The data type can be INT64, and the can be ND. Note that when the length of the TensorList in the output is 1, the last value in [object Object] constrains the valid portion of the output data. Any portion not specified in [object Object] will not be updated.
    • splitItem (int64_t, computation input): integer type, indicating whether tensor splitting is required for the output. [object Object] or [object Object] indicates multi-tensor, and [object Object] or [object Object] indicates single-tensor.
    • groupType (int64_t, computation input): integer type, indicating the axis to be grouped. For example, if the matrix multiplication is [object Object], [object Object] has the following options: [object Object] means no axis grouping, [object Object] indicates M-axis grouping, [object Object] indicates N-axis grouping, and [object Object] indicates K-axis grouping.
      • [object Object]Atlas A2 training products/Atlas A2 inference products[object Object]: Currently, N-axis grouping is not supported.
      • [object Object]Atlas inference products[object Object]: Currently, only M-axis grouping is supported.
    • y (aclTensorList*, computation output): aclTensorList on the device, yy in the formula. The can be ND, and the maximum length is 128.
      • [object Object]Atlas A2 training products/Atlas A2 inference products[object Object]: The data type can be FLOAT16, BFLOAT16, INT8, or FLOAT32.
      • [object Object]Atlas inference products[object Object]: The data type can be FLOAT16.
    • workspaceSize (uint64_t*, output): size of the workspace to be allocated on the device.
    • executor (aclOpExecutor**, output): operator executor, containing the operator computation process.
  • Return

    [object Object] status code. For details, see .

[object Object]
[object Object]
  • Parameters

    • workspace (void*, input): address of the workspace to be allocated on the device.
    • workspaceSize (uint64_t, input): size of the workspace to be allocated on the device, which is obtained by calling the first-phase API [object Object].
    • executor (aclOpExecutor*, input): operator executor, containing the operator computation process.
    • stream (aclrtStream, input): stream for executing the task.
  • Return

    [object Object] status code. For details, see .

[object Object]
  • Deterministic computation:
    • [object Object] defaults to deterministic implementation.
  • If [object Object] is passed, it must be a non-negative ascending array, and its length cannot be 1.
  • The size of each dimension for every tensor in [object Object] and [object Object], after 32-byte alignment, should be less than the maximum value of INT32 (2147483647).
  • [object Object]Atlas A2 training products/Atlas A2 inference products[object Object]:
    • The following input types are supported in non-quantization scenarios:

      • [object Object]: FLOAT16; [object Object]: FLOAT16; [object Object]: FLOAT16; [object Object]: null; [object Object]: null; [object Object]: null; [object Object]: null; [object Object]: FLOAT16
      • [object Object]: BFLOAT16; [object Object]: BFLOAT16; [object Object]: FLOAT32; [object Object]: null; [object Object]: null; [object Object]: null; [object Object]: null; [object Object]: BFLOAT16
      • [object Object]: FLOAT32; [object Object]: FLOAT32; [object Object]: FLOAT32; [object Object]: null; [object Object]: null; [object Object]: null; [object Object]: null; [object Object]: FLOAT32 (supported only when [object Object], [object Object], and [object Object] are all single-tensor)
    • The following input types are supported in quantization scenarios:

      • [object Object]: INT8; [object Object]: INT8; [object Object]: INT32; [object Object]: UINT64; [object Object]: null; [object Object]: null; [object Object]: null; [object Object]: INT8
    • The following input types are supported in fake-quantization scenarios:

      • [object Object]: FLOAT16; [object Object]: INT8; [object Object]: FLOAT16; [object Object]: null; [object Object]: null; [object Object]: FLOAT16; [object Object]: FLOAT16; [object Object]: FLOAT16
      • [object Object]: BFLOAT16; [object Object]: INT8; [object Object]: FLOAT32; [object Object]: null; [object Object]: null; [object Object]: BFLOAT16; [object Object]: BFLOAT16; [object Object]: BFLOAT16
    • Supported scenarios for different [object Object] values:

      • In quantization and fake-quantization scenarios, [object Object] can be either [object Object] or [object Object].

      • "S" stands for single-tensor, and "M" stands for multi-tensor, expressed in the sequence of [object Object], [object Object], [object Object]. For example, "SMS" indicates single-tensor [object Object], multi-tensor [object Object], and single-tensor [object Object].

        [object Object]undefined
    • The size of the last dimension for each tensor in [object Object] and [object Object] should be less than 65536. The last dimension of xix_i refers to the K-axis when [object Object] is not transposed or the M-axis when [object Object] is transposed. The last dimension of weightiweight_i refers to the N-axis when [object Object] is not transposed or the K-axis when [object Object] is transposed.

  • [object Object]Atlas inference products[object Object]:
    • The input and output support only the FLOAT16 type. The N-axis size of the output [object Object] must be a multiple of 16.

      [object Object]undefined
[object Object]

The following example is for reference only. For details, see .

[object Object]