aclnnIndexAddV2

📄 View Source Code

Applicable Products

ProductSupported
Ascend 950PR/Ascend 950DT×
Atlas A3 training products/Atlas A3 inference products√
Atlas A2 training products/Atlas A2 inference products√
Atlas 200I/500 A2 inference products×
Atlas inference products×
Atlas training products×

Function

Adds values from the source tensor to the corresponding positions in the input tensor along a specified dimension based on given indices.

Function Prototype

Each operator is divided into a two-phase API. The "aclnnIndexAddV2GetWorkspaceSize" API must be called first to obtain the workspace size required for computation and the executor that encapsulates the operator computation flow. Then, the "aclnnIndexAddV2" API is called to perform the computation.

aclnnStatus aclnnIndexAddV2GetWorkspaceSize(
 const aclTensor*  self,
 const int64_t     dim,
 const aclTensor*  index,
 const aclTensor*  source,
 const aclScalar*  alpha,
 int64_t           mode,
 aclTensor*        out,
 uint64_t*         workspaceSize,
 aclOpExecutor**   executor)
aclnnStatus aclnnIndexAddV2(
 void*          workspace,
 uint64_t       workspaceSize,
 aclOpExecutor* executor,
 aclrtStream    stream)

aclnnIndexAddV2GetWorkspaceSize

  • Parameters

    Parameter Input/Output Description Instruction Data Type Data Format Shape Non-contiguous Tensor
    self (aclTensor*) Input Input tensor. - FLOAT, FLOAT16, INT32, INT16, INT8, UINT8, DOUBLE, INT64, BOOL, BFLOAT16 ND 0-8 √
    dim(int64_t) Input Specifies the dimension. The value range is [-self.dim(), self.dim()-1]. INT64 - - -
    index (aclTensor*) Input Index. The shape of index must be equal to the shape of source along the dim dimension. INT64, INT32 ND 1 √
    source (aclTensor*) Input Source tensor. The shape of source must match the shape of self in all dimensions except the dim dimension. Same as self. ND Same as self. √
    alpha (aclScalar*) Input Scaling factor. The data type can be converted to the data type deduced from self and source.
    In non-deterministic computation scenarios where mode is set to 0, alpha does not participate in the computation.
    - Same as self. - -
    mode (int64_t) Input Computation mode. When the value is 0, it indicates high-performance mode; other values indicate non-high-performance mode. For high-performance mode restrictions, see Constraints.
    In deterministic computation scenarios, the mode value is ignored and the deterministic computation mode is used uniformly.
    INT64 - - -
    out (aclTensor*) Output Output aclTensor. - Same as self. ND Same as self. √
    workspaceSize (uint64_t*) Output Returns the workspace size to be allocated on the Device side. - - - - -
    executor (aclOpExecutor**) Output Returns the operator executor, which contains the operator computation flow. - - - - -
  • Return Value

    aclnnStatus: return code. For details, see aclnn Return Code.

    The first-phase API performs input parameter validation and returns an error in the following scenarios:

    Return Value Error Code Description
    ACLNN_ERR_PARAM_NULLPTR 161001 The input self, index, source, alpha, or out is a null pointer.
    ACLNN_ERR_PARAM_INVALID 161002 The data type of self, index, source, or out is not within the supported range.
    The data types of self, source, and out are inconsistent.
    The value of dim is greater than the shape size of self.
    The computed data type cannot be converted to the data type of the specified output out.
    self and source have unequal shape values in dimensions other than dim.
    index is not one-dimensional.
    The shape size of index is not equal to the shape value of source in dimension dim.
    The shape of out is not equal to the shape of self.
    The value range of index is not within [0, self.shape[dim]).
    When mode=0, the constraints for high-performance mode in Constraints are not met.

aclnnIndexAddV2

  • Parameters

    Parameter Input/Output Description
    workspace Input Address of the workspace memory applied for on the Device side.
    workspaceSize Input Size of the workspace applied for on the Device side, obtained through the first-phase API aclnnIndexAddV2GetWorkspaceSize.
    executor Input Operator executor that contains the operator computation flow.
    stream Input Specifies the stream for task execution.
  • Return Value

    aclnnStatus: return code. For details, see aclnn Return Code.

Constraints

  • Deterministic computation:

    • aclnnIndexAddV2 defaults to a non-deterministic implementation. You can enable deterministic computation through aclrtCtxSetSysParamOpt.
  • Index value range

    • The index input value range is [0, self.shape[dim]), meaning the index value input range is the shape size of self along the dim dimension. Negative indices and out-of-bounds indices are not supported.
  • High-performance mode (mode=0) has the following additional constraints:

    • dim must be 0 or -2.

    • The data type of self must be FLOAT, FLOAT16, INT32, INT16, or BFLOAT16.

    • self must be a two-dimensional tensor.

    • If mode=0 and the constraints in the parameter description are met but the above three constraints are not, the API throws an error.

Example

The following provides example code for reference only. For details about compilation and running, see Compile and Run Samples.

#include <iostream>
#include <vector>
#include "acl/acl.h"
#include "aclnnop/aclnn_index_add_v2.h"

#define CHECK_RET(cond, return_expr) \
  do {                               \
    if (!(cond)) {                   \
      return_expr;                   \
    }                                \
  } while (0)

#define LOG_PRINT(message, ...)     \
  do {                              \
    printf(message, ##__VA_ARGS__); \
  } while (0)

int64_t GetShapeSize(const std::vector<int64_t>& shape) {
  int64_t shapeSize = 1;
  for (auto i : shape) {
    shapeSize *= i;
  }
  return shapeSize;
}

int Init(int32_t deviceId, aclrtStream* stream) {
  // Boilerplate: resource initialization.
  auto ret = aclInit(nullptr);
  CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclInit failed. ERROR: %d\n", ret); return ret);
  ret = aclrtSetDevice(deviceId);
  CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclrtSetDevice failed. ERROR: %d\n", ret); return ret);
  ret = aclrtCreateStream(stream);
  CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclrtCreateStream failed. ERROR: %d\n", ret); return ret);
  return 0;
}

template <typename T>
int CreateAclTensor(const std::vector<T>& hostData, const std::vector<int64_t>& shape, void** deviceAddr,
                    aclDataType dataType, aclTensor** tensor) {
  auto size = GetShapeSize(shape) * sizeof(T);
  // Call aclrtMalloc to allocate device memory.
  auto ret = aclrtMalloc(deviceAddr, size, ACL_MEM_MALLOC_HUGE_FIRST);
  CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclrtMalloc failed. ERROR: %d\n", ret); return ret);
  // Call aclrtMemcpy to copy data from host memory to device memory.
  ret = aclrtMemcpy(*deviceAddr, size, hostData.data(), size, ACL_MEMCPY_HOST_TO_DEVICE);
  CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclrtMemcpy failed. ERROR: %d\n", ret); return ret);

  // Calculate the strides of a contiguous tensor.
  std::vector<int64_t> strides(shape.size(), 1);
  for (int64_t i = shape.size() - 2; i >= 0; i--) {
    strides[i] = shape[i + 1] * strides[i + 1];
  }

  // Call aclCreateTensor to create an aclTensor.
  *tensor = aclCreateTensor(shape.data(), shape.size(), dataType, strides.data(), 0, aclFormat::ACL_FORMAT_ND,
                            shape.data(), shape.size(), *deviceAddr);
  return 0;
}

int main() {
  // 1. (Boilerplate) Initialize the device/stream. Refer to the ACL API manual.
  // Fill in the deviceId based on your actual device.
  int32_t deviceId = 0;
  aclrtStream stream;
  auto ret = Init(deviceId, &stream);
  CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("Init acl failed. ERROR: %d\n", ret); return ret);

  // 2. Construct the input and output. Custom construction is required based on the API interface.
  std::vector<int64_t> selfShape = {4, 2};
  std::vector<int64_t> indexShape = {4};
  std::vector<int64_t> sourceShape = {4, 2};
  std::vector<int64_t> outShape = {4, 2};
  void* selfDeviceAddr = nullptr;
  void* indexDeviceAddr = nullptr;
  void* sourceDeviceAddr = nullptr;
  void* outDeviceAddr = nullptr;
  aclTensor* self = nullptr;
  aclTensor* index = nullptr;
  aclTensor* source = nullptr;
  aclScalar* alpha = nullptr;
  aclTensor* out = nullptr;
  std::vector<float> selfHostData = {0, 1, 2, 3, 4, 5, 6, 7};
  std::vector<int> indexHostData = {0, 1, 2, 3};
  std::vector<float> sourceHostData = {1, 1, 1, 2, 2, 2, 3, 3};
  std::vector<float> outHostData(8, 0);
  int64_t dim = 0;
  int64_t mode = 0;
  float alphaValue = 1.0f;
  // Create the self aclTensor.
  ret = CreateAclTensor(selfHostData, selfShape, &selfDeviceAddr, aclDataType::ACL_FLOAT, &self);
  CHECK_RET(ret == ACL_SUCCESS, return ret);
  // Create the index aclTensor.
  ret = CreateAclTensor(indexHostData, indexShape, &indexDeviceAddr, aclDataType::ACL_INT32, &index);
  CHECK_RET(ret == ACL_SUCCESS, return ret);
  // Create the source aclTensor.
  ret = CreateAclTensor(sourceHostData, sourceShape, &sourceDeviceAddr, aclDataType::ACL_FLOAT, &source);
  CHECK_RET(ret == ACL_SUCCESS, return ret);
  // Create the out aclTensor.
  ret = CreateAclTensor(outHostData, outShape, &outDeviceAddr, aclDataType::ACL_FLOAT, &out);
  CHECK_RET(ret == ACL_SUCCESS, return ret);
  // Create the alpha aclScalar.
  alpha = aclCreateScalar(&alphaValue, aclDataType::ACL_FLOAT);
  CHECK_RET(alpha != nullptr, return ret);

  // 3. Call the CANN operator library API. Replace with the specific API name.
  uint64_t workspaceSize = 0;
  aclOpExecutor* executor;
  // Call the first-phase API of aclnnIndexAddV2.
  ret = aclnnIndexAddV2GetWorkspaceSize(self, dim, index, source, alpha, mode, out, &workspaceSize, &executor);
  CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclnnIndexAddV2GetWorkspaceSize failed. ERROR: %d\n", ret); return ret);
  // Apply for device memory based on the workspaceSize calculated by the first-phase API.
  void* workspaceAddr = nullptr;
  if (workspaceSize > 0) {
    ret = aclrtMalloc(&workspaceAddr, workspaceSize, ACL_MEM_MALLOC_HUGE_FIRST);
    CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("allocate workspace failed. ERROR: %d\n", ret); return ret);
  }
  // Call the second-phase API of aclnnIndexAddV2.
  ret = aclnnIndexAddV2(workspaceAddr, workspaceSize, executor, stream);
  CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclnnIndexAddV2 failed. ERROR: %d\n", ret); return ret);

  // 4. (Boilerplate) Synchronously wait for task execution to complete.
  ret = aclrtSynchronizeStream(stream);
  CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclrtSynchronizeStream failed. ERROR: %d\n", ret); return ret);

  // 5. Obtain the output value and copy the result from the device memory to the host memory. Modify based on the specific API definition.
  auto size = GetShapeSize(outShape);
  std::vector<float> resultData(size, 0);
  ret = aclrtMemcpy(resultData.data(), resultData.size() * sizeof(resultData[0]), outDeviceAddr,
                    size * sizeof(resultData[0]), ACL_MEMCPY_DEVICE_TO_HOST);
  CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("copy result from device to host failed. ERROR: %d\n", ret); return ret);
  for (int64_t i = 0; i < size; i++) {
    LOG_PRINT("result[%ld] is: %f\n", i, resultData[i]);
  }

  // 6. Release aclTensor and aclScalar. Modify based on the specific API definition.
  aclDestroyTensor(self);
  aclDestroyTensor(index);
  aclDestroyTensor(source);
  aclDestroyScalar(alpha);
  aclDestroyTensor(out);

  // 7. Release device resources. Modify based on the specific API definition.
  aclrtFree(selfDeviceAddr);
  aclrtFree(indexDeviceAddr);
  aclrtFree(sourceDeviceAddr);
  aclrtFree(outDeviceAddr);
  if (workspaceSize > 0) {
    aclrtFree(workspaceAddr);
  }
  aclrtDestroyStream(stream);
  aclrtResetDevice(deviceId);
  aclFinalize();
  return 0;
}