aclnnIndexSelect

📄 View Source Code

Applicable Products

ProductSupported
Ascend 950PR/Ascend 950DT√
Atlas A3 training products/Atlas A3 inference products√
Atlas A2 training products/Atlas A2 inference products√
Atlas 200I/500 A2 inference products×
Atlas inference products√
Atlas training products√

Function

  • Description: Extracts elements from the input tensor along the specified dimension dim based on the index numbers in index, and saves them to the out tensor.

    For example, for the input tensor self=[123456789]self=\begin{bmatrix}1 & 2 & 3 \\ 4 & 5 & 6 \\ 7 & 8 & 9\end{bmatrix} and the index tensor index=[1, 0], the result of self.index_select(0, index): y=[456123]y=\begin{bmatrix}4 & 5 & 6 \\ 1 & 2 & 3\end{bmatrix};

    The result of x.index_select(1, index): y=[215487]y=\begin{bmatrix}2 & 1\\ 5 & 4\\8 & 7\end{bmatrix};

  • Formula:

    Taking a three-dimensional tensor as an example, the tensor self with shape (3,2,2) =[[[1,2],[3,4]],[[5,6],[7,8]],[[9,10],[11,12]]]\begin{bmatrix}[[1,&2],&[3,&4]], \\ [[5,&6],&[7,&8]], \\ [[9,&10],&[11,&12]]\end{bmatrix}, index=[1, 0], the subscripts corresponding to dim=0, 1, and 2 of the self tensor are ll, mm, and nn respectively, and index is one-dimensional.

    dim is 0, index_select(0, index): I=index[i];    out[i][m][n][i][m][n] = self[I][m][n][I][m][n]

    dim is 1, index_select(1, index): J=index[j];     out[l][j][n][l][j][n] = self[l][J][n][l][J][n]

    dim is 2, index_select(2, index): K=index[k];   out[l][m][k][l][m][k] = self[l][m][K][l][m][K]

Function Prototype

Each operator is divided into a two-phase API. The "aclnnIndexSelectGetWorkspaceSize" API must be called first to obtain the workspace size required for computation and the executor that encapsulates the operator computation flow. Then, the "aclnnIndexSelect" API is called to perform the computation.

aclnnStatus aclnnIndexSelectGetWorkspaceSize(
 const aclTensor *self,
 int64_t          dim,
 const aclTensor *index,
 aclTensor       *out,
 uint64_t        *workspaceSize,
 aclOpExecutor  **executor)
aclnnStatus aclnnIndexSelect(
 void             *workspace,
 uint64_t          workspaceSize,
 aclOpExecutor    *executor,
 const aclrtStream stream)

aclnnIndexSelectGetWorkspaceSize

  • Parameters

    Parameter Input/Output Description Instruction Data Type Data Format Dimension (Shape) Non-contiguous Tensor
    self Input Input tensor. - FLOAT、FLOAT16、BFLOAT16、INT64、INT32、INT16、INT8、UINT8、UINT16、UINT32、UINT64、BOOL、DOUBLE、COMPLEX64、COMPLEX128 ND、NCHW、NHWC、HWCN、NDHWC、NCDHW Not greater than 8. √
    dim Input The specified dimension. Range: [-self.dim(), self.dim() - 1]. INT64 - - -
    index Input Index. Must be 0D or 1D (a 0D tensor is treated as a 1D tensor of size 1). The index values in index range from 0 to self.shape[dim] (inclusive of 0, exclusive of self.shape[dim]). INT64, INT32 ND, NCHW, NHWC, HWCN, NDHWC, NCDHW - -
    out Output Output tensor. All dimensions except the dim dimension have the same length as the corresponding dimensions of self, while the dim dimension length equals the index length. Same as self. ND, NCHW, NHWC, HWCN, NDHWC, NCDHW Same as self. -
    workspaceSize Output Returns the workspace size to be requested on the Device side. - - - - -
    executor Output Returns the operator executor, which contains the operator computation flow. - - - - -
    • Atlas inference products, Atlas training products: The BFLOAT16 data type is not supported.
  • Return Value

    aclnnStatus: return code. For details, see aclnn Return Code.

    The first-phase API performs input parameter validation and returns an error in the following scenarios:

    Return Value Error Code Description
    ACLNN_ERR_PARAM_NULLPTR 161001 The self, index, or out parameter is a null pointer.
    ACLNN_ERR_PARAM_INVALID 161002 The data types of the parameters self and index are not within the supported range.
    dim >= self.dim() or dim < -self.dim().
    The dimension of index is greater than 1.
    The dimension of self is greater than 8.

aclnnIndexSelect

  • Parameters

    Parameter Input/Output Description
    workspace Input Workspace memory address applied for on the Device side.
    workspaceSize Input Workspace size applied for on the Device side, obtained from the first-phase API aclnnIndexSelectGetWorkspaceSize.
    executor Input Operator executor, which contains the operator computation process.
    stream Input Specifies the stream that executes the task.
  • Return Value

    aclnnStatus: return code. For details, see aclnn Return Code.

Constraints

  • When the shape of self is [], the shape of index can only be [1].

  • Deterministic computation:

    • aclnnIndexSelect defaults to a deterministic implementation.

Example

The following provides example code for reference only. For details about compilation and running, see Compile and Run Samples.

#include <iostream>
#include <vector>
#include "acl/acl.h"
#include "aclnnop/aclnn_index_select.h"

#define CHECK_RET(cond, return_expr) \
  do {                               \
    if (!(cond)) {                   \
      return_expr;                   \
    }                                \
  } while (0)

#define LOG_PRINT(message, ...)     \
  do {                              \
    printf(message, ##__VA_ARGS__); \
  } while (0)

int64_t GetShapeSize(const std::vector<int64_t>& shape) {
  int64_t shapeSize = 1;
  for (auto i : shape) {
    shapeSize *= i;
  }
  return shapeSize;
}

int Init(int32_t deviceId, aclrtStream* stream) {
  // Boilerplate. Initialize resources.
  auto ret = aclInit(nullptr);
  CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclInit failed. ERROR: %d\n", ret); return ret);
  ret = aclrtSetDevice(deviceId);
  CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclrtSetDevice failed. ERROR: %d\n", ret); return ret);
  ret = aclrtCreateStream(stream);
  CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclrtCreateStream failed. ERROR: %d\n", ret); return ret);
  return 0;
}

template <typename T>
int CreateAclTensor(const std::vector<T>& hostData, const std::vector<int64_t>& shape, void** deviceAddr,
                    aclDataType dataType, aclTensor** tensor) {
  auto size = GetShapeSize(shape) * sizeof(T);
  // Call aclrtMalloc to apply for device memory.
  auto ret = aclrtMalloc(deviceAddr, size, ACL_MEM_MALLOC_HUGE_FIRST);
  CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclrtMalloc failed. ERROR: %d\n", ret); return ret);
  // Call aclrtMemcpy to copy data from host to device memory.
  ret = aclrtMemcpy(*deviceAddr, size, hostData.data(), size, ACL_MEMCPY_HOST_TO_DEVICE);
  CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclrtMemcpy failed. ERROR: %d\n", ret); return ret);

  // Calculate strides of a contiguous tensor.
  std::vector<int64_t> strides(shape.size(), 1);
  for (int64_t i = shape.size() - 2; i >= 0; i--) {
    strides[i] = shape[i + 1] * strides[i + 1];
  }

  // Call aclCreateTensor to create an aclTensor.
  *tensor = aclCreateTensor(shape.data(), shape.size(), dataType, strides.data(), 0, aclFormat::ACL_FORMAT_ND,
                            shape.data(), shape.size(), *deviceAddr);
  return 0;
}

int main() {
  // 1. (Boilerplate) Initialize device/stream. See the ACL API manual.
  // Fill in the deviceId based on your actual device.
  int32_t deviceId = 0;
  aclrtStream stream;
  auto ret = Init(deviceId, &stream);
  CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("Init acl failed. ERROR: %d\n", ret); return ret);

  // 2. Construct the input and output. Customize the construction based on the API interface.
  std::vector<int64_t> selfShape = {4, 2};
  std::vector<int64_t> indexShape = {2};
  std::vector<int64_t> outShape = {2, 2};
  void* selfDeviceAddr = nullptr;
  void* indexDeviceAddr = nullptr;
  void* outDeviceAddr = nullptr;
  aclTensor* self = nullptr;
  aclTensor* index = nullptr;
  int64_t dim = 0;
  aclTensor* out = nullptr;
  std::vector<float> selfHostData = {0, 1, 2, 3, 4, 5, 6, 7};
  std::vector<int> indexHostData = {1, 0};
  std::vector<float> outHostData = {0, 0, 0, 0, 0, 0, 0, 0};

  // Create the self aclTensor.
  ret = CreateAclTensor(selfHostData, selfShape, &selfDeviceAddr, aclDataType::ACL_FLOAT, &self);
  CHECK_RET(ret == ACL_SUCCESS, return ret);
  // Create the index aclTensor.
  ret = CreateAclTensor(indexHostData, indexShape, &indexDeviceAddr, aclDataType::ACL_INT32, &index);
  CHECK_RET(ret == ACL_SUCCESS, return ret);
  // Create the out aclTensor.
  ret = CreateAclTensor(outHostData, outShape, &outDeviceAddr, aclDataType::ACL_FLOAT, &out);
  CHECK_RET(ret == ACL_SUCCESS, return ret);

  // 3. Call the CANN operator library API. Replace with the specific API name.
  uint64_t workspaceSize = 0;
  aclOpExecutor* executor;
  // Call the first-phase API of aclnnIndexSelect.
  ret = aclnnIndexSelectGetWorkspaceSize(self, dim, index, out, &workspaceSize, &executor);
  CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclnnIndexSelectGetWorkspaceSize failed. ERROR: %d\n", ret); return ret);
  // Apply for device memory based on the workspaceSize calculated by the first-phase API.
  void* workspaceAddr = nullptr;
  if (workspaceSize > 0) {
    ret = aclrtMalloc(&workspaceAddr, workspaceSize, ACL_MEM_MALLOC_HUGE_FIRST);
    CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("allocate workspace failed. ERROR: %d\n", ret); return ret);
  }
  // Call the second-phase API of aclnnIndexSelect.
  ret = aclnnIndexSelect(workspaceAddr, workspaceSize, executor, stream);
  CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclnnIndexSelect failed. ERROR: %d\n", ret); return ret);

  // 4. (Boilerplate) Synchronize and wait for task execution to complete.
  ret = aclrtSynchronizeStream(stream);
  CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclrtSynchronizeStream failed. ERROR: %d\n", ret); return ret);

  // 5. Obtain the output value and copy the result from the device memory to the host. Modify based on the specific API definition.
  auto size = GetShapeSize(outShape);
  std::vector<float> resultData(size, 0);
  ret = aclrtMemcpy(resultData.data(), resultData.size() * sizeof(resultData[0]), outDeviceAddr,
                    size * sizeof(resultData[0]), ACL_MEMCPY_DEVICE_TO_HOST);
  CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("copy result from device to host failed. ERROR: %d\n", ret); return ret);
  for (int64_t i = 0; i < size; i++) {
    LOG_PRINT("result[%ld] is: %f\n", i, resultData[i]);
  }

  // 6. Release aclTensor and aclScalar. Modify based on the specific API definition.
  aclDestroyTensor(self);
  aclDestroyTensor(index);
  aclDestroyTensor(out);

  // 7. Release device resources. Modify based on the specific API definition.
  aclrtFree(selfDeviceAddr);
  aclrtFree(indexDeviceAddr);
  aclrtFree(outDeviceAddr);
  if (workspaceSize > 0) {
    aclrtFree(workspaceAddr);
  }
  aclrtDestroyStream(stream);
  aclrtResetDevice(deviceId);
  aclFinalize();
  return 0;
}