aclSetOutputTensorAddr

Description

After aclOpExecutor reuse is enabled by calling aclSetAclOpExecutorRepeatable, if the output device memory address changes, the device memory address recorded in the output aclTensor needs to be updated.

Prototype

aclnnStatus aclSetOutputTensorAddr(aclOpExecutor *executor, const size_t index, aclTensor *tensor, void *addr)

Parameter Description

Parameter

Input/Output

Description

executor

Input

aclOpExecutor that is set to the reusable state.

index:

Input

Index of the output aclTensor to be updated. The value range is [0, total number of output tensors – 1].

tensor

Input

aclTensor pointer to be updated.

addr

Input

Device storage address to be updated to the specified aclTensor. The address must be 32-byte aligned. Otherwise, an undefined error may occur.

Return Value

0 on success; otherwise, failure. For details about the return codes, see Common API Return Codes.

Possible causes:

  • Error code 561103: executor or tensor is a null pointer.
  • Error code 161002: The index value is out of range.
  • Error code 161002: The input aclTensor when the first-phase API aclxxXxxGetWorkspaceSize is called for the first time is a null pointer. The address cannot be updated.

Constraints

None.

Example

The following sample code is for reference only. Do not copy and run it.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
// Create the input and output aclTensor and aclTensorList.
std::vector<int64_t> shape = {1, 2, 3};
aclTensor tensor1 = aclCreateTensor(shape.data(), shape.size(), aclDataType::ACL_FLOAT,
nullptr, 0, aclFormat::ACL_FORMAT_ND, shape.data(), shape.size(), nullptr);
aclTensor tensor2 = aclCreateTensor(shape.data(), shape.size(), aclDataType::ACL_FLOAT,
nullptr, 0, aclFormat::ACL_FORMAT_ND, shape.data(), shape.size(), nullptr);
aclTensor tensor3 = aclCreateTensor(shape.data(), shape.size(), aclDataType::ACL_FLOAT,
nullptr, 0, aclFormat::ACL_FORMAT_ND, shape.data(), shape.size(), nullptr);
aclTensor tensor4= aclCreateTensor(shape.data(), shape.size(), aclDataType::ACL_FLOAT,
nullptr, 0, aclFormat::ACL_FORMAT_ND, shape.data(), shape.size(), nullptr);
aclTensor output= aclCreateTensor(shape.data(), shape.size(), aclDataType::ACL_FLOAT,
nullptr, 0, aclFormat::ACL_FORMAT_ND, shape.data(), shape.size(), nullptr);
aclTensor *list[] = {tensor3, tensor4};
auto tensorList = aclCreateTensorList(list, 2);
uint64_t workspaceSize = 0;
aclOpExecutor *executor;
// The AddCustom operator has two inputs (aclTensor) and two outputs (aclTensor and aclTensorList).
// Call the first-phase API.
aclnnAddCustomGetWorkspaceSize(tensor1, tensor2, output, tensorList, &workspaceSize, &executor);
// Set the executor to be reusable.
aclSetAclOpExecutorRepeatable(executor); 
void *addr;
aclSetOutputTensorAddr(executor, 0, output, addr);  // Update the device address of the output aclTensor.
aclSetOutputTensorAddr(executor, 1, output, addr);  // Update the device address of the first aclTensor in the output tensor list.
aclSetOutputTensorAddr(executor, 2, output, addr);  // Update the device address of the second aclTensor in the output tensor list.
...
// Call the second-phase API.
aclnnAddCustom(workspace, workspaceSize, executor, stream);
// Clear the executor.
aclDestroyAclOpExecutor(executor);