aclnnLog1p&aclnnInplaceLog1p
Supported Products
| Product | Supported |
|---|---|
| √ | |
| √ | |
| × | |
| √ | |
| √ |
Function Description
- Description: Computes log1p for the input tensor.
- Formula:
Function Prototype
aclnnLog1pandaclnnInplaceLog1pimplement the same function in different ways. Select a proper operator based on your requirements.aclnnLog1p: An output tensor object needs to be created to store the computation result.aclnnInplaceLog1p: No output tensor object needs to be created, and the computation result is stored in the memory of the input tensor.
- Each operator has two-phase API calls. First,
aclnnLog1pGetWorkspaceSizeoraclnnInplaceLog1pGetWorkspaceSizeis called to obtain the workspace size required for computation and the executor that contains the operator computation process. Then,aclnnLog1poraclnnInplaceLog1pis called to perform computation.aclnnStatus aclnnLog1pGetWorkspaceSize(const aclTensor *self, aclTensor *out, uint64_t *workspaceSize, aclOpExecutor **executor)aclnnStatus aclnnLog1p(void *workspace, uint64_t workspaceSize, aclOpExecutor *executor, aclrtStream stream)aclnnStatus aclnnInplaceLog1pGetWorkspaceSize(aclTensor* selfRef, uint64_t* workspaceSize, aclOpExecutor** executor)aclnnStatus aclnnInplaceLog1p(void *workspace, uint64_t workspaceSize, aclOpExecutor *executor, aclrtStream stream)
aclnnLog1pGetWorkspaceSize
Parameters:
self(aclTensor*, compute input): self in the formula, aclTensor on the device. Non-contiguous tensors are supported. The data format can be ND. The shape cannot be greater than 8 dimensions, and must be the same as that of out. The data type must meet the type deduction rules with out.Atlas inference products andAtlas training products : The data type can be INT8, INT16, INT32, INT64, UINT8, BOOL, FLOAT, FLOAT16, or DOUBLE.Atlas A2 training products/Atlas A2 inference products andAtlas A3 training products/Atlas A3 inference products : The data type can be INT8, INT16, INT32, INT64, UINT8, BOOL, FLOAT, FLOAT16, DOUBLE, or BFLOAT16.
out(aclTensor *, compute output): out in the formula, which is an aclTensor on the device. Non-contiguous tensors are supported. The data format supports ND. The shape must be the same as that of self.Atlas inference products andAtlas training products : The data type can be FLOAT, FLOAT16, or DOUBLE.Atlas A2 training products/Atlas A2 inference products andAtlas A3 training products/Atlas A3 inference products : The data type can be FLOAT, FLOAT16, DOUBLE, or BFLOAT16.
workspaceSize(uint64_t *, output): size of the workspace to be allocated on the device.executor(aclOpExecutor **, output): operator executor, containing the operator computation process.
Returns:
aclnnStatus: status code. For details, see aclnn Return Codes.The first-phase API implements input parameter verification. The following errors may be thrown. 161001 (ACLNN_ERR_PARAM_NULLPTR): 1. The passed self or out is a null pointer. 161002 (ACLNN_ERR_PARAM_INVALID): 1. The data type or format of self or out is not supported. 2. The data types of self and out do not meet the type deduction rules. 3. The number of dimensions of Self or Out exceeds 8. 4. The shapes of self and out are inconsistent.
aclnnLog1p
Parameters:
workspace(void *, input): address of the workspace to be allocated on the device.workspaceSize(uint64_t, input): size of the workspace to be allocated on the device, which is obtained by calling first-phase APIaclnnLog1pGetWorkspaceSize.executor(aclOpExecutor*, input): operator executor, containing the operator computation process.stream(aclrtStream, input parameter): stream for executing the task.
Returns:
aclnnStatus: status code. For details, see aclnn Return Codes.
aclnnInplaceLog1pGetWorkspaceSize
Parameters:
selfRef(aclTensor*, computation input | computation output): aclTensor on the device. Non-contiguous tensors are supported. The data format supports ND. The shape cannot be greater than 8D.Atlas inference products andAtlas training products : The data type can be FLOAT, FLOAT16, or DOUBLE.Atlas A2 training products/Atlas A2 inference products andAtlas A3 training products/Atlas A3 inference products : The data type can be FLOAT, FLOAT16, DOUBLE, or BFLOAT16.
workspaceSize(uint64_t *, output): size of the workspace to be allocated on the device.executor(aclOpExecutor **, output): operator executor, containing the operator computation process.
Returns:
aclnnStatus: status code. For details, see aclnn Return Codes.
The first-phase API implements input parameter verification. The following errors may be thrown.
161001 (ACLNN_ERR_PARAM_NULLPTR): 1. The passed selfRef is a null pointer.
161002 (ACLNN_ERR_PARAM_INVALID): 1. The data type or format of selfRef is not supported.
2. The dimensions of selfRef are greater than 8.aclnnInplaceLog1p
Parameters:
workspace(void *, input): address of the workspace to be allocated on the device.workspaceSize(uint64_t, input): size of the workspace to be allocated on the device, which is obtained by calling first-phase APIaclnnInplaceLog1pGetWorkspaceSize.executor(aclOpExecutor*, input): operator executor, containing the operator computation process.stream(aclrtStream, input parameter): stream for executing the task.
Returns:
aclnnStatus: status code. For details, see aclnn Return Codes.
Constraints
- Deterministic computation:
aclnnLog1pandaclnnInplaceLog1pdefault to deterministic implementation.
Calling Examples
The following examples are for reference only. For details, see Compilation and Running Sample.
#include <iostream>
#include <vector>
#include "acl/acl.h"
#include "aclnnop/aclnn_log1p.h"
#define CHECK_RET(cond, return_expr) \
do { \
if (!(cond)) { \
return_expr; \
} \
} while (0)
#define LOG_PRINT(message, ...) \
do { \
printf(message, ##__VA_ARGS__); \
} while (0)
int64_t GetShapeSize(const std::vector<int64_t>& shape) {
int64_t shapeSize = 1;
for (auto i : shape) {
shapeSize *= i;
}
return shapeSize;
}
int Init(int32_t deviceId, aclrtStream* stream) {
// (Boilerplate) Initialize resources.
auto ret = aclInit(nullptr);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclInit failed. ERROR: %d\n", ret); return ret);
ret = aclrtSetDevice(deviceId);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclrtSetDevice failed. ERROR: %d\n", ret); return ret);
ret = aclrtCreateStream(stream);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclrtCreateStream failed. ERROR: %d\n", ret); return ret);
return 0;
}
template <typename T>
int CreateAclTensor(const std::vector<T>& hostData, const std::vector<int64_t>& shape, void** deviceAddr,
aclDataType dataType, aclTensor** tensor) {
auto size = GetShapeSize(shape) * sizeof(T);
// Call aclrtMalloc to allocate memory on the device.
auto ret = aclrtMalloc(deviceAddr, size, ACL_MEM_MALLOC_HUGE_FIRST);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclrtMalloc failed. ERROR: %d\n", ret); return ret);
// Call aclrtMemcpy to copy the data on the host to the memory on the device.
ret = aclrtMemcpy(*deviceAddr, size, hostData.data(), size, ACL_MEMCPY_HOST_TO_DEVICE);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclrtMemcpy failed. ERROR: %d\n", ret); return ret);
// Compute the strides of the contiguous tensor.
std::vector<int64_t> strides(shape.size(), 1);
for (int64_t i = shape.size() - 2; i >= 0; i--) {
strides[i] = shape[i + 1] * strides[i + 1];
}
// Call aclCreateTensor to create an aclTensor.
*tensor = aclCreateTensor(shape.data(), shape.size(), dataType, strides.data(), 0, aclFormat::ACL_FORMAT_ND, shape.data(), shape.size(), *deviceAddr);
return 0;
}
int main() {
// 1. (Boilerplate) Initialize the device and stream. For details, see the ACL API manual.
// Set the device ID in use.
int32_t deviceId = 0;
aclrtStream stream;
auto ret = Init(deviceId, &stream);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("Init acl failed. ERROR: %d\n", ret); return ret);
// 2. Construct the input and output based on the API.
std::vector<int64_t> selfShape = {4, 2};
std::vector<int64_t> outShape = {4, 2};
void* selfDeviceAddr = nullptr;
void* outDeviceAddr = nullptr;
aclTensor* self = nullptr;
aclTensor* out = nullptr;
std::vector<float> selfHostData = {0, 1, 2, 3, 4, 5, 6, 7};
std::vector<float> outHostData = {0, 0, 0, 0, 0, 0, 0, 0};
// Create a self aclTensor.
ret = CreateAclTensor(selfHostData, selfShape, &selfDeviceAddr, aclDataType::ACL_FLOAT, &self);
CHECK_RET(ret == ACL_SUCCESS, return ret);
// Create an out aclTensor.
ret = CreateAclTensor(outHostData, outShape, &outDeviceAddr, aclDataType::ACL_FLOAT, &out);
CHECK_RET(ret == ACL_SUCCESS, return ret);
// aclnnLog1p calling example
// 3. Call the CANN operator library API.
uint64_t workspaceSize = 0;
aclOpExecutor* executor;
// Call the first-phase API of aclnnLog1p.
ret = aclnnLog1pGetWorkspaceSize(self, out, &workspaceSize, &executor);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclnnLog1pGetWorkspaceSize failed. ERROR: %d\n", ret); return ret);
// Allocate device memory based on workspaceSize calculated by the first-phase API.
void* workspaceAddr = nullptr;
if (workspaceSize > 0) {
ret = aclrtMalloc(&workspaceAddr, workspaceSize, ACL_MEM_MALLOC_HUGE_FIRST);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("allocate workspace failed. ERROR: %d\n", ret); return ret);
}
// Call the second-phase API of aclnnLog1p.
ret = aclnnLog1p(workspaceAddr, workspaceSize, executor, stream);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclnnLog1p failed. ERROR: %d\n", ret); return ret);
// 4. (Boilerplate) Wait until the task execution is complete.
ret = aclrtSynchronizeStream(stream);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclrtSynchronizeStream failed. ERROR: %d\n", ret); return ret);
// 5. Obtain the output value and copy the result from the device memory to the host. Modify the code based on the API definition.
auto size = GetShapeSize(outShape);
std::vector<float> resultData(size, 0);
ret = aclrtMemcpy(resultData.data(), resultData.size() * sizeof(resultData[0]), outDeviceAddr,
size * sizeof(resultData[0]), ACL_MEMCPY_DEVICE_TO_HOST);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("copy result from device to host failed. ERROR: %d\n", ret); return ret);
for (int64_t i = 0; i < size; i++) {
LOG_PRINT("result[%ld] is: %f\n", i, resultData[i]);
}
// aclnnInplaceLog1p API calling example
// 3. Call the CANN operator library API.
LOG_PRINT("\ntest aclnnInplaceLog1p\n");
// Call the first-phase API of aclnnInplaceLog1p.
ret = aclnnInplaceLog1pGetWorkspaceSize(self, &workspaceSize, &executor);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclnnInplaceLog1pGetWorkspaceSize failed. ERROR: %d\n", ret); return ret);
// Allocate device memory based on workspaceSize calculated by the first-phase API.
if (workspaceSize > 0) {
ret = aclrtMalloc(&workspaceAddr, workspaceSize, ACL_MEM_MALLOC_HUGE_FIRST);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("allocate workspace failed. ERROR: %d\n", ret); return ret);
}
// Call the second-phase API of aclnnInplaceLog1p.
ret = aclnnInplaceLog1p(workspaceAddr, workspaceSize, executor, stream);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclnnInplaceLog1p failed. ERROR: %d\n", ret); return ret);
// 4. (Boilerplate) Wait until the task execution is complete.
ret = aclrtSynchronizeStream(stream);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclrtSynchronizeStream failed. ERROR: %d\n", ret); return ret);
// 5. Obtain the output value and copy the result from the device memory to the host. Modify the code based on the API definition.
ret = aclrtMemcpy(resultData.data(), resultData.size() * sizeof(resultData[0]), selfDeviceAddr,
size * sizeof(resultData[0]), ACL_MEMCPY_DEVICE_TO_HOST);
CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("copy result from device to host failed. ERROR: %d\n", ret); return ret);
for (int64_t i = 0; i < size; i++) {
LOG_PRINT("result[%ld] is: %f\n", i, resultData[i]);
}
// 6. Destroy aclTensor and aclScalar. Modify the code based on the API definition.
aclDestroyTensor(self);
aclDestroyTensor(out);
// 7. Destroy device resources. Modify the code based on the API definition.
aclrtFree(selfDeviceAddr);
aclrtFree(outDeviceAddr);
if (workspaceSize > 0) {
aclrtFree(workspaceAddr);
}
aclrtDestroyStream(stream);
aclrtResetDevice(deviceId);
aclFinalize();
return 0;
}