Packs one bit of Adam of the float16 or float32 type into uint8.
Each operator has calls. First, [object Object] is called to obtain the workspace size required for computation and the executor that contains the operator computation process. Then, [object Object] is called to perform computation.
[object Object][object Object]
Parameters:
[object Object](aclTensor*, compute input): 1D tensor used for computation, tensor on the device. Empty tensors are supported. The data type can be FLOAT16 or FLOAT. are supported. The can be ND.[object Object](int64_t, compute input): processing dimension, which is the first dimension of the output tensor during reshaping.[object Object](aclTensor*, compute output): output tensor. Only two-dimension tensors are supported. The data type can be UINT8. If the number of[object Object]elements is not exactly divided by 8, the total length of the output[object Object]is (Number of[object Object]elements/8) + 1. If the number of[object Object]elements is exactly divided by 8, the total length of the output[object Object]is (Number of[object Object]elements/8). The can be ND.[object Object](uint64_t*, output): size of the workspace to be allocated on the device.[object Object](aclOpExecutor**, output): operator executor, containing the operator computation process.
Returns:
Parameters:
[object Object](void *, input): address of the workspace to be allocated on the device.[object Object](uint64_t, input): size of the workspace to be allocated on the device, which is obtained by calling the first-phase API[object Object].[object Object](aclOpExecutor *, input): operator executor, containing the operator computation process.[object Object](aclrtStream, input): stream for executing the task.
Returns:
- Deterministic compute:
[object Object]defaults to a deterministic implementation.- The value of
[object Object]cannot be greater than the total length of the[object Object]output.
The following example is for reference only. For details, see .