Description: Performs matrix multiplication with float16 input tensor and int8 output tensor.
Formula:
Each operator has calls. First, aclnnBatchMatmulQuantGetWorkspaceSize is called to obtain the input parameters and compute the required workspace size based on the process. Then, aclnnBatchMatmulQuant is called to perform computation.
[object Object]
[object Object]
Parameters:
[object Object]Returns:
aclnnStatus: status code. For details, see .
The first-phase API implements input parameter verification. The following errors may be thrown.
[object Object]
- Determinism:
- [object Object]Atlas training series products[object Object] and [object Object]Atlas inference series products[object Object]: aclnnBatchMatmulQuant defaults to a deterministic implementation.
The following example is for reference only. For details, see .
[object Object]