--precision_mode_v2

Applicable Products

All processors

Description

Sets the precision mode of a model.

See Also

  • This option cannot be used together with --precision_mode. You are advised to use --precision_mode_v2. The --precision_mode_v2 option is added in the new version. The semantics of the option value is clearer and easier to understand.
  • When this option is set to mixed_float16, mixed_bfloat16, or mixed_hif8, to make adjustments based on the built-in tuning policy and manually specify which operators allow precision reduction and which operators disallow precision reduction, follow the instructions in --modify_mixlist.
  • In the inference scenario, --precision_mode_v2 can be used to set the global precision mode of a network model, but it may result in performance or accuracy problems on particular operators. Therefore, you can use --keep_dtype to keep the computation precision of these operators unchanged during the build of the original network model. However, --keep_dtype does not take effect when --precision_mode_v2 is set to origin.

Arguments

  • fp16 (default):

    Forces operators in the original graph to use float16, regardless of whether their original computation precision is float16, bfloat16, or float32.

  • origin:

    Retain the original precision.

    • If the precision of an operator in the original graph is float16, and the implementation of the operator in the AI Core does not support float16 but supports only float32 and bfloat16, the system automatically uses high-precision float32.
    • If the precision of an operator in the original graph is float16, and the implementation of the operator in the AI Core does not support float16 but supports only bfloat16, the AI CPU operator of float16 is used. If the AI CPU operator is not supported, an error is reported.
    • If the precision of an operator in the original graph is float32, and the implementation of the operator in the AI Core does not support float32 but supports only float16, the AI CPU operator of float32 is used. If the AI CPU operator is not supported, an error is reported.
  • cube_fp16in_fp32out:
    Selects different processing modes depending on the operator type when an operator supports both the float32 and float16 data types.
    • For cube operators, the system processes the computation based on the operator implementation.
      1. The preferred input data type is float16 and the output data type is float32.
      2. If the float16 input data and float32 output data types are not supported, set both the input and output data types to float32.
      3. If the float32 input and output data types are not supported, set both the input and output data types to float16.
      4. If the float16 input and output data types are not supported, an error is reported.
    • Forces vector compute operators in the original graph to use float32, regardless of whether their original computation precision is float16 or bfloat16.

      This option takes no effect if the original graph contains operators whose implementation on the AI Core does not support float32, for example, an operator that supports only float16. In this case, the supported float16 is used. If the operator implementation on the AI Core does not support float32 and the blocklist mode is enabled (by setting precision_reduce to false), the float32 AI CPU operator is used. If the AI CPU operator does not support float32 either, an error is reported.

  • mixed_float16:

    Mixed precision of float16, bfloat16, and float32 is used for neural network processing. For float32 and bfloat16 operators in the original graph, float16 is automatically used for certain float32 and bfloat16 operators based on the built-in tuning policy, which improves system performance and reduces memory usage with minimal accuracy loss.

    If this mode is configured, you can view the value of precision_reduce in the built-in tuning policy file ${INSTALL_DIR}/opp/built-in/op_impl/ai_core/tbe/config/xxx/aic-xxx-ops-info-*.json.

    • If it is set to true, the operator is on the trustlist and its precision will be reduced from float32 or bfloat16 to float16.
    • If it is set to false, the operator is on the blocklist and its precision will not be reduced from float32 or bfloat16 to float16. Such operators will continue to use their original precision (float32 or bfloat16).
    • If an operator in the network model does not have precision_reduce configured (that is, it is on the graylist), the mixed-precision handling mechanism for the current operator follows that of the previous operator. That is, if the previous operator supports precision reduction, the current operator also supports it; if the previous operator does not allow precision reduction, the current operator does not either.
  • mixed_bfloat16:

    Mixed precision of bfloat16 and float32 is used for neural network processing. In this mode, bfloat16 is automatically used for certain float32 operators in the original graph based on the built-in tuning policy, which improves system performance and reduces memory usage with minimal accuracy loss. If an operator supports neither bfloat16 nor float32, the AI CPU operator is used. If the AI CPU operator is also unsupported, an error is reported.

    If this mode is configured, you can view the value of precision_reduce in the built-in tuning policy file ${INSTALL_DIR}/opp/built-in/op_impl/ai_core/tbe/config/xxx/aic-xxx-ops-info-*.json.

    • If it is set to true, the operator is on trustlist and its precision will be reduced from float32 to bfloat16.
    • If it is set to false, the operator is on the blocklist and its precision will not be reduced from float32 to bfloat16.
    • If an operator in the network model does not have precision_reduce configured (that is, it is in the graylist), the mixed-precision handling mechanism for the current operator follows that of the previous operator. That is, if the previous operator supports precision reduction, the current operator also supports it; if the previous operator does not allow precision reduction, the current operator does not either.
  • mixed_hif8:

    Enables automatic mixed precision, indicating that hifloat8 (for details about this data type, click here), float16, bfloat16, and float32 are used together for neural network processing. In this mode, hifloat8 is automatically used for certain float16, bfloat16, and float32 operators in the original graph based on the built-in tuning policy, which improves system performance and reduces memory usage with minimal accuracy loss.

    If this mode is configured, you can view the values of precision_reduce in the built-in tuning policy file ${INSTALL_DIR}/opp/built-in/op_impl/ai_core/tbe/config/xxx/aic-xxx-ops-info-*.json.

    • If it is set to true, the operator is on the trustlist and its precision will be reduced from float16, bfloat16, or float32 to hifloat8.
    • If it is set to false, the operator is on the blocklist and its precision will not be reduced from float16, bfloat16, or float32 to hifloat8. In this case, the operator retains its original precision.
    • If an operator in the original graph does not have precision_reduce configured (that is, it is on the graylist), the mixed-precision handling mechanism for the current operator follows that of the previous operator. That is, if the previous operator supports precision reduction, the current operator also supports it; if the previous operator does not allow precision reduction, the current operator does not either.
  • cube_hif8:

    Forces Cube operators that support both hifloat8 and float16, bfloat16, or float32 in the original graph to use hifloat8.

Replace ${INSTALL_DIR} with the CANN component directory. For example, if the installation is performed by the root user, the default file storage path is /usr/local/Ascend/cann. Replace xxx in the preceding path based on the actual product.

Restrictions:

  • The bfloat16 data type supports only the following products:

    Atlas A2 training products / Atlas A2 inference products

    Atlas A3 training products / Atlas A3 inference products

    Atlas 200I/500 A2 inference products

    Ascend 950PR / Ascend 950DT

  • The hif8 data type supports only the following products:

    Ascend 950PR / Ascend 950DT

  • For this option, performance takes priority for the default value and accuracy overflow issues may occur during subsequent inference. If an accuracy issue occurs during inference, locate the fault by referring to Accuracy Improvement Suggestions for Model Inference.
  • To avoid accuracy issues, you can set the option to a value other than the default one, for example, origin.

Suggestions and Benefits

The accuracy and performance of the network model vary according to the configured precision mode.

Sort by precision: origin > mixed_float16 > fp16 > mixed_bfloat16 > mixed_hif8. Sort by performance: mixed_hif8 > mixed_bfloat16 > fp16 ≥ mixed_float16 > origin.

Example

--precision_mode_v2=fp16

Restrictions

In the mixed precision scenario, if the inference performance deteriorates after the version upgrade, you are advised to use the AOE tool to perform optimization again. After the optimization is complete, use the --op_bank_path option to load the path of the custom repository, and then convert the model again.

For details about operator tuning, see AOE Tuning Tool.