WeightQuantBatchMatmulV2TransposeFusionPass

Description

Offloads information about the Transpose nodes connected to WeightQuantBatchMatmulV2 to the transpose_x and transpose_weight attributes.

Restrictions

  • The Transpose nodes connected to antiquant_scale and antiquant_offset can be processed only when the weight node is connected to the Transpose node.
  • This fusion pattern is mandatory. Disabling it will result in functional errors.

Availability

Atlas A2 training product/Atlas A2 inference product

Atlas A3 training product/Atlas A3 inference product

Atlas 350 Accelerator Card