FusedInferAttentionScoreQuantFusionPass
Description
Fuses FusedInferAttentionScore+AscendQuant into the FusedInferAttentionScore operator in the quantization scenario. The scale and offset parameters of quant are converted into the input parameters quant_scale2 and quant_offset2 of the FIA interface.

Restrictions
- FusedInferAttentionScore supports only fp16 output.
- AscendQuant supports only fp16 input and int8 output.
- quant_scale2 and quant_offset2 of the FusedInferAttentionScore operator must be empty.
Availability
Atlas 350 Accelerator Card
Parent topic: Graph Fusion Patterns