[2025-11-06 19:37:11,840] [21108] [140473314031424] [llmmodels] [INFO] [run_pa.py-342] : pa_runner: PARunner(model_path=/models/Qwen2.5-VL-7B-Instruct, input_text=Explain the contents of the picture with more than 500 words and do not Answer the question using a single word or phrase., max_position_embeddings=None, max_input_length=4096, max_output_length=256, max_prefill_tokens=-1, load_tokenizer=True, enable_atb_torch=False, max_prefill_batch_size=None, max_batch_size=1, dtype=torch.float16, block_size=128, model_config=ModelConfig(num_heads=28, num_kv_heads=4, num_kv_heads_origin=4, head_size=128, k_head_size=128, v_head_size=128, num_layers=28, device=npu:0, dtype=torch.float16, soc_info=NPUSocInfo(soc_name='', soc_version=200, need_nz=True, matmul_nd_nz=False), kv_quant_type=None, fa_quant_type=None, mapping=Mapping(world_size=1, rank=0, num_nodes=1,pp_rank=0, pp_groups=[[0]], micro_batch_size=1, attn_dp_groups=[[0]], attn_tp_groups=[[0]], attn_inner_sp_groups=[[0]], attn_cp_groups=[[0]], attn_o_proj_tp_groups=[[0]], mlp_tp_groups=[[0]], moe_ep_groups=[[0]], moe_tp_groups=[[0]]), cla_share_factor=1, model_type=qwen2_5_vl, enable_nz=False), max_memory=93926195200,
[2025-11-06 19:37:11,886] [21108] [140473314031424] [llmmodels] [INFO] [run_pa.py-122] : ---------------Begin warm_up---------------
[2025-11-06 19:37:11,887] [21108] [140473314031424] [llmmodels] [INFO] [cache.py-154] : kv cache will allocate 0.232421875GB memory
[2025-11-06 19:37:11,979] [21108] [140473314031424] [llmmodels] [INFO] [generate.py-1139] : ------total req num: 1, infer start--------
[2025-11-06 19:37:12.218] [21108] [140472328595008] [llmmodels] [ERROR] [graph_operation.cpp:434] [MIE05E000002] Qwen25VL_VIT_graph nodes[15] atb operation execute fail, error:4
Traceback (most recent call last):
File "/usr/lib/python3.10/runpy.py", line 196, in _run_module_as_main
return _run_code(code, main_globals, None,
File "/usr/lib/python3.10/runpy.py", line 86, in _run_code
exec(code, run_globals)
File "/usr/local/Ascend/atb-models/examples/models/qwen2_vl/run_pa.py", line 343, in <module>
pa_runner.warm_up()
File "/usr/local/Ascend/atb-models/examples/models/qwen2_vl/run_pa.py", line 141, in warm_up
generate_req([single_req], self.model, self.max_batch_size, self.max_prefill_tokens, self.cache_manager)
File "/usr/local/Ascend/atb-models/examples/server/generate.py", line 1227, in generate_req
generate_token_with_clocking(model, cache_manager, batch, eplb_forwarder)
File "/usr/local/Ascend/atb-models/examples/server/generate.py", line 879, in generate_token_with_clocking
res = generate_token(model, cache_manager, input_batch_in, eplb_forwarder)
File "/usr/local/Ascend/atb-models/examples/server/generate.py", line 617, in generate_token
logits = model.forward(
File "/usr/local/Ascend/atb-models/atb_llm/runner/model_runner.py", line 310, in forward
res = self.model.forward(**kwargs)
File "/usr/local/Ascend/atb-models/atb_llm/models/qwen2_vl/flash_causal_qwen2_vl.py", line 118, in forward
inputs_embeds, image_grid_thw, video_grid_thw, second_per_grid_ts = self.prepare_prefill_token_service(
File "/usr/local/Ascend/atb-models/atb_llm/models/qwen2_vl/flash_causal_qwen2_vl.py", line 197, in prepare_prefill_token_service
image_features = self.vision_tower(image_pixel.to(self.vision_tower.dtype), image_grid_thw)
File "/usr/local/lib/python3.10/dist-packages/torch/nn/modules/module.py", line 1518, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
File "/usr/local/lib/python3.10/dist-packages/torch/nn/modules/module.py", line 1527, in _call_impl
return forward_call(*args, **kwargs)
File "/usr/local/Ascend/atb-models/atb_llm/models/qwen2_vl/modeling_qwen2_5_vl_vit_atb.py", line 872, in forward
vision_features = self.encoder(
File "/usr/local/lib/python3.10/dist-packages/torch/nn/modules/module.py", line 1518, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
File "/usr/local/lib/python3.10/dist-packages/torch/nn/modules/module.py", line 1527, in _call_impl
return forward_call(*args, **kwargs)
File "/usr/local/Ascend/atb-models/atb_llm/models/qwen2_vl/modeling_qwen2_5_vl_vit_atb.py", line 719, in forward
hidden_states = self.graph.forward(self.graph_inputs, self.graph_outputs, self.graph_param)
RuntimeError: The Inner error is reported as above. The process exits for this inner error, and the current working operator name is Qwen25VL_VIT_graph.
Since the operator is called asynchronously, the stacktrace may be inaccurate. If you want to get the accurate stacktrace, please set the environment variable ASCEND_LAUNCH_BLOCKING=1.
Note: ASCEND_LAUNCH_BLOCKING=1 will force ops to run in synchronous mode, resulting in performance degradation. Please unset ASCEND_LAUNCH_BLOCKING in time after debugging.
[ERROR] 2025-11-06-19:37:12 (PID:21108, Device:0, RankID:0) ERR00100 PTA call acl api failed.
EZ9999: Inner Error!
EZ9999: [PID: 21108] 2025-11-06-19:37:12.199.194 Kernel task happen error, retCode=0x26, [aicore exception].[FUNC:PreCheckTaskErr][FILE:davinci_kernel_task.cc][LINE:1539]
TraceBack (most recent call last):
The error from device(0), serial number is 2, there is an error of aicore, core id is 7, error code = 0x200000000, dump info: pc start: 0x8001240013f2a34, current: 0x1240013f2dc8, vec error info: 0xc6ef89f, mte error info: 0x4f, ifu error info: 0x146ffbffb7780, ccu error info: 0xdc076b8f004ffdd7, cube error info: 0xcb, biu error info: 0, aic error mask: 0x65000200d00028c, para base: 0x12c1400bc800, errorStr: The write address of the MTE instruction is out of range.[FUNC:PrintCoreErrorInfo][FILE:device_error_core_proc.cc][LINE:46]
The extend info from device(0), serial number is 2, there is aicore error, core id is 7, aicore int: 0x10, aicore error2: 0, axi clamp ctrl: 0, axi clamp state: 0x1717, biu status0: 0x101d14000000000, biu status1: 0x80000201020000, clk gate mask: 0, dbg addr: 0, ecc en: 0, mte ccu ecc 1bit error: 0x580000000000000, vector cube ecc 1bit error: 0, run stall: 0x1, dbg data0: 0, dbg data1: 0, dbg data2: 0, dbg data3: 0, dfx data: 0x8d[FUNC:PrintCoreErrorInfo][FILE:device_error_core_proc.cc][LINE:77]
The device(0), core list[0-0], error code is:[FUNC:PrintCoreInfoErrMsg][FILE:device_error_core_proc.cc][LINE:100]
coreId( 0): 0x200000000 [FUNC:PrintCoreInfoErrMsg][FILE:device_error_core_proc.cc][LINE:114]
[AIC_INFO] after execute:args print end[FUNC:GetError][FILE:stream.cc][LINE:1183]
Aicore kernel execute failed, device_id=0, stream_id=6, report_stream_id=6, task_id=1778, flip_num=0, fault kernel_name=flash_attention_prefill_0, fault kernel info ext=none, program id=193, hash=8941720987216483178.[FUNC:GetError][FILE:stream.cc][LINE:1183]
Failed to submit kernel task, retCode=0x7150026.[FUNC:LaunchKernelSubmit][FILE:context.cc][LINE:1206]
kernel launch submit failed.[FUNC:LaunchKernel][FILE:context.cc][LINE:1329]
rtKernelLaunchWithFlagV2 execute failed, reason=[aicore exception][FUNC:FuncErrorReason][FILE:error_message_manage.cc][LINE:53]
DEVICE[0] PID[21108]:
EXCEPTION STREAM:
Exception info:TGID=21108, model id=65535, stream id=6, stream phase=3
Message info[0]:RTS_HWTS: Aicore exception, slot_id=6, stream_id=6
Other info[0]:time=2025-11-06-19:37:11.960.615, function=process_hwts_error_exception, line=646, error code=0x26
[W NPUStream.cpp:526] Warning: NPU warning, error code is 507015[Error]:
[Error]: The aicore execution is abnormal.
Rectify the fault based on the error information in the ascend log.
EH9999: Inner Error!
rtDeviceSynchronizeWithTimeout execute failed, reason=[aicore exception][FUNC:FuncErrorReason][FILE:error_message_manage.cc][LINE:53]
EH9999: [PID: 21108] 2025-11-06-19:37:12.380.981 wait for compute device to finish failed, runtime result = 507015.[FUNC:ReportCallError][FILE:log_inner.cpp][LINE:161]
TraceBack (most recent call last):
(function npuSynchronizeUsedDevices)
[W NPUStream.cpp:508] Warning: NPU warning, error code is 507015[Error]:
[Error]: The aicore execution is abnormal.
Rectify the fault based on the error information in the ascend log.
EH9999: Inner Error!
rtDeviceSynchronizeWithTimeout execute failed, reason=[aicore exception][FUNC:FuncErrorReason][FILE:error_message_manage.cc][LINE:53]
EH9999: [PID: 21108] 2025-11-06-19:37:12.383.631 wait for compute device to finish failed, runtime result = 507015.[FUNC:ReportCallError][FILE:log_inner.cpp][LINE:161]
TraceBack (most recent call last):
(function npuSynchronizeDevice)
[W NPUWorkspaceAllocator.cpp:228] Warning: NPU warning, error code is 507015[Error]:
[Error]: The aicore execution is abnormal.
Rectify the fault based on the error information in the ascend log.
EH9999: Inner Error!
rtDeviceSynchronizeWithTimeout execute failed, reason=[aicore exception][FUNC:FuncErrorReason][FILE:error_message_manage.cc][LINE:53]
EH9999: [PID: 21108] 2025-11-06-19:37:12.385.117 wait for compute device to finish failed, runtime result = 507015.[FUNC:ReportCallError][FILE:log_inner.cpp][LINE:161]
TraceBack (most recent call last):
(function empty_cache)
/usr/lib/python3.10/multiprocessing/resource_tracker.py:224: UserWarning: resource_tracker: There appear to be 1 leaked shared_memory objects to clean up at shutdown
warnings.warn('resource_tracker: There appear to be %d '
[2025-11-06 19:37:25,110] torch.distributed.elastic.multiprocessing.api: [ERROR] failed (exitcode: 1) local_rank: 0 (pid: 21108) of binary: /usr/bin/python3
Traceback (most recent call last):
File "/usr/local/bin/torchrun", line 8, in <module>
sys.exit(main())
File "/usr/local/lib/python3.10/dist-packages/torch/distributed/elastic/multiprocessing/errors/__init__.py", line 346, in wrapper
return f(*args, **kwargs)
File "/usr/local/lib/python3.10/dist-packages/torch/distributed/run.py", line 806, in main
run(args)
File "/usr/local/lib/python3.10/dist-packages/torch/distributed/run.py", line 797, in run
elastic_launch(
File "/usr/local/lib/python3.10/dist-packages/torch/distributed/launcher/api.py", line 134, in __call__
return launch_agent(self._config, self._entrypoint, list(args))
File "/usr/local/lib/python3.10/dist-packages/torch/distributed/launcher/api.py", line 264, in launch_agent
raise ChildFailedError(
torch.distributed.elastic.multiprocessing.errors.ChildFailedError:
============================================================
examples.models.qwen2_vl.run_pa FAILED
------------------------------------------------------------
Failures:
<NO_OTHER_FAILURES>
------------------------------------------------------------
Root Cause (first observed failure):
[0]:
time : 2025-11-06_19:37:25
host : Ubuntu22.04.5
rank : 0 (local_rank: 0)
exitcode : 1 (pid: 21108)
error_file: <N/A>
traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html
============================================================
[2025-11-06 19:37:11,840] [21108] [140473314031424] [llmmodels] [INFO] [run_pa.py-342] : pa_runner: PARunner(model_path=/models/Qwen2.5-VL-7B-Instruct, input_text=Explain the contents of the picture with more than 500 words and do not Answer the question using a single word or phrase., max_position_embeddings=None, max_input_length=4096, max_output_length=256, max_prefill_tokens=-1, load_tokenizer=True, enable_atb_torch=False, max_prefill_batch_size=None, max_batch_size=1, dtype=torch.float16, block_size=128, model_config=ModelConfig(num_heads=28, num_kv_heads=4, num_kv_heads_origin=4, head_size=128, k_head_size=128, v_head_size=128, num_layers=28, device=npu:0, dtype=torch.float16, soc_info=NPUSocInfo(soc_name='', soc_version=200, need_nz=True, matmul_nd_nz=False), kv_quant_type=None, fa_quant_type=None, mapping=Mapping(world_size=1, rank=0, num_nodes=1,pp_rank=0, pp_groups=[[0]], micro_batch_size=1, attn_dp_groups=[[0]], attn_tp_groups=[[0]], attn_inner_sp_groups=[[0]], attn_cp_groups=[[0]], attn_o_proj_tp_groups=[[0]], mlp_tp_groups=[[0]], moe_ep_groups=[[0]], moe_tp_groups=[[0]]), cla_share_factor=1, model_type=qwen2_5_vl, enable_nz=False), max_memory=93926195200, [2025-11-06 19:37:11,886] [21108] [140473314031424] [llmmodels] [INFO] [run_pa.py-122] : ---------------Begin warm_up--------------- [2025-11-06 19:37:11,887] [21108] [140473314031424] [llmmodels] [INFO] [cache.py-154] : kv cache will allocate 0.232421875GB memory [2025-11-06 19:37:11,979] [21108] [140473314031424] [llmmodels] [INFO] [generate.py-1139] : ------total req num: 1, infer start-------- [2025-11-06 19:37:12.218] [21108] [140472328595008] [llmmodels] [ERROR] [graph_operation.cpp:434] [MIE05E000002] Qwen25VL_VIT_graph nodes[15] atb operation execute fail, error:4 Traceback (most recent call last): File "/usr/lib/python3.10/runpy.py", line 196, in _run_module_as_main return _run_code(code, main_globals, None, File "/usr/lib/python3.10/runpy.py", line 86, in _run_code exec(code, run_globals) File "/usr/local/Ascend/atb-models/examples/models/qwen2_vl/run_pa.py", line 343, in <module> pa_runner.warm_up() File "/usr/local/Ascend/atb-models/examples/models/qwen2_vl/run_pa.py", line 141, in warm_up generate_req([single_req], self.model, self.max_batch_size, self.max_prefill_tokens, self.cache_manager) File "/usr/local/Ascend/atb-models/examples/server/generate.py", line 1227, in generate_req generate_token_with_clocking(model, cache_manager, batch, eplb_forwarder) File "/usr/local/Ascend/atb-models/examples/server/generate.py", line 879, in generate_token_with_clocking res = generate_token(model, cache_manager, input_batch_in, eplb_forwarder) File "/usr/local/Ascend/atb-models/examples/server/generate.py", line 617, in generate_token logits = model.forward( File "/usr/local/Ascend/atb-models/atb_llm/runner/model_runner.py", line 310, in forward res = self.model.forward(**kwargs) File "/usr/local/Ascend/atb-models/atb_llm/models/qwen2_vl/flash_causal_qwen2_vl.py", line 118, in forward inputs_embeds, image_grid_thw, video_grid_thw, second_per_grid_ts = self.prepare_prefill_token_service( File "/usr/local/Ascend/atb-models/atb_llm/models/qwen2_vl/flash_causal_qwen2_vl.py", line 197, in prepare_prefill_token_service image_features = self.vision_tower(image_pixel.to(self.vision_tower.dtype), image_grid_thw) File "/usr/local/lib/python3.10/dist-packages/torch/nn/modules/module.py", line 1518, in _wrapped_call_impl return self._call_impl(*args, **kwargs) File "/usr/local/lib/python3.10/dist-packages/torch/nn/modules/module.py", line 1527, in _call_impl return forward_call(*args, **kwargs) File "/usr/local/Ascend/atb-models/atb_llm/models/qwen2_vl/modeling_qwen2_5_vl_vit_atb.py", line 872, in forward vision_features = self.encoder( File "/usr/local/lib/python3.10/dist-packages/torch/nn/modules/module.py", line 1518, in _wrapped_call_impl return self._call_impl(*args, **kwargs) File "/usr/local/lib/python3.10/dist-packages/torch/nn/modules/module.py", line 1527, in _call_impl return forward_call(*args, **kwargs) File "/usr/local/Ascend/atb-models/atb_llm/models/qwen2_vl/modeling_qwen2_5_vl_vit_atb.py", line 719, in forward hidden_states = self.graph.forward(self.graph_inputs, self.graph_outputs, self.graph_param) RuntimeError: The Inner error is reported as above. The process exits for this inner error, and the current working operator name is Qwen25VL_VIT_graph. Since the operator is called asynchronously, the stacktrace may be inaccurate. If you want to get the accurate stacktrace, please set the environment variable ASCEND_LAUNCH_BLOCKING=1. Note: ASCEND_LAUNCH_BLOCKING=1 will force ops to run in synchronous mode, resulting in performance degradation. Please unset ASCEND_LAUNCH_BLOCKING in time after debugging. [ERROR] 2025-11-06-19:37:12 (PID:21108, Device:0, RankID:0) ERR00100 PTA call acl api failed. EZ9999: Inner Error! EZ9999: [PID: 21108] 2025-11-06-19:37:12.199.194 Kernel task happen error, retCode=0x26, [aicore exception].[FUNC:PreCheckTaskErr][FILE:davinci_kernel_task.cc][LINE:1539] TraceBack (most recent call last): The error from device(0), serial number is 2, there is an error of aicore, core id is 7, error code = 0x200000000, dump info: pc start: 0x8001240013f2a34, current: 0x1240013f2dc8, vec error info: 0xc6ef89f, mte error info: 0x4f, ifu error info: 0x146ffbffb7780, ccu error info: 0xdc076b8f004ffdd7, cube error info: 0xcb, biu error info: 0, aic error mask: 0x65000200d00028c, para base: 0x12c1400bc800, errorStr: The write address of the MTE instruction is out of range.[FUNC:PrintCoreErrorInfo][FILE:device_error_core_proc.cc][LINE:46] The extend info from device(0), serial number is 2, there is aicore error, core id is 7, aicore int: 0x10, aicore error2: 0, axi clamp ctrl: 0, axi clamp state: 0x1717, biu status0: 0x101d14000000000, biu status1: 0x80000201020000, clk gate mask: 0, dbg addr: 0, ecc en: 0, mte ccu ecc 1bit error: 0x580000000000000, vector cube ecc 1bit error: 0, run stall: 0x1, dbg data0: 0, dbg data1: 0, dbg data2: 0, dbg data3: 0, dfx data: 0x8d[FUNC:PrintCoreErrorInfo][FILE:device_error_core_proc.cc][LINE:77] The device(0), core list[0-0], error code is:[FUNC:PrintCoreInfoErrMsg][FILE:device_error_core_proc.cc][LINE:100] coreId( 0): 0x200000000 [FUNC:PrintCoreInfoErrMsg][FILE:device_error_core_proc.cc][LINE:114] [AIC_INFO] after execute:args print end[FUNC:GetError][FILE:stream.cc][LINE:1183] Aicore kernel execute failed, device_id=0, stream_id=6, report_stream_id=6, task_id=1778, flip_num=0, fault kernel_name=flash_attention_prefill_0, fault kernel info ext=none, program id=193, hash=8941720987216483178.[FUNC:GetError][FILE:stream.cc][LINE:1183] Failed to submit kernel task, retCode=0x7150026.[FUNC:LaunchKernelSubmit][FILE:context.cc][LINE:1206] kernel launch submit failed.[FUNC:LaunchKernel][FILE:context.cc][LINE:1329] rtKernelLaunchWithFlagV2 execute failed, reason=[aicore exception][FUNC:FuncErrorReason][FILE:error_message_manage.cc][LINE:53] DEVICE[0] PID[21108]: EXCEPTION STREAM: Exception info:TGID=21108, model id=65535, stream id=6, stream phase=3 Message info[0]:RTS_HWTS: Aicore exception, slot_id=6, stream_id=6 Other info[0]:time=2025-11-06-19:37:11.960.615, function=process_hwts_error_exception, line=646, error code=0x26 [W NPUStream.cpp:526] Warning: NPU warning, error code is 507015[Error]: [Error]: The aicore execution is abnormal. Rectify the fault based on the error information in the ascend log. EH9999: Inner Error! rtDeviceSynchronizeWithTimeout execute failed, reason=[aicore exception][FUNC:FuncErrorReason][FILE:error_message_manage.cc][LINE:53] EH9999: [PID: 21108] 2025-11-06-19:37:12.380.981 wait for compute device to finish failed, runtime result = 507015.[FUNC:ReportCallError][FILE:log_inner.cpp][LINE:161] TraceBack (most recent call last): (function npuSynchronizeUsedDevices) [W NPUStream.cpp:508] Warning: NPU warning, error code is 507015[Error]: [Error]: The aicore execution is abnormal. Rectify the fault based on the error information in the ascend log. EH9999: Inner Error! rtDeviceSynchronizeWithTimeout execute failed, reason=[aicore exception][FUNC:FuncErrorReason][FILE:error_message_manage.cc][LINE:53] EH9999: [PID: 21108] 2025-11-06-19:37:12.383.631 wait for compute device to finish failed, runtime result = 507015.[FUNC:ReportCallError][FILE:log_inner.cpp][LINE:161] TraceBack (most recent call last): (function npuSynchronizeDevice) [W NPUWorkspaceAllocator.cpp:228] Warning: NPU warning, error code is 507015[Error]: [Error]: The aicore execution is abnormal. Rectify the fault based on the error information in the ascend log. EH9999: Inner Error! rtDeviceSynchronizeWithTimeout execute failed, reason=[aicore exception][FUNC:FuncErrorReason][FILE:error_message_manage.cc][LINE:53] EH9999: [PID: 21108] 2025-11-06-19:37:12.385.117 wait for compute device to finish failed, runtime result = 507015.[FUNC:ReportCallError][FILE:log_inner.cpp][LINE:161] TraceBack (most recent call last): (function empty_cache) /usr/lib/python3.10/multiprocessing/resource_tracker.py:224: UserWarning: resource_tracker: There appear to be 1 leaked shared_memory objects to clean up at shutdown warnings.warn('resource_tracker: There appear to be %d ' [2025-11-06 19:37:25,110] torch.distributed.elastic.multiprocessing.api: [ERROR] failed (exitcode: 1) local_rank: 0 (pid: 21108) of binary: /usr/bin/python3 Traceback (most recent call last): File "/usr/local/bin/torchrun", line 8, in <module> sys.exit(main()) File "/usr/local/lib/python3.10/dist-packages/torch/distributed/elastic/multiprocessing/errors/__init__.py", line 346, in wrapper return f(*args, **kwargs) File "/usr/local/lib/python3.10/dist-packages/torch/distributed/run.py", line 806, in main run(args) File "/usr/local/lib/python3.10/dist-packages/torch/distributed/run.py", line 797, in run elastic_launch( File "/usr/local/lib/python3.10/dist-packages/torch/distributed/launcher/api.py", line 134, in __call__ return launch_agent(self._config, self._entrypoint, list(args)) File "/usr/local/lib/python3.10/dist-packages/torch/distributed/launcher/api.py", line 264, in launch_agent raise ChildFailedError( torch.distributed.elastic.multiprocessing.errors.ChildFailedError: ============================================================ examples.models.qwen2_vl.run_pa FAILED ------------------------------------------------------------ Failures: <NO_OTHER_FAILURES> ------------------------------------------------------------ Root Cause (first observed failure): [0]: time : 2025-11-06_19:37:25 host : Ubuntu22.04.5 rank : 0 (local_rank: 0) exitcode : 1 (pid: 21108) error_file: <N/A> traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html ============================================================