版本信息:
MindeIE:1.0.0-300I-Duo-py311-openeuler24.03-lts
大模型:Qwen2.5-14B-Instruct
显卡:Atlas300I Duo
问题现象:
服务化拉起成功,使用curl命令进行单对话测试,响应为空
报错信息如下:
....[E compiler_depend.ts:416] call aclnnSWhere failed, detail:EL0004: [PID: 37513] 2025-04-30-17:08:33.471.645 Failed to allocate memory.
Possible Cause: Available memory is insufficient.
Solution: Close applications not in use.
TraceBack (most recent call last):
Check param failed, mdl can not be null.[FUNC:LaunchKernelPrepare][FILE:context.cc][LINE:769]
kernel launch prepare failed.[FUNC:LaunchKernelWithHandle][FILE:context.cc][LINE:985]
rtKernelLaunchWithHandleV2 execute failed, reason=[module new memory error][FUNC:FuncErrorReason][FILE:error_message_manage.cc][LINE:53]
rtKernelLaunchWithHandleV2 failed: 207001
#### KernelLaunch failed: /usr/local/Ascend/ascend-toolkit/8.0.0/opp/built-in/op_impl/ai_core/tbe//kernel/ascend310p/select_v2/SelectV2_bad39ec50a3854a750466c42b6ef8078_high_performance.o
Kernel Run failed. opType: 17, SelectV2
launch failed for SelectV2, errno:361001.
Init device error msg handler failed, retCode=0x7020016.[FUNC:GetDevErrMsg][FILE:api_impl.cc][LINE:5380]
rtGetDevMsg execute failed, reason=[driver error:out of memory][FUNC:FuncErrorReason][FILE:error_message_manage.cc][LINE:53]
[ERROR] 2025-04-30-17:08:33 (PID:37513, Device:6, RankID:-1) ERR01100 OPS call acl api failed
Exception raised from operator() at build/CMakeFiles/torch_npu.dir/compiler_depend.ts:146 (most recent call first):
frame #0: c10::Error::Error(c10::SourceLocation, std::string) + 0x68 (0xffff7a0fd898 in /usr/local/lib64/python3.11/site-packages/torch/lib/libc10.so)
frame #1: c10::detail::torchCheckFail(char const*, char const*, unsigned int, std::string const&) + 0x6c (0xffff7a0b62a8 in /usr/local/lib64/python3.11/site-packages/torch/lib/libc10.so)
frame #2: <unknown function> + 0x136f73c (0xffff6d9ff73c in /usr/local/lib64/python3.11/site-packages/torch_npu/lib/libtorch_npu.so)
frame #3: <unknown function> + 0x14f11b4 (0xffff6db811b4 in /usr/local/lib64/python3.11/site-packages/torch_npu/lib/libtorch_npu.so)
frame #4: <unknown function> + 0x783104 (0xffff6ce13104 in /usr/local/lib64/python3.11/site-packages/torch_npu/lib/libtorch_npu.so)
frame #5: <unknown function> + 0x783870 (0xffff6ce13870 in /usr/local/lib64/python3.11/site-packages/torch_npu/lib/libtorch_npu.so)
frame #6: <unknown function> + 0x780acc (0xffff6ce10acc in /usr/local/lib64/python3.11/site-packages/torch_npu/lib/libtorch_npu.so)
frame #7: <unknown function> + 0xdfee4 (0xaaaade0d8ee4 in /usr/local/Ascend/mindie/1.0.0/mindie-llm/bin//mindie_llm_backend_connector)
frame #8: <unknown function> + 0x80104 (0xffff965f0104 in /usr/lib64/libc.so.6)
frame #9: <unknown function> + 0xe7e8c (0xffff96657e8c in /usr/lib64/libc.so.6)
[W compiler_depend.ts:422] Warning: EL0004: [PID: 37513] 2025-04-30-17:08:33.515.579 Failed to allocate memory.
Possible Cause: Available memory is insufficient.
Solution: Close applications not in use.
TraceBack (most recent call last):
Init device error msg handler failed, retCode=0x7020016.[FUNC:GetDevErrMsg][FILE:api_impl.cc][LINE:5380]
rtGetDevMsg execute failed, reason=[driver error:out of memory][FUNC:FuncErrorReason][FILE:error_message_manage.cc][LINE:53]
EL0004: [PID: 37513] 2025-04-30-17:08:33.471.645 Failed to allocate memory.
Possible Cause: Available memory is insufficient.
Solution: Close applications not in use.
TraceBack (most recent call last):
Check param failed, mdl can not be null.[FUNC:LaunchKernelPrepare][FILE:context.cc][LINE:769]
kernel launch prepare failed.[FUNC:LaunchKernelWithHandle][FILE:context.cc][LINE:985]
rtKernelLaunchWithHandleV2 execute failed, reason=[module new memory error][FUNC:FuncErrorReason][FILE:error_message_manage.cc][LINE:53]
rtKernelLaunchWithHandleV2 failed: 207001
#### KernelLaunch failed: /usr/local/Ascend/ascend-toolkit/8.0.0/opp/built-in/op_impl/ai_core/tbe//kernel/ascend310p/select_v2/SelectV2_bad39ec50a3854a750466c42b6ef8078_high_performance.o
Kernel Run failed. opType: 17, SelectV2
launch failed for SelectV2, errno:361001.
Init device error msg handler failed, retCode=0x7020016.[FUNC:GetDevErrMsg][FILE:api_impl.cc][LINE:5380]
rtGetDevMsg execute failed, reason=[driver error:out of memory][FUNC:FuncErrorReason][FILE:error_message_manage.cc][LINE:53]
(function ExecFuncOpApi)
2025-04-30 17:08:33,517 [ERROR] standard_model.py:194 - [Model] >>>Execute type:1, Exception:The Inner error is reported as above. The process exits for this inner error, and the current working operator name is aclnnSWhere.
Since the operator is called asynchronously, the stacktrace may be inaccurate. If you want to get the accurate stacktrace, pleace set the environment variable ASCEND_LAUNCH_BLOCKING=1.
版本信息:
MindeIE:1.0.0-300I-Duo-py311-openeuler24.03-lts
大模型:Qwen2.5-14B-Instruct
显卡:Atlas300I Duo
问题现象:
服务化拉起成功,使用curl命令进行单对话测试,响应为空
报错信息如下:
....[E compiler_depend.ts:416] call aclnnSWhere failed, detail:EL0004: [PID: 37513] 2025-04-30-17:08:33.471.645 Failed to allocate memory.
Possible Cause: Available memory is insufficient.
Solution: Close applications not in use.
TraceBack (most recent call last):
Check param failed, mdl can not be null.[FUNC:LaunchKernelPrepare][FILE:context.cc][LINE:769]
kernel launch prepare failed.[FUNC:LaunchKernelWithHandle][FILE:context.cc][LINE:985]
rtKernelLaunchWithHandleV2 execute failed, reason=[module new memory error][FUNC:FuncErrorReason][FILE:error_message_manage.cc][LINE:53]
rtKernelLaunchWithHandleV2 failed: 207001
#### KernelLaunch failed: /usr/local/Ascend/ascend-toolkit/8.0.0/opp/built-in/op_impl/ai_core/tbe//kernel/ascend310p/select_v2/SelectV2_bad39ec50a3854a750466c42b6ef8078_high_performance.o
Kernel Run failed. opType: 17, SelectV2
launch failed for SelectV2, errno:361001.
Init device error msg handler failed, retCode=0x7020016.[FUNC:GetDevErrMsg][FILE:api_impl.cc][LINE:5380]
rtGetDevMsg execute failed, reason=[driver error:out of memory][FUNC:FuncErrorReason][FILE:error_message_manage.cc][LINE:53]
[ERROR] 2025-04-30-17:08:33 (PID:37513, Device:6, RankID:-1) ERR01100 OPS call acl api failed
Exception raised from operator() at build/CMakeFiles/torch_npu.dir/compiler_depend.ts:146 (most recent call first):
frame #0: c10::Error::Error(c10::SourceLocation, std::string) + 0x68 (0xffff7a0fd898 in /usr/local/lib64/python3.11/site-packages/torch/lib/libc10.so)
frame #1: c10::detail::torchCheckFail(char const*, char const*, unsigned int, std::string const&) + 0x6c (0xffff7a0b62a8 in /usr/local/lib64/python3.11/site-packages/torch/lib/libc10.so)
frame #2: <unknown function> + 0x136f73c (0xffff6d9ff73c in /usr/local/lib64/python3.11/site-packages/torch_npu/lib/libtorch_npu.so)
frame #3: <unknown function> + 0x14f11b4 (0xffff6db811b4 in /usr/local/lib64/python3.11/site-packages/torch_npu/lib/libtorch_npu.so)
frame #4: <unknown function> + 0x783104 (0xffff6ce13104 in /usr/local/lib64/python3.11/site-packages/torch_npu/lib/libtorch_npu.so)
frame #5: <unknown function> + 0x783870 (0xffff6ce13870 in /usr/local/lib64/python3.11/site-packages/torch_npu/lib/libtorch_npu.so)
frame #6: <unknown function> + 0x780acc (0xffff6ce10acc in /usr/local/lib64/python3.11/site-packages/torch_npu/lib/libtorch_npu.so)
frame #7: <unknown function> + 0xdfee4 (0xaaaade0d8ee4 in /usr/local/Ascend/mindie/1.0.0/mindie-llm/bin//mindie_llm_backend_connector)
frame #8: <unknown function> + 0x80104 (0xffff965f0104 in /usr/lib64/libc.so.6)
frame #9: <unknown function> + 0xe7e8c (0xffff96657e8c in /usr/lib64/libc.so.6)
[W compiler_depend.ts:422] Warning: EL0004: [PID: 37513] 2025-04-30-17:08:33.515.579 Failed to allocate memory.
Possible Cause: Available memory is insufficient.
Solution: Close applications not in use.
TraceBack (most recent call last):
Init device error msg handler failed, retCode=0x7020016.[FUNC:GetDevErrMsg][FILE:api_impl.cc][LINE:5380]
rtGetDevMsg execute failed, reason=[driver error:out of memory][FUNC:FuncErrorReason][FILE:error_message_manage.cc][LINE:53]
EL0004: [PID: 37513] 2025-04-30-17:08:33.471.645 Failed to allocate memory.
Possible Cause: Available memory is insufficient.
Solution: Close applications not in use.
TraceBack (most recent call last):
Check param failed, mdl can not be null.[FUNC:LaunchKernelPrepare][FILE:context.cc][LINE:769]
kernel launch prepare failed.[FUNC:LaunchKernelWithHandle][FILE:context.cc][LINE:985]
rtKernelLaunchWithHandleV2 execute failed, reason=[module new memory error][FUNC:FuncErrorReason][FILE:error_message_manage.cc][LINE:53]
rtKernelLaunchWithHandleV2 failed: 207001
#### KernelLaunch failed: /usr/local/Ascend/ascend-toolkit/8.0.0/opp/built-in/op_impl/ai_core/tbe//kernel/ascend310p/select_v2/SelectV2_bad39ec50a3854a750466c42b6ef8078_high_performance.o
Kernel Run failed. opType: 17, SelectV2
launch failed for SelectV2, errno:361001.
Init device error msg handler failed, retCode=0x7020016.[FUNC:GetDevErrMsg][FILE:api_impl.cc][LINE:5380]
rtGetDevMsg execute failed, reason=[driver error:out of memory][FUNC:FuncErrorReason][FILE:error_message_manage.cc][LINE:53]
(function ExecFuncOpApi)
2025-04-30 17:08:33,517 [ERROR] standard_model.py:194 - [Model] >>>Execute type:1, Exception:The Inner error is reported as above. The process exits for this inner error, and the current working operator name is aclnnSWhere.
Since the operator is called asynchronously, the stacktrace may be inaccurate. If you want to get the accurate stacktrace, pleace set the environment variable ASCEND_LAUNCH_BLOCKING=1.