使用1张Atlast300I Duo 96G卡,docker镜像方式运行DeepSeek-R1-Distill-Qwen-1.5B模型,稀疏量化报错
收藏回复举报
使用1张Atlast300I Duo 96G卡,docker镜像方式运行DeepSeek-R1-Distill-Qwen-1.5B模型,稀疏量化报错
t('forum.solved') 已解决
新人帖
发表于2025-07-09 16:51:39
0 查看
物理机驱动和固件版本:

Ascend-hdk-310p-npu-driver_24.1.0.1_linux-aarch64

Ascend-hdk-310p-npu-firmware_7.5.0.5.220

镜像使用的是

swr.cn-south-1.myhuaweicloud.com/ascendhub/mindie   2.0.RC2-300I-Duo-py311-openeuler24.03-lts 

步骤按照以下地址执行

https://gitee.com/ascend/ModelZoo-PyTorch/tree/master/MindIE/LLM/DeepSeek/DeepSeek-R1-Distill-Qwen-1.5B

执行到稀疏量化时,报错信息前段如下,感觉像OOM,应该怎么处理?:

[root@HwHiAiUser-pc Qwen]# python3 quant_qwen.py --model_path /data/DeepSeek-R1-Distill-Qwen-1.5B --save_directory /data/W8A8S --calib_file ../common/boolq.jsonl --w_bit 4 --a_bit 8 --fraction 0.011 --co_sparse True --device_type npu --use_sigma True --is_lowbit True
/usr/local/lib64/python3.11/site-packages/torch_npu/contrib/transfer_to_npu.py:295: ImportWarning: 
    *************************************************************************************************************
    The torch.Tensor.cuda and torch.nn.Module.cuda are replaced with torch.Tensor.npu and torch.nn.Module.npu now..
    The torch.cuda.DoubleTensor is replaced with torch.npu.FloatTensor cause the double type is not supported now..
    The backend in torch.distributed.init_process_group set to hccl now..
    The torch.cuda.* and torch.cuda.amp.* are replaced with torch.npu.* and torch.npu.amp.* now..
    The device parameters have been replaced with npu in the function below:
    torch.logspace, torch.randint, torch.hann_window, torch.rand, torch.full_like, torch.ones_like, torch.rand_like, torch.randperm, torch.arange, torch.frombuffer, torch.normal, torch._empty_per_channel_affine_quantized, torch.empty_strided, torch.empty_like, torch.scalar_tensor, torch.tril_indices, torch.bartlett_window, torch.ones, torch.sparse_coo_tensor, torch.randn, torch.kaiser_window, torch.tensor, torch.triu_indices, torch.as_tensor, torch.zeros, torch.randint_like, torch.full, torch.eye, torch._sparse_csr_tensor_unsafe, torch.empty, torch._sparse_coo_tensor_unsafe, torch.blackman_window, torch.zeros_like, torch.range, torch.sparse_csr_tensor, torch.randn_like, torch.from_file, torch._cudnn_init_dropout_state, torch._empty_affine_quantized, torch.linspace, torch.hamming_window, torch.empty_quantized, torch._pin_memory, torch.autocast, torch.load, torch.Generator, torch.set_default_device, torch.Tensor.new_empty, torch.Tensor.new_empty_strided, torch.Tensor.new_full, torch.Tensor.new_ones, torch.Tensor.new_tensor, torch.Tensor.new_zeros, torch.Tensor.to, torch.Tensor.pin_memory, torch.nn.Module.to, torch.nn.Module.to_empty
    *************************************************************************************************************
    
  warnings.warn(msg, ImportWarning)
/usr/local/lib64/python3.11/site-packages/torch_npu/contrib/transfer_to_npu.py:250: RuntimeWarning: torch.jit.script and torch.jit.script_method will be disabled by transfer_to_npu, which currently does not support them, if you need to enable them, please do not use transfer_to_npu.
  warnings.warn(msg, RuntimeWarning)
2025-07-09 16:01:21,358 - msmodelslim-logger - WARNING - The current CANN version does not support recall_window method.
2025-07-09 16:01:21,377 - msmodelslim-logger - WARNING - write directory not exists, creating directory '/data/W8A8S'
2025-07-09 16:01:47,660 - msmodelslim-logger - INFO - Automatically disabling the last linear layer: lm_head based on the `disable_last_linear` parameter setting.
feature process:   0%|                                                                                                                       | 0/5 [00:00<?, ?it/s]We detected that you are passing `past_key_values` as a tuple and this is deprecated and will be removed in v4.43. Please use an appropriate `Cache` class (https://huggingface.co/docs/transformers/v4.41.3/en/internal/generation_utils#transformers.Cache)
[E compiler_depend.ts:421] call failed, detail:EZ9903: [PID: 1161] 2025-07-09-16:01:48.283.588 rtKernelLaunchWithHandleV2 failed: 207001
        Solution: In this scenario, collect the plog when the fault occurs and locate the fault based on the plog.
        TraceBack (most recent call last):
        The error from device(0), serial number is 1, there is an aicore error, core id is 0, error code = 0x800000, dump info: pc start: 0x80012400004324c, current: 0x124000043308, vec error info: 0x3fdf389, mte error info: 0x30000bf, ifu error info: 0x273ff7f0edf80, ccu error info: 0xddb0cfe8007f8ffd, cube error info: 0xfc, biu error info: 0, aic error mask: 0x65000200d00028c, para base: 0x12c0003e5000, errorStr: The DDR address of the MTE instruction is out of range.[FUNC:PrintCoreErrorInfo][FILE:device_error_proc.cc][LINE:639]
        The extend info from device(0), serial number is 1, there is aicore error, core id is 0, aicore int: 0x10, aicore error2: 0, axi clamp ctrl: 0, axi clamp state: 0x1717, biu status0: 0x101d14000000000, biu status1: 0x80000201020000, clk gate mask: 0, dbg addr: 0, ecc en: 0, mte ccu ecc 1bit error: 0xd80000000000000, vector cube ecc 1bit error: 0, run stall: 0x1, dbg data0: 0, dbg data1: 0, dbg data2: 0, dbg data3: 0, dfx data: 0x55[FUNC:PrintCoreErrorInfo][FILE:device_error_proc.cc][LINE:670]
        The error from device(0), serial number is 1, there is an aicore error, core id is 1, error code = 0x800000, dump info: pc start: 0x80012400004324c, current: 0x124000043308, vec error info: 0x3cf3488, mte error info: 0x30000bf, ifu error info: 0x127faffd9ff80, ccu error info: 0x3f9073fb00179f6d, cube error info: 0xaf, biu error info: 0, aic error mask: 0x65000200d00028c, para base: 0x12c0003e5000, errorStr: The DDR address of the MTE instruction is out of range.[FUNC:PrintCoreErrorInfo][FILE:device_error_proc.cc][LINE:639]
        The extend info from device(0), serial number is 1, there is aicore error, core id is 1, aicore int: 0x10, aicore error2: 0, axi clamp ctrl: 0, axi clamp state: 0x1717, biu status0: 0x101d14000000000, biu status1: 0x80000201020000, clk gate mask: 0, dbg addr: 0, ecc en: 0, mte ccu ecc 1bit error: 0x2780000000000000, vector cube ecc 1bit error: 0, run stall: 0x1, dbg data0: 0, dbg data1: 0, dbg data2: 0, dbg data3: 0, dfx data: 0xef[FUNC:PrintCoreErrorInfo][FILE:device_error_proc.cc][LINE:670]
        The error from device(0), serial number is 1, there is an aicore error, core id is 2, error code = 0x800000, dump info: pc start: 0x80012400004324c, current: 0x124000043308, vec error info: 0x19b1f9c3, mte error info: 0x30000bf, ifu error info: 0x2cbd4debcdc00, ccu error info: 0xf49fb4c200736f5d, cube error info: 0xc3, biu error info: 0, aic error mask: 0x65000200d00028c, para base: 0x12c0003e5000, errorStr: The DDR address of the MTE instruction is out of range.[FUNC:PrintCoreErrorInfo][FILE:device_error_proc.cc][LINE:639]
        The extend info from device(0), serial number is 1, there is aicore error, core id is 2, aicore int: 0x10, aicore error2: 0, axi clamp ctrl: 0, axi clamp state: 0x1717, biu status0: 0x101d14000000000, biu status1: 0x80000201020000, clk gate mask: 0, dbg addr: 0, ecc en: 0, mte ccu ecc 1bit error: 0x2b80000000000000, vector cube ecc 1bit error: 0, run stall: 0x1, dbg data0: 0, dbg data1: 0, dbg data2: 0, dbg data3: 0, dfx data: 0x78[FUNC:PrintCoreErrorInfo][FILE:device_error_proc.cc][LINE:670]

我要发帖子