300IDuo 卡mindie 纯模型推理Qwen2-VL-7B-Instruct 报错
收藏回复举报
300IDuo 卡mindie 纯模型推理Qwen2-VL-7B-Instruct 报错
t('forum.solved') 已解决
发表于2025-01-22 17:31:34
0 查看

[2025-01-22 16:48:52,387] torch.distributed.run: [WARNING] [2025-01-22 16:48:52,387] torch.distributed.run: [WARNING] ***************************************** [2025-01-22 16:48:52,387] torch.distributed.run: [WARNING] Setting OMP_NUM_THREADS environment variable for each process to be 1 in default, to avoid your system being overloaded, please further tune the variable for optimal performance in your application as needed. [2025-01-22 16:48:52,387] torch.distributed.run: [WARNING] ***************************************** [2025-01-22 16:49:06,614] [24575] [281473438685760] [llm] [INFO] [cpu_binding.py-212] : rank_id: 7, device_id: 7, numa_id: 1, shard_devices: [4, 5, 6, 7], cpus: [32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63] [2025-01-22 16:49:06,617] [24575] [281473438685760] [llm] [INFO] [cpu_binding.py-238] : process 24575, new_affinity is [56, 57, 58, 59, 60, 61, 62, 63], cpu count 8 [2025-01-22 16:49:07,172] [24574] [281473355033152] [llm] [INFO] [cpu_binding.py-212] : rank_id: 6, device_id: 6, numa_id: 1, shard_devices: [4, 5, 6, 7], cpus: [32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63] [2025-01-22 16:49:07,175] [24574] [281473355033152] [llm] [INFO] [cpu_binding.py-238] : process 24574, new_affinity is [48, 49, 50, 51, 52, 53, 54, 55], cpu count 8 [2025-01-22 16:49:07,206] [24568] [281473026251328] [llm] [INFO] [cpu_binding.py-212] : rank_id: 0, device_id: 0, numa_id: 0, shard_devices: [0, 1, 2, 3], cpus: [0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31] [2025-01-22 16:49:07,208] [24568] [281473026251328] [llm] [INFO] [cpu_binding.py-238] : process 24568, new_affinity is [0, 1, 2, 3, 4, 5, 6, 7], cpu count 8 [2025-01-22 16:49:07,211] [24573] [281473652259392] [llm] [INFO] [cpu_binding.py-212] : rank_id: 5, device_id: 5, numa_id: 1, shard_devices: [4, 5, 6, 7], cpus: [32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63] [2025-01-22 16:49:07,213] [24573] [281473652259392] [llm] [INFO] [cpu_binding.py-238] : process 24573, new_affinity is [40, 41, 42, 43, 44, 45, 46, 47], cpu count 8 [2025-01-22 16:49:07,464] [24570] [281473225652800] [llm] [INFO] [cpu_binding.py-212] : rank_id: 2, device_id: 2, numa_id: 0, shard_devices: [0, 1, 2, 3], cpus: [0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31] [2025-01-22 16:49:07,465] [24569] [281472838982208] [llm] [INFO] [cpu_binding.py-212] : rank_id: 1, device_id: 1, numa_id: 0, shard_devices: [0, 1, 2, 3], cpus: [0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31] [2025-01-22 16:49:07,467] [24570] [281473225652800] [llm] [INFO] [cpu_binding.py-238] : process 24570, new_affinity is [16, 17, 18, 19, 20, 21, 22, 23], cpu count 8 [2025-01-22 16:49:07,467] [24569] [281472838982208] [llm] [INFO] [cpu_binding.py-238] : process 24569, new_affinity is [8, 9, 10, 11, 12, 13, 14, 15], cpu count 8 [2025-01-22 16:49:07,693] [24572] [281473779382848] [llm] [INFO] [cpu_binding.py-212] : rank_id: 4, device_id: 4, numa_id: 1, shard_devices: [4, 5, 6, 7], cpus: [32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63] [2025-01-22 16:49:07,693] [24571] [281472839043648] [llm] [INFO] [cpu_binding.py-212] : rank_id: 3, device_id: 3, numa_id: 0, shard_devices: [0, 1, 2, 3], cpus: [0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31] [2025-01-22 16:49:07,696] [24572] [281473779382848] [llm] [INFO] [cpu_binding.py-238] : process 24572, new_affinity is [32, 33, 34, 35, 36, 37, 38, 39], cpu count 8 [2025-01-22 16:49:07,696] [24571] [281472839043648] [llm] [INFO] [cpu_binding.py-238] : process 24571, new_affinity is [24, 25, 26, 27, 28, 29, 30, 31], cpu count 8 [2025-01-22 16:49:07,905] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : model_runner.quantize: None, model_runner.kv_quant_type: None, model_runner.fa_quant_type: None, model_runner.dtype: torch.float16 [2025-01-22 16:49:14,654] [24570] [281473225652800] [llm] [INFO] [dist.py-77] : initialize_distributed has been Set [2025-01-22 16:49:15,352] [24571] [281472839043648] [llm] [INFO] [dist.py-77] : initialize_distributed has been Set [2025-01-22 16:49:15,387] [24568] [281473026251328] [llm] [INFO] [dist.py-77] : initialize_distributed has been Set [2025-01-22 16:49:15,390] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : init tokenizer done: Qwen2TokenizerFast(name_or_path='/home/models/Qwen2-VL-7B-Instruct', vocab_size=151643, model_max_length=32768, is_fast=True, padding_side='left', truncation_side='right', special_tokens={'eos_token': '<|im_end|>', 'pad_token': '<|endoftext|>', 'additional_special_tokens': ['<|im_start|>', '<|im_end|>', '<|object_ref_start|>', '<|object_ref_end|>', '<|box_start|>', '<|box_end|>', '<|quad_start|>', '<|quad_end|>', '<|vision_start|>', '<|vision_end|>', '<|vision_pad|>', '<|image_pad|>', '<|video_pad|>']}, clean_up_tokenization_spaces=False), added_tokens_decoder={ 151643: AddedToken("<|endoftext|>", rstrip=False, lstrip=False, single_word=False, normalized=False, special=True), 151644: AddedToken("<|im_start|>", rstrip=False, lstrip=False, single_word=False, normalized=False, special=True), 151645: AddedToken("<|im_end|>", rstrip=False, lstrip=False, single_word=False, normalized=False, special=True), 151646: AddedToken("<|object_ref_start|>", rstrip=False, lstrip=False, single_word=False, normalized=False, special=True), 151647: AddedToken("<|object_ref_end|>", rstrip=False, lstrip=False, single_word=False, normalized=False, special=True), 151648: AddedToken("<|box_start|>", rstrip=False, lstrip=False, single_word=False, normalized=False, special=True), 151649: AddedToken("<|box_end|>", rstrip=False, lstrip=False, single_word=False, normalized=False, special=True), 151650: AddedToken("<|quad_start|>", rstrip=False, lstrip=False, single_word=False, normalized=False, special=True), 151651: AddedToken("<|quad_end|>", rstrip=False, lstrip=False, single_word=False, normalized=False, special=True), 151652: AddedToken("<|vision_start|>", rstrip=False, lstrip=False, single_word=False, normalized=False, special=True), 151653: AddedToken("<|vision_end|>", rstrip=False, lstrip=False, single_word=False, normalized=False, special=True), 151654: AddedToken("<|vision_pad|>", rstrip=False, lstrip=False, single_word=False, normalized=False, special=True), 151655: AddedToken("<|image_pad|>", rstrip=False, lstrip=False, single_word=False, normalized=False, special=True), 151656: AddedToken("<|video_pad|>", rstrip=False, lstrip=False, single_word=False, normalized=False, special=True), } [2025-01-22 16:49:15,811] [24569] [281472838982208] [llm] [INFO] [dist.py-77] : initialize_distributed has been Set [2025-01-22 16:49:15,805] [24570] [281473225652800] [llm] [INFO] [dist.py-98] : ProcessGroupHCCL has been Set [2025-01-22 16:49:15,818] [24570] [281473225652800] [llm] [INFO] [dist.py-112] : rank 2 init True, init_process_group has been activated [2025-01-22 16:49:15,951] [24575] [281473438685760] [llm] [INFO] [dist.py-77] : initialize_distributed has been Set [2025-01-22 16:49:16,523] [24571] [281472839043648] [llm] [INFO] [dist.py-98] : ProcessGroupHCCL has been Set [2025-01-22 16:49:16,539] [24571] [281472839043648] [llm] [INFO] [dist.py-112] : rank 3 init True, init_process_group has been activated [2025-01-22 16:49:16,560] [24568] [281473026251328] [llm] [INFO] [dist.py-98] : ProcessGroupHCCL has been Set [W compiler_depend.ts:714] Warning: The HCCL execution timeout 1800000ms is bigger than watchdog timeout 300000ms which is set by init_process_group! The plog may not be recorded. (function ProcessGroupHCCL) [2025-01-22 16:49:16,563] [24568] [281473026251328] [llm] [INFO] [dist.py-112] : rank 0 init True, init_process_group has been activated [2025-01-22 16:49:16,571] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : NPUSocInfo(soc_name='', soc_version=202, need_nz=True, matmul_nd_nz=False) [2025-01-22 16:49:16,647] [24572] [281473779382848] [llm] [INFO] [dist.py-77] : initialize_distributed has been Set [2025-01-22 16:49:16,644] [24574] [281473355033152] [llm] [INFO] [dist.py-77] : initialize_distributed has been Set [2025-01-22 16:49:16,911] [24573] [281473652259392] [llm] [INFO] [dist.py-77] : initialize_distributed has been Set [2025-01-22 16:49:16,955] [24569] [281472838982208] [llm] [INFO] [dist.py-98] : ProcessGroupHCCL has been Set [2025-01-22 16:49:16,958] [24569] [281472838982208] [llm] [INFO] [dist.py-112] : rank 1 init True, init_process_group has been activated [2025-01-22 16:49:17,227] [24575] [281473438685760] [llm] [INFO] [dist.py-98] : ProcessGroupHCCL has been Set [2025-01-22 16:49:17,231] [24575] [281473438685760] [llm] [INFO] [dist.py-112] : rank 7 init True, init_process_group has been activated [2025-01-22 16:49:17,983] [24574] [281473355033152] [llm] [INFO] [dist.py-98] : ProcessGroupHCCL has been Set [2025-01-22 16:49:17,987] [24574] [281473355033152] [llm] [INFO] [dist.py-112] : rank 6 init True, init_process_group has been activated [2025-01-22 16:49:18,041] [24572] [281473779382848] [llm] [INFO] [dist.py-98] : ProcessGroupHCCL has been Set [2025-01-22 16:49:18,044] [24572] [281473779382848] [llm] [INFO] [dist.py-112] : rank 4 init True, init_process_group has been activated [2025-01-22 16:49:18,113] [24573] [281473652259392] [llm] [INFO] [dist.py-98] : ProcessGroupHCCL has been Set [2025-01-22 16:49:18,117] [24573] [281473652259392] [llm] [INFO] [dist.py-112] : rank 5 init True, init_process_group has been activated [2025-01-22 16:49:19,992] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : NPUSocInfo(soc_name='', soc_version=202, need_nz=True, matmul_nd_nz=False) [2025-01-22 16:49:21,880] [24568] [281473026251328] [llm] [INFO] [flash_causal_qwen2.py-115] : >>>> qwen_QwenDecoderModel is called. [2025-01-22 16:49:21,935] [24570] [281473225652800] [llm] [INFO] [flash_causal_qwen2.py-115] : >>>> qwen_QwenDecoderModel is called. [2025-01-22 16:49:22,212] [24571] [281472839043648] [llm] [INFO] [flash_causal_qwen2.py-115] : >>>> qwen_QwenDecoderModel is called. /usr/local/Ascend/atb-models/atb_llm/utils/weights.py:383: UserWarning: torch.range is deprecated and will be removed in a future release because its behavior is inconsistent with Python's range builtin. Instead, use torch.arange, which produces values in [start, end). indices = torch.range(start, start + odd_rank_hidden_size - 1).to(torch.int32) /usr/local/Ascend/atb-models/atb_llm/utils/weights.py:383: UserWarning: torch.range is deprecated and will be removed in a future release because its behavior is inconsistent with Python's range builtin. Instead, use torch.arange, which produces values in [start, end). indices = torch.range(start, start + odd_rank_hidden_size - 1).to(torch.int32) /usr/local/Ascend/atb-models/atb_llm/utils/weights.py:356: UserWarning: torch.range is deprecated and will be removed in a future release because its behavior is inconsistent with Python's range builtin. Instead, use torch.arange, which produces values in [start, end). indices = torch.range(start, start + odd_rank_hidden_size - 1).to(torch.int32) [2025-01-22 16:49:22,415] [24569] [281472838982208] [llm] [INFO] [flash_causal_qwen2.py-115] : >>>> qwen_QwenDecoderModel is called. /usr/local/Ascend/atb-models/atb_llm/utils/weights.py:356: UserWarning: torch.range is deprecated and will be removed in a future release because its behavior is inconsistent with Python's range builtin. Instead, use torch.arange, which produces values in [start, end). indices = torch.range(start, start + odd_rank_hidden_size - 1).to(torch.int32) [2025-01-22 16:49:22,526] [24575] [281473438685760] [llm] [INFO] [flash_causal_qwen2.py-115] : >>>> qwen_QwenDecoderModel is called. /usr/local/Ascend/atb-models/atb_llm/utils/weights.py:390: UserWarning: torch.range is deprecated and will be removed in a future release because its behavior is inconsistent with Python's range builtin. Instead, use torch.arange, which produces values in [start, end). indices = torch.range(start, start + even_rank_hidden_size - 1).to(torch.int32) /usr/local/Ascend/atb-models/atb_llm/utils/weights.py:363: UserWarning: torch.range is deprecated and will be removed in a future release because its behavior is inconsistent with Python's range builtin. Instead, use torch.arange, which produces values in [start, end). indices = torch.range(start, start + even_rank_hidden_size - 1).to(torch.int32) /usr/local/Ascend/atb-models/atb_llm/utils/weights.py:390: UserWarning: torch.range is deprecated and will be removed in a future release because its behavior is inconsistent with Python's range builtin. Instead, use torch.arange, which produces values in [start, end). indices = torch.range(start, start + even_rank_hidden_size - 1).to(torch.int32) /usr/local/Ascend/atb-models/atb_llm/utils/weights.py:363: UserWarning: torch.range is deprecated and will be removed in a future release because its behavior is inconsistent with Python's range builtin. Instead, use torch.arange, which produces values in [start, end). indices = torch.range(start, start + even_rank_hidden_size - 1).to(torch.int32) [2025-01-22 16:49:23,041] [24574] [281473355033152] [llm] [INFO] [flash_causal_qwen2.py-115] : >>>> qwen_QwenDecoderModel is called. /usr/local/Ascend/atb-models/atb_llm/utils/weights.py:390: UserWarning: torch.range is deprecated and will be removed in a future release because its behavior is inconsistent with Python's range builtin. Instead, use torch.arange, which produces values in [start, end). indices = torch.range(start, start + even_rank_hidden_size - 1).to(torch.int32) [2025-01-22 16:49:23,260] [24573] [281473652259392] [llm] [INFO] [flash_causal_qwen2.py-115] : >>>> qwen_QwenDecoderModel is called. /usr/local/Ascend/atb-models/atb_llm/utils/weights.py:363: UserWarning: torch.range is deprecated and will be removed in a future release because its behavior is inconsistent with Python's range builtin. Instead, use torch.arange, which produces values in [start, end). indices = torch.range(start, start + even_rank_hidden_size - 1).to(torch.int32) [2025-01-22 16:49:23,296] [24572] [281473779382848] [llm] [INFO] [flash_causal_qwen2.py-115] : >>>> qwen_QwenDecoderModel is called. /usr/local/Ascend/atb-models/atb_llm/utils/weights.py:383: UserWarning: torch.range is deprecated and will be removed in a future release because its behavior is inconsistent with Python's range builtin. Instead, use torch.arange, which produces values in [start, end). indices = torch.range(start, start + odd_rank_hidden_size - 1).to(torch.int32) /usr/local/Ascend/atb-models/atb_llm/utils/weights.py:356: UserWarning: torch.range is deprecated and will be removed in a future release because its behavior is inconsistent with Python's range builtin. Instead, use torch.arange, which produces values in [start, end). indices = torch.range(start, start + odd_rank_hidden_size - 1).to(torch.int32) /usr/local/Ascend/atb-models/atb_llm/utils/weights.py:390: UserWarning: torch.range is deprecated and will be removed in a future release because its behavior is inconsistent with Python's range builtin. Instead, use torch.arange, which produces values in [start, end). indices = torch.range(start, start + even_rank_hidden_size - 1).to(torch.int32) /usr/local/Ascend/atb-models/atb_llm/utils/weights.py:383: UserWarning: torch.range is deprecated and will be removed in a future release because its behavior is inconsistent with Python's range builtin. Instead, use torch.arange, which produces values in [start, end). indices = torch.range(start, start + odd_rank_hidden_size - 1).to(torch.int32) /usr/local/Ascend/atb-models/atb_llm/utils/weights.py:356: UserWarning: torch.range is deprecated and will be removed in a future release because its behavior is inconsistent with Python's range builtin. Instead, use torch.arange, which produces values in [start, end). indices = torch.range(start, start + odd_rank_hidden_size - 1).to(torch.int32) /usr/local/Ascend/atb-models/atb_llm/utils/weights.py:363: UserWarning: torch.range is deprecated and will be removed in a future release because its behavior is inconsistent with Python's range builtin. Instead, use torch.arange, which produces values in [start, end). indices = torch.range(start, start + even_rank_hidden_size - 1).to(torch.int32) [2025-01-22 16:49:32,684] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : model: FlashQwen2vlForCausalLM( (rotary_embedding): PositionRotaryEmbedding() (attn_mask): AttentionMask() (vision_tower): Qwen2VisionTransformerPretrainedModelATB( (patch_embed): PatchEmbed( (proj): Conv3d(3, 1280, kernel_size=(2, 14, 14), stride=(2, 14, 14), bias=False) ) (rotary_pos_emb): VisionRotaryEmbedding() (encoder): Qwen2VLVisionEncoder( (layers): ModuleList( (0-31): 32 x Qwen2VLVisionBlock( (norm1): LayerNormATB() (attn): VisionAttention( (qkv): TensorParallelColumnLinear( (linear): FastLinear() ) (proj): TensorParallelRowLinear( (linear): FastLinear() ) ) (norm2): LayerNormATB() (mlp): VisionMlp( (fc1): TensorParallelColumnLinear( (linear): FastLinear() ) (fc2): TensorParallelRowLinear( (linear): FastLinear() ) ) ) ) ) (merger): PatchMerger( (ln_q): LayerNorm((1280,), eps=1e-06, elementwise_affine=True) (mlp): Sequential( (0): Linear(in_features=5120, out_features=5120, bias=True) (1): GELU(approximate='none') (2): Linear(in_features=5120, out_features=3584, bias=True) ) ) ) (language_model): FlashQwen2UsingMROPEForCausalLM( (rotary_embedding): PositionRotaryEmbedding() (attn_mask): AttentionMask() (transformer): FlashQwenModel( (wte): TensorEmbeddingWithoutChecking() (h): ModuleList( (0-27): 28 x FlashQwenLayer( (attn): FlashQwenAttention( (rotary_emb): PositionRotaryEmbedding() (c_attn): TensorParallelColumnLinear( (linear): FastLinear() ) (c_proj): TensorParallelColumnLinear( (linear): FastLinear() ) ) (mlp): QwenMLP( (act): SiLU() (w2_w1): TensorParallelColumnLinear( (linear): FastLinear() ) (c_proj): TensorParallelRowLinear( (linear): FastLinear() ) ) (ln_1): QwenRMSNorm() (ln_2): QwenRMSNorm() ) ) (ln_f): QwenRMSNorm() ) (lm_head): TensorParallelHead( (linear): FastLinear() ) ) (normalizer): Conv1d(3, 3, kernel_size=(1,), stride=(1,), groups=3) ) [2025-01-22 16:49:34,913] [24570] [281473225652800] [llm] [INFO] [cache.py-79] : kv cache will allocate 0.056396484375GB memory [2025-01-22 16:49:34,979] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : hbm_capacity(GB): 21.0859375, init_memory(GB): 5.271484375 [2025-01-22 16:49:34,979] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : pa_runner: PARunner(model_path=/home/models/Qwen2-VL-7B-Instruct/, input_text=Explain the details in the image., max_position_embeddings=None, max_input_length=4096, max_output_length=80, max_prefill_tokens=4176, enable_atb_torch=False, is_flash_model=True, max_batch_size=1, dtype=torch.float16, block_size=128, model_config=ModelConfig(num_heads=4, num_kv_heads=1, num_kv_heads_origin=1, head_size=128, k_head_size=128, v_head_size=128, num_layers=28, device=npu:0, dtype=torch.float16, soc_info=NPUSocInfo(soc_name='', soc_version=202, need_nz=True, matmul_nd_nz=False), kv_quant_type=None, fa_quant_type=None, , max_memory=22640852992, [2025-01-22 16:49:34,988] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : ---------------begin warm_up--------------- [2025-01-22 16:49:34,988] [24568] [281473026251328] [llm] [INFO] [cache.py-79] : kv cache will allocate 0.056396484375GB memory [2025-01-22 16:49:35,019] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : ------total req num: 1, infer start-------- ..[2025-01-22 16:49:36,096] [24571] [281472839043648] [llm] [INFO] [cache.py-79] : kv cache will allocate 0.056396484375GB memory .[2025-01-22 16:49:36,420] [24569] [281472838982208] [llm] [INFO] [cache.py-79] : kv cache will allocate 0.056396484375GB memory .[2025-01-22 16:49:38,136] [24574] [281473355033152] [llm] [INFO] [cache.py-79] : kv cache will allocate 0.056396484375GB memory [2025-01-22 16:49:38,139] [24572] [281473779382848] [llm] [INFO] [cache.py-79] : kv cache will allocate 0.056396484375GB memory .[2025-01-22 16:49:38,401] [24575] [281473438685760] [llm] [INFO] [cache.py-79] : kv cache will allocate 0.056396484375GB memory .[2025-01-22 16:49:38,449] [24573] [281473652259392] [llm] [INFO] [cache.py-79] : kv cache will allocate 0.056396484375GB memory ..[2025-01-22 16:49:40,034] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,034] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,035] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,036] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,036] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,037] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,038] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,038] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,039] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,039] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,040] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,040] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,041] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,041] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,042] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,042] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,043] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,043] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,044] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,044] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,045] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,045] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,046] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,046] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,047] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,047] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,048] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,048] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,048] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,049] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,049] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,050] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,050] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,051] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,051] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,052] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,052] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,053] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,053] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,054] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,054] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,055] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,055] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,056] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,056] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,057] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,057] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,058] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,058] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,058] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,059] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,059] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,060] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,060] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,061] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,061] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,061] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,062] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,062] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,063] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,063] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,063] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,064] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,064] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,065] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,065] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,066] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,066] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,067] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,067] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,068] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,068] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,069] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,069] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,070] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,070] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,071] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,071] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,071] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,072] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,072] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,073] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,073] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,073] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,074] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,075] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,075] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,075] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,076] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,076] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,077] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,077] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,078] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,078] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,079] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,079] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,080] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,080] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,080] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,081] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,081] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,082] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,082] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,083] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,083] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,083] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,084] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,084] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,085] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,085] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,086] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,086] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:40,087] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : trans to 29 [2025-01-22 16:49:41,790] [24569] [281472838982208] [llm] [INFO] [flash_causal_qwen2.py-394] : <<<<<<<after transdata k_caches[0].shape=torch.Size([33, 8, 128, 16]) [2025-01-22 16:49:41,806] [24571] [281472839043648] [llm] [INFO] [flash_causal_qwen2.py-394] : <<<<<<<after transdata k_caches[0].shape=torch.Size([33, 8, 128, 16]) [2025-01-22 16:49:41,823] [24570] [281473225652800] [llm] [INFO] [flash_causal_qwen2.py-394] : <<<<<<<after transdata k_caches[0].shape=torch.Size([33, 8, 128, 16]) [2025-01-22 16:49:41,839] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : <<<<<<< ori k_caches[0].shape=torch.Size([33, 8, 128, 16]) [2025-01-22 16:49:41,841] [24568] [281473026251328] [llm] [INFO] [flash_causal_qwen2.py-394] : <<<<<<<after transdata k_caches[0].shape=torch.Size([33, 8, 128, 16]) [2025-01-22 16:49:41,841] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : >>>>>>id of kcache is 281469620858864 id of vcache is 281469620858288 [2025-01-22 16:49:41,854] [24575] [281473438685760] [llm] [INFO] [flash_causal_qwen2.py-394] : <<<<<<<after transdata k_caches[0].shape=torch.Size([33, 8, 128, 16]) [2025-01-22 16:49:41,888] [24574] [281473355033152] [llm] [INFO] [flash_causal_qwen2.py-394] : <<<<<<<after transdata k_caches[0].shape=torch.Size([33, 8, 128, 16]) [2025-01-22 16:49:41,918] [24573] [281473652259392] [llm] [INFO] [flash_causal_qwen2.py-394] : <<<<<<<after transdata k_caches[0].shape=torch.Size([33, 8, 128, 16]) [2025-01-22 16:49:41,920] [24572] [281473779382848] [llm] [INFO] [flash_causal_qwen2.py-394] : <<<<<<<after transdata k_caches[0].shape=torch.Size([33, 8, 128, 16]) [2025-01-22 16:49:43,015] [24569] [281472838982208] [llm] [INFO] [generate.py-189] : Prefill time: 6464.224815368652ms, Decode token time: 97.58782386779785ms, E2E time: 6561.81263923645ms [2025-01-22 16:49:43,015] [24575] [281473438685760] [llm] [INFO] [generate.py-189] : Prefill time: 4480.690479278564ms, Decode token time: 97.28360176086426ms, E2E time: 4577.974081039429ms [2025-01-22 16:49:43,015] [24572] [281473779382848] [llm] [INFO] [generate.py-189] : Prefill time: 4746.464252471924ms, Decode token time: 97.28574752807617ms, E2E time: 4843.75ms [2025-01-22 16:49:43,015] [24571] [281472839043648] [llm] [INFO] [generate.py-189] : Prefill time: 6787.541151046753ms, Decode token time: 97.52225875854492ms, E2E time: 6885.063409805298ms [2025-01-22 16:49:43,015] [24573] [281473652259392] [llm] [INFO] [generate.py-189] : Prefill time: 4431.706190109253ms, Decode token time: 97.26762771606445ms, E2E time: 4528.973817825317ms [2025-01-22 16:49:43,015] [24574] [281473355033152] [llm] [INFO] [generate.py-189] : Prefill time: 4746.3860511779785ms, Decode token time: 97.28336334228516ms, E2E time: 4843.669414520264ms [2025-01-22 16:49:43,015] [24570] [281473225652800] [llm] [INFO] [generate.py-189] : Prefill time: 7971.807956695557ms, Decode token time: 97.58663177490234ms, E2E time: 8069.394588470459ms [2025-01-22 16:49:43,015] [24568] [281473026251328] [llm] [INFO] [generate.py-180] : max_generate_batch_size: 1 [2025-01-22 16:49:43,016] [24568] [281473026251328] [llm] [INFO] [generate.py-189] : Prefill time: 7894.780158996582ms, Decode token time: 98.7396240234375ms, E2E time: 7993.5197830200195ms [2025-01-22 16:49:43,022] [24568] [281473026251328] [llm] [INFO] [generate.py-221] : -------------------performance dumped------------------------

[2025-01-22 16:49:43,033] [24568] [281473026251328] [llm] [INFO] [generate.py-224] : | batch_size | input_seq_len | output_seq_len | e2e_time(ms) | prefill_time(ms) | decoder_token_time(ms) | prefill_count | max_generate_batch_size | |-------------:|----------------:|-----------------:|---------------:|-------------------:|-------------------------:|----------------:|--------------------------:| | 1 | 4096 | 2 | 7993.52 | 7894.78 | 98.74 | 1 | 1 | [2025-01-22 16:49:45,195] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : warmup_memory(GB): 6.96 [2025-01-22 16:49:45,195] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : ---------------end warm_up--------------- [2025-01-22 16:49:45,196] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : ---------------begin inference--------------- Traceback (most recent call last): File "", line 198, in _run_module_as_main File "", line 88, in _run_code File "/usr/local/Ascend/atb-models/examples/models/qwen2_vl/run_pa.py", line 530, in main() File "/usr/local/Ascend/atb-models/examples/models/qwen2_vl/run_pa.py", line 498, in main generate_texts, token_nums, latency = pa_runner.infer( ^^^^^^^^^^^^^^^^ File "/usr/local/Ascend/atb-models/examples/models/qwen2_vl/run_pa.py", line 299, in infer generate_req( File "/usr/local/Ascend/atb-models/examples/server/generate.py", line 108, in generate_req raise Exception(f"req: {req_idx} out of memory, need block:" + Exception: req: 0 out of memory, need block:84 is more than free block 33 Traceback (most recent call last): File "", line 198, in _run_module_as_main File "", line 88, in _run_code File "/usr/local/Ascend/atb-models/examples/models/qwen2_vl/run_pa.py", line 530, in main() File "/usr/local/Ascend/atb-models/examples/models/qwen2_vl/run_pa.py", line 498, in main generate_texts, token_nums, latency = pa_runner.infer( ^^^^^^^^^^^^^^^^ File "/usr/local/Ascend/atb-models/examples/models/qwen2_vl/run_pa.py", line 299, in infer generate_req( File "/usr/local/Ascend/atb-models/examples/server/generate.py", line 108, in generate_req raise Exception(f"req: {req_idx} out of memory, need block:" + Exception: req: 0 out of memory, need block:84 is more than free block 33 Traceback (most recent call last): File "", line 198, in _run_module_as_main File "", line 88, in _run_code File "/usr/local/Ascend/atb-models/examples/models/qwen2_vl/run_pa.py", line 530, in main() File "/usr/local/Ascend/atb-models/examples/models/qwen2_vl/run_pa.py", line 498, in main generate_texts, token_nums, latency = pa_runner.infer( ^^^^^^^^^^^^^^^^ File "/usr/local/Ascend/atb-models/examples/models/qwen2_vl/run_pa.py", line 299, in infer generate_req( File "/usr/local/Ascend/atb-models/examples/server/generate.py", line 108, in generate_req raise Exception(f"req: {req_idx} out of memory, need block:" + Exception: req: 0 out of memory, need block:84 is more than free block 33 Traceback (most recent call last): File "", line 198, in _run_module_as_main File "", line 88, in _run_code File "/usr/local/Ascend/atb-models/examples/models/qwen2_vl/run_pa.py", line 530, in main() File "/usr/local/Ascend/atb-models/examples/models/qwen2_vl/run_pa.py", line 498, in main generate_texts, token_nums, latency = pa_runner.infer( ^^^^^^^^^^^^^^^^ File "/usr/local/Ascend/atb-models/examples/models/qwen2_vl/run_pa.py", line 299, in infer generate_req( File "/usr/local/Ascend/atb-models/examples/server/generate.py", line 108, in generate_req raise Exception(f"req: {req_idx} out of memory, need block:" + Exception: req: 0 out of memory, need block:84 is more than free block 33 Traceback (most recent call last): File "", line 198, in _run_module_as_main File "", line 88, in _run_code File "/usr/local/Ascend/atb-models/examples/models/qwen2_vl/run_pa.py", line 530, in main() File "/usr/local/Ascend/atb-models/examples/models/qwen2_vl/run_pa.py", line 498, in main generate_texts, token_nums, latency = pa_runner.infer( ^^^^^^^^^^^^^^^^ File "/usr/local/Ascend/atb-models/examples/models/qwen2_vl/run_pa.py", line 299, in infer generate_req( File "/usr/local/Ascend/atb-models/examples/server/generate.py", line 108, in generate_req raise Exception(f"req: {req_idx} out of memory, need block:" + Exception: req: 0 out of memory, need block:84 is more than free block 33 Traceback (most recent call last): File "", line 198, in _run_module_as_main [2025-01-22 16:49:45,748] [24568] [281473026251328] [llm] [INFO] [logging.py-331] : ------total req num: 1, infer start-------- File "", line 88, in _run_code File "/usr/local/Ascend/atb-models/examples/models/qwen2_vl/run_pa.py", line 530, in main() File "/usr/local/Ascend/atb-models/examples/models/qwen2_vl/run_pa.py", line 498, in main generate_texts, token_nums, latency = pa_runner.infer( ^^^^^^^^^^^^^^^^ File "/usr/local/Ascend/atb-models/examples/models/qwen2_vl/run_pa.py", line 299, in infer generate_req( File "/usr/local/Ascend/atb-models/examples/server/generate.py", line 108, in generate_req raise Exception(f"req: {req_idx} out of memory, need block:" + Exception: req: 0 out of memory, need block:84 is more than free block 33 Traceback (most recent call last): File "", line 198, in _run_module_as_main File "", line 88, in _run_code File "/usr/local/Ascend/atb-models/examples/models/qwen2_vl/run_pa.py", line 530, in main() File "/usr/local/Ascend/atb-models/examples/models/qwen2_vl/run_pa.py", line 498, in main generate_texts, token_nums, latency = pa_runner.infer( ^^^^^^^^^^^^^^^^ File "/usr/local/Ascend/atb-models/examples/models/qwen2_vl/run_pa.py", line 299, in infer generate_req( File "/usr/local/Ascend/atb-models/examples/server/generate.py", line 108, in generate_req raise Exception(f"req: {req_idx} out of memory, need block:" + Exception: req: 0 out of memory, need block:84 is more than free block 33 Traceback (most recent call last): File "", line 198, in _run_module_as_main File "", line 88, in _run_code File "/usr/local/Ascend/atb-models/examples/models/qwen2_vl/run_pa.py", line 530, in main() File "/usr/local/Ascend/atb-models/examples/models/qwen2_vl/run_pa.py", line 498, in main generate_texts, token_nums, latency = pa_runner.infer( ^^^^^^^^^^^^^^^^ File "/usr/local/Ascend/atb-models/examples/models/qwen2_vl/run_pa.py", line 299, in infer generate_req( File "/usr/local/Ascend/atb-models/examples/server/generate.py", line 108, in generate_req raise Exception(f"req: {req_idx} out of memory, need block:" + Exception: req: 0 out of memory, need block:84 is more than free block 33 [ERROR] 2025-01-22-16:49:57 (PID:24570, Device:2, RankID:-1) ERR99999 UNKNOWN application exception [ERROR] 2025-01-22-16:49:57 (PID:24572, Device:4, RankID:-1) ERR99999 UNKNOWN application exception [ERROR] 2025-01-22-16:49:57 (PID:24575, Device:7, RankID:-1) ERR99999 UNKNOWN application exception [ERROR] 2025-01-22-16:49:57 (PID:24573, Device:5, RankID:-1) ERR99999 UNKNOWN application exception [ERROR] 2025-01-22-16:49:57 (PID:24574, Device:6, RankID:-1) ERR99999 UNKNOWN application exception [ERROR] 2025-01-22-16:49:57 (PID:24568, Device:0, RankID:-1) ERR99999 UNKNOWN application exception [ERROR] 2025-01-22-16:49:57 (PID:24571, Device:3, RankID:-1) ERR99999 UNKNOWN application exception [ERROR] 2025-01-22-16:49:58 (PID:24569, Device:1, RankID:-1) ERR99999 UNKNOWN application exception /usr/lib64/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 2 leaked shared_memory objects to clean up at shutdown warnings.warn('resource_tracker: There appear to be %d ' /usr/lib64/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 2 leaked shared_memory objects to clean up at shutdown warnings.warn('resource_tracker: There appear to be %d ' /usr/lib64/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 2 leaked shared_memory objects to clean up at shutdown warnings.warn('resource_tracker: There appear to be %d ' /usr/lib64/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 2 leaked shared_memory objects to clean up at shutdown warnings.warn('resource_tracker: There appear to be %d ' /usr/lib64/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 2 leaked shared_memory objects to clean up at shutdown warnings.warn('resource_tracker: There appear to be %d ' /usr/lib64/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 2 leaked shared_memory objects to clean up at shutdown warnings.warn('resource_tracker: There appear to be %d ' /usr/lib64/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 2 leaked shared_memory objects to clean up at shutdown warnings.warn('resource_tracker: There appear to be %d ' /usr/lib64/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 2 leaked shared_memory objects to clean up at shutdown warnings.warn('resource_tracker: There appear to be %d ' [2025-01-22 16:50:07,401] torch.distributed.elastic.multiprocessing.api: [ERROR] failed (exitcode: 1) local_rank: 0 (pid: 24568) of binary: /usr/bin/python3 Traceback (most recent call last): File "/usr/local/bin/torchrun", line 8, in sys.exit(main()) ^^^^^^ File "/usr/local/lib64/python3.11/site-packages/torch/distributed/elastic/multiprocessing/errors/init.py", line 346, in wrapper return f(*args, **kwargs) ^^^^^^^^^^^^^^^^^^ File "/usr/local/lib64/python3.11/site-packages/torch/distributed/run.py", line 806, in main run(args) File "/usr/local/lib64/python3.11/site-packages/torch/distributed/run.py", line 797, in run elastic_launch( File "/usr/local/lib64/python3.11/site-packages/torch/distributed/launcher/api.py", line 134, in call return launch_agent(self._config, self._entrypoint, list(args)) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/usr/local/lib64/python3.11/site-packages/torch/distributed/launcher/api.py", line 264, in launch_agent raise ChildFailedError( torch.distributed.elastic.multiprocessing.errors.ChildFailedError: ============================================================ examples.models.qwen2_vl.run_pa FAILED


Failures: [1]: time : 2025-01-22_16:50:07 host : localhost.localdomain rank : 1 (local_rank: 1) exitcode : 1 (pid: 24569) error_file: <N/A> traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html [2]: time : 2025-01-22_16:50:07 host : localhost.localdomain rank : 2 (local_rank: 2) exitcode : 1 (pid: 24570) error_file: <N/A> traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html [3]: time : 2025-01-22_16:50:07 host : localhost.localdomain rank : 3 (local_rank: 3) exitcode : 1 (pid: 24571) error_file: <N/A> traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html [4]: time : 2025-01-22_16:50:07 host : localhost.localdomain rank : 4 (local_rank: 4) exitcode : 1 (pid: 24572) error_file: <N/A> traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html [5]: time : 2025-01-22_16:50:07 host : localhost.localdomain rank : 5 (local_rank: 5) exitcode : 1 (pid: 24573) error_file: <N/A> traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html [6]: time : 2025-01-22_16:50:07 host : localhost.localdomain rank : 6 (local_rank: 6) exitcode : 1 (pid: 24574) error_file: <N/A> traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html [7]: time : 2025-01-22_16:50:07 host : localhost.localdomain rank : 7 (local_rank: 7) exitcode : 1 (pid: 24575) error_file: <N/A> traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html


Root Cause (first observed failure): [0]: time : 2025-01-22_16:50:07 host : localhost.localdomain rank : 0 (local_rank: 0) exitcode : 1 (pid: 24568) error_file: <N/A> traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html ============================================================

我要发帖子