(APIServer pid=351099) INFO 06-22 22:40:47 [utils.py:299] (APIServer pid=351099) INFO 06-22 22:40:47 [utils.py:299] █ █ █▄ ▄█ (APIServer pid=351099) INFO 06-22 22:40:47 [utils.py:299] ▄▄ ▄█ █ █ █ ▀▄▀ █ version 0.20.2rc1.dev13+g9557d9108.d20260620 (APIServer pid=351099) INFO 06-22 22:40:47 [utils.py:299] █▄█▀ █ █ █ █ model /mnt/fast-ai/llm-cache/hf/models--nameistoken--Qwen3.6-35B-A3B-Quark-W8A8-INT8/snapshots/cced56592e8c8935f8220836b4baa04dfd389118 (APIServer pid=351099) INFO 06-22 22:40:47 [utils.py:299] ▀▀ ▀▀▀▀▀ ▀▀▀▀▀ ▀ ▀ (APIServer pid=351099) INFO 06-22 22:40:47 [utils.py:299] (APIServer pid=351099) INFO 06-22 22:40:47 [utils.py:233] non-default args: {'model_tag': '/mnt/fast-ai/llm-cache/hf/models--nameistoken--Qwen3.6-35B-A3B-Quark-W8A8-INT8/snapshots/cced56592e8c8935f8220836b4baa04dfd389118', 'host': '127.0.0.1', 'port': 18080, 'model': '/mnt/fast-ai/llm-cache/hf/models--nameistoken--Qwen3.6-35B-A3B-Quark-W8A8-INT8/snapshots/cced56592e8c8935f8220836b4baa04dfd389118', 'trust_remote_code': True, 'max_model_len': 32768, 'quantization': 'quark', 'served_model_name': ['qwen36-35b-a3b-fp8'], 'generation_config': 'vllm', 'distributed_executor_backend': 'mp', 'tensor_parallel_size': 4, 'gpu_memory_utilization': 0.95, 'enable_prefix_caching': False, 'mamba_cache_mode': 'align', 'language_model_only': True, 'max_num_batched_tokens': 8192, 'max_num_seqs': 48, 'async_scheduling': False, 'speculative_config': {'method': 'mtp', 'num_speculative_tokens': 3, 'max_model_len': 32768}, 'compilation_config': {'mode': None, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': [], 'ir_enable_torch_wrap': None, 'splitting_ops': None, 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': None, 'compile_ranges_endpoints': None, 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': , 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': None, 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': None, 'pass_config': {}, 'max_cudagraph_capture_size': 8, 'dynamic_shapes_config': {'type': , 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': None, 'static_all_moe_layers': []}} (APIServer pid=351099) WARNING 06-22 22:40:47 [envs.py:1941] Unknown vLLM environment variable detected: VLLM_XPU_ZERO_FRESH_GDN_STATE (APIServer pid=351099) WARNING 06-22 22:40:47 [envs.py:1941] Unknown vLLM environment variable detected: VLLM_XPU_GDN_REPLAYSSM_SPEC (APIServer pid=351099) WARNING 06-22 22:40:47 [envs.py:1941] Unknown vLLM environment variable detected: VLLM_XPU_SPEC_DECODE_RESTORE_DRAFT_FULL_ACCEPT_GDN_STATE (APIServer pid=351099) WARNING 06-22 22:40:47 [envs.py:1941] Unknown vLLM environment variable detected: VLLM_XPU_FORCE_QUARK_REPACK (APIServer pid=351099) WARNING 06-22 22:40:47 [envs.py:1941] Unknown vLLM environment variable detected: VLLM_EXTRA_ARGS (APIServer pid=351099) WARNING 06-22 22:40:47 [envs.py:1941] Unknown vLLM environment variable detected: VLLM_XPU_GDN_NATIVE_FALLBACK (APIServer pid=351099) WARNING 06-22 22:40:47 [envs.py:1941] Unknown vLLM environment variable detected: VLLM_XPU_FORCE_GRAPH_WITH_COMM (APIServer pid=351099) WARNING 06-22 22:40:47 [envs.py:1941] Unknown vLLM environment variable detected: VLLM_XPU_CUSTOM_ALLREDUCE_CLONE_INPUT (APIServer pid=351099) WARNING 06-22 22:40:47 [envs.py:1941] Unknown vLLM environment variable detected: VLLM_XPU_COMPILE_ALLREDUCE_CUSTOM_OP (APIServer pid=351099) WARNING 06-22 22:40:47 [envs.py:1941] Unknown vLLM environment variable detected: VLLM_XPU_USE_CUSTOM_OP_COLLECTIVES (APIServer pid=351099) WARNING 06-22 22:40:47 [envs.py:1941] Unknown vLLM environment variable detected: VLLM_XPU_GRAPH_NOOP_COMM_CAPTURE (APIServer pid=351099) WARNING 06-22 22:40:47 [envs.py:1941] Unknown vLLM environment variable detected: VLLM_XPU_GDN_PREFILL_RECURRENT_FALLBACK (APIServer pid=351099) WARNING 06-22 22:40:47 [envs.py:1941] Unknown vLLM environment variable detected: VLLM_XPU_CUSTOM_ALLREDUCE_GRAPH_CLONE_INPUT (APIServer pid=351099) WARNING 06-22 22:40:47 [envs.py:1941] Unknown vLLM environment variable detected: VLLM_XPU_SPEC_DECODE_RESTORE_DRAFT_PARTIAL_REJECT_GDN_STATE (APIServer pid=351099) WARNING 06-22 22:40:47 [envs.py:1941] Unknown vLLM environment variable detected: VLLM_XPU_QUARK_W8A8_MOE (APIServer pid=351099) WARNING 06-22 22:40:47 [envs.py:1941] Unknown vLLM environment variable detected: VLLM_XPU_DISABLE_PREFILL_CUDAGRAPH_REPLAY (APIServer pid=351099) WARNING 06-22 22:40:47 [envs.py:1941] Unknown vLLM environment variable detected: VLLM_XPU_GREEDY_SAMPLE_TOPK_FALLBACK (APIServer pid=351099) WARNING 06-22 22:40:47 [envs.py:1941] Unknown vLLM environment variable detected: VLLM_USE_V1 (APIServer pid=351099) INFO 06-22 22:40:55 [nixl_utils.py:20] Setting UCX_RCACHE_MAX_UNRELEASED to '1024' to avoid a rare memory leak in UCX when using NIXL. (APIServer pid=351099) WARNING 06-22 22:40:55 [nixl_utils.py:34] NIXL is not available (APIServer pid=351099) WARNING 06-22 22:40:55 [nixl_utils.py:44] NIXL agent config is not available (APIServer pid=351099) INFO 06-22 22:40:55 [model.py:563] Resolved architecture: Qwen3_5MoeForConditionalGeneration (APIServer pid=351099) INFO 06-22 22:40:55 [model.py:1692] Using max model len 32768 (APIServer pid=351099) INFO 06-22 22:41:01 [model.py:563] Resolved architecture: Qwen3_5MoeMTP (APIServer pid=351099) INFO 06-22 22:41:01 [model.py:1692] Using max model len 262144 (APIServer pid=351099) WARNING 06-22 22:41:01 [speculative.py:659] Enabling num_speculative_tokens > 1 will run multiple times of forward on same MTP layer,which may result in lower acceptance rate (APIServer pid=351099) INFO 06-22 22:41:01 [scheduler.py:239] Chunked prefill is enabled with max_num_batched_tokens=8192. (APIServer pid=351099) WARNING 06-22 22:41:01 [config.py:402] Mamba cache mode is set to 'none' when prefix caching is disabled (APIServer pid=351099) INFO 06-22 22:41:01 [vllm.py:844] Asynchronous scheduling is disabled. (APIServer pid=351099) INFO 06-22 22:41:01 [kernel.py:210] Final IR op priority after setting platform defaults: IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native']) (APIServer pid=351099) WARNING 06-22 22:41:01 [xpu.py:208] Forcing XPU Graph with communication ops because VLLM_XPU_FORCE_GRAPH_WITH_COMM=1. (APIServer pid=351099) [transformers] `Qwen2VLImageProcessorFast` is deprecated. The `Fast` suffix for image processors has been removed; use `Qwen2VLImageProcessor` instead. (APIServer pid=351099) INFO 06-22 22:41:02 [registry.py:126] All limits of multimodal modalities supported by the model are set to 0, running in text-only mode. WARNING 06-22 22:41:08 [nixl_utils.py:34] NIXL is not available WARNING 06-22 22:41:08 [nixl_utils.py:44] NIXL agent config is not available (EngineCore pid=351909) INFO 06-22 22:41:08 [core.py:246] Initializing a V1 LLM engine (v0.20.2rc1.dev13+g9557d9108.d20260620) with config: model='/mnt/fast-ai/llm-cache/hf/models--nameistoken--Qwen3.6-35B-A3B-Quark-W8A8-INT8/snapshots/cced56592e8c8935f8220836b4baa04dfd389118', speculative_config=SpeculativeConfig(method='mtp', model='/mnt/fast-ai/llm-cache/hf/models--nameistoken--Qwen3.6-35B-A3B-Quark-W8A8-INT8/snapshots/cced56592e8c8935f8220836b4baa04dfd389118', num_spec_tokens=3), tokenizer='/mnt/fast-ai/llm-cache/hf/models--nameistoken--Qwen3.6-35B-A3B-Quark-W8A8-INT8/snapshots/cced56592e8c8935f8220836b4baa04dfd389118', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=4, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=True, quantization=quark, quantization_config=None, enforce_eager=False, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=xpu, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=0, served_model_name=qwen36-35b-a3b-fp8, enable_prefix_caching=False, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': , 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['none'], 'ir_enable_torch_wrap': True, 'splitting_ops': ['vllm::unified_attention_with_output', 'vllm::unified_mla_attention_with_output', 'vllm::mamba_mixer2', 'vllm::mamba_mixer', 'vllm::short_conv', 'vllm::linear_attention', 'vllm::plamo2_mamba_mixer', 'vllm::gdn_attention_core', 'vllm::gdn_attention_core_xpu', 'vllm::olmo_hybrid_gdn_full_forward', 'vllm::kda_attention', 'vllm::sparse_attn_indexer', 'vllm::rocm_aiter_sparse_attn_indexer', 'vllm::deepseek_v4_attention', 'vllm::unified_kv_cache_update', 'vllm::unified_mla_kv_cache_update'], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [8192], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': , 'cudagraph_num_of_warmups': 1, 'cudagraph_capture_sizes': [1, 2, 4, 8], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 8, 'dynamic_shapes_config': {'type': , 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native']), enable_flashinfer_autotune=False, moe_backend='auto') (EngineCore pid=351909) WARNING 06-22 22:41:08 [multiproc_executor.py:1367] Reducing Torch parallelism from 16 threads to 1 to avoid unnecessary CPU contention. Set OMP_NUM_THREADS in the external environment to tune this value as needed. (EngineCore pid=351909) INFO 06-22 22:41:08 [multiproc_executor.py:243] DP group leader: node_rank=0, node_rank_within_dp=0, master_addr=127.0.0.1, mq_connect_ip=10.0.0.65 (local), world_size=4, local_world_size=4 WARNING 06-22 22:41:13 [nixl_utils.py:34] NIXL is not available WARNING 06-22 22:41:13 [nixl_utils.py:44] NIXL agent config is not available WARNING 06-22 22:41:13 [nixl_utils.py:34] NIXL is not available WARNING 06-22 22:41:13 [nixl_utils.py:44] NIXL agent config is not available WARNING 06-22 22:41:13 [nixl_utils.py:34] NIXL is not available WARNING 06-22 22:41:13 [nixl_utils.py:44] NIXL agent config is not available WARNING 06-22 22:41:13 [nixl_utils.py:34] NIXL is not available WARNING 06-22 22:41:13 [nixl_utils.py:44] NIXL agent config is not available [transformers] `Qwen2VLImageProcessorFast` is deprecated. The `Fast` suffix for image processors has been removed; use `Qwen2VLImageProcessor` instead. [transformers] `Qwen2VLImageProcessorFast` is deprecated. The `Fast` suffix for image processors has been removed; use `Qwen2VLImageProcessor` instead. [transformers] `Qwen2VLImageProcessorFast` is deprecated. The `Fast` suffix for image processors has been removed; use `Qwen2VLImageProcessor` instead. [transformers] `Qwen2VLImageProcessorFast` is deprecated. The `Fast` suffix for image processors has been removed; use `Qwen2VLImageProcessor` instead. INFO 06-22 22:41:15 [registry.py:126] All limits of multimodal modalities supported by the model are set to 0, running in text-only mode. INFO 06-22 22:41:15 [registry.py:126] All limits of multimodal modalities supported by the model are set to 0, running in text-only mode. INFO 06-22 22:41:15 [registry.py:126] All limits of multimodal modalities supported by the model are set to 0, running in text-only mode. INFO 06-22 22:41:15 [registry.py:126] All limits of multimodal modalities supported by the model are set to 0, running in text-only mode. (Worker pid=352026) INFO 06-22 22:41:18 [parallel_state.py:1542] world_size=4 rank=3 local_rank=3 distributed_init_method=tcp://127.0.0.1:41525 backend=xccl (Worker pid=352025) INFO 06-22 22:41:18 [parallel_state.py:1542] world_size=4 rank=2 local_rank=2 distributed_init_method=tcp://127.0.0.1:41525 backend=xccl (Worker pid=352024) INFO 06-22 22:41:18 [parallel_state.py:1542] world_size=4 rank=1 local_rank=1 distributed_init_method=tcp://127.0.0.1:41525 backend=xccl (Worker pid=352023) INFO 06-22 22:41:18 [parallel_state.py:1542] world_size=4 rank=0 local_rank=0 distributed_init_method=tcp://127.0.0.1:41525 backend=xccl (Worker pid=352023) INFO 06-22 22:41:18 [parallel_state.py:1855] rank 0 in world size 4 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0, EPLB rank N/A 2026:06:22-22:41:18:352025 |CCL_WARN| value of CCL_ATL_TRANSPORT changed to be ofi (default:mpi) 2026:06:22-22:41:18:352025 |CCL_WARN| value of CCL_TOPO_P2P_ACCESS changed to be 1 (default:-1) 2026:06:22-22:41:18:352024 |CCL_WARN| value of CCL_ATL_TRANSPORT changed to be ofi (default:mpi) 2026:06:22-22:41:18:352024 |CCL_WARN| value of CCL_TOPO_P2P_ACCESS changed to be 1 (default:-1) 2026:06:22-22:41:18:352026 |CCL_WARN| value of CCL_ATL_TRANSPORT changed to be ofi (default:mpi) 2026:06:22-22:41:18:352026 |CCL_WARN| value of CCL_TOPO_P2P_ACCESS changed to be 1 (default:-1) 2026:06:22-22:41:18:352023 |CCL_WARN| value of CCL_ATL_TRANSPORT changed to be ofi (default:mpi) 2026:06:22-22:41:18:352023 |CCL_WARN| value of CCL_TOPO_P2P_ACCESS changed to be 1 (default:-1) 2026:06:22-22:41:18:352025 |CCL_WARN| could not get local_idx/count from environment variables, trying to get them from ATL 2026:06:22-22:41:18:352024 |CCL_WARN| could not get local_idx/count from environment variables, trying to get them from ATL 2026:06:22-22:41:18:352026 |CCL_WARN| could not get local_idx/count from environment variables, trying to get them from ATL 2026:06:22-22:41:18:352023 |CCL_WARN| could not get local_idx/count from environment variables, trying to get them from ATL 2026:06:22-22:41:19:352025:[2] |CCL_WARN| topology recognition shows PCIe connection between devices. If this is not correct, you can disable topology recognition, with CCL_TOPO_FABRIC_VERTEX_CONNECTION_CHECK=0. This will assume XeLinks across devices 2026:06:22-22:41:19:352023:[0] |CCL_WARN| topology recognition shows PCIe connection between devices. If this is not correct, you can disable topology recognition, with CCL_TOPO_FABRIC_VERTEX_CONNECTION_CHECK=0. This will assume XeLinks across devices 2026:06:22-22:41:19:352024:[1] |CCL_WARN| topology recognition shows PCIe connection between devices. If this is not correct, you can disable topology recognition, with CCL_TOPO_FABRIC_VERTEX_CONNECTION_CHECK=0. This will assume XeLinks across devices 2026:06:22-22:41:19:352026:[3] |CCL_WARN| topology recognition shows PCIe connection between devices. If this is not correct, you can disable topology recognition, with CCL_TOPO_FABRIC_VERTEX_CONNECTION_CHECK=0. This will assume XeLinks across devices (Worker pid=352023) WARNING 06-22 22:41:20 [__init__.py:204] min_p and logit_bias parameters won't work with speculative decoding. (Worker_TP0 pid=352023) INFO 06-22 22:41:20 [gpu_model_runner.py:10757] Starting to load model /mnt/fast-ai/llm-cache/hf/models--nameistoken--Qwen3.6-35B-A3B-Quark-W8A8-INT8/snapshots/cced56592e8c8935f8220836b4baa04dfd389118... (Worker pid=352026) WARNING 06-22 22:41:20 [__init__.py:204] min_p and logit_bias parameters won't work with speculative decoding. (Worker pid=352025) WARNING 06-22 22:41:20 [__init__.py:204] min_p and logit_bias parameters won't work with speculative decoding. (Worker pid=352024) WARNING 06-22 22:41:20 [__init__.py:204] min_p and logit_bias parameters won't work with speculative decoding. (Worker_TP0 pid=352023) INFO 06-22 22:41:20 [xpu.py:122] Using backend AttentionBackendEnum.FLASH_ATTN for vit attention (Worker_TP0 pid=352023) INFO 06-22 22:41:20 [mm_encoder_attention.py:372] Using AttentionBackendEnum.FLASH_ATTN for MMEncoderAttention. (Worker_TP0 pid=352023) INFO 06-22 22:41:20 [__init__.py:463] Selected XPUInt8ScaledMMLinearKernel for QuarkW8A8Int8 (Worker_TP0 pid=352023) INFO 06-22 22:41:20 [gdn_linear_attn.py:1454] Using Triton/FLA GDN prefill kernel (Worker_TP0 pid=352023) INFO 06-22 22:41:20 [int8.py:152] Using XPU Int8 MoE backend out of potential backends: ['XPU', 'TRITON']. (Worker_TP0 pid=352023) INFO 06-22 22:41:21 [xpu.py:59] Setting VLLM_KV_CACHE_LAYOUT to 'NHD' for XPU; only NHD layout is supported by XPU attention kernels. (Worker_TP0 pid=352023) INFO 06-22 22:41:21 [xpu.py:87] Using Flash Attention backend. (Worker_TP0 pid=352023) INFO 06-22 22:41:21 [flash_attn.py:649] Using FlashAttention version 2 (Worker_TP1 pid=352024) INFO 06-22 22:41:21 [xpu.py:59] Setting VLLM_KV_CACHE_LAYOUT to 'NHD' for XPU; only NHD layout is supported by XPU attention kernels. (Worker_TP2 pid=352025) INFO 06-22 22:41:21 [xpu.py:59] Setting VLLM_KV_CACHE_LAYOUT to 'NHD' for XPU; only NHD layout is supported by XPU attention kernels. (Worker_TP3 pid=352026) INFO 06-22 22:41:21 [xpu.py:59] Setting VLLM_KV_CACHE_LAYOUT to 'NHD' for XPU; only NHD layout is supported by XPU attention kernels. (Worker_TP0 pid=352023) INFO 06-22 22:41:21 [weight_utils.py:904] Filesystem type for checkpoints: EXT4. Checkpoint size: 34.15 GiB. Available RAM: 110.69 GiB. (Worker_TP0 pid=352023) INFO 06-22 22:41:21 [weight_utils.py:927] Auto-prefetch is disabled because the filesystem (EXT4) is not a recognized network FS (NFS/Lustre). If you want to force prefetching, start vLLM with --safetensors-load-strategy=prefetch. (Worker_TP0 pid=352023) Loading safetensors checkpoint shards: 0% Completed | 0/7 [00:00= mamba page size. (Worker_TP1 pid=352024) INFO 06-22 22:41:34 [interface.py:633] Padding mamba page size by 7.46% to ensure that mamba page size and attention page size are exactly equal. (Worker_TP0 pid=352023) INFO 06-22 22:41:34 [weight_utils.py:904] Filesystem type for checkpoints: EXT4. Checkpoint size: 34.15 GiB. Available RAM: 109.91 GiB. (Worker_TP0 pid=352023) Loading safetensors checkpoint shards: 0% Completed | 0/7 [00:00= mamba page size. (Worker_TP3 pid=352026) INFO 06-22 22:41:34 [interface.py:633] Padding mamba page size by 7.46% to ensure that mamba page size and attention page size are exactly equal. (Worker_TP2 pid=352025) INFO 06-22 22:41:34 [interface.py:492] Setting kv cache block size to 64 for FLASH_ATTN backend. (Worker_TP2 pid=352025) INFO 06-22 22:41:34 [interface.py:609] Setting attention block size to 576 tokens to ensure that attention page size is >= mamba page size. (Worker_TP2 pid=352025) INFO 06-22 22:41:34 [interface.py:633] Padding mamba page size by 7.46% to ensure that mamba page size and attention page size are exactly equal. (Worker_TP0 pid=352023) Loading safetensors checkpoint shards: 14% Completed | 1/7 [00:00<00:00, 7.76it/s] (Worker_TP0 pid=352023) Loading safetensors checkpoint shards: 29% Completed | 2/7 [00:00<00:00, 7.44it/s] (Worker_TP0 pid=352023) Loading safetensors checkpoint shards: 43% Completed | 3/7 [00:00<00:00, 7.05it/s] (Worker_TP0 pid=352023) Loading safetensors checkpoint shards: 57% Completed | 4/7 [00:00<00:00, 6.93it/s] (Worker_TP0 pid=352023) Loading safetensors checkpoint shards: 71% Completed | 5/7 [00:00<00:00, 6.83it/s] (Worker_TP0 pid=352023) Loading safetensors checkpoint shards: 86% Completed | 6/7 [00:00<00:00, 6.80it/s] (Worker_TP0 pid=352023) Loading safetensors checkpoint shards: 100% Completed | 7/7 [00:00<00:00, 7.46it/s] (Worker_TP0 pid=352023) Loading safetensors checkpoint shards: 100% Completed | 7/7 [00:00<00:00, 7.20it/s] (Worker_TP0 pid=352023) (Worker_TP0 pid=352023) INFO 06-22 22:41:35 [default_loader.py:391] Loading weights took 0.99 seconds (Worker_TP0 pid=352023) INFO 06-22 22:41:35 [fp8.py:578] Using MoEPrepareAndFinalizeNoDPEPModular (Worker_TP0 pid=352023) INFO 06-22 22:41:35 [llm_base_proposer.py:1484] Detected MTP model. Sharing target model embedding weights with the draft model. (Worker_TP0 pid=352023) INFO 06-22 22:41:35 [llm_base_proposer.py:1540] Detected MTP model. Sharing target model lm_head weights with the draft model. (Worker_TP0 pid=352023) INFO 06-22 22:41:35 [gpu_model_runner.py:10859] Model loading took 8.79 GiB memory and 14.930704 seconds (Worker_TP0 pid=352023) INFO 06-22 22:41:35 [interface.py:492] Setting kv cache block size to 64 for FLASH_ATTN backend. (Worker_TP0 pid=352023) INFO 06-22 22:41:35 [interface.py:609] Setting attention block size to 576 tokens to ensure that attention page size is >= mamba page size. (Worker_TP0 pid=352023) INFO 06-22 22:41:35 [interface.py:633] Padding mamba page size by 7.46% to ensure that mamba page size and attention page size are exactly equal. (Worker_TP1 pid=352024) WARNING 06-22 22:41:44 [backends.py:1048] Failed to read file (Worker_TP0 pid=352023) WARNING 06-22 22:41:44 [backends.py:1048] Failed to read file (Worker_TP2 pid=352025) WARNING 06-22 22:41:44 [backends.py:1048] Failed to read file (Worker_TP0 pid=352023) INFO 06-22 22:41:44 [backends.py:1089] Using cache directory: /mnt/fast-ai/vllm-cache-exp/qwen36-ablation-tp4-mtp-k3-graph-throughput-probe/vllm/torch_compile_cache/7b543d29da/rank_0_0/backbone for vLLM's torch.compile (Worker_TP0 pid=352023) INFO 06-22 22:41:44 [backends.py:1148] Dynamo bytecode transform time: 7.90 s (Worker_TP3 pid=352026) WARNING 06-22 22:41:44 [backends.py:1048] Failed to read file (Worker_TP0 pid=352023) INFO 06-22 22:41:48 [backends.py:378] Cache the graph of compile range (1, 8192) for later use (Worker_TP0 pid=352023) INFO 06-22 22:42:26 [backends.py:393] Compiling a graph for compile range (1, 8192) takes 41.88 s (Worker_TP0 pid=352023) INFO 06-22 22:42:33 [decorators.py:708] saved AOT compiled function to /mnt/fast-ai/vllm-cache-exp/qwen36-ablation-tp4-mtp-k3-graph-throughput-probe/vllm/torch_compile_cache/torch_aot_compile/6daee6364651b8d02175ef8d4bae18c86eccc87559c7674e9d3fe9af57e89612/rank_0_0/model (Worker_TP0 pid=352023) INFO 06-22 22:42:33 [monitor.py:53] torch.compile took 57.01 s in total (EngineCore pid=351909) INFO 06-22 22:42:37 [shm_broadcast.py:681] No available shared memory broadcast block found in 60 seconds. This typically happens when some processes are hanging or doing some time-consuming work (e.g. compilation, weight/kv cache quantization). (Worker_TP0 pid=352023) INFO 06-22 22:42:38 [monitor.py:81] Initial profiling/warmup run took 4.90 s (Worker_TP0 pid=352023) WARNING 06-22 22:42:39 [backends.py:1048] Failed to read file (Worker_TP0 pid=352023) INFO 06-22 22:42:39 [backends.py:1089] Using cache directory: /mnt/fast-ai/vllm-cache-exp/qwen36-ablation-tp4-mtp-k3-graph-throughput-probe/vllm/torch_compile_cache/7b543d29da/rank_0_0/eagle_head for vLLM's torch.compile (Worker_TP0 pid=352023) INFO 06-22 22:42:39 [backends.py:1148] Dynamo bytecode transform time: 0.98 s (Worker_TP1 pid=352024) WARNING 06-22 22:42:39 [backends.py:1048] Failed to read file (Worker_TP3 pid=352026) WARNING 06-22 22:42:39 [backends.py:1048] Failed to read file (Worker_TP2 pid=352025) WARNING 06-22 22:42:39 [backends.py:1048] Failed to read file (Worker_TP0 pid=352023) INFO 06-22 22:42:54 [backends.py:393] Compiling a graph for compile range (1, 8192) takes 15.26 s (Worker_TP0 pid=352023) INFO 06-22 22:42:54 [decorators.py:708] saved AOT compiled function to /mnt/fast-ai/vllm-cache-exp/qwen36-ablation-tp4-mtp-k3-graph-throughput-probe/vllm/torch_compile_cache/torch_aot_compile/1b79d13c67134b00c4adb04abce932bab253710c9701e8c14a663709d7cbf6d1/rank_0_0/model (Worker_TP0 pid=352023) INFO 06-22 22:42:54 [monitor.py:53] torch.compile took 16.67 s in total (Worker_TP0 pid=352023) INFO 06-22 22:43:00 [fused_moe.py:1078] Using configuration from /home/steve/src/vllm/vllm/model_executor/layers/fused_moe/configs/E=256,N=128,device_name=Intel(R)_Arc(TM)_Pro_B70_Graphics,dtype=fp8_w8a8,block_shape=[128,128].json for MoE layer. (Worker_TP0 pid=352023) INFO 06-22 22:43:06 [monitor.py:81] Initial profiling/warmup run took 11.31 s (Worker_TP0 pid=352023) INFO 06-22 22:43:13 [gpu_model_runner.py:2256] Cleared ReplaySSM spec dummy state for 30 GDN layers (profile run) (Worker_TP0 pid=352023) INFO 06-22 22:43:14 [gpu_worker.py:487] Available KV cache memory: 20.15 GiB (EngineCore pid=351909) WARNING 06-22 22:43:14 [kv_cache_utils.py:1155] Add 3 padding layers, may waste at most 10.00% KV cache memory (EngineCore pid=351909) INFO 06-22 22:43:14 [kv_cache_utils.py:1334] Accounting for XPU ReplaySSM spec scratch: 1.88 MiB per KV block; planned KV blocks 3335 -> 2557 (EngineCore pid=351909) INFO 06-22 22:43:14 [kv_cache_utils.py:1334] Accounting for XPU ReplaySSM spec scratch: 1.88 MiB per KV block; planned KV blocks 3333 -> 2555 (EngineCore pid=351909) INFO 06-22 22:43:14 [kv_cache_utils.py:1855] GPU KV cache size: 1,213,365 tokens (EngineCore pid=351909) INFO 06-22 22:43:14 [kv_cache_utils.py:1856] Maximum concurrency for 32,768 tokens per request: 37.03x (Worker_TP0 pid=352023) INFO 06-22 22:43:14 [utils.py:60] `_KV_CACHE_LAYOUT_OVERRIDE` variable detected. Setting KV cache layout to NHD. (Worker_TP0 pid=352023) INFO 06-22 22:43:14 [gpu_model_runner.py:2231] Preallocated ReplaySSM spec rings for 30 GDN layers (max_spec_len=4) (Worker_TP0 pid=352023) INFO 06-22 22:43:14 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled. (Worker_TP3 pid=352026) INFO 06-22 22:43:14 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled. (Worker_TP1 pid=352024) INFO 06-22 22:43:14 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled. (Worker_TP2 pid=352025) INFO 06-22 22:43:14 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled. (Worker_TP0 pid=352023) WARNING 06-22 22:43:15 [parallel_state.py:566] Skipping communicator graph-capture context for XpuCommunicator under VLLM_XPU_GRAPH_NOOP_COMM_CAPTURE=1. (Worker_TP0 pid=352023) Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 0%| | 0/4 [00:00, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['none'], 'ir_enable_torch_wrap': True, 'splitting_ops': ['vllm::unified_attention_with_output', 'vllm::unified_mla_attention_with_output', 'vllm::mamba_mixer2', 'vllm::mamba_mixer', 'vllm::short_conv', 'vllm::linear_attention', 'vllm::plamo2_mamba_mixer', 'vllm::gdn_attention_core', 'vllm::gdn_attention_core_xpu', 'vllm::olmo_hybrid_gdn_full_forward', 'vllm::kda_attention', 'vllm::sparse_attn_indexer', 'vllm::rocm_aiter_sparse_attn_indexer', 'vllm::deepseek_v4_attention', 'vllm::unified_kv_cache_update', 'vllm::unified_mla_kv_cache_update'], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_vision_items_per_batch': 0, 'encoder_cudagraph_max_frames_per_batch': None, 'compile_sizes': [], 'compile_ranges_endpoints': [8192], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': , 'cudagraph_num_of_warmups': 1, 'cudagraph_capture_sizes': [1, 2, 4, 8], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 8, 'dynamic_shapes_config': {'type': , 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': False, 'static_all_moe_layers': []}, kernel_config=KernelConfig(ir_op_priority=IrOpPriorityConfig(rms_norm=['native'], fused_add_rms_norm=['native']), enable_flashinfer_autotune=False, moe_backend='auto'), (EngineCore pid=351909) ERROR 06-22 22:47:45 [dump_input.py:79] Dumping scheduler output for model execution: SchedulerOutput(scheduled_new_reqs=[], scheduled_cached_reqs=CachedRequestData(req_ids=['cmpl-b8c0097e20849331-0-a558a7e9'],resumed_req_ids=set(),new_token_ids_lens=[],all_token_ids_lens={},new_block_ids=[None],num_computed_tokens=[931],num_output_tokens=[434]), num_scheduled_tokens={cmpl-b8c0097e20849331-0-a558a7e9: 4}, total_num_scheduled_tokens=4, scheduled_spec_decode_tokens={cmpl-b8c0097e20849331-0-a558a7e9: [0, 0, 0]}, scheduled_encoder_inputs={}, num_common_prefix_blocks=[0, 0, 0, 0], finished_req_ids=[], free_encoder_mm_hashes=[], preempted_req_ids=[], has_structured_output_requests=false, pending_structured_output_tokens=false, num_invalid_spec_tokens=null, kv_connector_metadata=null, ec_connector_metadata=null, new_block_ids_to_zero=null) (EngineCore pid=351909) ERROR 06-22 22:47:45 [dump_input.py:81] Dumping scheduler stats: SchedulerStats(num_running_reqs=1, num_waiting_reqs=0, num_skipped_waiting_reqs=0, step_counter=0, current_wave=0, kv_cache_usage=0.005481597494126911, prefix_cache_stats=PrefixCacheStats(reset=False, requests=0, queries=0, hits=0, preempted_requests=0, preempted_queries=0, preempted_hits=0), connector_prefix_cache_stats=None, kv_cache_eviction_events=[], spec_decoding_stats=None, kv_connector_stats=None, waiting_lora_adapters={}, running_lora_adapters={}, cudagraph_stats=None, perf_stats=None) (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] EngineCore encountered a fatal error. (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] Traceback (most recent call last): (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] File "/home/steve/src/vllm/vllm/v1/engine/core.py", line 1356, in run_engine_core (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] engine_core.run_busy_loop() (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] File "/home/steve/src/vllm/vllm/v1/engine/core.py", line 1397, in run_busy_loop (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] self._process_engine_step() (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] File "/home/steve/src/vllm/vllm/v1/engine/core.py", line 1436, in _process_engine_step (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] outputs, model_executed = self.step_fn() (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] ^^^^^^^^^^^^^^ (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] File "/home/steve/src/vllm/vllm/v1/engine/core.py", line 574, in step (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] model_output = future.result() (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] ^^^^^^^^^^^^^^^ (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] File "/home/steve/src/vllm/vllm/v1/executor/multiproc_executor.py", line 194, in result (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] return super().result() (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] ^^^^^^^^^^^^^^^^ (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] File "/home/steve/.local/share/uv/python/cpython-3.12.13-linux-x86_64-gnu/lib/python3.12/concurrent/futures/_base.py", line 449, in result (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] return self.__get_result() (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] ^^^^^^^^^^^^^^^^^^^ (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] File "/home/steve/.local/share/uv/python/cpython-3.12.13-linux-x86_64-gnu/lib/python3.12/concurrent/futures/_base.py", line 401, in __get_result (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] raise self._exception (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] File "/home/steve/src/vllm/vllm/v1/executor/multiproc_executor.py", line 198, in _wait_for_response (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] response = self.aggregate(self.get_response()) (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] ^^^^^^^^^^^^^^^^^^^ (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] File "/home/steve/src/vllm/vllm/v1/executor/multiproc_executor.py", line 514, in get_response (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] status, result = mq.dequeue(timeout=dequeue_timeout) (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] File "/home/steve/src/vllm/vllm/distributed/device_communicators/shm_broadcast.py", line 755, in dequeue (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] with self.acquire_read(timeout, indefinite) as buf: (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] File "/home/steve/.local/share/uv/python/cpython-3.12.13-linux-x86_64-gnu/lib/python3.12/contextlib.py", line 137, in __enter__ (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] return next(self.gen) (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] ^^^^^^^^^^^^^^ (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] File "/home/steve/src/vllm/vllm/distributed/device_communicators/shm_broadcast.py", line 677, in acquire_read (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] raise RuntimeError("cancelled") (EngineCore pid=351909) ERROR 06-22 22:47:45 [core.py:1365] RuntimeError: cancelled (EngineCore pid=351909) Process EngineCore: (EngineCore pid=351909) Traceback (most recent call last): (EngineCore pid=351909) File "/home/steve/.local/share/uv/python/cpython-3.12.13-linux-x86_64-gnu/lib/python3.12/multiprocessing/process.py", line 314, in _bootstrap (EngineCore pid=351909) self.run() (EngineCore pid=351909) File "/home/steve/.local/share/uv/python/cpython-3.12.13-linux-x86_64-gnu/lib/python3.12/multiprocessing/process.py", line 108, in run (EngineCore pid=351909) self._target(*self._args, **self._kwargs) (EngineCore pid=351909) File "/home/steve/src/vllm/vllm/v1/engine/core.py", line 1367, in run_engine_core (EngineCore pid=351909) raise e (EngineCore pid=351909) File "/home/steve/src/vllm/vllm/v1/engine/core.py", line 1356, in run_engine_core (EngineCore pid=351909) engine_core.run_busy_loop() (EngineCore pid=351909) File "/home/steve/src/vllm/vllm/v1/engine/core.py", line 1397, in run_busy_loop (EngineCore pid=351909) self._process_engine_step() (EngineCore pid=351909) File "/home/steve/src/vllm/vllm/v1/engine/core.py", line 1436, in _process_engine_step (EngineCore pid=351909) outputs, model_executed = self.step_fn() (EngineCore pid=351909) ^^^^^^^^^^^^^^ (EngineCore pid=351909) File "/home/steve/src/vllm/vllm/v1/engine/core.py", line 574, in step (EngineCore pid=351909) model_output = future.result() (EngineCore pid=351909) ^^^^^^^^^^^^^^^ (EngineCore pid=351909) File "/home/steve/src/vllm/vllm/v1/executor/multiproc_executor.py", line 194, in result (EngineCore pid=351909) return super().result() (EngineCore pid=351909) ^^^^^^^^^^^^^^^^ (EngineCore pid=351909) File "/home/steve/.local/share/uv/python/cpython-3.12.13-linux-x86_64-gnu/lib/python3.12/concurrent/futures/_base.py", line 449, in result (EngineCore pid=351909) return self.__get_result() (EngineCore pid=351909) ^^^^^^^^^^^^^^^^^^^ (EngineCore pid=351909) File "/home/steve/.local/share/uv/python/cpython-3.12.13-linux-x86_64-gnu/lib/python3.12/concurrent/futures/_base.py", line 401, in __get_result (EngineCore pid=351909) raise self._exception (EngineCore pid=351909) File "/home/steve/src/vllm/vllm/v1/executor/multiproc_executor.py", line 198, in _wait_for_response (EngineCore pid=351909) response = self.aggregate(self.get_response()) (EngineCore pid=351909) ^^^^^^^^^^^^^^^^^^^ (EngineCore pid=351909) File "/home/steve/src/vllm/vllm/v1/executor/multiproc_executor.py", line 514, in get_response (EngineCore pid=351909) status, result = mq.dequeue(timeout=dequeue_timeout) (EngineCore pid=351909) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (EngineCore pid=351909) File "/home/steve/src/vllm/vllm/distributed/device_communicators/shm_broadcast.py", line 755, in dequeue (EngineCore pid=351909) with self.acquire_read(timeout, indefinite) as buf: (EngineCore pid=351909) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ (EngineCore pid=351909) File "/home/steve/.local/share/uv/python/cpython-3.12.13-linux-x86_64-gnu/lib/python3.12/contextlib.py", line 137, in __enter__ (EngineCore pid=351909) return next(self.gen) (EngineCore pid=351909) ^^^^^^^^^^^^^^ (EngineCore pid=351909) File "/home/steve/src/vllm/vllm/distributed/device_communicators/shm_broadcast.py", line 677, in acquire_read (EngineCore pid=351909) raise RuntimeError("cancelled") (EngineCore pid=351909) RuntimeError: cancelled /home/steve/.local/share/uv/python/cpython-3.12.13-linux-x86_64-gnu/lib/python3.12/multiprocessing/resource_tracker.py:279: UserWarning: resource_tracker: There appear to be 4 leaked shared_memory objects to clean up at shutdown warnings.warn('resource_tracker: There appear to be %d '