Qwen 3.6이 그냥 멈춰버림
qwen3.6 just stops
핵심 요약
Qwen 3.6 27B 모델이 작업 도중 멈추는 현상에 대해 vLLM 설정 및 파라미터 조정을 통한 해결 방법을 논의함.
- 작업 중단 현상 — Qwen 3.6 27B 모델이 CLI 및 OpenCode 환경에서 작업 도중 갑자기 멈추는 문제가 보고됨.
- 설정 최적화 — preserve thinking 옵션을 활성화하고 dflash를 비활성화하여 안정성을 높이는 방법이 제안됨.
- 컨텍스트 오버필링 — 모델의 컨텍스트가 가득 찼을 때 발생하는 문제일 수 있다는 분석이 나옴.
- 커스텀 솔루션 — 중단된 모델을 다시 자극해 실행하게 만드는 'cattleprod' 플러그인 사례가 언급됨.
Sometimes qwen 3.6 just stops at the middle of a task, is there a way to avoid it?
This is qwen-code CLI, but also happens on opencode.
Running with vLLM with docker compose:
services:
vllm-qwen36-27b-dual-dflash-noviz:
image: vllm/vllm-openai:nightly-1acd67a795ebccdf9b9db7697ae9082058301657
container_name: vllm-qwen36-27b-dual-dflash-noviz
restart: on-failure
ports:
- "${BIND_HOST:-0.0.0.0}:${PORT:-8080}:8000"
volumes:
- ${MODEL_DIR:-/home/ai/models/vllm}:/root/.cache/huggingface
- /home/ai/club-3090/models/qwen3.6-27b/vllm/cache/torch_compile:/root/.cache/vllm/torch_compile_cache
- /home/ai/club-3090/models/qwen3.6-27b/vllm/cache/triton:/root/.triton/cache
- /home/ai/club-3090/models/qwen3.6-27b/vllm/patches/vllm-marlin-pad/marlin.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/marlin.py:ro
- /home/ai/club-3090/models/qwen3.6-27b/vllm/patches/vllm-marlin-pad/MPLinearKernel.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mixed_precision/MPLinearKernel.py:ro
environment:
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- CUDA_DEVICE_ORDER=PCI_BUS_ID
- VLLM_WORKER_MULTIPROC_METHOD=spawn
- NCCL_CUMEM_ENABLE=0
- NCCL_P2P_DISABLE=1
- VLLM_NO_USAGE_STATS=1
- VLLM_USE_FLASHINFER_SAMPLER=1
- OMP_NUM_THREADS=1
- PYTORCH_CUDA_ALLOC_CONF=${PYTORCH_CUDA_ALLOC_CONF:-expandable_segments:True,max_split_size_mb:512}
shm_size: "16gb"
ipc: host
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids: ["0", "2"]
capabilities: [gpu]
entrypoint:
- /bin/bash
- -c
- |
exec vllm serve ${VLLM_ENFORCE_EAGER:+--enforce-eager} "$@"
- --
command:
- --model
- /root/.cache/huggingface/qwen3.6-27b-autoround-int4
- --served-model-name
- qwen
- --quantization
- auto_round
- --dtype
- bfloat16
- --tensor-parallel-size
- "2"
- --disable-custom-all-reduce
- --max-model-len
- "${MAX_MODEL_LEN:-185000}"
- --gpu-memory-utilization
- "${GPU_MEMORY_UTILIZATION:-0.95}"
- --max-num-seqs
- "${MAX_NUM_SEQS:-2}"
- --max-num-batched-tokens
- "8192"
- --language-model-only
- --trust-remote-code
- --reasoning-parser
- qwen3
- --default-chat-template-kwargs
- '{"enable_thinking": true}'
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
- --enable-prefix-caching
- --enable-chunked-prefill
- --speculative-config
- '{"method":"dflash","model":"/root/.cache/huggingface/qwen3.6-27b-dflash","num_speculative_tokens":5}'
- --host
- 0.0.0.0
- --port
- "8000"

