vLLM에서 Qwen3.6 27B + 35B 구동하기 (단일 R9700, gfx1201)
Qwen3.6 27B + 35B on vLLM, single R9700 (gfx1201)
핵심 요약
단일 Radeon R9700 GPU를 사용하여 vLLM 환경에서 Qwen3.6 모델을 최적화하고 벤치마크한 결과입니다.
- vLLM 최적화 — 단일 R9700 GPU 환경에서 FP8 및 INT4 가중치를 활용한 설정 공유함.
- 벤치마크 결과 — 35B MoE 및 27B Dense 모델의 컨텍스트 깊이별 토큰 처리 속도를 측정함.
- 설정 팁 — vLLM 참조 설정 대비 텐서 병렬 크기 조정 및 speculative tokens 최적화 수행함.
- 토크나이저 수정 — Avesed 리포지토리의 고정된 패딩/절단 설정이 비전 모델에 미치는 오류를 해결함.
I've been tuning my new Radeon AI Pro R9700, and figured that this would be useful information for people who are trying to optimise their setups. I'm pretty happy with these results and looking forward to Qwen3.8..
Summary below provided by Claude (which helped me configure it to run on my system via podman).
Setup: stilldeadcode/vllm-radiance:0.5.8. Single (not dual) card.
https://hub.docker.com/r/stilldeadcode/vllm-radiance/
https://codeberg.org/StillDeadcode/vllm-radiance/
The reference config shipped with the image is tuned for FP8 weights on 2× R9700 (TP=2). Most of its defaults (AITER attention backend, FP8 KV, --no-async-scheduling, --mamba-cache-mode align, all RADIANCE_* toggles) are correct as-is and don't need touching. Here's what actually differs when running one card with INT4:
Config differences vs. reference
- --tensor-parallel-size 1 (no second card)
- --gpu-memory-utilization 0.98 (reference band is 0.90–0.97 on dual cards)
- num_speculative_tokens=4 on the 27B. Ladder-tested 2/3/4/8 directly against the container (4 arms × 2 loads × 2 reps × 4 depths); 4 wins at every depth by 17–48% over 8.
Model Weights:
Weights: Avesed/Qwen3.6-{27B,35B}-INT4-W4A16 (compressed-tensors, group_size 32). The 35B at FP8 simply won't fit one 32GB card at any useful context length.
Checkpoint fix (not an image issue): tokenizer.json in the Avesed INT4 repo ships truncation.max_length: 512 / padding: Fixed(512) baked in from calibration — breaks vision above ~672px. Set both to null.
Model notes
27B: Dense (no MoE), MTP on, num_speculative_tokens=4, 131,072 ctx.
35B: MoE (A3B), MTP off, 262,144 ctx.
Benchmark Results

