1800달러(GPU 비용)로 Qwen/Qwen3.6-27b-FP8 구동: P2P, 262K 컨텍스트, BF16 KV 캐시, 55 tok/s
$1800 (in GPU cost running with P2P running Qwen/Qwen3.6-27b-FP8 with 262K context and BF16 KV cache at 55 tok/s
핵심 요약
1,800달러 규모의 5060 Ti 4개 구성으로 262K 컨텍스트의 대규모 모델을 효율적으로 추론하는 환경을 공유함.
- 가성비 추론 환경 — 1,800달러의 GPU 비용으로 262K 컨텍스트의 대규모 모델 구동 가능함.
- 하드웨어 구성 — 5060 Ti 16GB 4개를 P2P로 연결하여 추론 전용 시스템 구축함.
- 성능 지표 — 초당 55 토큰의 생성 속도와 1,000 토큰/s 수준의 프롬프트 처리 성능을 기록함.
- 제약 사항 — 추론 전용으로만 사용 가능하며, PCIe 대역폭과 메모리 오버클럭이 성능에 영향을 줌.
안녕하세요 여러분, 1,700달러의 GPU 비용으로 추론 전용 단일 사용자 환경을 구축할 수 있다는 것을 공유하고 싶습니다.
설정: P2P를 사용한 4x 5060 ti (16GB)
미국에 거주하신다면 페이스북 마켓플레이스나 Slickdeals 같은 곳을 잘 살펴보세요. 중고 5060 ti 16GB 모델을 425~475달러에 찾을 수 있습니다.
한 가지 큰 주의점은 이런 구성은 오직 추론에만 관심이 있는 경우에만 유효하다는 것입니다.
사용한 VLLM 명령어:
export VLLM_SLEEP_WHEN_IDLE=1
export VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=1
export VLLM_WORKER_MULTIPROC_METHOD=spawn
export SAFETENSORS_FAST_GPU=1
export NCCL_P2P_DISABLE=0
export NCCL_CUMEM_ENABLE=1
export CUDA_DEVICE_ORDER=PCI_BUS_ID
export TORCH_FLOAT32_MATMUL_PRECISION=high
export PYTORCH_ALLOC_CONF=expandable_segments:True
# 제외됨: VLLM_USE_FLASHINFER_MOE_FP8 (dense 모델), VLLM_TEST_FORCE_FP8_MARLIN (네이티브 FP8 우선 테스트)
vllm serve /data/models/Qwen/Qwen3.6-27B-FP8 \
--host 0.0.0.0 --port 8080 \
--tensor-parallel-size 4 \
--performance-mode interactivity \
--trust-remote-code \
--language-model-only \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--reasoning-parser qwen3 \
--max-model-len 262144 \
--kv-cache-dtype bfloat16 \
--max-num-seqs 4 \
--gpu-memory-utilization 0.92 \
--speculative-config '{"method":"qwen3_next_mtp","num_speculative_tokens":3}' \
--compilation-config '{"max_cudagraph_capture_size":16,"mode":"VLLM_COMPILE"}' \
--async-scheduling \
--attention-backend flashinfer \
--enable-prefix-caching
벤치마크 명령어:
vllm bench serve --backend vllm --base-url http://localhost:8080 --endpoint /v1/completions --model /data/models/Qwen/Qwen3.6-27B-FP8 --dataset-name random --random-input-len 4096 --random-output-len 1024 --num-prompts 40 --max-concurrency 1 --num-warmups 5 --ignore-eos --seed 1234 --percentile-metrics ttft,tpot,itl,e2el --save-result --result-filename qwen36_c1_4k.json
============ Serving Benchmark Result ============
Successful requests: 40
Failed requests: 0
Maximum request concurrency: 1
Benchmark duration (s): 735.75
Total input tokens: 163840
Total generated tokens: 40960
Request throughput (req/s): 0.05
Output token throughput (tok/s): 55.67
Peak output token throughput (tok/s): 25.00
Peak concurrent requests: 2.00
Total token throughput (tok/s): 278.36
---------------Time to First Token----------------
Mean TTFT (ms): 4226.91
Median TTFT (ms): 4315.47
P99 TTFT (ms): 4320.32
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 13.85
Median TPOT (ms): 13.44
P99 TPOT (ms): 25.61
---------------Inter-token Latency----------------
Mean ITL (ms): 40.91
Median ITL (ms): 40.84
P99 ITL (ms): 41.59
----------------End-to-end Latency----------------
Mean E2EL (ms): 18393.49
Median E2EL (ms): 17991.18
P99 E2EL (ms): 30508.70
---------------Speculative Decoding---------------
Acceptance rate (%): 65.25
Acceptance length: 2.96
Drafts: 13853
Draft tokens: 41559
Accepted tokens: 27116
Per-position acceptance (%):
Position 0: 78.29
Position 1: 64.14
Position 2: 53.31
==================================================


