Llama.cpp 8% 속도 향상 PR
Llama.cpp PR 8% speed boost
핵심 요약
Llama.cpp의 샘플링 과정을 GPU로 이전하여 추론 속도를 높이는 PR 소개 및 벤치마크 결과 공유.
- GPU 샘플링 — 기존 CPU 기반 샘플링을 GPU로 이전하여 성능 개선함.
- 속도 향상 — 5090에서 8%, P40에서 4%의 토큰 처리 속도 향상 확인됨.
- 메모리 대역폭 — P40은 메모리 대역폭 제한으로 인해 상대적으로 개선 폭이 작음.
- 벤치마크 결과 — CPU와 GPU 샘플링 간의 처리 속도 비교 데이터 제공됨.
Llama.cpp는 현재 MTP가 활성화된 사용자를 위해 CPU 기반 샘플링을 사용함. 이 PR은 샘플링을 GPU로 옮기며, 5090에서 Qwen3.6:35b 모델 기준 tok/s가 8% 증가함. P40에서 테스트했을 때 4%의 추론 속도 향상을 관찰함.
Nvidia P40에서 최대 84 tok/s를 보게 되어 매우 흥분됨.
리눅스 + Tesla P40 (sm_61, Pascal) 환경에서 백엔드 샘플링 시 약 4% 향상됨:
CPU 샘플링: llama-server -m Qwen3.6-35B-A3B-UD-IQ4_NL.gguf --spec-type draft-mtp --seed 42
python3 mtp-bench.py code_python pred= 192 draft= 132 acc= 124 rate=0.939 tok/s=73.1 code_cpp pred= 113 draft= 76 acc= 74 rate=0.974 tok/s=75.9 explain_concept pred= 192 draft= 159 acc= 111 rate=0.698 tok/s=62.4 summarize pred= 192 draft= 167 acc= 107 rate=0.641 tok/s=59.6 qa_factual pred= 192 draft= 159 acc= 111 rate=0.698 tok/s=62.4 translation pred= 119 draft= 92 acc= 73 rate=0.793 tok/s=67.0 creative_short pred= 192 draft= 197 acc= 92 rate=0.467 tok/s=50.7 stepwise_math pred= 192 draft= 133 acc= 124 rate=0.932 tok/s=73.6 long_code_review pred= 192 draft= 155 acc= 113 rate=0.729 tok/s=63.8
백엔드 샘플링: llama-server -m Qwen3.6-35B-A3B-UD-IQ4_NL.gguf --spec-type draft-mtp --seed 42 -bs
python3 mtp-bench.py code_python pred= 192 draft= 132 acc= 124 rate=0.939 tok/s=76.2 code_cpp pred= 113 draft= 76 acc= 74 rate=0.974 tok/s=79.4 explain_concept pred= 192 draft= 159 acc= 111 rate=0.698 tok/s=64.6 summarize pred= 192 draft= 167 acc= 107 rate=0.641 tok/s=61.6 qa_factual pred= 192 draft= 159 acc= 111 rate=0.698 tok/s=64.6 translation pred= 119 draft= 92 acc= 73 rate=0.793 tok/s=69.6 creative_short pred= 192 draft= 197 acc= 92 rate=0.467 tok/s=52.1 stepwise_math pred= 192 draft= 133 acc= 124 rate=0.932 tok/s=76.6 long_code_review pred= 192 draft= 155 acc= 113 rate=0.729 tok/s=65.7
수락률(Acceptance ratio)은 백엔드와 CPU 샘플링 모두 정확히 동일함. RTX 5090(12%)보다 향상 폭이 작은 것은 예상된 결과임. P40은 메모리 대역폭 제한(sm_61, 580 GB/s vs RTX 5090의 1,792 GB/s)을 받기 때문에, CPU↔GPU 로짓 왕복 시간이 전체 디코드 시간에서 차지하는 비중이 작음. 하지만 여전히 최근 본 것 중 가장 큰 tok/s 향상임. (~+2 t/s).
