[3090] Gemma4 QAT + MTP 빠른 TPS 수치 [요약: 1.2~1.8배 향상]
[3090] Gemma4 QAT + MTP quick TPS numbers [TLDR 1.2-1.8x better]
핵심 요약
RTX 3090에서 Gemma 4 모델에 QAT와 MTP를 적용해 토큰 생성 속도를 획기적으로 높인 벤치마크 결과입니다.
- 성능 향상 — QAT와 MTP 기술을 통해 24GB VRAM 환경에서 토큰 생성 속도가 최대 1.8배 빨라짐
- 하드웨어 가성비 — 3090 같은 24GB GPU 사용자들에게 최적의 AI 모델 구동 환경 제공
- 벤치마크 공유 — 12B 및 31B 모델에 대한 구체적인 설정값과 속도 향상 데이터 공개
- 기술적 한계 — 긴 컨텍스트 처리 시 속도 저하 문제와 특정 설정에서의 호환성 이슈 존재
지난 몇 주는 24GB(이하) GPU 쓰는 거지들에겐 그야말로 축복이었다.
- 개쩌는 모델들 출시 (Gemma 4 / Qwen 3.6)
- QAT로 공짜 지능 획득
- MTP로 속도 보너스까지
이제 GPU 거지(24GB 이하)들이 더 이상 거지가 아닌 티핑 포인트에 도달했다.
Gemma 4 31b 돌리면서 40tok/s 나오는 것도 이미 만족스러웠는데, 이제는 70-80tok/s가 찍힌다.
3090 가격이 오르는 게 당연하지.
참고용:
- limit=1, OSL=192, concurrency 1, temp=1.0/top_k=64/top_p=0.95, ctx=40960, q8_0 KV cache, parallel=1
- 12b 모델은 TEXT 전용이랑 mmproj 멀티모달 둘 다 테스트해 봤는데, 속도 향상 폭은 똑같음.
(모델이랑 진짜로 '대화'가 가능하고, 질문하자마자 0.1초 만에 답변 쏟아내는 거 진짜 미쳤음. TTS는 아직 안 되지만.)
• 하드웨어
- CPU: Intel Core i9-13900H, 14 cores / 20 threads
- RAM: 62 GiB system RAM, 8 GiB swap
- GPU: NVIDIA GeForce RTX 3090, 24 GiB VRAM
- Driver/CUDA: NVIDIA driver 595.71.05, CUDA 13.2
- OS/kernel: Ubuntu 24.04-ish, Linux 6.17.0-35-generic
Startup config:
llama-server \
-m gemma-4-12B-it-qat-UD-Q4_K_XL.gguf \
--model-draft gemma-4-12B-it-qat-assistant-MTP-Q8_0.gguf \
--spec-type draft-mtp \
--spec-draft-n-max 4 \
--parallel 1 \
--ctx-size 40960 \
--temp 1.0 \
--top-p 0.95 \
--top-k 64 \
--spec-draft-ngl all \
--spec-draft-type-k q8_0 \
--spec-draft-type-v q8_0 \
UPDATE:
for 26b, turns out best N-max is 1, which gives a 1.26x speedup:
setting tok/s speedup accept
━━━━━━━━━ ━━━━━━━━ ━━━━━━━━━ ━━━━━━━━
no MTP 143.01 1.00x -
───────── ──────── ───────── ────────
n-max 1 180.01 1.26x 0.765
───────── ──────── ───────── ────────
n-max 2 175.77 1.23x 0.654
───────── ──────── ───────── ────────
n-max 3 170.37 1.19x 0.576
───────── ──────── ───────── ────────
n-max 4 165.90 1.16x 0.492
───────── ──────── ───────── ────────
n-max 5 155.51 1.09x 0.444
NOTE: These are Temp 1.0, so there is some stochastic voltatility to the numbers, but i think they are directionalyl correct.
Also what are the deets on this quick test?
11 requests, one each for coding, humanities, math, QA, RAG, reasoning, STEM, writing, multilingual, summarization, roleplay. Context allocated is 40960, but prompt lengths were only about 22 to 1578 tokens, average about 280. Output target is --osl 192 per turn; some samples are multi-turn, so max full-length total is 15 turns * 192 = 2880 generated tokens, but stop tokens can end samples early.
This is meant to be a quick and dirty benchmark to get a rough idea of potential impact of QAT + MTP on Gemma4 (on a 3090 GPU) A full proper grid of context + depth will be done separately.


