RTX 5080 16GB: 128k 컨텍스트에서 Qwen3.6 35B MoE 벤치마크 — 56 tok/s 및 MTP가 효과 없는 이유
RTX 5080 16GB: Qwen3.6 35B MoE at 128k context — 56 tok/s, and why MTP doesn't help
핵심 요약
128k 컨텍스트 환경에서 RTX 5080 16GB로 Qwen3.6 35B 모델을 돌릴 때 MTP가 성능에 미치는 영향 분석.
MTP 성능 분석 — 128k 컨텍스트에서는 MTP가 속도 향상에 도움을 주지 않음.
VRAM 최적화 — 35B 모델 사용 시 --fit-target 1536 설정으로 KV 캐시 부족 방지.
모델별 차이 — 27B 모델은 GPU에 완전히 올라가 MTP 효과가 크지만 35B는 오히려 느려짐.
코딩 에이전트 환경 — 128k 컨텍스트에서 56 tok/s의 안정적인 생성 속도 확보 가능.
MTP (Multi-Token Prediction) just merged into mainline llama.cpp at b9190. I promised u/WarthogConfident4039 a Qwen3.6 benchmarking round. Three configs, tested at real coding-agent context lengths (not just 512 tokens). The main finding surprised me.
TL;DR: 35B Q4_K_XL, no MTP,--fit-target 1536**, 131k context. That's the config.** 56 tok/s generation, 1,584 tok/s prompt processing at 128k context. MTP doesn't help at 128k — both converge to the same speed. Skip the complexity. The 27B IQ3 is worth considering if 56k context is enough for you (or if you have a 12 GB card where the 35B won't fit).
MTP requires --fit-target 1536 to reserve ~1.5 GB for the MTP compute buffer
That 1.5 GB pushes ~3 more MoE expert layers from GPU to CPU
CPU-bound expert layers are the bottleneck for MoE inference
MTP's multi-token speculation (~79% acceptance) doesn't compensate for the slower per-step speed
For the 27B, MTP helps because the model fits entirely on GPU (12.45 GB) — --fit-target 0 works with and without MTP, so there's no VRAM penalty. The 27B goes from ~56 tok/s (no MTP, older builds) to 73 tok/s with MTP.
Rule of thumb: MTP helps when your model fits on GPU. It hurts when the MTP compute buffer forces more layers to CPU.
Speed at Coding-Agent Context Lengths (the real test)
Everyone runs coding agents at 128k. Here's what actually happens as you fill the context window. Tested with synthetic prompts (Python classes, architecture docs, error stack traces — varied enough to prevent tokenizer compression), prompt cache disabled, 35B Q4_K_XL with --fit-target 1536:
Context
PP (no MTP)
PP (MTP)
TG (no MTP)
TG (MTP)
~8k
1,855 tok/s
1,712 tok/s
73 tok/s
79 tok/s
~32k
1,810 tok/s
1,674 tok/s
74 tok/s
70 tok/s
~64k
1,723 tok/s
1,583 tok/s
67 tok/s
76 tok/s
~128k
1,584 tok/s
1,437 tok/s
56 tok/s
56 tok/s
8k/32k TG measured in a separate run from 64k/128k — expect ~5-10% variance between rows from measurement noise.
At 128k context, MTP and no-MTP converge to the same TG speed (~56 tok/s). The KV cache fills VRAM at long context regardless of MTP, so the offload split ends up identical. MTP's multi-token speculation is offset by its compute overhead.
PP degrades gracefully: 1,855 → 1,584 tok/s from 8k to 128k (~15% decline). A 128k prompt processes in ~81 seconds.
The "97 tok/s" only exists at short context with --fit-target 0. At 64k+, --fit-target 0 OOMs because there's no headroom for KV cache growth. You must use --fit-target 1536 for long-context work, which brings speed down to ~73 tok/s at short context and ~56 tok/s at 128k.
Bottom line for coding agents: expect ~56 tok/s TG and ~1,500 tok/s PP at 128k context on 16GB. MTP is a wash — doesn't help or hurt at full context.
VRAM Usage
Config
VRAM used
VRAM free
Notes
A (27B IQ3+MTP)
14,803 MiB
1,039 MiB
Fully on GPU, fit-target 0
B (35B Q4_K_XL+MTP)
14,623 MiB
1,219 MiB
Partial offload, fit-target 1536
B (35B Q4_K_XL, no MTP)
15,815 MiB
27 MiB
Maximum GPU layers, fit-target 0
C (35B Q8_0+MTP)
14,567 MiB
1,275 MiB
Heavy offload, fit-target 1536
Context Limits (push to OOM)
Limit
27B IQ3
35B Q4_K_XL
35B Q8_0
Max ctx (q8_0 KV)
56k
131k+
131k+
Max ctx (q4_0 KV)
110k
131k+
131k+
Speed at max ctx
80.5 / 57.2
56
45
This is the biggest differentiator. The 35B MoE handles 131k context easily because its hybrid architecture (Gated DeltaNet + Attention) only has ~10 full-attention layers that need KV cache. The remaining SSM layers use a tiny recurrent state. The 27B dense model has KV on every layer, so it maxes out at 56k with q8_0 KV.
Tip for 27B users: switching from -ctk q8_0 -ctv q8_0 to -ctk q4_0 -ctv q4_0 extends your max context from 56k → 110k. Quality cost is minimal: q4_0 KV at 56k scores 218/220 CodeNeedle vs 220/220 with q8_0 KV (q4_0 at regular context: 219/220 — so most of the 2-line drop is from q4_0 itself, not the longer context).
The OOM at higher contexts is the MTP compute buffer (529 MiB fixed allocation), not the KV cache itself. This is a llama.cpp implementation detail that may improve in future versions.
The 27B IQ3 has a perfect score — every line exact, zero hallucinations. The 35B models are close but not quite there. Interesting that Q8_0 doesn't beat Q4_K_XL here.
Quality — GSM8K (grade school math, 100 cases)
Metric
27B IQ3
35B Q4_K_XL
35B Q8_0
Accuracy
89%
91%
90%
CI (95%, excl. truncated)
[86.9%, 97.1%]
[84.9%, 95.8%]
[85.8%, 96.5%]
Truncated
5
1
3
Wall time
106 min
67 min
114 min
All three overlap in confidence intervals — the quality difference is negligible. But the 35B Q4_K_XL is 37% faster to evaluate (67 vs 106 min) with fewer truncations.
Note: AIME2025 was also tested on the 27B — 50% overall but100% on non-truncated cases*. Every failure was context exhaustion at 32k, not wrong reasoning. The 35B MoE with 131k context would likely score higher.*
Ubatch PP Trick (coder543, May 18)
u/coder543 discovered that increasing -ub from 512→8192 gives 5.5x prompt processing speedup for --n-cpu-moe partially offloaded models. I tested this on the 35B:
Result: doesn't apply with--fit on**.** The -ub 2048+ OOMs because --fit on already maximizes VRAM for model layers — no headroom for larger batch buffers. If you use --n-cpu-moe manual offload instead, the trick works. But --fit on is simpler and handles the split automatically.
Concurrency (-np sweep)
Tested -np 1/2/4 on 10 GSM8K cases:
-np
27B tok/s
27B throughput
35B tok/s
35B throughput
1
83.3
0.6 cases/min
70.7
0.8 cases/min
2
57.7
1.3 cases/min
49.7
1.1 cases/min
4
10.0 (CPU overflow)
0.6 cases/min
28
failed
-np 2doubles batch throughput at 30% slower per-request speed. -np 4 pushes layers to CPU — 27B drops to 10 tok/s, 35B partially fails. Use -np 1 for interactive chat, -np 2 for batch evaluation.
MTP Reference (for 27B / fully-on-GPU setups)
MTP is worth it when the model fits entirely on GPU (no offload penalty). For the 27B IQ3 on 12GB: 73 tok/s with MTP vs ~56 without. For the 35B on 16GB: skip it (see speed table above).
If you do use MTP:
--spec-type draft-mtp — not mtp. Mainline renamed it.
-np 1 — b9204 defaults to 4 slots which pushes layers to CPU.
No MTP. No special flags. --fit-target 1536 is the key — it reserves VRAM headroom so the KV cache doesn't OOM at 128k. Load it, leave it running, point your coding agent at localhost:8080/v1/chat/completions.
What you get: 56 tok/s generation at 128k context. 1,584 tok/s prompt processing (81s to ingest 128k tokens). 131k max context. GSM8K 91%. Stable.
Why no MTP? At 128k context both MTP and no-MTP give the same 56 tok/s — the KV cache dominates VRAM either way. MTP adds 5 gotchas for zero benefit. Skip the complexity.
Other VRAM budgets (community data, not tested by us)
Everything above was tested on our RTX 5080 16GB. These estimates for other GPUs are from community reports:
VRAM
Model
Speed
Source
8 GB
35B MoE Q2_K_XL+MTP
~50 tok/s (est.)
u/Still-Notice8155 (GTX 1070, -fit off --n-cpu-moe 32)
12 GB
35B MoE Q4_K_XL+MTP
~73-80 tok/s
u/janvitos (RTX 4070 Super 12GB)
16 GB
35B Q4_K_XL
56 tok/s @ 128k
This post (RTX 5080)
24 GB
35B Q4_K_XL (no MTP)
~90+ tok/s (est.)
Model is ~22 GB, fits fully on GPU with headroom for KV
The 27B IQ3+MTP needs the MTP head grafted — graft-mtp.py in the repo.
Why not the others?
27B IQ3 — We tested it on our 16GB card where it fits fully on GPU (12.45 GB model). Perfect CodeNeedle (220/220), 73 tok/s with MTP (GGUF). But it caps at 56k context (110k with q4_0 KV). If your coding agent needs 128k, it's out. Better suited for 12 GB cards where the 35B won't fit.
35B Q8_0 — 38% slower (46 tok/s with MTP), negligible quality gain (GSM8K 90% vs 91%, overlapping CIs). Not worth the VRAM on 16 GB.
Credits
This post exists because of the community:
am17an — original MTP implementation (PR #22673), merged mainline b9190
havenoammo — MTP GGUF variants + graft script
u/janvitos — 80 tok/s MTP config on 12GB (635 upvotes), documented the flags
u/coder543 — ubatch PP trick for --n-cpu-moe (May 18)
u/OsmanthusBloom — earlier ubatch discovery
u/Still-Notice8155 — GTX 1070 8GB MTP benchmarks proving it works everywhere
EDIT 2: u/FusionX correctly points out that --fit-target 1536 is too conservative for headless setups. My machine runs a desktop compositor + terminal that eats ~1 GB VRAM before the model loads. If you're running headless, --fit-target 128 keeps more expert layers on GPU. FusionX reports 70-80 tok/s at 131k context on the same GPU with this setting. I'll re-benchmark with a lower fit-target and update. The recommended config is adjust --fit-target down if you're headless.
EDIT 3: Hey thanks everyone for commenting, and for the ones who really skeptical of the results because the post was AI generated. u/the__storm u/Special_Animal2049kevin_1994 I really appreciate your criticisms, and I should have been more upfront about this. So to remedy this I have posted the scripts that produced these results and the raw data themselves, you can find them here: https://github.com/gaztrabisme/llm-server/tree/main/docs/dev
EDIT 4: u/OsmanthusBloom caught that the community VRAM table incorrectly listed the 27B dense model for the 8 GB and 12 GB rows. Both sources actually ran the 35B MoE with CPU offload.
주요 댓글
r/localllama
MTP가 35B 모델 성능에 미치는 영향에 대한 기술적 논의와 함께, AI로 작성된 글의 가독성 문제에 대한 피드백이 오가는 분위기입니다.
65
모델이 VRAM에 안 들어가면 속도 느려진다는 말을 하려고 4283 토큰이나 쓴 거임?
10
경제적으로는 맞는 말임. 근데 난 개인적으로 같은 시스템으로 게임도 해서 이 GPU를 고른 거임 ㅋㅋ
24
아 젠장, 내가 좀 대충 했네. 수정해줘서 고마워, 너한테 공을 돌릴게!
16
맞아, 이 서브레딧은 이제 포기하고 쓰레기 같은 글들을 받아들이기로 한 것 같네(아니면 레딧 알고리즘이 그냥 이런 걸 좋아하는 걸지도). MTP가 메모리를 잡아먹고 느린 메모리로 오프로딩하면 추론 속도가 떨어진다는 결론은 딱히 놀라운 사실도 아닌데 말이야.
5
21
이거 AI가 쓴 거임? 초당 토큰 속도 너무 낮은데. 그리고 헤드리스 모드라고 가정하면 1536MB VRAM을 예약할 필요가 전혀 없음. KV 캐시는 이미 llama의 피팅 휴리스틱에 계산되어 있음. OOM 안 날 정도로 최대한 낮게 설정해. 난 128MB로도 잘 돌아감. 나도 똑같은 사양인데 램만 훨씬 구린 DDR4임. 131k 컨텍스트 사이즈에서 같은 GPU, 모델, llama 파라미터로 70 tk/s 나옴. 사실 --fit-target=128MB로 설정하면 80 tk/s까지 찍힘. 256k 컨텍스트에서도 65 tk/s 나오고.
15
경제적으로는 맞는 말이죠. 개인적으로는 같은 시스템으로 게임도 해서 이 GPU를 선택한 거예요 ㅋㅋ
2
16
아이고, 내가 너무 대충 확인했네. 지적해줘서 고맙고, 본문에 네 이름 언급해서 크레딧 남겨놨어!