Unsloth Qwen3.6-35B-A3B UD XL 모델에 MTP를 이식한 결과 공유
Uploaded Unsloth Qwen3.6-35B-A3B UD XL models with MTP grafted, here are the results
핵심 요약
MTP를 적용한 Qwen3.6-35B-A3B 모델을 테스트했으나, MoE 구조의 특성상 속도 향상이 기대보다 낮게 나타남.
- MTP 적용 결과 — Q4/Q8 양자화 모델에서 속도 향상이 2.5~6% 수준으로 낮게 나타남.
- 하드웨어 영향 — 사용자 환경에 따라 속도 향상 폭이 크게 달라지는 양상을 보임.
- MoE 구조 한계 — MTP가 대역폭 병목은 해결하지만 MoE 모델의 연산 효율에는 제한적임.
- 커뮤니티 테스트 — 일부 사용자는 특정 환경에서 유의미한 속도 개선을 보고함.
Following my previous post https://www.reddit.com/r/LocalLLaMA/comments/1t5ageq, a few people asked for the 35B A3B version.
The model is up on HuggingFace at https://huggingface.co/havenoammo/Qwen3.6-35B-A3B-MTP-GGUF if anyone wants to check it out. It includes the isolated MTP layers and convert.py as well.
The results are not great though. Q4 only got a 6% speed increase and Q8 only 2.5%. On the 27B it was a 2-2.5x gain, so this could be related to the MTP implementation of llama.cpp and the qwen35moe architecture or just a limitation of the model. Results are preliminary and might change in future. Either way, wanted to report back for anyone who was wondering.
Edit: u/AdamDhahabi reported:
2x 5070 Ti + 3090: Q8 went from 110 t/s to 165 t/s.
27B dense model runs at 2-2.5x speed.
So the gain might depend on your setup. Worth giving it a try!
Here is my own tests:
Tested with the prompt hello can you tell me a story on Q4.
Hardware: 5090 FE
Without MTP: 215 t/s
prompt eval time = 24.12 ms / 17 tokens ( 1.42 ms per token, 704.84 tokens per second)
eval time = 6872.43 ms / 1478 tokens ( 4.65 ms per token, 215.06 tokens per second)
total time = 6896.55 ms / 1495 tokens
With MTP: 228.83 t/s
prompt eval time = 30.08 ms / 17 tokens ( 1.77 ms per token, 565.10 tokens per second)
eval time = 8552.05 ms / 1957 tokens ( 4.37 ms per token, 228.83 tokens per second)
total time = 8582.13 ms / 1974 tokens
draft acceptance rate = 0.61434 ( 1268 accepted / 2064 generated)
Same prompt on Q8.
Hardware: 5090 FE + 3090
Without MTP: 148.20 t/s
prompt eval time = 25.80 ms / 17 tokens ( 1.52 ms per token, 658.97 tokens per second)
eval time = 11525.23 ms / 1708 tokens ( 6.75 ms per token, 148.20 tokens per second)
total time = 11551.03 ms / 1725 tokens

