b9200 업데이트 벤치마크: RTX 3090에서 Qwen 3.6 27B mtp 및 Hermes Agent 최적화
Benchmarking the new b9200 update: Optimizing Qwen 3.6 27B mtp for Hermes Agent on a single RTX 3090
핵심 요약
RTX 3090에서 Qwen 3.6 27B 모델의 MTP 설정을 최적화하고 b9200 업데이트를 적용해 에이전트 성능을 대폭 개선함.
- MTP 설정 최적화 — 룩어헤드와 병렬 슬롯을 조정해 에이전트 워크플로우의 병목 현상을 해결함.
- b9200 업데이트 효과 — 메모리 트래픽 오버헤드를 줄여 프롬프트 처리와 토큰 생성 속도를 크게 높임.
- 에이전트 성능 향상 — 27B 모델에서 초당 13~27 토큰의 실사용 가능한 속도를 확보함.
- 하드웨어 효율성 — VRAM 버스 부하를 최소화하여 언더볼팅된 3090에서도 안정적인 성능을 유지함.
UPDATED (POST b9200)
Okay, here is the updated version using the new Qwen 3.6 27B mtp gguf from Unsloth, running it as the backend for the hermes agent. While dialing it in, I noticed that the currently recommended Unsloth mtp flags actually bottleneck performance and tank draft acceptance rates for strict, multi-turn agentic workflows. Pairing a custom config with today's brand new llama.cpp b9200 release — which specifically fixes mtp memory traffic overhead — completely turns that around.
Hardware/Software
* RTX 3090 (24GB VRAM) — currently undervolted to keep temps down
* Ryzen 7 5700G / 64GB
* Qwen3.6-27B-IQ4_NL.gguf
* llama-server (b9200+ compiled from source, commit #23234)
* hermes agent (64K context) max to limit spillover
The problem with default mtp settings
Running the standard recommended mtp flags (--spec-draft-n-max 6 and --spec-draft-p-min 0.75) gave poor results for agentic loops. Generation speeds sat around 7–8 t/s, and the mtp draft acceptance rate hovered around 22–26%.
Agent workflows are rigid. A 6-token lookahead frequently guesses the wrong punctuation, the main model rejects the draft, and the GPU throws out the math and recalculates — completely negating the mtp speed boost. Without explicitly declaring parallel slots, llama-server also defaults to 4, eating up memory bandwidth managing unused context slots.
The fix and the b9200 boost
For agent workflows on a 24GB card, limit to a single slot, drop the lookahead to 3, and remove the p-min threshold so it doesn't hesitate on rigid syntax. Combined with the b9200 release — which stops copying the full logits for every token in the batch during prompt processing — the optimized launch command looks like this:
.\build\bin\Release\llama-server.exe ^
-m D:\models\Qwen3.6-27B-IQ4_NL.gguf ^
--spec-type draft-mtp ^
--spec-draft-n-max 3 ^
--ctx-size 65536 ^
--parallel 1 ^
--flash-attn on ^
--cache-type-k q8_0 ^
--cache-type-v q8_0 ^
--port 8081
Results (Prior to the update vs. Post-b9200)


