mudler/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-APEX-MTP-GGUF 출시!
mudler/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-APEX-MTP-GGUF just released !
핵심 요약
MTP 헤드가 포함된 Qwen3.6-35B-A3B 모델의 APEX 양자화 버전이 공개되었습니다.
- APEX 양자화 — MoE 모델을 위한 혼합 정밀도 최적화 기법 적용
- MTP 헤드 통합 — 별도 모델 없이 자체 추론 가속(self-speculative decoding) 지원
- 하드웨어 효율성 — 소비자용 하드웨어에서도 빠른 추론 속도 보장
- imatrix 미사용 — MTP 헤드에 정적 Q8_0 양자화를 적용하여 정확도 유지
모듈 설명:
개인 연구 차원에서 <strong>30개 이상의 무료 APEX MoE 양자화</strong> 모델을 호스팅하고 있다. 내 로컬 장비는 <strong>NVIDIA DGX Spark</strong>(통합 메모리 122GB)가 전부인데, 30~50B급 MoE 모델까지는 어떻게든 돌리지만 <strong>200B가 넘어가는 큰 모델들은 H100/H200/Blackwell 같은 클라우드 컴퓨팅을 빌려야 해서</strong> 양자화 한 번 할 때마다 보통 20~100달러씩 깨진다.
APEX 양자화 모델이 쓸만하다고 생각되면, 후원을 통해 더 큰 모델들을 돌릴 수 있게 힘을 보태달라.
Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled — APEX-MTP GGUF
<a href="https://huggingface.co/lordx64/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled">lordx64/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled</a> 모델을 <strong>APEX(Adaptive Precision for EXpert Models)</strong> 방식으로 양자화했고, 자체 추측 디코딩(self-speculative decoding)을 바로 쓸 수 있게 <strong>MTP(multi-token prediction) 헤드</strong>를 포함시켰다.
일반 APEX 저장소랑 뭐가 다르냐고?
이 GGUF 파일들은 <a href="https://github.com/ggml-org/llama.cpp/pull/22673">llama.cpp PR #22673</a> 덕분에 모델의 <strong>MTP 헤드</strong>를 본체와 함께 하나의 파일로 묶어놨다. 최신 llama.cpp(커밋 255582687 이상)를 쓰면 별도의 드래프트 모델 없이 이 파일 하나만으로 바로 자체 추측 디코딩을 돌릴 수 있다:
llama-server -m Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-APEX-MTP-I-Balanced.gguf --draft-mtp
MTP가 없는 버전도 <a href="https://huggingface.co/mudler/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-APEX-GGUF">mudler/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-APEX-GGUF</a>에서 받을 수 있다. 용량은 조금 더 작지만 자체 추측 디코딩은 안 된다.
파일 크기
각 양자화 파일은 MTP가 없는 버전보다 약 2.5% 정도 더 크다(트랜스포머 블록 가중치가 하나 더 들어갔기 때문. MTP가 본체의 embed_tokens를 공유하므로 임베딩 중복은 없다).
MTP 드래프트 헤드 정밀도
번들로 포함된 MTP 헤드(<code>blk.40.</code>, <code>nextn.</code> 프로젝션 + 노름 포함)는 <strong>I-Nano 티어를 제외한 모든 티어</strong>에서 <strong>Q8_0</strong>(거의 무손실)로 양자화했다. I-Nano는 MTP 블록에 본체 티어와 동일한 정밀도를 유지하되(Q3_K 라우팅 전문가, Q4_K 어텐션), <code>blk.40.nextn.eh_proj</code>는 Q4_K로 고정했다. <a href="https://huggingface.co/mudler/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-APEX-MTP-GGUF#why-the-mtp-head-doesnt-use-imatrix">아래 설명</a>을 참고해라.
이렇게 하면 본체 티어 정밀도 대비 파일당 약 1GB 정도의 추가 비용으로 드래프트 정확도(추측 디코딩 수락률에 중요함)를 높게 유지할 수 있다.
MTP 헤드는 왜 imatrix를 안 쓰냐고?
<code>llama-imatrix</code>는 본체(<code>blk.0..blk.39</code>)만 활성화하는 일반적인 포워드 패스만 돌린다. MTP 헤드는 <code>--draft-mtp</code> 추측 디코딩 중에만 작동하기 때문에 imatrix 활성화 데이터가 남지 않는다. 그래서 imatrix가 필요 없는 정적 K-quant / Q8_0 방식으로 MTP 헤드를 양자화해서 해결했다.


