DiffusionGemma를 비난하는 대신 해킹해보는 건 어떨까?
Can we stop dunking on DiffusionGemma and hack it instead?
핵심 요약
DiffusionGemma의 환각 문제를 해결하기 위한 추론 최적화 방법과 워크플로우 개선안을 제안하는 포스트입니다.
- 추론 최적화 — 엔트로피 기반 샘플러와 작업별 튜닝을 통해 환각을 줄이고 속도를 개선함
- 구조화된 출력 — JSON 스키마 스캐폴딩을 활용해 모델의 출력 정확도와 충실도를 높임
- 사고 모드 활용 — 추론 및 도구 호출 시 'Thinking Mode'를 활성화하여 성능을 극대화함
- 워크플로우 개선 — 추론 엔진과 컨텍스트 관리 방식을 재설계하여 dLLM의 한계를 극복함
DiffusionGemma 나온 지 이제 겨우 일주일 됐는데, 다들 "순진하게" 추론 돌리니까 환각 현상 너무 심하다고 난리네. 벌써 이 문제 해결하려고 나온 논문들이 꽤 있어서, AI한테 dLLM이 완전히 맛탱이 가지 않게 만들 방법들 표로 정리 좀 해보라고 시켜봤어 (Mercury가 이미 비슷한 짓을 하긴 했는데, 그건 폐쇄형 모델 쪽이라). 그러니까 AI가 뽑아준 결과물 보고 llama.cpp나 vLLM 같은 애들이 추론 속도 3배 펌핑하는 데 부족해 보이면 나를 아주 그냥 갈궈줘.
범례: ⚙️ = Drop-in (지금 당장 프롬프트/설정으로 적용 가능) | 🛠️ = Wrapper (오케스트레이션/검증/검색) | 🔧 = Decoder (최대 성능 향상을 위한 커스텀 샘플러/런타임).
| # | Method | Type | Concise Action | Expected Benefit (vs Naive 256-Token Rendering) | Citation Cluster |
|---|---|---|---|---|---|
| Tier 0: Foundational Official Settings (Must-Use Baseline – Fixes ~80% of Complaints) | |||||
| 1 | Entropy-Bounded Sampler + Adaptive Stopping | ⚙️ Drop-in | Commit lowest-entropy tokens until accumulated entropy exceeds bound (0.1); stop when argmax stable (2+ steps) and mean entropy < 0.005 | Prevents premature termination/over-refinement hallucinations; dynamic steps by task complexity; 2–3× effective speedup; core path to match Qwen-level quality | Google model card & HF config (2026); Ben-Hamu et al. (EB-Sampler, NeurIPS 2025, arXiv:2505.24857) |
| 2 | Canvas Cap + Task-Tuned Entropy | ⚙️ Drop-in | Keep 256-token canvas but set max_new_tokens short for tool calls (64–128); lower bound (0.03–0.05) for tools/deterministic, higher (0.15–0.2) for factual/reasoning | Reduces noise/waste on short structured outputs; deterministic tool selection; preserves candidate diversity to cut premature hallucination and improve reasoning | Google serving examples (2026); EB-Sampler family + hallucination-mode papers (2026) |
| 3 | Thinking Mode + Clean History | ⚙️ Drop-in | Add enable_thinking=True for reasoning/tool selection; retain only final (non-thinking) response in multi-turn history |


