Ternary Bonsai 2 (27B) 허깅페이스 출시. 6GB 미만 용량으로 WebGPU에서 로컬 구동 가능
Ternary Bonsai 2 (27B) just released on Hugging Face. At <6GB in size, it can even run locally in-browser on WebGPU.
핵심 요약
Qwen 기반의 27B 모델을 6GB 미만으로 압축한 Ternary Bonsai 2가 출시되어 성능 유지 여부를 두고 논의가 활발합니다.
- 모델 경량화 — 27B 모델을 3진법 가중치로 압축하여 6GB 미만으로 구현함
- 성능 논란 — FP16 대비 98.2% 성능 유지라는 수치에 대해 사용자들의 의구심이 제기됨
- WebGPU 지원 — 브라우저 환경에서 로컬로 구동 가능한 효율성을 확보함
- 버전 업데이트 — 기존 Qwen 3.6 기반 모델에서 3.8로 업그레이드됨

The model is derived from Qwen3.8-27B, a 27B hybrid-attention causal language model (architecture unchanged), but uses ternary weights to shrink model size down to <6GB in size. According to the model card, it's 9x smaller than FP16 while retaining 98.2% of the intelligence.
- Collection: https://huggingface.co/collections/prism-ml/bonsai-2
- Demo: https://huggingface.co/spaces/webml-community/ternary-bonsai-2-webgpu-kernels

