Bonsai-27B 및 Ternary-Bonsai-27B - 업데이트 (PR 관련)
Bonsai-27B & Ternary-Bonsai-27B - Updates (on PRs)
핵심 요약
Bonsai-27B 모델의 1비트 및 3진법(Ternary) 양자화 모델의 최신 llama.cpp 지원 현황과 성능 한계를 공유합니다.
- 모델 업데이트 — 1비트 및 3진법 Bonsai 모델의 llama.cpp 업스트림 통합 현황 공유됨
- 성능 한계 — 에이전트 코딩 작업에는 아직 최적화되지 않았으며, 3진법 모델은 모바일 기기 메모리 제한을 초과함
- 기술적 진전 — CPU, Metal, Vulkan 등 다양한 백엔드에서 지원이 확대되고 있으며 최적화 PR이 진행 중임
- 사용자 피드백 — 일반 사무용 PC에서도 구동 가능한 점은 고무적이나, 3진법 모델의 CUDA 지원은 아직 포크 버전에서만 원활함
아래 업스트림 상태 섹션은 https://github.com/PrismML-Eng/Bonsai-demo 에서 가져왔음.
Binary(이진) 업스트림 상태
Q1_0은 업스트림 llama.cpp의 다양한 백엔드(CPU(일반, NEON, 최적화된 x86), Metal, CUDA, Vulkan)에서 기본적으로 지원됨.
| Runtime | Status |
|---|---|
| llama.cpp (CPU, Metal, CUDA, Vulkan) | ✅ Merged upstream, works out of the box |
| MLX (1-bit) | ⏳ Pending upstream: mlx#3161; until it merges, use PrismML-Eng/mlx (branch prism, built automatically by setup.sh) |
Ternary(삼진) 업스트림 상태
Ternary 지원은 메인라인 llama.cpp로 넘어가는 중임. 백엔드들이 하나씩 들어오고 있어서, 현재는 메인라인이랑 우리 포크 버전이 섞여 있는 상태임. 일단 알아둬야 할 실질적인 내용은 이거임: 현재 세 가지 ternary GGUF 버전을 배포 중인데, 각각 맞는 환경에서 돌려야 함.
| File | Format | Runs on |
|---|---|---|
*-Q2_0.gguf | Group size 128. The format this demo uses, compatible with our fork. Once the llama.cpp migration completes, these files will be deprecated and replaced by the PQ2_0 ggufs | This demo / the fork binaries. Will not load on mainline (same type id, different block size) |
*-Q2_0_g64.gguf | Group size 64 (2.25 bpw). The official llama.cpp format; these will be renamed to plain Q2_0, replacing the current ones | Mainline llama.cpp (CPU and Metal so far) |

