llama.cpp DeepSeek v4 Flash 실험적 추론 지원
llama.cpp DeepSeek v4 Flash experimental inference
핵심 요약
Redis 개발자 antirez가 llama.cpp에서 DeepSeek v4 모델을 2비트 양자화로 구동하는 실험적 지원을 공개함.
- DeepSeek v4 지원 — llama.cpp를 통해 DeepSeek v4 모델을 실험적으로 구동할 수 있는 환경을 제공함.
- 2비트 양자화 — 모델의 routed expert를 2비트로 양자화하여 128GB RAM 환경에서 효율적인 추론을 가능하게 함.
- 성능 최적화 — 맥북 M3 Max 환경에서 최적화 후 초당 21토큰의 추론 속도를 달성함.
- 버그 수정 — 초기 배포 시 발생했던 CMake 오류와 긴 문맥 처리 관련 버그를 수정함.
Hi, here you can find experimental llama.cpp support for DeepSeek v4, and here there is the GGUF you can use to run the inference with "just" (lol) 128GB of RAM. The model, even quantized at 2 bit, looks very solid in my limited testing, and the speed of 17 t/s in my MacBook M3 Max is quite interesting, I would say we are into the usable zone.
What I did was to heavily quantize the routed experts to 2 bits using two different 2 bit quants to balance error and size. All the rest of the model, including the shared expert for each layer, is Q8: it is not worth it to play with the most sensible parts of the model if the bulk of the weights are in the routed experts.
I have the feeling that even 2 bit quantized this will prove to be a stronger model than Qwen 3.6 27B, but this is only a feeling based on the quality of the replies I get chatting with it. There is to experiment more and run benchmarks.
EDIT sorry for the CMake error, I produced the GGUF using a tool that I decided not to ship (not ready for prime time..., mostly a hack) instead of using the standard quantizer of llama.cpp. Now the problem is fixed. Also the inference in Metal is now 21 token/sec after some optimization.
EDIT2 also fixed the long context bug.


