RTX 4090 + DDR5 환경에서 DeepSeek v4 Flash 구동 경험
DeepSeek v4 Flash on 4090 + DDR5, my experience
핵심 요약
RTX 4090과 DDR5 환경에서 DeepSeek v4 Flash를 구동하며 얻은 성능 최적화 경험과 Qwen 3.6 27B와의 비교 분석.
- 성능 최적화 — CPU 코어 고정 및 플래시 어텐션 설정으로 추론 속도 향상함
- 모델 비교 — DeepSeek v4가 추론 능력은 좋으나 에이전트 작업에는 Qwen 3.6 27B가 더 효율적임
- 하드웨어 제약 — 24GB VRAM 환경에서 컨텍스트 크기와 배치 사이즈 조절이 필수적임
- 기술적 팁 — KV 오프로드 설정 및 배치 사이즈 최적화를 통해 성능 개선 가능함
Disclosure: No AI was used to write this
My specs are:
- RTX 4090
- 128 GB DDR5 5600 MT/s
- Intel Core Ultra 7 270k
Running nvidia-595 on ubuntu 26.04 with latest llama.cpp build (pulled and rebuilt this morning).
Tried a lot of things, ended up running unsloth's UD-Q2_K_XL quant with command:
taskset -c 0-7 /home/kevin/ai/llama.cpp/build/bin/llama-server -lv 4 -m /home/kevin/ai/models/DeepSeek-V4-Flash-UD-Q2_K_XL-00001-of-00003.gguf --temp 1.0 --top-p 1.0 --min-p 0.0 -t 8 -fitc 64000 -fa off -np 1
Speed: [ Prompt: 132.5 t/s | Generation: 10.9 t/s ]
Some notes:
- On Intel Core Ultra 7 270k (I recently bought this CPU), pinning pcores makes a big difference. Like 2x, from 6.8 tok/s to 11 tok/s
--no-mmapis much slower- using
-ctk q8_0or-ctv q8_0crashed the llama.cpp process - adjusting
-bor-ubto > 4096 with context > 32k seems to explode the CUDA buffer to 90 GB+ - with llama-server, for some reason,
-fa offis necessary, otherwise it also explodes the CUDA buffer
Overall, seems smarter compared to Qwen 3.6 27B Q4_K_XL. It runs slower, but reasons less, meaning tasks still complete in a reasonable amount of time. However, for agentic, Qwen 3.6 27B is still much better because it runs so much faster, and also qwen models don't seem to "over-reason" too much when doing agentic tasks.
With a couple fixes (flash attention, microbatch/batch adjust, context quantisation, etc.) I think this model could be pretty decent on 4090/3090. If we could get these numbers up to like ~20 tg/s and ~300 pp/s it might replace qwen 3.6 27b for me.
Also tried running IQ4_NL quant but ended up being too slow and couldn't fit enough context (only about 10k):
taskset -c 0-7 /home/kevin/ai/llama.cpp/build/bin/llama-cli -m /home/kevin/ai/models/DeepSeek-V4-Flash-UD-IQ4_NL-00001-of-00004.gguf --temp 1.0 --top-p 1.0 --min-p 0.0 -t 8
It was at speed: [ Prompt: 50.7 t/s | Generation: 8.1 t/s ]

