[Deepseek-V4-Flash-0731] Full 1M context on a single RTX5090 + DDR5 Desktop Setup with VLLM CPU/Ram Offloading, ~800 tps pp & 15+ tps decode [Agentic Coding]
핵심 요약
RTX 5090과 256GB RAM을 사용하여 DeepSeek-V4-Flash 모델의 1M 컨텍스트를 효율적으로 구동하는 설정과 최적화 방법을 공유합니다.
하드웨어 구성 — RTX 5090 32GB와 256GB DDR5 RAM을 활용한 로컬 환경 구축
소프트웨어 최적화 — vLLM 포크 버전과 FlashInfer 라이브러리 패치를 통한 실행 오류 해결
성능 최적화 — DSpark 추론 단계에 따른 동적 추론 깊이 조절로 생성 속도 개선 시도
실행 설정 — MoE 레이어의 GPU/RAM 분할 배치를 통한 1M 컨텍스트 처리
Update — long-context SM120 fallback workaround validated to ~500k
I reduced it to a deterministic reproducer: a ~71.9k-token seed request succeeded, but a second request with a near-full prefix-cache hit plus a tiny suffix reliably killed the engine in the sparse-indexer MQA prefill path. Instrumentation on the failing request showed q=(221,64,128), kv=(17975,128), and only ~15.15 MiB of full FP32 logits, so this was not simply a giant-logits allocation problem.
The SM12x Triton MQA kernel changes its M tile once indexer KV crosses 16K entries. The workaround keeps long-KV top-k prefills on the existing chunked path, caps each inner KV chunk at 16K, and dynamically bounds the temporary FP32 chunk logits to 64 MiB.
Validation so far: the exact ~71.9k reproducer that previously crashed now passes, the same patched server passed ~149.9k seed + cache-hit testing, and it has now also passed a 500,084-token seed request followed by a 500,101-token near-full prefix-cache-hit request. The 500k seed completed in ~40m30s and the cached follow-up in ~3.9s.
I am treating this as a validated workaround through ~500k on this configuration, not as a blanket 1M-context guarantee.
During startup, FlashInfer's CUDA IPC helper could accidentally find TileLang's:
libcudart_stub.so
instead of the real loaded CUDA runtime.
That eventually caused:
undefined symbol: cudaDeviceReset
The problem was FlashInfer's find_loaded_library("libcudart") doing a substring search over /proc/self/maps.
I patched:
flashinfer/comm/cuda_ipc.py
so it checks the actual filename instead:
def find_loaded_library(lib_name):
with open("/proc/self/maps") as f:
for line in f:
if "/" not in line:
continue
start = line.index("/")
path = line[start:].strip()
filename = path.split("/")[-1]
if (
filename.startswith(lib_name + ".so")
or filename.startswith(lib_name + "-")
):
return path
return None
After that, FlashInfer correctly resolves the real libcudart instead of the TileLang stub.
This is a local patch and obviously needs to be reapplied if the package gets replaced.
For this workload I suspect the ideal behavior would be approximately:
reasoning/thinking: 1 speculative token
normal/final decoding: 2 speculative tokens
The second draft token often isn't worth computing while the model is doing difficult reasoning, but becomes very valuable when it transitions into more predictable code/text generation.
vLLM does not currently give me a simple runtime switch for this, so I may patch the speculative decoding path later and experiment with changing the draft depth based on whether the model is currently emitting reasoning or final output.
That looks like one of the biggest remaining decode optimizations.
---non AI comment section begins---
Stay tuned, I am working on a if/else block to fix that stupid behavior slowing down during reasoning and squeeze even more tps out of this stack.
---non AI comment section ends---
주요 댓글
r/localllama
사용자들은 vllmds4의 출처와 DSpark 지원 여부에 대해 논의하며 기술적인 세부 사항을 확인하고 있습니다.
14
공유해 줘서 고마워요. 링크해 준 포스트를 보면서 시도해 볼 가치가 있을지 고민 중이었거든요. 지금까지 본 바로는 디코드 속도는 (DSpark 없는) llama.cpp와 비슷하고, 프리필 속도에서 눈에 띄는 향상이 있는 것 같은데 동의하시나요?
다른 작성자는 4소켓 시스템을 사용해서 이 엔진이 NUMA를 훨씬 잘 처리하기 때문에 성능 향상을 얻은 것 같은데, 단일 소켓 시스템에서도 유의미할까요?
혹시...
5
네, 거의 맞아요. llamacpp에서의 프리필은 정말 골칫거리였고 에이전트 코딩 작업 속도를 너무 많이 떨어뜨렸거든요.
예전에는 llamacpp를 돌렸었는데(당신과 비슷한 tps, pp 약 100, 디코드 약 15 tps 정도 나왔음), u-ub 값을 8192로 올려보려고 했었죠(다른 포스트에서 u-ub를 8192로 올리면 pp 성능이 개선된다고 했거든요). 그러다 당신이 눈여겨보던 그 포스트를 보고 시도해 본 겁니다.
당분간은 이 설정을 유지할 생각입니다. 그 엄청난 프리필 성능 덕분에 DeepSeek가 에이전트 작업에 쓸만해졌거든요...