16GB VRAM의 굴레에 대한 토론 스레드
16 GB VRAM purgatory discussion thread
핵심 요약
16GB VRAM 환경에서 LLM을 구동하기 위한 최적의 설정과 모델 구성에 대한 사용자들의 노하우 공유.
- VRAM 최적화 — 윈도우 환경의 VRAM 점유를 피하기 위한 헤드리스 설정 및 파라미터 튜닝 공유함.
- 모델 추천 — Qwen 3.8 27B 모델을 중심으로 다양한 양자화 버전과 실행 환경 비교함.
- 성능 개선 — exllamav2나 exl3 등 GGUF 외의 대안을 활용한 토큰 처리 속도 향상 논의함.
- 컨텍스트 관리 — VRAM 부족을 해결하기 위한 KV 캐시 오프로딩 및 체크포인트 설정 최적화함.
어떤 모델이랑 설정 쓰고 있는지 여기 좀 공유해 봐.
윈도우 쓰는 중이면 난 이거 쓰고 있음. https://huggingface.co/Bucoid/Qwen3.8-27B-Uncensored-IQ4-XS-MTP-16GB-VRAM-GGUF MTP는 끄고, q4 k/q4 v mmproj는 CPU/RAM으로 유배 보냈음. VRAM 최대한 아끼려고 ub 좀 조절해서 컨텍스트 90k~100k 정도 확보해서 VRAM 안에 다 구겨 넣는 중이다.
리눅스 쓰거나 내장 그래픽 있으면 윈도우가 VRAM 1.5GB씩 처먹는 꼴 안 봐도 되니까 14.5GB 넘게 쓸 수 있을 거다. 너네는 지옥까진 아니겠네.
@echo off
.\\ikllama\\llama-server.exe ^
-m "D:\\AI models\\qwen3.8\\Qwen3.8-27B-Uncensored-IQ4-XS-MTP-16GB-VRAM-GGUF.gguf" ^
:: gpu offload all layers (99 is more than the max which means it will offload everything)
-ngl 99 ^
:: this depends on your cpu
-t 8 ^
:: literally can't fit in anything at higher q to save vram
--cache-type-k q4\_0 ^
--cache-type-v q4\_0 ^
:: check your max context size with fit, at around >100k context rot sets in
-c 100100 ^
:: this flag should always be on to optimise speed and memory use
-fa on ^
:: You need a fixed jinja file to prevent it rambling forever, I believe this deefaults to xhigh
--chat-template-file chat-template.jinja ^
--chat-template-kwargs "{\\"preserve\_thinking\\": true, \\"enable\_thinking\\": true}" ^
:: you ain't getting more than this
-np 1 ^
:: use the mmproj and banish it to CPU/RAM land to save vram (probably ~800-900 mb vram saving)
--mmproj mmproj-F16.gguf ^
--no-mmproj-offload ^
:: we need reasoning
--reasoning on ^
:: the image mmproj needs this line
--image-min-tokens 1024 ^
--metrics ^
--port 8080 ^
:: reduce vram spikes saving some vram
--batch-size 1024 ^
--ubatch-size 256 ^
:: allows more caching in RAM. According to Claude it's mostly for your context slot checkpoints that there is literally no room for
--cache-ram 24576 ^
--ctx-checkpoints 32 ^
:: Delta net architecture apparently has a bug where it just stalls forever saving and shifting contexts this is apparently supposed to help with this according to Cl\*ude
--no-context-shift ^
:: force mtp header into the CPU/RAM (cl\*ude estimates ~200 mb savings)
--override-tensor nextn=CPU ^
--jinja
pause

