100달러로 맞추는 20GB VRAM, 448GB/s 대역폭의 초저가 가성비 세팅
Ultra budget 20GB vram with 448GB/s for $100 bucks.
핵심 요약
100달러짜리 중고 P102-100 그래픽카드 2개로 20GB VRAM을 확보해 LLM 추론을 효율적으로 돌리는 가성비 빌드입니다.
- 가성비 빌드 — 100달러로 20GB VRAM과 448GB/s 대역폭 확보함
- 성능 검증 — Qwen 3.6 35B 모델로 70tk/s 속도와 32K 컨텍스트 지원함
- 전력 효율 — 카드당 150W 제한으로 성능 손실 최소화 및 효율 최적화함
- 지속 가능성 — llama.cpp의 폭넓은 지원으로 구형 파스칼 아키텍처도 장기간 사용 가능함
100달러짜리 그래픽 카드로 뽑아낼 수 있는 성능의 한계는 이 정도다.
VRAM은 쥐꼬리만큼 주면서 가격은 4배 넘게 처받는 카드들 여러 개 쓰는 것보다, 훨씬 쾌적한 속도에 컨텍스트도 넉넉하게 챙기면서 동시 사용자 3명까지 돌릴 수 있다.
0.00.008.388 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
0.00.008.391 I device_info:
0.00.089.439 I - CUDA0 : NVIDIA P102-100 (10144 MiB, 10013 MiB free)
0.00.197.645 I - CUDA1 : NVIDIA P102-100 (10144 MiB, 10013 MiB free)
0.00.197.656 I - CPU : Intel(R) Xeon(R) W-2135 CPU @ 3.70GHz (128396 MiB, 128396 MiB free)
0.00.197.728 I system_info: n_threads = 6 (n_threads_batch = 6) / 12 | CUDA : ARCHS = 600,610,750,860,890 | USE_GRAPHS = 1 | PEER_MAX_BATCH_SIZE = 128 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
0.00.197.764 I srv init: running without SSL
0.00.197.849 I srv init: using 11 threads for HTTP server
0.00.198.515 I srv start: binding port with default address family
0.00.199.823 I srv llama_server: loading model
0.00.199.902 I srv load_model: loading model '/models/Qwen3.6-35B-A3B-UD-IQ4_XS.gguf'
0.00.199.906 I common_init_result: fitting params to device memory ...
0.00.199.907 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
0.00.987.288 W common_fit_params: failed to fit params to free device memory: n_gpu_layers already set by user to 99, abort
0.23.223.625 W llama_context: n_ctx_seq (32768) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
0.23.481.073 I common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
0.23.570.914 I srv load_model: initializing slots, n_slots = 3
0.23.598.842 W srv load_model: speculative decoding will use checkpoints
0.23.598.851 W common_speculative_init: no implementations specified for speculative decoding
0.23.598.852 I slot load_model: id 0 | task -1 | new slot, n_ctx = 32768
0.23.598.854 I slot load_model: id 1 | task -1 | new slot, n_ctx = 32768
0.23.598.854 I slot load_model: id 2 | task -1 | new slot, n_ctx = 32768
0.23.598.961 I srv load_model: prompt cache is enabled, size limit: 8192 MiB
0.23.598.963 I srv load_model: use `--cache-ram 0` to disable the prompt cache
0.23.598.964 I srv load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
0.23.598.965 I srv load_model: context checkpoints enabled, max = 32, min spacing = 8192
0.23.598.985 I srv init: idle slots will be saved to prompt cache upon starting a new task
0.23.628.848 I init: chat template, example_format: '<|im_start|>system
You are a helpful assistant<|im_end|>
<|im_start|>user
Hello<|im_end|>
<|im_start|>assistant
Hi there<|im_end|>
<|im_start|>user
How are you?<|im_end|>
<|im_start|>assistant
<think>
</think>
'
0.23.666.546 I srv init: init: chat template, thinking = 0
0.23.666.572 I srv llama_server: model loaded
0.23.666.575 I srv llama_server: server is listening on http://127.0.0.1:5802
0.23.666.579 I srv update_slots: all slots are idle
0.48.181.695 I srv operator(): Chat format: peg-native
0.48.182.094 I slot get_availabl: id 2 | task -1 | selected slot by LRU, t_last = -1
0.48.182.101 I srv get_availabl: updating prompt cache
0.48.182.111 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000
0.48.182.123 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 98304 tokens, 8589934592 est)
0.48.182.128 I srv get_availabl: prompt cache update took 0.02 ms
0.48.182.302 I slot launch_slot_: id 2 | task 0 | processing task, is_child = 0
0.48.182.309 I slot process_sing: id 0 | task -1 | saving idle slot to prompt cache
0.48.182.311 I slot process_sing: id 1 | task -1 | saving idle slot to prompt cache
0.48.186.009 I srv operator(): Chat format: peg-native
0.48.189.081 I srv operator(): Chat format: peg-native
0.49.483.103 I slot get_availabl: id 1 | task -1 | selected slot by LRU, t_last = -1
0.49.483.111 I srv get_availabl: updating prompt cache
0.49.483.116 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000
0.49.483.119 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 98304 tokens, 8589934592 est)
0.49.483.120 I srv get_availabl: prompt cache update took 0.01 ms
0.49.483.178 I slot launch_slot_: id 1 | task 2 | processing task, is_child = 0
0.49.483.179 I slot process_sing: id 0 | task -1 | saving idle slot to prompt cache
0.49.483.181 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = -1
0.49.483.181 I srv get_availabl: updating prompt cache
0.49.483.182 I srv load: - looking for better prompt, base f_keep = -1.000, sim = 0.000
0.49.483.183 I srv update: - cache state: 0 prompts, 0.000 MiB (limits: 8192.000 MiB, 98304 tokens, 8589934592 est)
0.49.483.183 I srv get_availabl: prompt cache update took 0.00 ms
0.49.483.215 I slot launch_slot_: id 0 | task 3 | processing task, is_child = 0
0.51.242.275 I slot create_check: id 0 | task 3 | created context checkpoint 1 of 32 (pos_min = 1376, pos_max = 1376, n_tokens = 1377, size = 62.813 MiB)
0.51.367.765 I slot create_check: id 1 | task 2 | created context checkpoint 1 of 32 (pos_min = 670, pos_max = 670, n_tokens = 671, size = 62.813 MiB)
0.51.367.773 I slot print_timing: id 2 | task 0 | prompt processing, n_tokens = 1377, progress = 1.00, t = 3.19 s / 432.28 tokens per second
0.51.480.037 I slot create_check: id 2 | task 0 | created context checkpoint 1 of 32 (pos_min = 1376, pos_max = 1376, n_tokens = 1377, size = 62.813 MiB)
0.56.647.801 I slot print_timing: id 0 | task 3 | n_decoded = 100, tg = 23.30 t/s, tg_3s = 23.30 t/s
0.56.653.219 I slot print_timing: id 2 | task 0 | n_decoded = 100, tg = 23.30 t/s, tg_3s = 23.30 t/s
0.56.692.218 I slot print_timing: id 1 | task 2 | n_decoded = 100, tg = 23.54 t/s, tg_3s = 23.54 t/s
0.59.655.679 I slot print_timing: id 0 | task 3 | n_decoded = 171, tg = 23.43 t/s, tg_3s = 23.60 t/s
0.59.661.606 I slot print_timing: id 2 | task 0 | n_decoded = 171, tg = 23.42 t/s, tg_3s = 23.60 t/s
0.59.702.608 I slot print_timing: id 1 | task 2 | n_decoded = 171, tg = 23.56 t/s, tg_3s = 23.58 t/s
1.02.659.591 I slot print_timing: id 0 | task 3 | n_decoded = 242, tg = 23.49 t/s, tg_3s = 23.64 t/s
1.02.665.066 I slot print_timing: id 2 | task 0 | n_decoded = 242, tg = 23.49 t/s, tg_3s = 23.64 t/s
1.02.705.486 I slot print_timing: id 1 | task 2 | n_decoded = 242, tg = 23.58 t/s, tg_3s = 23.64 t/s
1.03.253.784 I slot print_timing: id 0 | task 3 | prompt eval time = 2873.48 ms / 1381 tokens ( 2.08 ms per token, 480.60 tokens per second)
1.03.253.789 I slot print_timing: id 0 | task 3 | eval time = 10897.06 ms / 256 tokens ( 42.57 ms per token, 23.49 tokens per second)
1.03.253.791 I slot print_timing: id 0 | task 3 | total time = 13770.54 ms / 1637 tokens
1.03.253.792 I slot print_timing: id 0 | task 3 | graphs reused = 253
1.03.253.924 I slot release: id 0 | task 3 | stop processing: n_tokens = 1636, truncated = 0
1.03.259.600 I slot print_timing: id 2 | task 0 | prompt eval time = 4178.32 ms / 1381 tokens ( 3.03 ms per token, 330.52 tokens per second)
1.03.259.605 I slot print_timing: id 2 | task 0 | eval time = 10898.93 ms / 256 tokens ( 42.57 ms per token, 23.49 tokens per second)
1.03.259.606 I slot print_timing: id 2 | task 0 | total time = 15077.26 ms / 1637 tokens
1.03.259.607 I slot print_timing: id 2 | task 0 | graphs reused = 253
1.03.259.741 I slot release: id 2 | task 0 | stop processing: n_tokens = 1636, truncated = 0
1.03.288.482 I slot print_timing: id 1 | task 2 | prompt eval time = 2960.66 ms / 1381 tokens ( 2.14 ms per token, 466.45 tokens per second)
1.03.288.486 I slot print_timing: id 1 | task 2 | eval time = 10844.49 ms / 256 tokens ( 42.36 ms per token, 23.61 tokens per second)
1.03.288.487 I slot print_timing: id 1 | task 2 | total time = 13805.15 ms / 1637 tokens
1.03.288.488 I slot print_timing: id 1 | task 2 | graphs reused = 253
1.03.288.614 I slot release: id 1 | task 2 | stop processing: n_tokens = 1636, truncated = 0
1.03.288.625 I srv update_slots: all slots are idle

