PSA: DO NOT use Intel consumer platforms for multi-GPU setups
핵심 요약
인텔 소비자용 플랫폼은 PCIe P2P 제한으로 인해 멀티 GPU AI 작업에 부적합하며 성능 저하를 유발합니다.
PCIe P2P 제한 — 인텔 소비자용 CPU는 GPU 간 직접 통신을 방해하는 하드웨어/펌웨어 제한이 있음
성능 저하 — P2P가 차단되면 대역폭이 반토막 나고 추론 성능이 크게 떨어짐
드라이버 문제 — 패치된 커널 드라이버를 사용해도 VLLM 등에서 오류가 발생할 수 있음
대안 플랫폼 — AMD AM5나 EPYC 플랫폼이 멀티 GPU 구성 및 PCIe P2P 지원에 훨씬 유리함
요즘 멀티 GPU 머신 직접 맞추려는 사람들 많길래, 다들 흔히 하는 실수 하나 좀 짚어주려고 글 쓴다. 바로 Z890 같은 인텔 소비자용 플랫폼을 멀티 GPU 구성에 쓰는 짓거리 말이야.
CPU가 PCIe 5.0 레인을 24개나 제공하고, 하이엔드 보드에선 PCIe x16 슬롯 두 개를 8x8x로 쪼개 쓸 수 있다고는 하지만, GPU 간 P2P 통신이 필수인 AI 추론이나 학습 작업에선 이거 완전 무용지물이다.
테스트한다고 내 오버클럭 테스트용 시스템인 Asus Z890 Apex 보드에 Intel Core Ultra 7 270K Plus 꽂고 최신 BIOS 3202 버전 올려서 돌려봤거든. 원래는 Epyc 기반 서버보다 싱글 코어 성능 좋은 CPU 쓰면 GPU 추론 성능이 좀 나아질까 싶어서 비교해 보려고 했던 거임. 요즘 물가도 미쳐 돌아가니까 내가 가진 GPU들 처리량이라도 최대한 뽑아내 보려고 발악 중이거든.
근데 인텔 데스크탑 CPU의 빠른 싱글 코어 성능을 챙기려면 PCIe 스위치 보드를 따로 달아야 할 것 같더라. 애초에 인텔 플랫폼은 메인 PCIe 슬롯에서 8x4x4x 분할밖에 안 되게 인위적으로 막혀 있기도 하고.
테스트해보니 Arrow Lake CPU의 PCIe 루트 컴플렉스 하에서 PCIe P2P가 제대로 작동하지 않게 만드는 하드웨어/펌웨어상의 제약이 있는 것 같아.
이게 GPU가 REBAR를 지원 안 해서 생기는 문제는 아님. lspci -v 쳐보면 GPU가 BAR 사이즈 64G로 아주 잘 잡히거든. 이론상으론 PCIe P2P 돌아가는 데 아무 문제 없어야 함. BIOS에서 REBAR도 켜놨고, grub 설정이랑 IOMMU도 다 꺼놨는데도 이 모양임:
GRUB_CMDLINE_LINUX_DEFAULT="quiet splash pcie_aspm=off intel_iommu=on iommu=pt"GRUB_CMDLINE_LINUX_DEFAULT="quiet splash pcie_aspm=off intel_iommu=on iommu=pt"
02:00.0 VGA compatible controller: NVIDIA Corporation GA102GL [RTX A6000] (rev a1) (prog-if 00 [VGA controller])
Subsystem: NVIDIA Corporation GA102GL [RTX A6000]
Flags: bus master, fast devsel, latency 0, IRQ 219
Memory at 8f000000 (32-bit, non-prefetchable) [size=16M]
Memory at c000000000 (64-bit, prefetchable) [size=64G]
Memory at d000000000 (64-bit, prefetchable) [size=32M]
I/O ports at a000 [size=128]
Expansion ROM at 90000000 [virtual] [disabled] [size=512K]
Capabilities: <access denied>
Kernel driver in use: nvidia
Kernel modules: nvidiafb, nouveau, nvidia_drm, nvidia
02:00.1 Audio device: NVIDIA Corporation GA102 High Definition Audio Controller (rev a1)
Subsystem: NVIDIA Corporation GA102 High Definition Audio Controller
Flags: bus master, fast devsel, latency 0, IRQ 17
Memory at 90080000 (32-bit, non-prefetchable) [size=16K]
Capabilities: <access denied>
Kernel driver in use: snd_hda_intel
Kernel modules: snd_hda_intel
03:00.0 VGA compatible controller: NVIDIA Corporation GA102GL [RTX A6000] (rev a1) (prog-if 00 [VGA controller])
Subsystem: NVIDIA Corporation GA102GL [RTX A6000]
Flags: bus master, fast devsel, latency 0, IRQ 222
Memory at 8d000000 (32-bit, non-prefetchable) [size=16M]
Memory at a000000000 (64-bit, prefetchable) [size=64G]
Memory at b000000000 (64-bit, prefetchable) [size=32M]
I/O ports at 9000 [size=128]
Expansion ROM at 8e000000 [virtual] [disabled] [size=512K]
Capabilities: <access denied>
Kernel driver in use: nvidia
Kernel modules: nvidiafb, nouveau, nvidia_drm, nvidia
03:00.1 Audio device: NVIDIA Corporation GA102 High Definition Audio Controller (rev a1)
Subsystem: NVIDIA Corporation GA102 High Definition Audio Controller
Flags: bus master, fast devsel, latency 0, IRQ 18
Memory at 8e080000 (32-bit, non-prefetchable) [size=16K]
Capabilities: <access denied>
Kernel driver in use: snd_hda_intel
Kernel modules: snd_hda_intel
nvidia-smi 출력값만 보면 PCIe P2P가 잘 될 것처럼 나오는데 말이지:
GPU0 GPU1 CPU Affinity NUMA Affinity GPU NUMA ID
GPU0 X PHB 0-23 0 N/A
GPU1 PHB X 0-23 0 N/A
Legend:
X = Self
SYS = Connection traversing PCIe as well as the SMP interconnect between NUMA nodes (e.g., QPI/UPI)
NODE = Connection traversing PCIe as well as the interconnect between PCIe Host Bridges within a NUMA node
PHB = Connection traversing PCIe as well as a PCIe Host Bridge (typically the CPU)
PXB = Connection traversing multiple PCIe bridges (without traversing the PCIe Host Bridge)
PIX = Connection traversing at most a single PCIe bridge
NV# = Connection traversing a bonded set of # NVLinks
인텔 소비자용 플랫폼에서 PCIe P2P를 막아버리는 순정 엔비디아 드라이버를 쓰면, 원래 지원해야 할 RTX A6000에서도 PCIe P2P가 비활성화된 걸 확인할 수 있음:
[P2P (Peer-to-Peer) GPU Bandwidth Latency Test]
Device: 0, NVIDIA RTX A6000, pciBusID: 2, pciDeviceID: 0, pciDomainID:0
Device: 1, NVIDIA RTX A6000, pciBusID: 3, pciDeviceID: 0, pciDomainID:0
Device=0 CANNOT Access Peer Device=1
Device=1 CANNOT Access Peer Device=0
***NOTE: In case a device doesn't have P2P access to other one, it falls back to normal memcopy proce
dure.
So you can see lesser Bandwidth (GB/s) and unstable Latency (us) in those cases.
P2P Connectivity Matrix
D\D 0 1
0 1 0
1 0 1
Unidirectional P2P=Disabled Bandwidth Matrix (GB/s)
D\D 0 1
0 675.24 11.74
1 11.79 676.71
Unidirectional P2P=Enabled Bandwidth (P2P Writes) Matrix (GB/s)
D\D 0 1
0 618.81 11.78
1 11.72 676.41
Bidirectional P2P=Disabled Bandwidth Matrix (GB/s)
D\D 0 1
0 649.82 16.19
1 13.66 608.81
Bidirectional P2P=Enabled Bandwidth Matrix (GB/s)
D\D 0 1
0 591.97 14.85
1 16.57 679.47
P2P=Disabled Latency Matrix (us)
GPU 0 1
0 1.61 16.40
1 17.65 1.66
CPU 0 1
0 1.38 4.55
1 4.43 1.27
P2P=Enabled Latency (P2P Writes) Matrix (us)
GPU 0 1
0 1.60 17.15
1 16.71 1.66
CPU 0 1
0 1.29 4.59
1 4.50 1.26
NOTE: The CUDA Samples are not meant for performance measurements. Results may vary when GPU Boost is
enabled.
심지어 인텔 서버 플랫폼인 Ice Lake Xeon까지 써봤는데, 이것도 AMD Epyc SP3 플랫폼보다 대역폭이랑 지연 시간 면에서 영 별로더라.
AMD가 인텔보다 PCIe 컨트롤러 구현을 훨씬 잘해놓은 것 같음. 적어도 내가 테스트해 본 플랫폼들에서는 멀티 GPU 구성할 때 AMD가 훨씬 잘 돌아갔거든. 아쉽게도 DDR5 RDIMM 가격이 미쳐 돌아가는 바람에 최신 AMD SP5나 Xeon 6 플랫폼은 테스트 못 해봤는데, 거기서도 AMD가 인텔보다는 나을 것 같음.
아래는 내 컴퓨터에서 RTX Pro 6000 GPU 2개 꽂고 돌려본 PCIe P2P 테스트 결과임. 처음에는 Supermicro X12SPa-TF 보드에 인텔 Ice Lake Xeon W-3365로 빌드했다가, 나중에 Asrock ROMED8-2T 보드에 AMD Epyc 7V73X로 옮겨서 테스트했음.
인텔 Xeon Ice Lake:
[P2P (Peer-to-Peer) GPU Bandwidth Latency Test]
Device: 0, NVIDIA RTX PRO 6000 Blackwell Workstation Edition, pciBusID: 51, pciDeviceID: 0, pciDomainID:0
Device: 1, NVIDIA RTX PRO 6000 Blackwell Workstation Edition, pciBusID: c3, pciDeviceID: 0, pciDomainID:0
Device=0 CAN Access Peer Device=1
Device=1 CAN Access Peer Device=0
***NOTE: In case a device doesn't have P2P access to other one, it falls back to normal memcopy procedure.
So you can see lesser Bandwidth (GB/s) and unstable Latency (us) in those cases.
P2P Connectivity Matrix
D\D 0 1
0 1 1
1 1 1
Unidirectional P2P=Disabled Bandwidth Matrix (GB/s)
D\D 0 1
0 1575.10 24.07
1 23.97 1600.97
Unidirectional P2P=Enabled Bandwidth (P2P Writes) Matrix (GB/s)
D\D 0 1
0 1576.29 18.37
1 20.78 1581.48
Bidirectional P2P=Disabled Bandwidth Matrix (GB/s)
D\D 0 1
0 1550.05 30.78
1 30.57 1562.45
Bidirectional P2P=Enabled Bandwidth Matrix (GB/s)
D\D 0 1
0 1550.05 39.75
1 39.75 1557.00
P2P=Disabled Latency Matrix (us)
GPU 0 1
0 2.06 14.33
1 144.15 2.07
CPU 0 1
0 2.51 5.58
1 5.55 2.32
P2P=Enabled Latency (P2P Writes) Matrix (us)
GPU 0 1
0 2.06 0.45
1 0.37 2.07
CPU 0 1
0 2.40 1.68
1 1.69 2.45
NOTE: The CUDA Samples are not meant for performance measurements. Results may vary when GPU Boost is enabled.
AMD Epyc SP3 7003:
[P2P (Peer-to-Peer) GPU Bandwidth Latency Test]
Device: 0, NVIDIA RTX PRO 6000 Blackwell Workstation Edition, pciBusID: 81, pciDeviceID: 0, pciDomainID:0
Device: 1, NVIDIA RTX PRO 6000 Blackwell Workstation Edition, pciBusID: c1, pciDeviceID: 0, pciDomainID:0
Device=0 CAN Access Peer Device=1
Device=1 CAN Access Peer Device=0
***NOTE: In case a device doesn't have P2P access to other one, it falls back to normal memcopy procedure.
So you can see lesser Bandwidth (GB/s) and unstable Latency (us) in those cases.
P2P Connectivity Matrix
D\D 0 1
0 1 1
1 1 1
Unidirectional P2P=Disabled Bandwidth Matrix (GB/s)
D\D 0 1
0 1860.12 23.87
1 23.98 1903.17
Unidirectional P2P=Enabled Bandwidth (P2P Writes) Matrix (GB/s)
D\D 0 1
0 1855.70 27.95
1 27.91 1900.92
Bidirectional P2P=Disabled Bandwidth Matrix (GB/s)
D\D 0 1
0 1831.70 30.55
1 30.67 1854.57
Bidirectional P2P=Enabled Bandwidth Matrix (GB/s)
D\D 0 1
0 1836.01 48.74
1 48.86 1854.53
P2P=Disabled Latency Matrix (us)
GPU 0 1
0 0.99 14.31
1 14.30 1.00
CPU 0 1
0 2.65 7.34
1 7.29 2.47
P2P=Enabled Latency (P2P Writes) Matrix (us)
GPU 0 1
0 0.99 0.37
1 0.36 1.00
CPU 0 1
0 2.56 2.03
1 2.09 2.57
NOTE: The CUDA Samples are not meant for performance measurements. Results may vary when GPU Boost is enabled.
주요 댓글
r/localllama
인텔 소비자용 플랫폼의 멀티 GPU 구성 시 발생하는 PCIe P2P 제한 문제에 대해 논의하며, AMD 플랫폼이 대안으로 언급되고 보드별 레인 공유 문제에 대한 정보가 공유됨.
38
Z890에서 확인했는데:
cuda-samples의 simpleP2P 테스트(p2platencytest 말고)를 실행해서 GPU 간 직접 메모리 읽기/쓰기가 가능한지 확인해야 함(안 됨) 아니면 nvbandwidth를 돌려보셈. 둘 다 엔비디아 깃허브에 있음.
그게 안 되더라도 NCCL/vLLM/SGLang은 CPU를 거쳐 복사할 수 있음. 프리필/프롬프트 처리에는 영향이 있지만(10~20% 정도 손실 가능), 단일 쿼리 기준 디코드/토큰 생성에는 영향 없음.
이건 NVME 슬롯 1, 2를 둘 다 사용할 때만 발생하는 거지? NVME 슬롯 3, 4는 칩셋에 연결되어 있고 4.0 속도로 제한되잖아. NVME 1, 3, 4를 쓰면서 PCIe 5.0 x8x8 구성을 할 수 있지 않아?
3
응 맞아
1
휴! 다행이다. 그럼 내 장비는 제대로 세팅한 거네, 고마워!
2
X870E ProArt에서 PCIe를 8x 8x 모드로 설정하면 Gen. 5 NVMe 드라이브 2개를 동시에 사용할 수 없다는 걸 알았을 때 진짜 짜증 났음.
0
Asus Pro Art X870e에서 96GB BAR 지원할까?
1
ReBAR는 문제가 아님.
문제는 P2P 그 자체임. PCIe 장치들이 CPU를 거치지 않고 서로 직접 통신할 수 있느냐가 관건이지.
지금까지는 Asus ProArt X670e나 X870e(AMD AM5)에서는 되는 것 같은데, Asus ProArt Z890(Intel Arrow Lake)에서는 안 되는 것 같음.
0
아, Z890 문제가 플랫폼 제한이라는 건 이해했음. 깃 스레드 파보니까 NIC <-> GPU 관련이긴 했지만 direct writes는 가능한데 direct reads는 안 되는 것 같더라.
AM5 플랫폼은 P2P 문제없이 잘 돌아가는 것 같고, 그냥 해당 보드에서 Resizable BAR가 궁금했음. 전문가가 아니라서 잘 모르지만, P2P 성능을 제대로 뽑으려면 256MB 제한에 걸리면 안 되니까. 2의 거듭제곱이니까 카드당 96GB면 128GB가 필요하겠네.
0
응, Z890에서도 131072 MiB ~ 128GiB로 뜨더라.
-1
AsRock Taichi 870은 2개의 PCI 포트에 각각 8x 전용 레인을 할당해서 이 문제를 해결해. Asus에 비해 플래시나 기능은 좀 빠지지만, 내 보드는 아주 안정적이었어.
두 번째 답글 - 많은 글을 읽고, 네 링크랑 추가 검색, 그리고 AI 쿼리 2개를 돌려본 결과... 매뉴얼이 맞고 구리 배선이 PCI 슬롯과 분리되어 있다는 걸 95% 확신해. 다른 제조사들은 공유 회로랑 소프트웨어 패치로 기능을 쑤셔 넣으려고 하는데, 이건 Asrock만의 특징이야.
다음 주에 돌아가면 직접 확인해 볼 거야.
1
음 — 이제 잘 모르겠네 - 매뉴얼 15페이지를 보면 GPU 2개가 8x, 8x로 작동하고... USB랑 드라이브는 독립적이라고 아주 명확하게 나와 있어.
방금 보드를 샀으니까 이제 직접 파헤쳐 볼 거야. 고마워.
1
나 Taichi 쓰는데 M.2 1번이랑 2번 다 채웠고, USB 4도 쓰고 있고, GPU 2개 다 x8x8로 작동 중이야.