RTX Pro 4500(Oculink 연결)에서 PrismaQuant, INT4 Autoround 및 NVFP4 W4A4 양자화 모델 테스트
Some testing on RTX Pro 4500 (With Oculink) on PrismaQuant, INT4 Autoround and NVFP4 W4A4 quantized model
핵심 요약
미니 PC와 Oculink로 연결된 RTX Pro 4500에서 다양한 양자화 모델의 성능과 vLLM 환경에서의 문제 해결 과정을 공유합니다.
- 하드웨어 구성 — Beelink SER 8 미니 PC와 Oculink로 연결된 RTX Pro 4500 32GB 사용
- 문제 해결 — vLLM에서 발생하는 툴 콜 오류와 무한 루프 문제를 MTP 비활성화로 해결
- 양자화 비교 — PrismaQuant(PrismaSCOUT, PrismaAURA) 모델의 KLD 및 생성 속도 벤치마크 수행
- 실사용 사례 — 로컬 호스팅 서비스 관리 및 Opencode/Cline의 계획 실행 에이전트로 27B 모델 활용
This little beast has been around for a while after asking about whether it's possible to set up in this subreddit.
Beelink SER 8 8745 HS, AooStar eg01, and RTX Pro 4500 32GB
The Sakamakismile model I've been running was throwing tool call errors and getting stuck in thinking loops in both Opencode and Cline across vLLM 0.22, 0.23, and 0.24 for some reasons, and even swapping to the Froggeric Chat Template didn't seem to improve things (it was relatively OK for roughly 2 days of intense using, then the problem surfaced again). I went looking to see if there were any NVFP4 quantized models available.
Then I found a new quantization method that's been getting some discussion on the official Nvidia DGX Spark forum, called PrismaQuant. Simply put, it selects the most suitable format for each linear layer to maximize the model's capabilities at a specific bit.
Note that the PrismaQuant quantization method is currently only usable in vLLM for the Blackwell architecture (50 series, RTX Pro series). Also, because it's so new, GGUF is basically completely unsupported right now.
| Model Name | Quantization Bits | Weight Size | Base/Source Model |
|---|


