GLM 5.2 Q1_S vs Qwen 27B Q8 성능 비교
GLM 5.2 Q1_S vs Qwen 27B Q8
핵심 요약
거대 모델의 초저양자화(Q1_S)가 사고 과정을 통해 중형 모델의 고양자화보다 뛰어난 코딩 성능을 보여줌.
- 양자화 효율성 — 대형 모델의 낮은 양자화가 소형 모델의 높은 양자화보다 코딩에서 유리함을 증명함.
- 사고 과정의 힘 — GLM 5.2의 사고 트레이스가 복잡한 코딩 작업을 한 번에 성공시키는 핵심 요인임.
- 하드웨어 제약 — 2x 3090 환경에서 거대 모델 구동 시 발생하는 극심한 발열과 속도 저하를 언급함.
- 모델 특화 논쟁 — 코딩에 특화된 REAP 모델의 범용성 부족과 특정 지식 망각 현상에 대해 토론함.
TL;DR; GLM-5.2 Q1_S beats Qwen 3.6 27B Q8, both run at KV Q8
edit: GLM run a K & V Q8, Qwen run with KV cache at full FP16., with preserve thinking on.
Disclaimer: This is a hobby/amateur comparison with n=1, so go easy on it. I just thought it would be fun to share.
The Context and The Task
Some time ago there were quite a few discussions on what's better: a lower quant of a larger model, or a higher quant of a smaller model. We got quite a few benchmarks and in-house tests, which were mostly consistent — the larger model at a lower quant was better.
Nowadays I often see claims of anything lower than Q3 being 'braindead' regardless of the actual size. I've also noticed some comments belittling people who share how they've managed to run huge models on their consumer-grade hardware, just because it was a low quant.
So, I did a little test. Beloved Qwen 27B at Q8 vs 'braindead' GLM 5.2 at Q1_S. The Q1_S is the smallest quant I could find, but I really wouldn't be able to run Q2 anyway.
My hardware is 2 x RTX 3090, 24GB VRAM each (limited to 200W power) and 192 GB DDR5 RAM. I run Qwen at ~60 tps gen, and GLM at ~6 tps at low context down to 3 tps nearing 100k context.
I picked a simple tech stack and clear instructions, so that there would be as little variance due to instruction ambiguity as possible.
Both models were run under the pi harness, with the exact same config and prompts. The instruction was to build a simple 3D game in Three.js (HTML/CSS/JS); the full content is attached at the end.
This is the second attempt at this test. The first one was not documented and used a different tech stack, but the results were practically the same.
Qwen 3.6 27B
It went quick, that's for sure — just a couple of minutes and ~20k tokens. But it failed to build a working product. After instructing it to fix it, it was 'working' but still not playable; it required another 2 prompts to make it 'done'. So in total: 1 initial + 3 follow-ups, with a total of ~42k tokens.


