Gemma 4 12B(통합 오디오)로 긴 시스템 프롬프트에서 음성 인식을 제대로 시킨 사람 있음?
Anyone gotten Gemma 4 12B (unified audio) to actually attend to speech with a large system prompt?
핵심 요약
Gemma 4 12B 모델에서 긴 시스템 프롬프트 사용 시 오디오 인식이 무시되는 현상에 대한 해결책을 찾는 중입니다.
- 모델 성능 저하 — 긴 시스템 프롬프트 사용 시 오디오 인식 기능이 제대로 작동하지 않음
- 다양한 스택 테스트 — vLLM, llama.cpp, LiteRT-LM 등에서 동일한 문제 발생
- 잠재적 원인 — 오디오와 방대한 텍스트 컨텍스트 간의 어텐션 포화 문제로 추정
- 해결책 모색 — 프롬프트 구조 조정이나 어텐션 설정 변경 등 우회 방법 질문
I'm trying to use Gemma 4 12B — the new encoder-free unified model (audio/vision/text in one) — for a one-pass audio → response voice assistant: feed the recorded WAV + system prompt and get the reply back as text directly, collapsing the separate ASR + LLM steps into a single model (TTS still happens afterward).
Works great with a minimal prompt — the model clearly hears and responds to the audio. But once the text prompt gets large/dense (mine is ~21k tokens: detailed instructions + tool definitions), it basically stops attending to the audio — replies as if the audio weren't there (generic/hallucinated) or only weakly transcribes. Trim the prompt back down and audio attention returns.
Same behavior across three stacks, so it doesn't look stack-specific:
- vLLM (gemma4-unified image + pip install av), audio as base64 audio_url
- llama.cpp (--mmproj, input_audio content, chat_template_kwargs {enable_thinking:false})
- LiteRT-LM (gemma4-12b,gpu)
Feels like an inherent attention/saturation limit when audio competes with a long dense text context. (Notably, E4B with a tiny prompt keeps audio attention fine — so I'm using it as a small audio front-end instead.)
Questions for anyone who's tried:
1. Has anyone gotten 12B unified audio to reliably attend to speech with a big system prompt (lots of instructions/tools)?
-
Known limitation of the unified arch, or a serving/config thing (audio placement in the sequence, attention settings, chat template, sampling)?
-
Workarounds — audio-first vs audio-last ordering, prompt structuring, attention/RoPE tweaks?
Served on an NVIDIA GB10 (Blackwell).

