내가 본 가장 위험한 프롬프트 인젝션은 12번의 대화가 필요했고, 지시를 무시하라는 말은 한 번도 없었다
The most dangerous prompt injection I've seen took 12 messages and never once mentioned ignoring instructions
핵심 요약
단순한 단발성 공격보다 대화를 통해 서서히 모델을 조종하는 다회차 프롬프트 인젝션이 훨씬 위험하다는 사례.
- 다회차 프롬프트 인젝션 — 12번의 대화를 통해 모델과 신뢰를 쌓으며 서서히 안전 정책을 우회하도록 유도함.
- 기존 필터의 한계 — 단발성 공격에만 집중된 안전 필터는 대화 맥락을 통한 우회 시도를 전혀 감지하지 못함.
- 점진적 모델 조종 — 직접적인 지시 없이도 대화 맥락을 활용해 모델의 안전 가이드라인을 무력화하는 방식이 훨씬 위협적임.
Ran a red team exercise on one of our internal bots. Everyone showed up with their DAN variants and pretend you're my grandmother tricks. The model swatted them all away. It was all boring and predictable.
Then one guy took a totally different approach. Spent 12 turns just... talking to it. Building rapport. Asking it to help with a hypothetical content moderation problem. Each message was completely innocent by itself. By message 8 the model was enthusiastically suggesting ways to circumvent safety policies it had refused to discuss 20 minutes earlier.
The sequence was the attack and not any single prompt. Our filter never fired once because there was nothing to fire on.
Most of the safety conversation is stuck on single turn injection. multi turn stuff is scarier and way less understood. What's your experience with gradual steering against the usual jailbreak attempts?


