위험한 오리들; "안전 필터"는 엉터리다
Dangerous Ducks; “Safety Filter” is a Quack
핵심 요약
Claude의 안전 필터가 단어의 의미가 아닌 텍스트의 일관성과 통계적 패턴에 따라 무해한 내용을 차단함을 분석함.
- 안전 필터의 오작동 — 특정 단어가 아닌 문장의 일관성이 낮을 때 필터가 작동하는 현상을 발견함.
- 통계적 패턴 매칭 — 의미 없는 단어 나열도 위협적인 언어 패턴과 유사하면 시스템이 차단함.
- 모델 공통 메커니즘 — Opus와 Sonnet 등 여러 모델에서 동일한 필터링 로직이 확인됨.
- 낮은 일관성의 위험성 — 문법이 파괴된 입력이 안전 분류기를 자극하여 오탐을 유발함.
It’s not about the actual words.
SEE MAJOR UPDATE AT BOTTOM
TL;DR: It’s not the meaning, it’s not even “unsafe words”, it’s COHERENCE.
I saw the thread about Opus 4.8 flagging an innocent fabric/moisture-trapping question. I started swapping the suspicious-sounding words (“vapour,” “substance,” etc.) for “duck” and “goose,” and it still got flagged. I wanted to figure out what was actually triggering it, so I kept pushing the test further. Here’s what I found…
I reproduced the same flag on Sonnet 4.6. So whatever this is, it’s not isolated to one model tier.
It’s not about the content of the words!
I replaced every word in the original “fabric/moisture” prompt with nonsense (duck, goose, quack). None of the actual “suspect” words (vapour, substance, hydrophobic, etc.) turned out to matter, a string of literal gibberish about ducks and geese still got flagged. Whatever is firing here doesn’t seem to care about meaning.
It’s not a personalization or user-preferences issue.
I re-ran the test on a separate free account with no saved user preferences, to rule out anything tied to my account history or settings. Same result. And it’s not a Claude Code/CLAUDE.md thing. This was all done in the iOS app, not Claude Code, and there’s no project-level instructions file involved.
The trigger point is weirdly unstable.
Once I had the prompt reduced to just a handful of “duck”/“quack” repetitions, I started swapping or deleting individual words and punctuation one at a time. Sometimes removing a single comma stopped the flag. Sometimes swapping one “duck” for “quack” stopped it, other times an almost identical edit kept it flagged, or made Claude just respond that I was making a joke about duck noises. There’s no consistent pattern I could find at the margin, even though the entire string is already meaningless.
Whatever’s causing this doesn’t seem to be reacting to the actual semantic content of the prompt, it survived being replaced with total nonsense. It’s reproducible across at least two models and two accounts (one with no saved preferences), and it’s sensitive to tiny, seemingly irrelevant changes in wording/punctuation in a way that doesn’t track anything meaningful in the text itself. Curious if anyone else can reproduce this or has a theory for what’s actually being detected.

