IMPORTANT:This isn’t just a problem with the German language. It affects all languages!
Claude Code systematically translates, renames, paraphrases, and sometimes invents variants of established technical and architectural terminology when the conversation/output language is German.
This is NOT merely a UI localization issue and NOT merely a matter of translation quality.
It directly changes the semantic identifiers Claude uses to reason about the architecture of a software project.
Observed examples from real long-running Claude Code projects:
- `Gate` → `Zaun` ("fence")
- `Gates` → `Zäune` ("fences")
- `Policy` / `Policies` → `Politik` ("politics")
- `Root` → `Wurzel` ("root" literally translated as the botanical/anatomical German word)
- canonical internal component/process names → newly invented German variants
- previously established architecture labels → inconsistent translated synonyms
The `Root` case is particularly important because this demonstrates that the problem is broader than obviously absurd translations.
In German technical prose, words such as "Wurzelverzeichnis" may sometimes be linguistically understandable as a translation of "root directory".
That is NOT the issue.
If `Root` is the canonical name of an architectural component, filesystem concept, process, scope, node, registry entry, or project-defined identifier, then changing `Root` to `Wurzel` is semantically destructive even if the translated word could be understood by a human.
The same applies to all canonical project terminology.
A model must distinguish between:
ordinary natural-language prose that may be translated
Claude currently fails to preserve this distinction reliably.
The most severe observed examples are:
`Gate` → `Zaun`
In an agent architecture, a Gate is a control/validation boundary. A "Zaun" is literally a physical fence. These are not semantically interchangeable concepts.
`Policy` / `Policies` → `Politik`
In a software architecture, a Policy is a rule, policy object, enforcement definition, permission definition, or behavioral contract. "Politik" means politics/political policy in German and completely changes the semantic domain.
`Root` → `Wurzel`
In a software architecture, Root may identify a canonical project root, hierarchy root, root node, filesystem root, or explicitly named architecture component. Translating it creates a different identifier and breaks one-to-one terminology mapping.
The core problem is therefore not simply "bad German".
The core problem is:
CANONICAL CONCEPT A
→ Claude translates/renames it
→ Claude now represents it internally/output-wise as CONCEPT NAME B
→ the new name enters documentation, plans, memory, summaries, or agent communication
→ later Claude treats B as an established project concept
→ more aliases and translations are produced
→ the architecture vocabulary diverges from the actual architecture.
This becomes especially destructive during long-running agent work because the mistranslated or newly invented terminology does not remain conversational wording.
Claude can propagate it into:
- project documentation
- architecture documentation
- implementation plans
- task definitions
- persistent memory
- auto-memory
- external project memory files
- process descriptions
- internal labels
- proposed filenames
- directory descriptions
- rule descriptions
- hook feedback
- Sentinel feedback
- summaries
- context-compaction summaries
- subagent prompts
- subagent responses
- subsequent reasoning
- implementation decisions
Once an incorrect translated term enters persistent project state, Claude can retrieve it later as if it were a valid canonical concept.
This creates a self-reinforcing feedback loop:
Canonical English technical term exists.
Claude translates or renames it while communicating in German.
The translation changes the identifier and potentially its meaning.
Claude stores/references the translated version.
Context is compacted or the session changes.
The translated version survives in memory/documentation.
Claude retrieves the translated variant.
Claude assumes the translated term is legitimate.
Claude creates additional synonyms or variations.
Project terminology and architectural reasoning drift further away from the source architecture.
This is effectively persistent semantic-memory contamination.
I have extensively attempted to prevent this behavior using:
- `CLAUDE.md` instructions
- canonical terminology indexes
- terminology registries
- explicit glossaries
- exact-name lists
- blocklists for previously invented names
- memory definitions
- explicit "DO NOT TRANSLATE" rules
- explicit "DO NOT INVENT ALIASES" rules
- Sentinels
- hooks that inject terminology requirements immediately before operations
- filesystem names as canonical identifiers
- process names as canonical identifiers
- repeated corrections during sessions
These mechanisms can reduce individual occurrences but do not reliably solve the problem.
Claude eventually starts translating, paraphrasing, renaming, or inventing terminology again.
The issue becomes dramatically worse when Claude communicates in German while operating on a project whose technical terminology is primarily English.
Equivalent workflows performed entirely in English show substantially less terminology corruption.
This makes German-language operation unreliable for large agentic software projects.
Concrete observed transformations:
- `Gate` → `Zaun`
- `Gates` → `Zäune`
- `Policy` / `Policies` → `Politik`
- `Root` → `Wurzel`
These are not isolated vocabulary problems. The same translation mechanism is dangerous for a large class of overloaded English software-engineering terms.
Representative examples of the same HIGH-RISK failure class (illustrative examples, not all claimed as individually observed):
- `Hook` → `Haken`
- `Branch` → `Zweig`
- `Fork` → `Gabel`
- `Socket` → `Steckdose`
- `Port` → `Hafen`
- `Thread` → `Faden`
- `Pipeline` → `Rohrleitung`
- `Worker` → `Arbeiter`
- `Shell` → `Schale` / `Muschel`
- `Handle` → `Griff`
- `Container` → `Behälter`
- `Registry` → `Registrierung`
- `Controller` → `Steuerung` / `Kontrolleur`
- `Contract` → `Vertrag`
- `Sentinel` → `Wächter`
- `Validator` → `Prüfer`
- `Capability` → `Fähigkeit`
- `State` → `Zustand`
- `Scope` → `Bereich`
- `Artifact` → `Artefakt`
Some of these translations may be linguistically acceptable in ordinary German prose.
That is precisely why this bug is dangerous.
The problem is NOT whether a German translation exists.
The problem is that a canonical identifier MUST NOT be translated at all.
For example, suppose an architecture explicitly defines:
`PolicyGate`
`RootGate`
`ContractRegistry`
`CapabilityRegistry`
`Sentinel`
`WorkerPool`
If Claude internally or externally starts referring to these as:
`Politik-Zaun`
`Wurzel-Zaun`
`Vertragsregistrierung`
`Fähigkeitsregistrierung`
`Wächter`
`Arbeiter-Pool`
then the deterministic one-to-one relationship between project terminology and model terminology has been destroyed.
This can affect reasoning even before it affects code because Claude begins reasoning about the system using vocabulary that is no longer identical to the vocabulary used by the architecture.
In an agentic system, terminology is not decoration.
Terminology is part of the system model.
주요 댓글
r/anthropic
사용자들은 Claude Code의 강제 번역이 프로젝트의 기술적 일관성을 해치는 심각한 문제라고 지적하며, 단순한 용어집 설정으로는 해결이 불가능해 Anthropic의 시스템 프롬프트 개선이 필수적이라고 주장함.
1
독일인이신가요? 독일 사람들은 실제로 어떤 단어를 쓰나요? 번역하지 않고 그냥 영어 단어를 그대로 쓰나요?
0
네, 맞습니다. 하지만 이건 단순히 독일어만의 문제가 아니에요. 모든 번역과 관련된 문제입니다. 그게 바로 핵심이죠. 기술 용어는 단어 그대로 채택되어야 합니다. 이건 모델들이 일반적으로 저지르는 근본적인 실수인데, 그렇지 않으면 기능이 왜곡되기 때문입니다. 독일어로 다시 작성할 수는 있겠죠. 하지만 그걸 다시 영어로 번역하면 기능이 또 잘못 설명될 겁니다. 그래서 Claude가 이 문제를 더 지능적으로 해결하고 기술 용어를...
0
LLM의 본질은 토큰이 벡터로 존재한다는 것이고, 그 벡터는 주변 토큰들과 비교하며 의미를 정의함. 이건 거의 피할 수 없는 문제지만, 커스텀 용어집을 정의해서 Claude.md 안에 넣거나 Claude.md가 그걸 참조하게 만들 수는 있음.
0
Claude.md는 이 문제를 해결하기에 완전히 잘못된 장소임. Claude.md는 세션 시작 시 딱 한 번 로드되고, 그 이후에는 거대한 컨텍스트 안의 그냥 또 다른 컨텍스트일 뿐임. Anthropic은 모든 프롬프트마다 시스템 프롬프트를 가지고 있으니 이 문제를 아무 문제 없이 해결할 수 있음. 완벽하게 고칠 수 있다는 뜻임. 그렇지 않으면 우리가 훅(hook)을 써서 직접 해결해야 함. 문제는 애초에 이런 이상한 번역을 유발하는 게 Anthropic의 시스템 프롬프트라는 거임. 그들이 번역을 수정해야…
용어집을 쓰고 싶지 않다면, 주로 기술 용어를 사용한다는 걸 인식하도록 프로필을 설정해보는 건 어때?
0
맞음, 용어집은 다른 상황에서는 별 도움이 안 됨. 왜냐면 다른 주제에서도 실수를 하거든. 예를 들어 이상한 라벨을 쓰거나, 사실을 지어내거나, 함수 이름을 틀리기도 함. 훅을 써도 도대체 무슨 소리를 하는 건지 알 수 없는 이상한 재표현들을 만들어냄. 그래서 용어집으로는 절대 해결이 안 됨. 그런 방식으로는 절대 작업 못 함. 결국 나중에 또 문제를 일으킬 수 있는 끝없는 데이터를 집어넣게 되는 꼴임…
저도 독일인이지만, 프롬프트와 문서화에는 항상 영어만 사용해서 한 번도 문제를 겪은 적이 없습니다. 직장에서 동료가 독일어를 쓰는 걸 봤는데 결과물이 정말 오글거렸어요. 주석이랑 문서가 독일어랑 영어가 섞여 있고, 단어 하나나 문장 중간에 언어가 바뀌기도 하더라고요. 영어만 고수하시면 아마 이런 문제는 절대 겪지 않으실 겁니다. 제 착각일 수도 있는데, 독일어로 프롬프트하면 어차피 처리되기 전에 영어로 번역되는 거 아닌가요?
0
네, 지금 제 시스템 전체를 재구축하고 있어요. 모든 걸 다시 쓰고 있고, 메모리 파일 같은 것도 전부요. 모델이 엉망진창인 실수를 너무 많이 해서 엄청난 작업량이었습니다. 정말 재앙이었어요, 장담합니다. 아마 처음에 모델한테 그런 용어를 쓰라고 시켰을 수도 있고, 그게 메모리에 기록되면서 계속 사용했을 수도 있겠죠. 하지만 처음부터 그렇게 시작하고 명시하지 않으면, 그 뒤에 오는 모든 게 완전히 엉망이...
저도 같은 현상을 겪었고 아직 좋은 해결책을 못 찾았습니다. 제 애플리케이션의 번역 파일을 조정하게 시키면 'Version 1.0'을 'Fassung 1.0'으로 바꾸는 등 이상한 결과가 나올 때가 많아요. Claude가 영어로만 작업하고 응답을 번역 레이어로 감싸는 것 같은 느낌이 듭니다.
0
시도해 볼 수는 있겠죠. 저는 연구와 전문가들의 의견을 바탕으로 다 해결했습니다. 지금 제 시스템 전체를 메모리 파일까지 포함해서 업데이트하고 있어요. 모든 명령어와 규칙도 영어로 작성해야 합니다. 모델이 바로 그런 것들을 가지고 작업하고 생각하기 때문이죠. 모든 내부 프로세스는 영어로 기술되어야 합니다. 그리고 모델에게 영어로 생각하라고 명시적으로 지시하는 것도 중요합니다.
0
지금 제 시스템을 전환 중인데, 점점 더 많은 문제들이 눈에 띄네요. 에이전트는 환경에 적응하거든요. 즉, 시스템 내부의 모든 주제가 영어로 되어 있어야 한다는 뜻입니다. 그렇지 않으면 다시 드리프트가 발생해서 독일어로 말하기로 결정해버릴 테니까요. 모델이 작업하는 모든 파일이 영어로 되어 있는 것이 필수적입니다. 그렇지 않으면 모델이 적응해서 계속...