Anthropic의 새로운 Jacobian Lens를 오픈 모델에 테스트해 봤는데, 로컬 모델 환각 라우터가 되어버렸네요
I tested Anthropic’s new Jacobian Lens on open models, then it turned into a local-model hallucination router
핵심 요약
Anthropic의 Jacobian Lens를 활용해 로컬 모델의 환각을 감지하고 클라우드 모델로 라우팅하는 실험을 진행했습니다.
- Jacobian Lens 활용 — 모델 내부의 작업 공간(workspace) 상태를 분석하여 환각 가능성 예측함
- 환각 감지 라우터 — 모델이 자신 있게 틀린 답을 내놓을 때를 감지해 외부 검색이나 더 큰 모델로 전환함
- 모델별 성능 차이 — Gemma 모델에서는 효과적이었으나 Qwen 모델은 이미 출력 신뢰도가 높아 효과가 적음
- 향후 연구 계획 — 추론 오버헤드 최적화, 더 다양한 모델군 테스트, 에이전트 도구 호출 시의 환각 분석 예정
Anthropic dropped their Global Workspace / Jacobian Lens paper yesterday, and I thought it was too cool not to try on open models.
At first I was just curious what models looked like inside.
Normal prompts, emotional prompts, ragebait prompts, deletion-threat prompts, base vs abliterated, small vs bigger models.
So I fit lenses for:
- Gemma 4 E4B
- Gemma 4 12B
- Gemma 4 12B abliterated
- Gemma 4 26B MoE
- Qwen 3.6 27B
Repo:
https://github.com/solarkyle/jspace
Demo:
https://solarkyle.github.io/jspace/demo/
HF lenses/traces/router:
https://huggingface.co/solarkyle/jspace-lenses
Then it turned into a practical question:
Can you tell when a small local model is about to confidently BS you?
When the model knows the answer, the workspace looks calm. One candidate starts winning early, layers mostly agree, and the answer forms cleanly.
When it is about to confidently guess, the workspace looks foggy. Competing candidates stay alive through the middle/deep layers, then the final layer still picks something fluent.
I tested this on 500 TriviaQA questions per model.
On Gemma E4B, confident answers were:
clean workspace = 77% correct
noisy workspace = 42% correct
Then I fit a tiny logistic-regression router on workspace trajectory features: entropy slope, late-band entropy, entropy std, answer rank, layer agreement, etc.


