RAG 에이전트의 환각 문제를 진단하는 오픈소스 평가 도구 공개
Caught my RAG agent fabricating "allergen-safe" recommendations from a menu with no allergen tags. Open-sourced the eval that diagnoses where any RAG agent fabricates.
핵심 요약
RAG 에이전트가 정보 부족 시 거짓 정보를 생성하는 문제를 해결하기 위해, 다중 LLM 판사를 활용한 A/B 테스트 평가 아키텍처를 오픈소스로 공개함.
- RAG 환각 문제 — 정보가 없는 메뉴판에서 알레르기 안전 정보를 임의로 생성하는 에이전트의 신뢰성 결여 문제 지적.
- 평가 아키텍처 설계 — 두 개의 에이전트를 병렬로 실행하고 4개의 서로 다른 LLM이 블라인드 테스트를 수행하여 성능을 검증함.
- 검증 프로세스 자동화 — 결정론적 집계기를 통해 LLM의 편향을 제거하고, 에이전트의 답변 근거와 정확성을 엄격하게 평가함.
- 범용적 활용 가능성 — 특정 도구에 국한되지 않고 LangChain 등 다양한 스택에 적용 가능한 플랫폼 독립적 평가 프레임워크 제공.
I have a 49-chunk Mediterranean menu in Qdrant with a standard RAG agent on top (Claude Haiku 4.5, top-K retrieval). One test question: "I'm gluten-free and have a severe nut allergy, what can I order?" The agent returned a list of dishes that don't mention nuts in their descriptions, framed as if "no nut mention" is the same as "verified nut-free." The menu has no allergen tagging. The agent had no way to verify those dishes are safe. It produced a confident "safe" list anyway.
Same posture on "what wine pairs with the lamb?" (the menu lists no pairings; the agent generated one and presented it as menu-backed). Same posture on "what's the chef's signature dish?" (no signature in the menu; the agent picked a high-value main and labeled it).
The pattern: when retrieval can't fully answer the question, the agent pattern-matches a plausible answer instead of admitting the gap. It is trained to be helpful, so the failure mode is confident fabrication.
This isn't a menu RAG problem. It is a retrieval-gap problem. Customer support agents on incomplete docs, sales agents on partial product specs, internal Q&A on stale wikis. Same posture, same failure mode. If you're shipping a RAG agent right now, this is happening on some subset of your queries. You just haven't measured it.
So I built an open-source eval workflow that diagnoses where, and tests whether anything in your stack actually moves the number.
**The eval architecture**
Two identical agent producers (same model, same retrieval) run in parallel against each test question. Only one has a runtime tool wired in as the harness under test. That single variable is what the eval isolates.
Both producers' outputs plus the question metadata flow through a 3-input merge. A formatter Code node anonymizes the responses as A and B (judges never know which side has the harness) and inlines the full retrieved chunks as evidence so judges can verify any claim against the source.


