프로덕션 환경에서 원인을 알 수 없는 토큰 사용 패턴이 발생함
Seeing unexpected token usage patterns in production hard to attribute where it’s coming from
핵심 요약
프로덕션 환경에서 발생하는 예상치 못한 토큰 사용량 증가의 원인을 파악하고 디버깅하는 방법을 공유함.
- 토큰 사용 패턴 — 시스템 프롬프트 반복 및 에이전트 재시도로 인한 토큰 소모 증가를 확인함
- 컨텍스트 관리 — 멀티 턴 환경에서 컨텍스트가 제대로 정리되지 않고 계속 늘어나는 문제를 겪음
- 디버깅 도구 — LangGraphics를 통한 실행 그래프 시각화 및 Helicone을 활용한 토큰 사용량 추적을 제안함
Ran into something recently while looking at our LLM usage.
At an aggregate level, costs look fine. But when breaking it down, some patterns were surprising:
- A single workflow accounted for a large % of token usage, mostly due to repeated system prompts
- Some agent-style flows were making significantly more tool calls than expected (likely retries / edge cases)
- Context length was growing over time in certain flows without much pruning
Individually none of this looked alarming, but combined it was a noticeable chunk of usage.
What made it tricky is that this wasn’t obvious from standard logs or dashboards had to dig a bit to even notice it.
Curious if others have seen similar patterns, especially around:
- attributing token usage at a workflow/feature level
- debugging agent loops or unexpected tool calls
- managing context growth in multi-turn setups
Would be interested in how people are instrumenting or debugging this.


