더 이상 토큰 부족에 시달리지 마세요
Never Run Out of Tokens Again
핵심 요약
Luna API의 저렴한 가격을 활용해 토큰 비용을 획기적으로 줄이는 워크플로우를 소개합니다.
- API 비용 절감 — Luna API의 저렴한 입력/출력 토큰 가격을 활용함.
- 워크플로우 최적화 — SOL을 오케스트레이터로, Luna를 에이전트로 사용하는 구조임.
- 컨텍스트 압축 — 비용 효율성을 위해 계획 수립 후 컨텍스트를 압축함.
- 에이전트 설정 — Luna 서브 에이전트 활용을 위한 프롬프트 및 설정 가이드를 제공함.
Luna의 새로운 API 가격인 입력 토큰 100만 개당 $0.20, 출력 토큰 100만 개당 $1.20이면 그냥 거저먹기지.
내가 지금까지 아주 잘 써먹고 있는 워크플로우는 이거야:
1. SOL xhigh로 구현 계획 짜기
SOL한테 작업을 분석하게 시키고 상세한 구현 계획을 뽑아내.
<선택 사항> 나는 "Grill Me"랑 비슷한 스킬을 하나 만들었는데, 이게 작업 구현에 필요한 도메인 지식에 대해 스스로 확신이 들 때까지 날카로운 질문을 던져서 충분히 이해하게 만들어.
2. 컨텍스트 압축하기
계획이 끝나면 대화 내용을 압축해. 그래야 오케스트레이터가 비싼 컨텍스트 토큰을 낭비 안 하지.
3. 오케스트레이터로 SOL high 돌리기
이런 식의 프롬프트를 써봐:
TASK
Your job is to orchestrate and review the Luna max-thinking agent.
Focus especially on:
- Code quality
- Simple and understandable implementations
- Useful comments and documentation
- Idiomatic framework-specific best practices
- Meaningful tests
Tests should not cover only the happy path when additional edge cases or failure scenarios would be useful.
After reviewing Luna’s work, decide whether to:
1. Call Luna max-thinking again with the full context required to resolve the identified issues, or
2. Fix the issues yourself when doing so would require substantially fewer tokens.
START THE LUNA AGENT WITH:
codex exec \
-m gpt-5.6-luna \
-c 'model_reasoning_effort="max"' \
--ephemeral \
-s workspace-write \
-a never \
'PLAN'
---- OPTIONAL IF YOU WANT TO SEE SOME RESULTS FRIENDO ----
ADD THIS TO YOUR PROMPT
Summarize the cost generated by the Luna agent using the new API prices:
- $0.20 per million input tokens
- $1.20 per million output tokens
Show Luna’s cost separately from your own cost as the orchestrator.
Then estimate what the total cost would have been if SOL xhigh had completed the entire task alone without Luna.
UPDATE: You may be able to spawn native Luna sub-agents, which would be easier and potentially even more cost-efficient because they require less context.


