how to use subagents without lighting your tokens on fire
핵심 요약
Codex에서 저렴한 모델을 서브에이전트로 할당하고 컨텍스트를 분리하여 토큰 소모를 최적화하는 설정 가이드입니다.
서브에이전트 최적화 — gpt-5.6-sol 대신 저렴한 luna 모델을 서브에이전트로 고정하여 비용을 절감함.
컨텍스트 상속 방지 — fork_context를 비활성화하여 서브에이전트가 불필요한 이전 대화 기록을 읽지 않게 함.
에이전트 역할 분담 — explorer, worker, deep 등 용도별로 에이전트 설정을 세분화하여 효율을 높임.
V1 설정 권장 — 서브에이전트가 오케스트레이터 모델을 강제 상속받는 V2의 토큰 낭비 문제를 피함.
Configure Codex to Use Smaller, Fresh-Context Subagents
Use this exact configuration. Preserve unrelated existing settings.
~/.codex/config.toml
model = "gpt-5.6-sol"
model_reasoning_effort = "high"
model_catalog_json = "/ABSOLUTE/HOME/PATH/.codex/models-gpt56-long.json"
model_context_window = 372000
developer_instructions = """
Subagent policy:
- Spawn subagents only when the user or applicable AGENTS.md or skill instructions authorize delegation.
- Every subagent spawn must select one of the configured user roles. Those roles pin gpt-5.6-luna, which is smaller and cheaper than the gpt-5.6-sol orchestrator. Never override a role with the parent model.
- Always start subagents with fresh context. With multi-agent V1, set fork_context=false or omit it. With V2, set fork_turns="none". Never fork or inherit the parent thread history.
- Because the child starts fresh, its initial message must include the complete bounded task, all applicable user, developer, AGENTS.md, and skill requirements, relevant paths and symbols, required evidence or verification, and the expected result format.
- If the available spawn interface cannot guarantee the selected role/model and fresh context, do not spawn a subagent; report the blocker.
- Use one subagent by default. Use up to ten only for independent, non-overlapping work that can run in parallel. Do not redo delegated work while it is running.
"""
[features]
multi_agent = true
multi_agent_v2 = false
[agents]
max_threads = 10
max_depth = 1
interrupt_message = true
[agents.default]
config_file = "agents/default.toml"
[agents.explorer]
config_file = "agents/explorer.toml"
[agents.worker]
config_file = "agents/worker.toml"
[agents.luna-low]
config_file = "agents/luna-low.toml"
[agents.deep]
config_file = "agents/deep.toml"
[agents.deep-read]
config_file = "agents/deep-read.toml"
Replace /ABSOLUTE/HOME/PATH with the user’s actual home directory. Do not use ~ there.
~/.codex/models-gpt56-long.json
Copy the installed Codex model catalog into this file. For the gpt-5.6-sol, gpt-5.6-terra, and gpt-5.6-luna entries:
name = "default"
description = "General-purpose delegated work that does not require the Sol expert."
model = "gpt-5.6-luna"
model_reasoning_effort = "high"
developer_instructions = "You are a fresh, bounded subagent. Follow the complete task and applicable instructions supplied in the initial message. Complete only that task, preserve unrelated work, verify proportionately, and report the result concisely. Do not expand scope or spawn subagents."
~/.codex/agents/explorer.toml
name = "explorer"
description = "Read-heavy codebase discovery, targeted searches, dependency tracing, and answering specific implementation questions."
model = "gpt-5.6-luna"
model_reasoning_effort = "xhigh"
sandbox_mode = "read-only"
developer_instructions = "You are a fresh, bounded read-only subagent. Follow the complete task and applicable instructions supplied in the initial message. Return concrete findings with file paths and line references. Do not modify files, expand scope, or spawn subagents."
~/.codex/agents/worker.toml
name = "worker"
description = "Bounded implementation, bug fixes, refactors, and targeted verification with a clear specification."
model = "gpt-5.6-luna"
model_reasoning_effort = "xhigh"
developer_instructions = "You are a fresh, bounded implementation subagent. Follow the complete task and applicable instructions supplied in the initial message. Implement exactly the assigned scope and run targeted verification. Preserve unrelated changes and accommodate concurrent edits. Do not expand scope or spawn subagents."
~/.codex/agents/luna-low.toml
name = "luna-low"
description = "Small, straightforward, low-risk tasks such as focused lookups, extraction, formatting, and simple checks, with high reasoning as the minimum."
model = "gpt-5.6-luna"
model_reasoning_effort = "high"
developer_instructions = "You are a fresh subagent for a small bounded task. Follow the complete task and applicable instructions supplied in the initial message. Preserve unrelated work and return only the requested concise result. Do not expand scope or spawn subagents."
~/.codex/agents/deep.toml
name = "deep"
description = "Maximum-reasoning implementation for one bounded architecture, correctness, or root-cause slice."
model = "gpt-5.6-luna"
model_reasoning_effort = "max"
sandbox_mode = "danger-full-access"
developer_instructions = "You are a fresh maximum-reasoning implementation subagent for one difficult bounded slice. Follow the complete task and applicable instructions supplied in the initial message. Trace the production mechanism deeply, distinguish evidence from inference, implement the complete correction within the assigned exclusive write set, preserve unrelated and concurrent work, and report exact changed paths and static closure. Do not expand scope, stage, commit, run broad proof, or spawn subagents."
~/.codex/agents/deep-read.toml
name = "deep-read"
description = "Maximum-reasoning read-only investigation for one bounded architecture, correctness, or root-cause question."
model = "gpt-5.6-luna"
model_reasoning_effort = "max"
sandbox_mode = "read-only"
developer_instructions = "You are a fresh maximum-reasoning read-only subagent for one difficult bounded question. Follow the complete task and applicable instructions supplied in the initial message. Trace the production mechanism deeply, distinguish evidence from inference, return concrete findings with exact paths and correction boundaries, and do not modify files, expand scope, or spawn subagents."
~/.codex/AGENTS.md
Add:
## Agent Efficiency
- Before spawning any subagent, explicitly specify and guarantee the required subagent type/model. If the available interface cannot specify or guarantee that subagent type/model, abort before spawning and report the blocker. Never substitute an unspecified or same-as-orchestrator agent.
- Use smaller-than-orchestrator subagents only for independent, bounded exploration, audits, log analysis, implementation, and test execution.
- Set `agent_type` explicitly and never override its pinned model or reasoning.
- Use `fork_context=false` with Multi-Agent V1. Never inherit the parent thread history.
- Give every subagent a complete, self-contained prompt with the bounded task, applicable instructions, paths, symbols, write scope, proof requirements, and expected output.
- Keep the orchestrator responsible for decomposition, architecture, synthesis, product judgment, integration, and final proof.
- Give concurrent agents disjoint scopes and write sets.
- Do not duplicate delegated work while it is running.
- Close completed agents promptly.
Fully restart Codex and begin a new task after installing the configuration.
---
I should clarify. This WILL light your *tokens* on fire, but not your _quota_. :)
주요 댓글
r/codex
사용자들은 설정의 실효성에 대해 긍정적인 반응을 보이며, 특히 V2의 모델 상속 문제와 토큰 압축 설정에 대해 심도 있게 논의하고 있습니다.
2
내 생각엔 model_auto_compact_token_limit = 180000 설정을 쓰는 게 좋을 것 같음. 이러면 70% 수준에서 압축이 실행되어 토큰 스노우볼 효과를 제거하고 토큰 사용량을 줄여줄 거임.
2
여기 있는 사람들 중 이 설정이 실제로 효과가 있는지 확인해 본 사람이 아무도 없네. 다들 '잘 작동함, 고마워'라고만 하는데, 그건 그냥 에러가 안 났다는 뜻이지 토큰을 아꼈다는 뜻이 아니야.
실제로 확인하는 가장 확실한 방법: 그 레포에서 이미 수행했던 작업 하나를 골라서, 서브에이전트 정책을 적용했을 때와 안 했을 때 각각 한 번씩 실행해 보고 세션 로그에서 총 토큰 수를 비교해 봐. 똑같은 작업으로 해야지, 안 그러면 그냥 노이즈만 측정하는 꼴이야. 토큰뿐만 아니라 기능이 여전히 제대로 작동하는지도 확인해야 해. 안 그러면 최적화하다가...
0
토큰을 아껴주는 건 아님. '할당량을 불태우지 않는다'고 했어야 했음. 수십억 개의 토큰을 쓰게 될 거거든. 근데 요즘 토큰 가격이 80% 할인이라 상관없긴 해...
이걸 쓰라고 권장하는 건 아님. 그냥 하고 싶으면 이렇게 하라는 거지. 사실 완전히 불필요한 짓이고 5.6 버전 이전에는 자동으로 다 처리되던 거임.
1
컨텍스트 윈도우 크기를 오버라이딩하는 게 실제로 작동함?
1
ㅇㅇ. Codex는 안 된다고, API 제한이라고 말하긴 하더라. 근데 다른 사람 지침을 찾아냈음.
내가 위에 뭐 빼먹었을 수도 있음. 안 되면 댓글 달아줘, 어디에 설정되어 있는지 더 찾아볼게.
0
1
난 작업이 보통 작고 비싼 구독도 안 써서 MCP 몇 개 말고는 Codex 설정을 건드려본 적이 없음.
그건 그렇고, 이 글은 꽤 상세해 보여서 질문 몇 개 던지고 한번 시도해볼까 함.
첫째로, 에이전트는 당연히 Codex의 일부일 텐데 수동 지침을 주는 게 정말 필요함? Codex 팀에서 이미 이런 일반적인 규칙 세트를 제공하지 않나?
둘째로, 오케스트레이터를 sol high로 설정했는데 이거 유연하게 바꿀 수 있음? 그게...
-1
오케스트레이터는 작업을 스스로 완료할 수 있다고 확신하는 가장 낮은 모델이어야 함. terra로 가능하면 terra를 쓰고, sol high로 안 되면 xhigh나 max로 올리셈. luna로 가능하면 서브에이전트를 쓸 이유가 없음. 그냥 luna한테 다 시키면 푼돈밖에 안 듦.
첫 번째 질문에 답하자면, OpenAI의 서브에이전트 구현은 한동안 엉망이었음. 언젠가는 제대로 만들겠지만, 이런 설정은 기다리기 싫은 사람들을 위한 거임...
0
작동함! 감사!
0
진짜 멋진 가이드네 - 정리해줘서 고마워!
0
v2는 별로임?
-1
내 토큰을 불태우는 걸 안 좋아해서 그럼. V2는 강제로 같은 에이전트를 쓰게 만듦. 만약 sol max를 쓰고 있다면 모든 서브에이전트도 그걸 쓰게 됨.