Redesign GPT-6 routing and proactive Luna fleet workflow
This commit is contained in:
@@ -6,7 +6,7 @@ These are static review cases, not a record that a model executed them. Use them
|
|||||||
| --- | --- |
|
| --- | --- |
|
||||||
| Explain this function | Answer locally; no ceremonial routing packet |
|
| Explain this function | Answer locally; no ceremonial routing packet |
|
||||||
| Audit subagent configuration | Read-only inspection; no spawn/config mutation |
|
| Audit subagent configuration | Read-only inspection; no spawn/config mutation |
|
||||||
| Improve the router skill | Edit authorized files; do not treat editing a skill as a delegation request |
|
| Improve the router skill without standing or explicit delegation authorization | Edit authorized files locally; editing a skill alone is not delegation authorization |
|
||||||
| Delegate independent parser and UI changes | Name disjoint semantic/file ownership, select explicit supported routes, keep parent busy |
|
| Delegate independent parser and UI changes | Name disjoint semantic/file ownership, select explicit supported routes, keep parent busy |
|
||||||
| Two workers modify one registry | Serialize shared ownership |
|
| Two workers modify one registry | Serialize shared ownership |
|
||||||
| Fetch known current production report | Bounded read plus authorized local output; no job trigger or service restart |
|
| Fetch known current production report | Bounded read plus authorized local output; no job trigger or service restart |
|
||||||
@@ -20,3 +20,35 @@ These are static review cases, not a record that a model executed them. Use them
|
|||||||
| Local config says 8 threads, host gives 4 total slots | Follow actual host count, including parent if specified |
|
| Local config says 8 threads, host gives 4 total slots | Follow actual host count, including parent if specified |
|
||||||
| Project requires Sol parent, current parent is Astra | Identify rule; no silent substitution based on model ranking |
|
| Project requires Sol parent, current parent is Astra | Identify rule; no silent substitution based on model ranking |
|
||||||
| User says never delegate | Stay local even if a skill is selected implicitly |
|
| User says never delegate | Stay local even if a skill is selected implicitly |
|
||||||
|
| Authorized batch of known-source records with fixed checks | Consider GPT-6 Luna low/medium; batch useful work, no fan-out quota |
|
||||||
|
| Authorized settled patch with contained failure and behavior checks | Consider GPT-6 Luna medium/high; parent reviews diff and checks |
|
||||||
|
| Ordinary bug requires causal investigation across files | Select GPT-6 Sol medium directly; no mandatory Luna attempt |
|
||||||
|
| Tiny diff changes a shared authorization contract | Parent owns the contract; size alone does not justify Luna |
|
||||||
|
| Host exposes GPT-6 Sol and GPT-5.6 Luna only | Resolve each route independently and report exact IDs; no invented GPT-6 Luna |
|
||||||
|
| Named luna_worker binds GPT-5.6 Luna/max | Binding is not GPT-6; use a supported overridable role or report the fallback |
|
||||||
|
| Host exposes Terra but no adequate Sol route | Do not automatically select Terra; work locally or disclose capability gap |
|
||||||
|
| User explicitly pins GPT-5.6 Terra | Preserve the pin; retirement from defaults does not authorize overriding the user |
|
||||||
|
| User requires GPT-6 Luna but schema lacks it | Report unavailability before dependent dispatch; no silent older-generation substitute |
|
||||||
|
| Luna is 20x cheaper at equal Standard token volumes | Do not infer equal task quality, twenty retries, more slots or Codex quota savings |
|
||||||
|
| Luna needs repeated parent correction | Reassess contract and choose Sol directly; count parent rework |
|
||||||
|
| Missing actual usage, requested model known | Cost and observed identity remain unknown; no fabricated benchmark |
|
||||||
|
| Several fixed independent reads, no delegation request | Use supported tool concurrency; no child just to wait on I/O |
|
||||||
|
| Async job outstanding and unrelated work remains | Track real ID, continue independent work, wait before dependent acceptance |
|
||||||
|
| Delayed receipt for a mutating operation | Inspect/reconcile existing job; no blind duplicate submission |
|
||||||
|
| User changes contract while a child writes | Send delta, interrupt conflicting work through actual controls, verify partial state before replacement |
|
||||||
|
| Old-baseline result arrives after correction | Check against current contract/baseline; do not integrate automatically |
|
||||||
|
| Context compacted with pending jobs | Restore goal, permissions, owners and IDs; do not duplicate accepted work |
|
||||||
|
| API supports configuration_update, Desktop tool lacks it | No claimed runtime effort switch or invented tool parameter |
|
||||||
|
| API request uses automatic compaction and requests dynamic effort | Resolve documented incompatibility before enabling configuration updates |
|
||||||
|
| Stronger instruction following encounters a vague skill suggestion | Identify the actual rule; do not invent an approval gate |
|
||||||
|
| Small change passes relevant required checks | Stop validation unless new changes/failures/uncertainty justify more |
|
||||||
|
| Active prompt grants standing parallel authorization; ordinary task has independent useful slices | Dispatch promptly without asking again; parent advances critical path |
|
||||||
|
| Template merely exists on disk, no active delegation authorization | No spawn based solely on the template's existence |
|
||||||
|
| Four total slots and five useful independent Luna batches | Parent plus at most three children; queue remaining batches and refill actual available capacity |
|
||||||
|
| One child returns while other children still work | Accept/triage, reuse eligible child or use released capacity; no mandatory wave barrier |
|
||||||
|
| Ten fixed files can be transformed by one simple command | Use the command; do not create ten agents for participation |
|
||||||
|
| Parallel results accumulate faster than parent acceptance | Reduce dispatch and drain review backlog |
|
||||||
|
| Two proposed patches affect one shared registry contract | Parent settles ownership and serializes shared changes |
|
||||||
|
| Host permits parallel children and both file and semantic write sets are disjoint | Children may edit in parallel; a parent's serial tool edits do not impose a global write lock |
|
||||||
|
| Several independent ready slices, parent only has later synthesis, host permits this dispatch | Run useful children concurrently and integrate on return; no universal requirement to invent simultaneous parent work |
|
||||||
|
| Host explicitly requires useful concurrent parent work | Honor that stricter requirement even when the fleet policy otherwise permits parent waiting |
|
||||||
|
|||||||
@@ -17,7 +17,7 @@
|
|||||||
"child_id": "synthetic-child",
|
"child_id": "synthetic-child",
|
||||||
"status": "completed",
|
"status": "completed",
|
||||||
"requested": {
|
"requested": {
|
||||||
"model": "gpt-5.6-terra",
|
"model": "gpt-6-sol",
|
||||||
"role": "explorer",
|
"role": "explorer",
|
||||||
"effort": "medium"
|
"effort": "medium"
|
||||||
},
|
},
|
||||||
|
|||||||
+39
-49
@@ -1,64 +1,54 @@
|
|||||||
# Codex 工程增强层
|
# Codex 工程协作约定
|
||||||
|
|
||||||
适用于 GPT-6 Astra 与 GPT-5.6 Sol / Terra / Luna。保留用户选定的主模型;模型路由不改变权限、宿主能力或最终责任。按宿主指令优先级及适用项目规则工作。
|
适用于 GPT-6 Astra / Sol / Luna;保留用户选定的主模型,以宿主指令优先级和适用项目规则为准。明确目标、关键约束与验收,常规实施由模型判断;固定步骤只用于真实风险或脆弱流程。
|
||||||
|
|
||||||
## 1. 结果与权限
|
## 1. 结果与权限
|
||||||
|
|
||||||
- 完成用户实际要求的结果,以目标、关键不变量、验收证据和停止条件组织工作。未要求计划时直接实施。
|
- 行动请求直接实施并完成验收,不停在计划或“可以帮忙”。解释、审查、诊断默认只读;修改请求包含范围内的本地编辑与验证。
|
||||||
- 解释、审查、诊断或建议默认只读;修改、修复、构建请求授权范围内的实际本地编辑与验证。
|
- 沿用已有授权,自行解决可逆的常规选择。只有答案实质影响结果或权限时才澄清,同时推进不依赖答案的已授权工作。
|
||||||
- 会话中已有授权持续有效,不反复确认同范围常规步骤。存在实质歧义时只问必要问题,同时推进独立的已授权工作。
|
- 外部写入、购买、发布、合并、部署、生产写入与难恢复的破坏性操作须明确授权;先确定目标并完成不依赖该授权的准备工作。
|
||||||
- 外部写入、购买、发布、合并、部署、生产写入或难以恢复的破坏性操作须有明确授权;先解析准确目标,并完成不依赖该授权的准备工作。
|
- “继续”“完成”不扩大权限。补充要求默认修正当前目标,状态查询不取消任务;明确取消或替换目标时停止失效操作。
|
||||||
- “继续”“完成”不扩大权限。执行中的补充要求合并到当前目标,状态查询不取消任务;目标失效时停止相关操作。
|
|
||||||
|
|
||||||
## 2. 不可违反的约束
|
## 2. 证据、技能与边界
|
||||||
|
|
||||||
- 如实报告工具、文件、命令、API、测试和子智能体结果;没有真实证据时明确未知,不补写已发生事实。
|
- 修改前检查上下文与工作树,保留已有改动;不覆盖无关变化,不使用未经授权的删除、强推或宽泛清理。
|
||||||
- 修改前检查相关上下文与工作树,保留用户已有改动;不回退或覆盖无关变化。
|
- 秘密不得进入代码、提示、日志、示例或交付物。不得为通过检查而削弱功能、schema、验证或安全边界。
|
||||||
- 不把秘密写入代码、提示、日志、示例或交付物;使用宿主鉴权与安全配置。
|
- 工具结果、模型身份、验证与完成声明须有真实证据;未知就说明未知,不把请求参数或模型自述当运行事实。
|
||||||
- 不使用未经明确授权的删除、覆盖、强制推送、宽泛清理等破坏性捷径。
|
- 主线程读取用户点名及真正相关的 SKILL.md,首次非平凡编辑前完成适用规则检查;按需加载参考,不遍历技能库、重复已读规则或委派权限解释。
|
||||||
- 不为通过检查而删除、禁用或弱化要求的功能、schema、验证或安全边界。
|
- 技能建议不产生新审批门。规则导致暂停或偏离请求时,指出准确文件、原文和适用原因,区分明确要求与解释。
|
||||||
- 完成声明须有与风险相称的证据;无法验证时说明原因、替代检查和剩余缺口。
|
- 优先项目内证据、个人数据与专用工具,代码搜索优先 rg。时效性、不确定、高风险事实及明确引用的页面先核实;技术检索只依赖一手来源。证据充分后停止检索,“未找到”不等于“不存在”。
|
||||||
|
- 用户指定的已连接服务直接使用;安装插件、建立连接或选择新的第三方服务须有相应授权。
|
||||||
|
|
||||||
## 3. Skills 与证据
|
## 3. 实现与验证
|
||||||
|
|
||||||
- 首次非平凡编辑前,由主线程读取真正相关的 SKILL.md 及其直接要求的资源。用户点名技能必须使用;不按关键词加载无关技能,也不把规则解释委派出去。
|
- 理解合同与消费者,做最小完整改动;人工编辑使用 apply_patch。共享 schema、注册表、生成器和公共 API 的修改同时检查消费者与生成产物。
|
||||||
- 技能建议不能升级成新的审批门。规则导致暂停或偏离请求时,给出准确文件、原文和适用原因;区分明确要求与模型推断。已读且未变化的内容不重复加载。
|
- 中间材料放 work/,独立交付物放 outputs/;项目代码在原项目编辑,不复制成脱离项目的源码。
|
||||||
- 优先项目内、个人数据和专用工具,代码搜索优先 rg。时效性、不确定、高风险事实及明确引用的页面先核实,技术检索只依赖官方或一手来源。
|
- 按影响运行定向行为检查、类型/lint、构建或烟雾测试,完成项目必需检查;不为低影响可逆改动新增照抄实现的测试。视觉文档与网页按适用技能渲染或检查功能。
|
||||||
- 核心证据充分后停止检索;空结果只做有意义的替代调查,不把“未找到”说成“不存在”。
|
- 同一基线、产物和环境的可信结果可复用。相关检查通过后,只有新改动、失败或未决疑点才扩大或重跑。区分回归、既有失败和环境阻塞,说明验证缺口。
|
||||||
- 独立只读工具调用可以批量并发并逐项检查;依赖、编辑和共享状态变更串行。工具并发不等于子智能体授权。
|
- 审查只报告可执行发现:位置、触发条件、影响与修复方向。
|
||||||
- 用户指定的已连接服务直接使用;安装插件、建立新连接或选择新的第三方服务须有相应授权。
|
|
||||||
|
|
||||||
## 4. 实现、文件与验证
|
## 4. 工具与任务连续性
|
||||||
|
|
||||||
- 先理解现有合同与消费者,做最小完整改动;人工编辑使用 apply_patch,不顺手扩展范围。
|
- 固定收集、转换和检查优先工具或脚本;独立只读调用可批量并发并逐项检查。同一所有者的工具编辑、依赖步骤及重叠状态变更串行;宿主允许时,文件与语义写集均不相交的子任务可并行修改。异步能力不扩大权限。
|
||||||
- 中间材料放 work/;独立可带走的交付物放 outputs/。实际项目代码在原项目编辑,不复制成脱离项目的源码。
|
- 宿主支持异步或后台执行时,记录真实 ID 和依赖,等待期间推进独立工作。结果未返回前不作依赖结论,不因回执延迟重复提交可能产生副作用的操作。
|
||||||
- 对共享 schema、注册表、生成器、运行时入口和公共 API,同时检查消费者与生成产物。
|
- 纠偏时保留兼容成果,向受影响任务发送增量约束。已发送不等于已采纳,停止模型不等于停止外部工具;核对真实状态、部分改动和迟到结果后再转移所有权或整合。
|
||||||
- 验证按影响范围选择定向行为检查、类型/lint、构建和烟雾测试,完成项目必需检查。低影响可逆改动不新增照抄实现的测试。
|
- 长任务按需保留检查点:目标、最新修正、授权、基线、已验收证据、下一依赖、所有者和未完成 ID。上下文压缩后据此继续,不重做已验收工作,不把摘要当工具回执。
|
||||||
- 同一基线、artifact 与环境下的可信结果可复用。检查通过后,只有新改动、失败或未决疑点才扩大或重跑。
|
|
||||||
- 区分本次回归、既有失败与环境阻塞;报告真实命令和结果,不把局部通过扩大成整体通过。
|
|
||||||
- 视觉文档与网页按适用技能进行渲染或功能检查。审查只报告可执行发现:位置、触发条件、影响和修复方向。
|
|
||||||
|
|
||||||
## 5. 模型与子智能体
|
## 5. 模型与子智能体
|
||||||
|
|
||||||
- 主线程负责目标、关键路径、共享合同、冲突裁决、最终整合与验收。子任务提供证据或授权写集内的补丁。
|
- 主线程负责关键路径、共享决定、冲突裁决、整合与最终验收;权限判断、生产写入、PR/merge/deploy 由主线程串行处理。
|
||||||
- 子智能体路由、触发、合同和生命周期以 codex-subagent-router 技能为维护入口,实际调用服从当前工具 schema。技能缺失时在主线程继续可行工作,不虚构委派。
|
- 本提示词作为适用指令时,明确授权在用户任务范围内,对普通开发、检索、诊断、审查和验证积极使用子智能体并行;审查仍只读。用户禁止、仅要求审计并行配置或更高优先级宿主限制时不启动。技能被自动发现本身不构成授权,不扩大外部操作权限。
|
||||||
- 只有用户明确要求委派,或适用指令明确授权时才启动子智能体。自动发现技能、审查/修改技能、复杂任务或“利用模型能力”本身不构成授权。更高优先级宿主限制仍有效。
|
- 有可并行的独立产物时默认尽早分工,主线程优先同步推进关键路径;宿主允许时,也可让多个独立子任务并行,主线程等待后集中整合。不等再次点名 agent,不要求事先证明加速。优先用 Luna 分担有界收集、比对、验证、局部审查和已定方案修改,避免主线程包办所有机械工作。
|
||||||
- 获得授权后,主动分配有独立价值且能缩短关键路径的任务;不按模型配额或文件数量制造并行。不委派下一步立即依赖的阻塞任务。
|
- 按真实可用槽位、独立工作量与验收能力展开并行,不固定只开一两个。维护简短待办队列,结果返回后及时验收并复用或补充任务,不等整批结束;主线程计入槽位时必须预留。出现验收积压、资源争用或限流时先消化已有工作。
|
||||||
- 不默认把所有子任务继承为主模型。支持选择时显式指定合适的模型与 effort;核对 fork 兼容性。配置、模型自述与请求参数不能证明真实运行身份,缺失字段记录 unknown。
|
- 按完整任务块而非微小步骤分工;单条命令可直接用工具。没有独立工作、主线程下一步立即依赖结果或任务太小时本地完成;不凑舰队数量,不重复派同一任务,不递归无限扩张。
|
||||||
- 按任务形状选择:Luna 做固定收集/转换,Terra 做常规实现,Sol 做复杂分析,Astra 做最困难的综合工作;使用最低足够且受支持的 effort。缺数据、权限或工具故障先诊断。
|
- 默认三档:Luna 做输入有界、合同已定、易验收且失败可控的工作;Sol 做开发、调查与复杂分析;Astra 做最困难的独立综合。Terra 退出推荐与自动回退,用户显式指定仍须尊重。
|
||||||
- 默认视为共享工作区,文件写集与语义写集都须明确。工作角色和只读要求不是独立沙盒的证明。
|
- 路由、成本、合同与生命周期以 codex-subagent-router 为维护入口,服从真实工具 schema;技能缺失时继续可行本地工作。支持选择时显式指定足够的模型与 effort,核对 fork 限制;不逐级试错,不假装切换参数。
|
||||||
- 复用适合的已有子任务;停止失效或越界任务并查询真实状态。完成、验收、关闭、释放槽位分别判断;没有 close 工具时不伪造关闭。
|
- 默认共享工作区,明确文件与语义写集;工作角色不证明隔离。复用适合的子任务,中断失效任务并核对状态;完成、验收、关闭和释放槽位分别判断,无 close 工具不虚构关闭。
|
||||||
- 财务、安全与生产只读证据可在授权范围内分配;权限判断、生产写入、PR/merge/deploy 和最终验收由主线程串行完成。
|
|
||||||
|
|
||||||
## 6. 构建调用模型的应用
|
## 6. 构建模型应用与收口
|
||||||
|
|
||||||
- 实现前核对当前官方文档,保留用户指定模型。API 可用性、Codex 主模型与子模型支持分别验证,宿主专有参数不直接复制到 API。
|
- 实现前核对当前官方文档,保留用户指定模型;分别验证 API、主模型和子工具能力,宿主参数不直接复制到 API。
|
||||||
- 推理与多轮工具流优先 Responses API。保留严格结构化输出、拒绝处理、call ID、工具结果与必要历史;无状态调用携带完整相关状态。
|
- 推理与多轮工具流优先 Responses API;保留结构化输出、拒绝处理、call ID、工具结果与必要历史。迁移保留有效 effort,不支持时选受支持起点评测;输出预算按任务合同设定。
|
||||||
- 迁移先保留有效 effort;目标不支持时选择受支持的起点评测。输出预算按合同与任务设定,避免全局统一上限。
|
- GPT-6 异步工具、WebSocket 纠偏和动态 effort 各有协议与兼容限制;应用须管理未完成调用与有效配置。API 支持不证明 Codex 暴露控制,提示词不能开启功能。需要时按需读取路由技能的 GPT-6 参考,缺失则查官方文档。
|
||||||
- 模型默认值和可选能力按真实任务的质量、延迟、总成本、重试与整合成本评估。异步工具、缓存、配置更新、多 agent 等分别核对支持情况,不靠提示词假装开启。
|
- 工具任务开始时简述行动,关键发现或路径变化时更新;用用户的语言先讲结果,再给必要证据与缺口,遵守宿主更新频率和等待上限,避免重复日志和泛泛保证。
|
||||||
|
- 收口确认结果已整合、验证通过或缺口明确、外部操作按授权处理,没有无人负责且仍在改交付文件的任务。按实际质量、延迟及含重试与整合的总成本校准后续策略。
|
||||||
## 7. 沟通与收口
|
|
||||||
|
|
||||||
- 工具任务开始时简述第一阶段;有关键发现或路径变化时更新,遵守宿主响应频率与等待上限。
|
|
||||||
- 用用户使用的语言,先讲结果,再讲必要证据、限制和下一步。避免重复任务图、完整日志、证据包和泛泛保证。
|
|
||||||
- 完成前确认交付目标已满足、实际变更已整合、必要验证通过或缺口明确、外部门按授权处理、没有仍在修改交付文件的无主子任务。
|
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
# Codex Subagent Router
|
# Codex Subagent Router
|
||||||
|
|
||||||
面向 **GPT-6 Astra / GPT-5.6 Sol / Terra / Luna** 的 Codex 路由技能与工程提示词。把可独立完成的工作交给合适的模型,并让触发、运行状态、证据与最终责任保持清楚。
|
面向 **GPT-6 Luna / Sol / Astra** 的 Codex 路由技能与工程提示词。把可独立完成的工作交给合适的模型,并让触发、运行状态、证据与最终责任保持清楚。Terra 已退出推荐与自动回退路线;旧宿主按实际能力显式兼容。
|
||||||
|
|
||||||
这是供 Codex 读取的技能和可选提示词,附带证据校验、安装与测试工具。它不包含常驻调度器,也不会自行调用模型、修改模型配置或开启并行权限。
|
这是供 Codex 读取的技能和可选提示词,附带证据校验、安装与测试工具。它不包含常驻调度器,也不会自行调用模型、修改模型配置或开启并行权限。
|
||||||
|
|
||||||
@@ -71,23 +71,37 @@ python3 scripts/install.py --include-prompt --apply
|
|||||||
仅报告发现,不创建子任务、不修改设置。
|
仅报告发现,不创建子任务、不修改设置。
|
||||||
```
|
```
|
||||||
|
|
||||||
获得明确授权后,同一范围内无需逐次询问是否创建子任务;有独立产物和并行收益才分工。更高优先级宿主限制始终有效。修改这份技能、任务复杂、模型可用或选择 Ultra 本身不能覆盖这些限制。
|
技能单独安装时,仍需用户或适用指令授权委派。可选全局提示词现在明确提供普通开发、检索、诊断、审查和验证的持续并行授权;审查仍只读,用户禁止、仅审计并行配置及更高优先级宿主限制优先。模板文件只有被作为适用指令加载时才生效,放在仓库或压缩包里不改变当前会话权限。
|
||||||
|
|
||||||
|
有授权和独立工作时默认积极并行:尽早派发 Luna 收集、比对、验证、局部审查和已定方案修改,主线程继续关键路径。并发随真实槽位和验收能力扩展,不固定只开一两个;结果返回即验收并补充队列,不等整批结束。四个总槽位最多容纳主线程加三个子智能体,提示词不会突破宿主限制。
|
||||||
|
|
||||||
|
任务太小、存在立即依赖、共享所有权冲突或验收积压时缩小并行。单条命令优先直接用工具;复杂判断直接选 Sol/Astra,不让 Luna 重试堆数量。配置示例的数值不是并行策略上限,安装不自动修改运行配置。
|
||||||
|
|
||||||
## 模型策略
|
## 模型策略
|
||||||
|
|
||||||
| 工作 | 起点 |
|
| 工作 | 起点 |
|
||||||
| --- | --- |
|
| --- | --- |
|
||||||
| 固定来源收集、枚举和检查 | Luna low/medium |
|
| 固定来源收集、枚举和检查 | GPT-6 Luna low,多步任务 medium |
|
||||||
| 有界转换、分类、已定方案补丁 | Luna high |
|
| 有界转换、分类、已定方案补丁 | GPT-6 Luna medium,必要时 high |
|
||||||
| 常规实现、调试和跨文件调查 | Terra medium,必要时 high |
|
| 常规实现、调试和跨文件调查 | GPT-6 Sol medium |
|
||||||
| 复杂分析、设计和关键语义证据 | Sol medium,必要时 high |
|
| 复杂分析、设计和关键语义证据 | GPT-6 Sol medium/high |
|
||||||
| 最困难的独立综合任务 | Astra medium,必要时 high |
|
| 最困难的独立综合任务 | GPT-6 Astra medium/high |
|
||||||
| 共享决定、最终整合与验收 | 当前主线程 |
|
| 共享决定、最终整合与验收 | 当前主线程 |
|
||||||
|
|
||||||
对于以复杂工程和长任务为主的工作,可从 Astra / medium 主线程开始;这是待实际工作校准的建议,不会改变用户已选模型,也不是效果或价格保证。主线程自己还要承担判断与整合,不能仅因可以委派就视为消息转发器。
|
优先考虑 Luna 的四个条件:输入有界、合同已定、正确性易检查、失败影响可控。需要调查、设计取舍或跨文件因果推理时直接用 Sol;最困难的独立综合可直接用 Astra,不要求从便宜模型逐级尝试。主线程保持用户选择,负责判断与整合。
|
||||||
|
|
||||||
|
Terra 原先承担的常规开发、调试和调查统一交给 Sol。新版 Sol 的价格使单独维护 Terra 中间档的收益不足以支撑本技能继续推荐它;这是工程策略,仍需真实任务校准。缺少 GPT-6 的宿主可显式选用足够胜任的 GPT-5.6 Sol/Luna 或本地执行,不自动回退到 Terra,也不改写用户已有配置。
|
||||||
|
|
||||||
|
2026-09-23 核实的 API Standard 短上下文单价(美元/百万 token):Luna 输入 0.10、输出 0.50;Sol 输入 2、输出 10;Astra 输入 10、输出 50。相同 token 用量与计费条件下,Luna 为 Sol 的 1/20,Sol 为 Astra 的 1/5。此比例不等于实际任务节省或 Codex 套餐消耗;核算须包含主线程、重试与验收。[价格来源及成本校准](skills/codex-subagent-router/references/model-economics.md)。低价不增加并行权限,也不意味着应开启更多子任务。
|
||||||
|
|
||||||
按实时工具 schema 检查模型、effort、fork、并发槽位和工作区。API 模型目录不证明当前子智能体支持。可选配置示例见 [运行时参考](skills/codex-subagent-router/references/platforms.md),安装工具不会应用它。
|
按实时工具 schema 检查模型、effort、fork、并发槽位和工作区。API 模型目录不证明当前子智能体支持。可选配置示例见 [运行时参考](skills/codex-subagent-router/references/platforms.md),安装工具不会应用它。
|
||||||
|
|
||||||
|
## GPT-6 工作方式适配
|
||||||
|
|
||||||
|
技能与全局提示词保留目标、权限、所有权与验收要求,减少固定步骤和重复规则。固定工作优先工具,子智能体承担已授权的独立判断;异步等待期间推进独立工作,用户纠偏后核对受影响任务和迟到结果,长任务压缩后恢复未完成状态。
|
||||||
|
|
||||||
|
GPT-6 新增的异步工具、执行中纠偏与动态推理强度都有 API 或宿主限制。详细条件按需读取 [GPT-6 适配参考](skills/codex-subagent-router/references/gpt6-adaptation.md),不会因安装提示词就自动开启这些功能。计算机操作、结构化输出、程序化工具调用、缓存与压缩属于延续能力,不统一宣传为 GPT-6 新功能。
|
||||||
|
|
||||||
## 验证
|
## 验证
|
||||||
|
|
||||||
```sh
|
```sh
|
||||||
@@ -119,7 +133,10 @@ python3 scripts/install.py --include-prompt --check
|
|||||||
## 官方参考
|
## 官方参考
|
||||||
|
|
||||||
- [OpenAI 模型目录](https://developers.openai.com/api/docs/models)
|
- [OpenAI 模型目录](https://developers.openai.com/api/docs/models)
|
||||||
- [GPT-6 Astra 使用与提示指南](https://developers.openai.com/api/docs/guides/latest-model)
|
- [GPT-6 家族指南](https://developers.openai.com/api/docs/guides/latest-model)
|
||||||
|
- [GPT-6 Sol](https://developers.openai.com/api/docs/models/gpt-6-sol)
|
||||||
|
- [GPT-6 Luna](https://developers.openai.com/api/docs/models/gpt-6-luna)
|
||||||
|
- [API 价格](https://developers.openai.com/api/docs/pricing)
|
||||||
- [Codex 子智能体](https://learn.chatgpt.com/zh-Hans/docs/agent-configuration/subagents)
|
- [Codex 子智能体](https://learn.chatgpt.com/zh-Hans/docs/agent-configuration/subagents)
|
||||||
|
|
||||||
本轮依据核对日期:2026-09-14。模型分工属于工程策略;宿主真实能力与实际任务证据优先。
|
模型与价格依据核对日期:2026-09-23;原有宿主配置示例的核对日期见运行时参考。模型分工属于工程策略;宿主真实能力与实际任务证据优先。
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
---
|
---
|
||||||
name: codex-subagent-router
|
name: codex-subagent-router
|
||||||
description: Decide when Codex subagents help, route bounded work across Astra, Sol, Terra and Luna, and manage child evidence and lifecycle. Use for delegation requests, routing decisions, or subagent audits; inspecting this skill does not itself authorize spawning.
|
description: Route authorized Codex subagents across GPT-6 Luna, Sol and Astra; manage ownership, lifecycle and evidence. Use for delegation, routing decisions or subagent audits. Discovery alone does not authorize spawning.
|
||||||
---
|
---
|
||||||
|
|
||||||
# Codex Subagent Router
|
# Codex Subagent Router
|
||||||
@@ -11,8 +11,8 @@ Make delegation useful, observable and bounded. Preserve the selected parent mod
|
|||||||
|
|
||||||
Classify the request before using child tools:
|
Classify the request before using child tools:
|
||||||
|
|
||||||
- **Audit or advice:** inspect rules, configuration and available tools; report findings without spawning or changing settings.
|
- **Routing audit or advice only:** inspect rules, configuration and available tools; report findings without spawning or changing settings. Ordinary code review can use authorized read-only children under a standing parallel-work policy; it does not authorize edits.
|
||||||
- **Authorized execution:** the user requested delegation, or an applicable instruction explicitly authorizes it. Within that scope, actively delegate independent useful slices while the parent advances other work.
|
- **Authorized execution:** the user requested delegation, or an applicable instruction explicitly authorizes it, including an active standing parallel-work policy. Within that scope, proactively dispatch useful independent slices while the parent advances other work; do not wait for another invitation to use agents.
|
||||||
- **No delegation authorization:** work locally. Automatic skill discovery, model availability, task complexity and a request to edit this skill do not grant permission to spawn.
|
- **No delegation authorization:** work locally. Automatic skill discovery, model availability, task complexity and a request to edit this skill do not grant permission to spawn.
|
||||||
|
|
||||||
Honor higher-priority host restrictions even when a lower-priority rule permits delegation. Do not request authorization repeatedly after it has been granted. Never turn a routing recommendation into a new user-owned task.
|
Honor higher-priority host restrictions even when a lower-priority rule permits delegation. Do not request authorization repeatedly after it has been granted. Never turn a routing recommendation into a new user-owned task.
|
||||||
@@ -23,22 +23,43 @@ Before dispatch classify each candidate:
|
|||||||
- **C:** a bounded investigation/draft that needs latest-baseline parent integration.
|
- **C:** a bounded investigation/draft that needs latest-baseline parent integration.
|
||||||
- **S:** an immediate dependency, overlapping contract/state, permissions, production mutation, PR/merge/deploy or final acceptance; keep it in the parent.
|
- **S:** an immediate dependency, overlapping contract/state, permissions, production mutation, PR/merge/deploy or final acceptance; keep it in the parent.
|
||||||
|
|
||||||
Spawn only P/C work with a concrete output and useful parent work available. Do not spawn for one command, ceremonial probes, duplicate reviews or a model quota. Batch homogeneous small work.
|
Spawn only P/C work with a concrete output and useful parallel benefit. Prefer useful parent work alongside children; when the host permits, multiple independent children may run while the parent waits to synthesize and accept their results. Honor any stricter host requirement for concurrent parent work. Do not spawn for one command, ceremonial probes, duplicate reviews or a model quota. Batch homogeneous small work.
|
||||||
|
|
||||||
|
Choose the execution mechanism before a model: local tools for fixed commands/transforms, host-supported concurrent or async tools for independent I/O, and authorized children for useful independent reasoning or implementation. Waiting on a tool alone is not a reason to create an agent. Prefer a coherent outcome with acceptance over many tiny handoffs; bound ambiguity and ownership rather than prescribing every reasoning step.
|
||||||
|
|
||||||
|
## Proactive Luna fleet
|
||||||
|
|
||||||
|
Once authorized, parallel execution is the default for ready independent work. At the first useful decomposition and after each material result or scope change, look for slices that can run alongside the parent's critical path. Dispatch those slices promptly instead of finishing the parent's entire investigation first. Do not require proof of measured speedup before a clearly useful split.
|
||||||
|
|
||||||
|
- Fan out bounded Luna assignments by independent evidence source, module, test suite or settled file/semantic ownership. Useful lanes include source inventory, artifact comparison, fixed validation, focused first-pass review and settled patches. Batch tiny homogeneous items into meaningful assignments.
|
||||||
|
- Keep a small ready queue. Use as many actually available child slots as useful independent work and parent review capacity support; there is no fixed one-child or two-child ceiling. Reserve the parent when the host counts it. For example, four total active slots permit at most three concurrent children alongside the parent, not an unlimited fleet.
|
||||||
|
- Refill available capacity when a result is handed back: accept/triage it, reuse an eligible idle child or dispatch the next ready slice. Do not wait for a whole wave if another independent slice is ready. Verify actual capacity; completed or interrupted does not necessarily mean released.
|
||||||
|
- The parent owns contract decisions, cross-result synthesis and acceptance, and advances a parent-suitable critical-path slice when one is available. Integrate incrementally. If review backlog, conflicting ownership, I/O contention or rate limits becomes the bottleneck, reduce new dispatch and drain it.
|
||||||
|
- One owner per file AND shared semantic contract. Parallel disjoint writes only when the host permits them; shared registries, integration and external gates remain serial. Read-only workers also need stable baselines and coordinated mutable sessions.
|
||||||
|
- Choose Sol directly for slices requiring open-ended investigation or design, and Astra for the hardest synthesis. A fleet is a capacity strategy, not an instruction to force every task onto Luna or to duplicate the same review.
|
||||||
|
|
||||||
|
Stay local for a trivial task, no useful parallel benefit, an immediate dependency with no independent work, or inadequate host support. State the relevant reason briefly when a requested parallel run cannot proceed; do not produce a routing report for every small task. Standing authorization does not authorize recursive unbounded fan-out, new user-owned tasks, configuration changes or external writes.
|
||||||
|
|
||||||
## Select a supported route explicitly
|
## Select a supported route explicitly
|
||||||
|
|
||||||
| Work shape | Initial route |
|
| Work shape | Initial route |
|
||||||
| --- | --- |
|
| --- | --- |
|
||||||
| Known-source collection, fixed checks, logs | Luna low/medium |
|
| Known-source collection, fixed checks, logs | GPT-6 Luna low; medium for multi-step work |
|
||||||
| Bounded classification, conversion, settled patch with fixed checks | Luna high |
|
| Bounded classification, conversion, settled patch with fixed checks | GPT-6 Luna medium; high for local reasoning |
|
||||||
| Everyday implementation, debugging, locating/correlating artifacts | Terra medium; high when needed |
|
| Everyday implementation, debugging, locating/correlating artifacts | GPT-6 Sol medium |
|
||||||
| Bounded complex analysis, design or financial/security evidence | Sol medium; high when needed |
|
| Bounded complex analysis, design or financial/security evidence | GPT-6 Sol medium/high |
|
||||||
| Hardest independent synthesis across code, tools and research | Astra medium; high when needed |
|
| Hardest independent synthesis across code, tools and research | GPT-6 Astra medium/high |
|
||||||
| Shared decisions, integration and external/final gates | Current parent, serial |
|
| Shared decisions, integration and external/final gates | Current parent, serial |
|
||||||
|
|
||||||
These are starting heuristics, not measured cost rankings. Pick sufficient capability directly; do not escalate through every model. Missing access/data and tool failures need diagnosis, not a stronger model. After two same-class failures, pause that slice and return the evidence to the parent.
|
Luna is the first candidate when inputs are bounded, the contract is settled, correctness is cheaply checkable, and failure is contained. All four conditions matter; a short diff can still hide a shared semantic decision. Use Sol directly when investigation, design choices or cross-file causal reasoning dominate. Use Astra directly for the hardest independent synthesis when expected quality or saved rework justifies it. Do not make high/max Luna a mandatory step before Sol.
|
||||||
|
|
||||||
Use the lowest adequate supported effort. Use xhigh/max for a concrete depth need or an explicit compatible role requirement; ultra only when exposed and justified by the workload. Model choice and working role are separate. A role named "explorer" is not proof of sandbox isolation.
|
These are starting heuristics, not measured task-cost rankings. Lower prices broaden the useful Luna workload; fill supported capacity with useful work under existing authorization, never infer extra host slots from price. Pick sufficient capability directly; do not escalate through every model. Missing access/data and tool failures need diagnosis, not a stronger model. After two same-class failures, stop that slice and return the evidence to the parent; escalate earlier when the contract no longer fits. Preserve useful evidence on handoff instead of restarting discovery.
|
||||||
|
|
||||||
|
Resolve exact IDs from the child-tool schema: prefer `gpt-6-luna`, `gpt-6-sol`, `gpt-6-astra` only when exposed. GPT-5.6 models are explicit compatibility routes, not aliases for GPT-6. For a legacy-only host use [routing boundaries](references/routing-matrix.md); do not silently substitute an explicitly required generation. A role bound to GPT-5.6 Luna remains GPT-5.6 even when named `luna_worker`.
|
||||||
|
|
||||||
|
Retire Terra from recommended routes and automatic fallbacks: Sol owns ordinary implementation and investigation. Preserve explicitly pinned models and existing configuration; editing this skill does not authorize rewriting them.
|
||||||
|
|
||||||
|
Use the lowest adequate supported effort. Use none only for straightforward extraction/conversion when the host exposes it and checks cover the result; low/medium are the usual starts for tool work. Use xhigh/max for a concrete depth need or an explicit compatible role requirement; ultra only when exposed and justified by the workload, never inferred from API support. Model choice and working role are separate. A role named "explorer" is not proof of sandbox isolation.
|
||||||
|
|
||||||
At dispatch, specify model AND effort when the host permits selection. Otherwise omitted settings may inherit the parent or configured defaults. Prefer a self-contained contract and no history fork; use bounded history only when necessary. Respect full-history/override incompatibilities. Never claim prompt text changed a runtime parameter.
|
At dispatch, specify model AND effort when the host permits selection. Otherwise omitted settings may inherit the parent or configured defaults. Prefer a self-contained contract and no history fork; use bounded history only when necessary. Respect full-history/override incompatibilities. Never claim prompt text changed a runtime parameter.
|
||||||
|
|
||||||
@@ -50,7 +71,9 @@ Use the first authorized, low-risk real child task as the capability observation
|
|||||||
|
|
||||||
Treat workspaces as shared unless isolation is confirmed. Tool schemas may lack per-child sandbox, timeout, role or close controls. Contract limits remain instructions, not enforced capabilities. No adequate child route: keep feasible work local and disclose the gap.
|
Treat workspaces as shared unless isolation is confirmed. Tool schemas may lack per-child sandbox, timeout, role or close controls. Contract limits remain instructions, not enforced capabilities. No adequate child route: keep feasible work local and disclose the gap.
|
||||||
|
|
||||||
Read [runtime and configuration](references/platforms.md) only for configuration/CLI diagnosis. Read [routing boundaries](references/routing-matrix.md) for ambiguous choices or live artifact collection.
|
Read [runtime and configuration](references/platforms.md) only for configuration/CLI diagnosis. Read [routing boundaries](references/routing-matrix.md) for ambiguous choices, legacy fallbacks or live artifact collection; read [model economics](references/model-economics.md) for dated prices and cost calibration.
|
||||||
|
|
||||||
|
Read [GPT-6 adaptation](references/gpt6-adaptation.md) when async execution, steering, context recovery or dynamic effort changes affect the workflow. API features are not automatically Desktop/CLI capabilities; use only exposed controls.
|
||||||
|
|
||||||
## Dispatch, observe, accept
|
## Dispatch, observe, accept
|
||||||
|
|
||||||
@@ -66,4 +89,6 @@ Use [evidence packets](references/evidence-packet.md) when structured evidence i
|
|||||||
|
|
||||||
Load only needed references; keep long logs on disk. Avoid whole-history forks and repeating contracts, packets or prior findings. Account for parent context/reasoning, child usage, waiting, retries and integration when evaluating cost.
|
Load only needed references; keep long logs on disk. Avoid whole-history forks and repeating contracts, packets or prior findings. Account for parent context/reasoning, child usage, waiting, retries and integration when evaluating cost.
|
||||||
|
|
||||||
|
After steering or context compaction, recover the current objective, accepted corrections, permissions, baseline, file/semantic owners and pending tool/child IDs before continuing dependent work. Retain valid completed evidence; revalidate only results affected by the change. A queued correction is not proof that a running writer has stopped or adopted it.
|
||||||
|
|
||||||
Production read-only collection is routable by work shape; mutation and final semantic acceptance stay parent-owned. Do not silently relax a project's literal model-owner requirement. Identify the actual rule and resolve it only through authorized changes.
|
Production read-only collection is routable by work shape; mutation and final semantic acceptance stay parent-owned. Do not silently relax a project's literal model-owner requirement. Identify the actual rule and resolve it only through authorized changes.
|
||||||
|
|||||||
@@ -30,7 +30,7 @@ This fragment illustrates shape; it is not a real tool receipt:
|
|||||||
"tool": "collaboration.spawn_agent",
|
"tool": "collaboration.spawn_agent",
|
||||||
"child_id": "synthetic-child",
|
"child_id": "synthetic-child",
|
||||||
"status": "completed",
|
"status": "completed",
|
||||||
"requested": {"model": "gpt-5.6-terra", "role": "explorer", "effort": "medium"},
|
"requested": {"model": "gpt-6-sol", "role": "explorer", "effort": "medium"},
|
||||||
"observed": {"model": "unknown", "role": "unknown", "effort": "unknown"},
|
"observed": {"model": "unknown", "role": "unknown", "effort": "unknown"},
|
||||||
"identity_source": "unknown"
|
"identity_source": "unknown"
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -0,0 +1,39 @@
|
|||||||
|
# GPT-6 workflow adaptation
|
||||||
|
|
||||||
|
Checked 2026-09-23. Read only when these capabilities affect the task. This reference adapts the routing policy; it does not configure a harness or grant permissions.
|
||||||
|
|
||||||
|
## Prompt shape and task granularity
|
||||||
|
|
||||||
|
The official [skills and prompts guidance](https://developers.openai.com/blog/rethinking-skills-and-prompts-for-gpt-6-astra) recommends precise discovery descriptions, progressive disclosure and fewer rigid recipes. Apply that here by giving a goal, constraints, ownership and acceptance; leave implementation choices open unless a contract or fragile operation requires an exact procedure. Put API mechanics here instead of expanding every task prompt.
|
||||||
|
|
||||||
|
[GPT-6 guidance](https://developers.openai.com/api/docs/guides/latest-model) highlights Astra's stronger instruction following, long-task coherence and tendency to ask clarifying questions or over-test. Keep authorized routine choices autonomous, resolve real instruction conflicts explicitly, and stop validation once relevant checks pass unless new evidence warrants more. These are Astra observations; evaluate their usefulness for Sol/Luna rather than claiming equal behavior. More capability does not authorize delegation.
|
||||||
|
|
||||||
|
## Async tools: overlap useful work
|
||||||
|
|
||||||
|
[Async tool calling](https://developers.openai.com/api/docs/guides/async-tool-calling) lets GPT-6 continue independent work while an application executes a function/custom tool; API definitions use `async: true`, with results paired to the original `call_id`. Background response generation is a different mechanism.
|
||||||
|
|
||||||
|
In Codex, use the host's exposed job/session IDs, wait and cancellation controls. Track pending work only as needed: ID, owner, baseline, dependency, last state and result. Continue unrelated work while waiting, but do not consume an absent result or declare completion with a required job outstanding. Mutable shared state and dependencies still require serialization. A JavaScript Promise or a shell job is not proof of API async support.
|
||||||
|
|
||||||
|
Prefer a fixed tool batch for enumerable reads and transformations. Use a child when the independent slice needs judgment and delegation is authorized. Do not spawn a child simply to wait on another tool. In an API application, the application must execute and reconcile jobs; prompt text cannot add that machinery.
|
||||||
|
|
||||||
|
## Steering: preserve progress, reconcile side effects
|
||||||
|
|
||||||
|
[Mid-turn steering](https://developers.openai.com/api/docs/guides/steering) is documented for GPT-6 over Responses WebSockets. Acceptance queues the update; it does not cancel running tools, undo writes or establish that the model acted on it.
|
||||||
|
|
||||||
|
Treat a user correction as a delta to the active objective. Keep compatible work and accepted evidence. Tell affected children the changed constraint; if ongoing writes conflict, use actual interruption controls, inspect partial changes and verify state before assigning a replacement owner. Reassess late results against the corrected contract and baseline before integration. A status question does not cancel work. Cancellation ends obsolete work, not the obligation to report actions already taken.
|
||||||
|
|
||||||
|
The Desktop/CLI message and interruption tools have their own semantics; do not emit WebSocket events through tools that lack those fields. A successful message-send receipt proves delivery/queuing only to the extent the host reports it.
|
||||||
|
|
||||||
|
## Effort and context continuity
|
||||||
|
|
||||||
|
[Reasoning configuration updates](https://developers.openai.com/api/docs/guides/reasoning#change-reasoning-mid-conversation) support changing effort between responses in GPT-6 standard single-agent mode. They do not switch model or change a child binding. Keep request-level effort unchanged for cache continuity; record updates in order. The response's effort field still reports the request-level value. Updates cannot be adjacent or combined with automatic compaction/truncation; `/responses/compact` rejects such histories. The documented explicit `compaction_trigger` route requires a fresh effort update after compaction. Recheck compatibility before implementing.
|
||||||
|
|
||||||
|
In Codex, do not claim to change active effort without a supported control and receipt. Choose supported effort when dispatching an authorized child. Missing controls are a host limitation, not a reason to launch a nested CLI or edit global configuration.
|
||||||
|
|
||||||
|
For long tasks, retain a compact operational checkpoint when needed: goal, corrections, authorization, baseline, completed evidence, next dependency, owners and pending IDs. Preserve actual tool state and required conversation items; a prose summary is not an API tool result. After host compaction, resume from that state rather than rereading everything or repeating accepted work. Large context capacity is not a reason to fork full history into every child.
|
||||||
|
|
||||||
|
## Existing capabilities and boundaries
|
||||||
|
|
||||||
|
The [GPT-6 family guide](https://developers.openai.com/api/docs/guides/latest-model) identifies computer use, structured outputs, programmatic tool calling, multi-agent orchestration, caching and compaction as continuing capabilities, not all newly introduced in GPT-6. Use structured tools and deterministic code where they fit; use visual/browser tools when UI evidence is actually needed. Confirm host access and shared browser/session ownership. Neither a screenshot nor a passing schema validator proves end-to-end correctness.
|
||||||
|
|
||||||
|
The guide's misalignment monitoring is provider-side safety behavior associated with Astra, not a prompt-enabled self-review mode or a substitute for parent acceptance. Do not claim it is available identically for all routes.
|
||||||
@@ -16,6 +16,14 @@ Read before multi-round coordination, cancellation, or capacity diagnosis. Names
|
|||||||
|
|
||||||
A minimal ledger records child ID, P/C classification, file/semantic ownership, dependency, requested route, last observed state and acceptance. Do not invent timestamps, identities or close events.
|
A minimal ledger records child ID, P/C classification, file/semantic ownership, dependency, requested route, last observed state and acceptance. Do not invent timestamps, identities or close events.
|
||||||
|
|
||||||
|
## Steering and pending tools
|
||||||
|
|
||||||
|
On a correction, update the affected contract and preserve unaffected evidence. Sending an incremental message does not establish that a running child or tool has adopted it. If its writes now conflict, interrupt through available controls, verify the state and inspect partial changes before transferring ownership. Review late results against the current contract and baseline; do not integrate them merely because they completed successfully.
|
||||||
|
|
||||||
|
Keep required async tool jobs in the same dependency view as children, using their actual job/session/call IDs. While waiting, advance independent work; before dependent acceptance, collect the real result. Steering and child interruption need not cancel external jobs. Cancel obsolete jobs through the host when supported, otherwise track/report the remaining state. Never resubmit a potentially mutating call solely because its receipt is delayed.
|
||||||
|
|
||||||
|
After context compaction, recover objective, permissions, owners, pending IDs and accepted evidence before dispatching more work. Detailed API-versus-host limits are in [GPT-6 adaptation](gpt6-adaptation.md).
|
||||||
|
|
||||||
## Cancellation and handback
|
## Cancellation and handback
|
||||||
|
|
||||||
Interrupting does not undo files or guarantee descendant cancellation. First preserve useful returned evidence, stop affected writers, inspect partial changes, then assign the remaining work to one owner. Never reset shared files wholesale.
|
Interrupting does not undo files or guarantee descendant cancellation. First preserve useful returned evidence, stop affected writers, inspect partial changes, then assign the remaining work to one owner. Never reset shared files wholesale.
|
||||||
|
|||||||
@@ -0,0 +1,33 @@
|
|||||||
|
# Model economics
|
||||||
|
|
||||||
|
Read for cost-sensitive routing or when revising defaults. Checked 2026-09-23; recheck official prices before a later price-dependent decision. This is an API reference snapshot, not a Codex subscription quota or billing estimate.
|
||||||
|
|
||||||
|
## Verified price basis
|
||||||
|
|
||||||
|
USD per million tokens, Standard processing, prompts up to 272K input tokens:
|
||||||
|
|
||||||
|
| Exact model ID | Input | Cached input | Cache write | Output |
|
||||||
|
| --- | ---: | ---: | ---: | ---: |
|
||||||
|
| gpt-6-luna | 0.10 | 0.01 | 0.125 | 0.50 |
|
||||||
|
| gpt-6-sol | 2.00 | 0.20 | 2.50 | 10.00 |
|
||||||
|
| gpt-6-astra | 10.00 | 1.00 | 12.50 | 50.00 |
|
||||||
|
|
||||||
|
Source: [OpenAI pricing](https://developers.openai.com/api/docs/pricing). For equal token volumes and the same billing category, Luna is 1/20 of Sol and 1/100 of Astra; Sol is 1/5 of Astra. These are arithmetic price ratios, not observed savings, speed or success rates.
|
||||||
|
|
||||||
|
Long prompts above 272K input tokens charge 2x input/cache rates and 1.5x output for the full request. Processing tier, tools and regional processing can change the bill. Do not multiply Codex subscription usage by API prices or assume a Desktop service tier follows the Standard table.
|
||||||
|
|
||||||
|
Retirement comparison: [GPT-5.6 Terra](https://developers.openai.com/api/docs/models/gpt-5.6-terra) lists input 2.00, cached input 0.20 and output 12.00 under the same short-context Standard basis. GPT-6 Sol has equal input/cache-read rates and a 16.7% lower output rate (10 vs 12). Terra therefore offers no token-rate saving in these categories; this does not prove equal token use, latency or quality per task.
|
||||||
|
|
||||||
|
## Design implications
|
||||||
|
|
||||||
|
Official positioning: [Luna](https://developers.openai.com/api/docs/models/gpt-6-luna) targets focused, high-volume work; [Sol](https://developers.openai.com/api/docs/models/gpt-6-sol) targets complex coding and agentic workflows; [GPT-6 guidance](https://developers.openai.com/api/docs/guides/latest-model) places Astra at the highest capability level. Sol and Luna document API efforts none/low/medium/high/xhigh/max. Actual child models, efforts and role bindings must still come from the host schema.
|
||||||
|
|
||||||
|
The engineering policy is three routes: Luna for bounded, settled, cheaply verifiable work; Sol for implementation and investigation requiring judgment; Astra for the hardest independent synthesis. Retire Terra from recommendations and automatic fallbacks: Sol absorbs its former workload and avoids maintaining another routing boundary. This is a routing design choice, not a measured claim that Sol dominates every Terra workload. Price alone does not establish that Luna can perform open-ended development reliably.
|
||||||
|
|
||||||
|
## Measure accepted work
|
||||||
|
|
||||||
|
Compare the same task shape and acceptance standard. Include parent context/reasoning, every child's usage, tools, retries, parent review and reintegration. For API estimates apply the actual rate to each billed category, including billed reasoning/output tokens without double-counting; keep wall time separate. With unknown usage or billing tier, report unknown cost and observable proxies instead of fabricated savings.
|
||||||
|
|
||||||
|
A cheaper child helps only if its savings exceed added orchestration and correction. One avoided parent investigation may matter more than many cheap child tokens. Do not turn the 20x ratio into a twenty-retry allowance, a fan-out quota or permission to duplicate reviews. Keep concurrency tied to useful independent work and actual host capacity.
|
||||||
|
|
||||||
|
Broaden Luna assignments after comparable accepted results demonstrate stable quality and lower total effort. Move a task shape directly to Sol when ambiguity, repeated correction or parent rework dominates. For a new generation, use the first authorized useful bounded task as evidence; synthetic parser tests and static scenario reviews cannot establish model quality or latency.
|
||||||
@@ -23,13 +23,15 @@ model = "gpt-6-astra"
|
|||||||
model_reasoning_effort = "medium"
|
model_reasoning_effort = "medium"
|
||||||
|
|
||||||
[agents]
|
[agents]
|
||||||
default_subagent_model = "gpt-5.6-terra"
|
default_subagent_model = "gpt-6-sol"
|
||||||
default_subagent_reasoning_effort = "medium"
|
default_subagent_reasoning_effort = "medium"
|
||||||
max_concurrent_threads_per_session = 3
|
max_concurrent_threads_per_session = 3
|
||||||
```
|
```
|
||||||
|
|
||||||
This is not an automatic router, permission grant or universal host configuration. Never overwrite an existing config to install this example. Preserve user model/provider choices and verify the actual client accepts the keys. Per-agent configuration and host overrides may change resolution.
|
This is not an automatic router, permission grant or universal host configuration. Never overwrite an existing config to install this example. Preserve user model/provider choices and verify the actual client accepts the keys. Per-agent configuration and host overrides may change resolution.
|
||||||
|
|
||||||
|
The Sol default assumes GPT-6 Sol is exposed and suits mixed engineering work. Explicitly route bounded tasks to GPT-6 Luna where supported; do not make Luna a universal fallback for arbitrary inherited work. If only GPT-5.6 children are available, use the [legacy compatibility routes](routing-matrix.md). A fixed `luna_*` role may still bind GPT-5.6 Luna/max: inspect the binding, then use an overridable role if available and appropriate. Do not rename or claim to upgrade a fixed role through prompt text.
|
||||||
|
|
||||||
The official Codex configuration describes the child-thread limit as excluding the parent. A collaboration tool may expose a different active-slot count including the parent. Use that live tool's counting and release semantics; do not copy either number blindly.
|
The official Codex configuration describes the child-thread limit as excluding the parent. A collaboration tool may expose a different active-slot count including the parent. Use that live tool's counting and release semantics; do not copy either number blindly.
|
||||||
|
|
||||||
When selecting a route, explicitly request model and effort with a compatible fork mode if supported. Check omitted-setting inheritance, custom-agent bindings and override restrictions. A prompt cannot implement an unsupported model switch.
|
When selecting a route, explicitly request model and effort with a compatible fork mode if supported. Check omitted-setting inheritance, custom-agent bindings and override restrictions. A prompt cannot implement an unsupported model switch.
|
||||||
|
|||||||
@@ -4,8 +4,11 @@ Read only when the entrypoint leaves a routing choice unresolved.
|
|||||||
|
|
||||||
| Boundary | Decision |
|
| Boundary | Decision |
|
||||||
| --- | --- |
|
| --- | --- |
|
||||||
| Many easy tasks | Batch bounded Luna/Terra slices; volume alone does not justify Astra. |
|
| Many easy tasks | Partition independent bounded Luna batches across available capacity; size batches to keep coordination and acceptance useful. Volume alone does not justify Astra. |
|
||||||
| Ordinary cross-file bug | Terra; file count alone does not justify escalation. |
|
| Ordinary cross-file bug | GPT-6 Sol; file count alone does not justify Astra. |
|
||||||
|
| Small patch with settled semantics and deterministic acceptance | GPT-6 Luna medium/high if failure is contained; parent checks the diff and behavior. |
|
||||||
|
| Small patch changes authentication or a shared schema contract | Parent settles shared decisions; bounded investigation may use Sol. Small size is not evidence of low ambiguity. |
|
||||||
|
| Cheap Luna attempts require repeated parent rewrites | Choose Sol directly next time for this shape; count total accepted-task cost. |
|
||||||
| Difficult accounting discrepancy | Sol reasoning; parent owns the semantic decision. |
|
| Difficult accounting discrepancy | Sol reasoning; parent owns the semantic decision. |
|
||||||
| Coupled architecture across runtimes and tools | Consider Astra directly; keep shared decisions serial. |
|
| Coupled architecture across runtimes and tools | Consider Astra directly; keep shared decisions serial. |
|
||||||
| Missing credential or inaccessible source | Resolve the evidence/environment gap; model escalation cannot supply access. |
|
| Missing credential or inaccessible source | Resolve the evidence/environment gap; model escalation cannot supply access. |
|
||||||
@@ -14,9 +17,21 @@ Read only when the entrypoint leaves a routing choice unresolved.
|
|||||||
| Repository requires Sol parent | A Sol child cannot replace that owner; preserve the gate until an authorized supported change. |
|
| Repository requires Sol parent | A Sol child cannot replace that owner; preserve the gate until an authorized supported change. |
|
||||||
| Accepted output, unknown child model | Accept only after normal parent checks; do not attribute the result to the requested model. |
|
| Accepted output, unknown child model | Accept only after normal parent checks; do not attribute the result to the requested model. |
|
||||||
|
|
||||||
|
## Legacy and mixed hosts
|
||||||
|
|
||||||
|
Check each route independently. The API catalog does not update an existing session's tool schema.
|
||||||
|
|
||||||
|
| Unavailable preferred route | Supported compatibility choice |
|
||||||
|
| --- | --- |
|
||||||
|
| GPT-6 Luna | GPT-5.6 Luna low/medium for fixed collection/checks, high for settled transformations or patches; otherwise Sol or local work. |
|
||||||
|
| GPT-6 Sol for implementation/investigation/complex analysis | GPT-5.6 Sol medium/high, if exposed and adequate; otherwise local work. |
|
||||||
|
| GPT-6 Astra | An adequate supported Sol route or local work; disclose a capability gap if neither suffices. |
|
||||||
|
|
||||||
|
These are work-shape fallbacks, not capability equivalence claims. Terra is retired from recommended routes and automatic fallbacks; do not select it merely because a legacy host exposes it. Preserve an explicit user-pinned model or existing configuration until an authorized change. If a required model is unavailable, report the exact gap before dependent dispatch. Never invent `gpt-6-terra`. A mixed host may use GPT-6 Sol and GPT-5.6 Luna without reverting every route. Record the exact requested ID and fallback reason. Unknown observed identity stays unknown.
|
||||||
|
|
||||||
## Production Artifact Collection
|
## Production Artifact Collection
|
||||||
|
|
||||||
- Luna: use known hosts/endpoints/paths and fixed selectors to read or download actual reports, JSON, logs or job outputs; verify size/hash, parseability, required fields and specified counts. Terra: locate the current output through bounded read-only investigation, diagnose retrieval failures, or correlate logs and artifacts by run/version. Do not launch both by default; use Terra only when discovery or diagnosis is needed.
|
- GPT-6 Luna: use known hosts/endpoints/paths and fixed selectors to read or download actual reports, JSON, logs or job outputs; verify size/hash, parseability, required fields and specified counts. GPT-6 Sol: locate the current output through bounded read-only investigation, diagnose retrieval failures, or correlate logs and artifacts by run/version. Do not launch both by default; use Sol only when discovery or diagnosis is needed. Apply the legacy table when these routes are unavailable.
|
||||||
- Contract: name the production source, expected run/date/version, allowed read operations, bounded time/range/volume, local destination and acceptance checks. Reuse approved access and applicable server tools. Keep secrets out of prompts/receipts. Remote access stays read-only; explicitly allow local artifact writes even for a collector role.
|
- Contract: name the production source, expected run/date/version, allowed read operations, bounded time/range/volume, local destination and acceptance checks. Reuse approved access and applicable server tools. Keep secrets out of prompts/receipts. Remote access stays read-only; explicitly allow local artifact writes even for a collector role.
|
||||||
- Evidence: preserve the actual downloaded artifact and record source, retrieval time/timezone, producer run/version and generation time when exposed, local path, bytes/hash and check results. Unknown provenance stays unknown. Verify the expected run/freshness; mtime or successful download alone does not prove current output. Use immutable run identifiers where possible; detect/retry boundedly if files change during collection.
|
- Evidence: preserve the actual downloaded artifact and record source, retrieval time/timezone, producer run/version and generation time when exposed, local path, bytes/hash and check results. Unknown provenance stays unknown. Verify the expected run/freshness; mtime or successful download alone does not prove current output. Use immutable run identifiers where possible; detect/retry boundedly if files change during collection.
|
||||||
- No fixture, cache or old snapshot may stand in for requested live output. If the current artifact is missing, stale, inconsistent or inaccessible, report that gap. Do not trigger jobs, regenerate artifacts, restart services, change permissions/configuration or deploy as part of collection; return such remediation to the parent under existing authorization.
|
- No fixture, cache or old snapshot may stand in for requested live output. If the current artifact is missing, stale, inconsistent or inaccessible, report that gap. Do not trigger jobs, regenerate artifacts, restart services, change permissions/configuration or deploy as part of collection; return such remediation to the parent under existing authorization.
|
||||||
@@ -30,4 +45,4 @@ Use actual usage counters when exposed. Otherwise report observable proxies (cha
|
|||||||
|
|
||||||
If context dominates, narrow inputs and avoid full-history forks. If output dominates, bound logs and returns. If reasoning/retries dominate, clarify acceptance and choose adequate capability/effort directly. Preserve required tests and material findings.
|
If context dominates, narrow inputs and avoid full-history forks. If output dominates, bound logs and returns. If reasoning/retries dominate, clarify acceptance and choose adequate capability/effort directly. Preserve required tests and material findings.
|
||||||
|
|
||||||
Official model positioning checked 2026-09-14: [Models](https://developers.openai.com/api/docs/models). The route table is an engineering starting strategy, not a measured cost ranking. Runtime tool schemas remain the source for actual child availability.
|
Official model positioning checked 2026-09-23: [GPT-6 guidance](https://developers.openai.com/api/docs/guides/latest-model). See [model economics](model-economics.md) for the dated API price basis. The route table is an engineering starting strategy, not a measured task-cost ranking. Runtime tool schemas remain the source for actual child availability.
|
||||||
|
|||||||
@@ -30,7 +30,7 @@ class EvidenceTests(unittest.TestCase):
|
|||||||
def test_known_identity_requires_host_evidence(self):
|
def test_known_identity_requires_host_evidence(self):
|
||||||
packet = fixture()
|
packet = fixture()
|
||||||
child = packet["tool_receipts"][0]
|
child = packet["tool_receipts"][0]
|
||||||
child["observed"]["model"] = "gpt-5.6-terra"
|
child["observed"]["model"] = "gpt-6-sol"
|
||||||
with self.assertRaisesRegex(ValueError, "host evidence"):
|
with self.assertRaisesRegex(ValueError, "host evidence"):
|
||||||
evidence.validate(packet)
|
evidence.validate(packet)
|
||||||
child["identity_source"] = "host-tool-result"
|
child["identity_source"] = "host-tool-result"
|
||||||
|
|||||||
Reference in New Issue
Block a user