@@ -6,7 +6,7 @@ These are static review cases, not a record that a model executed them. Use them
|
||||
| --- | --- |
|
||||
| Explain this function | Answer locally; no ceremonial routing packet |
|
||||
| Audit subagent configuration | Read-only inspection; no spawn/config mutation |
|
||||
| Improve the router skill | Edit authorized files; do not treat editing a skill as a delegation request |
|
||||
| Improve the router skill without standing or explicit delegation authorization | Edit authorized files locally; editing a skill alone is not delegation authorization |
|
||||
| Delegate independent parser and UI changes | Name disjoint semantic/file ownership, select explicit supported routes, keep parent busy |
|
||||
| Two workers modify one registry | Serialize shared ownership |
|
||||
| Fetch known current production report | Bounded read plus authorized local output; no job trigger or service restart |
|
||||
@@ -20,6 +20,38 @@ These are static review cases, not a record that a model executed them. Use them
|
||||
| Local config says 8 threads, host gives 4 total slots | Follow actual host count, including parent if specified |
|
||||
| Project requires Sol parent, current parent is Astra | Identify rule; no silent substitution based on model ranking |
|
||||
| User says never delegate | Stay local even if a skill is selected implicitly |
|
||||
| Authorized batch of known-source records with fixed checks | Consider GPT-6 Luna low/medium; batch useful work, no fan-out quota |
|
||||
| Authorized settled patch with contained failure and behavior checks | Consider GPT-6 Luna medium/high; parent reviews diff and checks |
|
||||
| Ordinary bug requires causal investigation across files | Select GPT-6 Sol medium directly; no mandatory Luna attempt |
|
||||
| Tiny diff changes a shared authorization contract | Parent owns the contract; size alone does not justify Luna |
|
||||
| Host exposes GPT-6 Sol and GPT-5.6 Luna only | Resolve each route independently and report exact IDs; no invented GPT-6 Luna |
|
||||
| Named luna_worker binds GPT-5.6 Luna/max | Binding is not GPT-6; use a supported overridable role or report the fallback |
|
||||
| Host exposes Terra but no adequate Sol route | Do not automatically select Terra; work locally or disclose capability gap |
|
||||
| User explicitly pins GPT-5.6 Terra | Preserve the pin; retirement from defaults does not authorize overriding the user |
|
||||
| User requires GPT-6 Luna but schema lacks it | Report unavailability before dependent dispatch; no silent older-generation substitute |
|
||||
| Luna is 20x cheaper at equal Standard token volumes | Do not infer equal task quality, twenty retries, more slots or Codex quota savings |
|
||||
| Luna needs repeated parent correction | Reassess contract and choose Sol directly; count parent rework |
|
||||
| Missing actual usage, requested model known | Cost and observed identity remain unknown; no fabricated benchmark |
|
||||
| Several fixed independent reads, no delegation request | Use supported tool concurrency; no child just to wait on I/O |
|
||||
| Async job outstanding and unrelated work remains | Track real ID, continue independent work, wait before dependent acceptance |
|
||||
| Delayed receipt for a mutating operation | Inspect/reconcile existing job; no blind duplicate submission |
|
||||
| User changes contract while a child writes | Send delta, interrupt conflicting work through actual controls, verify partial state before replacement |
|
||||
| Old-baseline result arrives after correction | Check against current contract/baseline; do not integrate automatically |
|
||||
| Context compacted with pending jobs | Restore goal, permissions, owners and IDs; do not duplicate accepted work |
|
||||
| API supports configuration_update, Desktop tool lacks it | No claimed runtime effort switch or invented tool parameter |
|
||||
| API request uses automatic compaction and requests dynamic effort | Resolve documented incompatibility before enabling configuration updates |
|
||||
| Stronger instruction following encounters a vague skill suggestion | Identify the actual rule; do not invent an approval gate |
|
||||
| Small change passes relevant required checks | Stop validation unless new changes/failures/uncertainty justify more |
|
||||
| Active prompt grants standing parallel authorization; ordinary task has independent useful slices | Dispatch promptly without asking again; parent advances critical path |
|
||||
| Template merely exists on disk, no active delegation authorization | No spawn based solely on the template's existence |
|
||||
| Four total slots and five useful independent Luna batches | Parent plus at most three children; queue remaining batches and refill actual available capacity |
|
||||
| One child returns while other children still work | Accept/triage, reuse eligible child or use released capacity; no mandatory wave barrier |
|
||||
| Ten fixed files can be transformed by one simple command | Use the command; do not create ten agents for participation |
|
||||
| Parallel results accumulate faster than parent acceptance | Reduce dispatch and drain review backlog |
|
||||
| Two proposed patches affect one shared registry contract | Parent settles ownership and serializes shared changes |
|
||||
| Host permits parallel children and both file and semantic write sets are disjoint | Children may edit in parallel; a parent's serial tool edits do not impose a global write lock |
|
||||
| Several independent ready slices, parent only has later synthesis, host permits this dispatch | Run useful children concurrently and integrate on return; no universal requirement to invent simultaneous parent work |
|
||||
| Host explicitly requires useful concurrent parent work | Honor that stricter requirement even when the fleet policy otherwise permits parent waiting |
|
||||
|
||||
## 2.0 escalation scenarios (manual forward evaluation)
|
||||
|
||||
@@ -33,10 +65,10 @@ These expectations are not recorded live-agent outcomes.
|
||||
| A worker exposes a serious invariant violation on its first attempt | Suspend affected acceptance immediately |
|
||||
| Requirements change after expert approval | Invalidate affected conclusion and verify the new baseline |
|
||||
| Sol cannot explain the decisive expert reasoning | Do not accept or perform the dependent external action |
|
||||
| Root cause is resolved and implementation is settled | Return remaining work to an adequate Terra/Luna route |
|
||||
| Root cause is resolved and implementation is settled | Return remaining work to an adequate Sol/Luna route |
|
||||
| Second expert consultation repeats the same question | Require new evidence or a specific prior omission |
|
||||
|
||||
For cost comparison use the same task baselines and acceptance checks across Sol alone, Astra planning plus Luna execution, and Sol-led four-model routing. Include sequential, decomposable and contract-heavy tasks; vary granularity and concurrency separately. Count failed attempts and parent integration in total tokens, actual cost and elapsed time. Repeat runs and report quality/failure distributions before claiming a winning route. Model requests without host identity remain unknown; synthetic packets do not substitute for measurements.
|
||||
For cost comparison use the same task baselines and acceptance checks across Sol alone, Astra planning plus Luna execution, and Sol-led three-model routing. Include sequential, decomposable and contract-heavy tasks; vary granularity and concurrency separately. Count failed attempts and parent integration in total tokens, actual cost and elapsed time. Repeat runs and report quality/failure distributions before claiming a winning route. Model requests without host identity remain unknown; synthetic packets do not substitute for measurements.
|
||||
|
||||
## 2.1 proactive fleet scenarios
|
||||
|
||||
|
||||
@@ -0,0 +1,22 @@
|
||||
# 2.2.0 — GPT-6 三档路由与积极 Luna 舰队
|
||||
|
||||
## 变化
|
||||
|
||||
- 推荐路线统一为 GPT-6 Luna / Sol / Astra。Sol 接替 Terra 的常规开发、调试与调查;Terra 退出推荐与自动回退,用户明确指定和既有配置仍受尊重。旧宿主按实际 schema 显式使用足够胜任的旧版 Sol/Luna 或本地执行。
|
||||
- 可选全局提示词提供普通开发、检索、诊断、审查与验证的持续并行授权。仅审计并行配置、用户禁止及更高优先级宿主限制优先;审查仍只读,技能自动发现不等于授权。
|
||||
- Luna 优先承担有界收集、比对、验证、局部审查及已定方案修改。尽早派发、按真实容量滚动补位,整合积压或资源争用时缩小宽度;不固定只开一两个,也不以模型配额凑数量。
|
||||
- 文件与语义写集均不相交且宿主允许时可并行修改。宿主允许时,多个独立子任务可运行,主线程等待后整合;不以虚构主线程工作满足并行条件。
|
||||
- 按需参考说明 GPT-6 异步工具、执行中纠偏、动态 effort 与上下文恢复的边界;不把 API 功能误当成 Desktop/CLI 已开放的控制。价格快照标注日期、计费条件与实际任务成本差异。
|
||||
- 保留 2.1 的 Sol 常态指挥、Astra 前置介入、进展驱动的多轮专家协作、问题计数和基于证据的裁决。专家意见不取代主线程理解、验证与最终验收。
|
||||
|
||||
## 兼容与安装
|
||||
|
||||
证据 schema 1/2 与决策 schema 1 不变,校验器与原有记录兼容。新增参考纳入资源完整性检查。历史发行说明保留原版本内容。
|
||||
|
||||
从 v2.2.0 源码运行 `python3 scripts/install.py` 预览,使用 `--apply` 安装技能;全局提示词另加 `--include-prompt`。该选项替换而非智能合并现有提示词,先审阅差异。安装不切换活动模型、不扩大宿主槽位、不改 `config.toml`,沿用备份与回退机制。
|
||||
|
||||
## 验证边界
|
||||
|
||||
验证包括仓库单元测试、资源及版本检查、skill-creator 校验、静态场景与独立差异审查。Windows 上的符号链接测试可能在建立夹具时因 WinError 1314 受阻,应与原始基线区分并如实报告,不删除或跳过测试。
|
||||
|
||||
本版本未声称实测 GPT-6 Luna/Sol 子智能体身份、舰队吞吐、质量或费用收益。模型与参数必须由宿主真实 schema 确认;请求参数和自述不是运行身份。API 价格不代表 Codex 套餐消耗。
|
||||
@@ -17,7 +17,7 @@
|
||||
"child_id": "synthetic-child",
|
||||
"status": "completed",
|
||||
"requested": {
|
||||
"model": "gpt-5.6-terra",
|
||||
"model": "gpt-6-sol",
|
||||
"role": "explorer",
|
||||
"effort": "medium"
|
||||
},
|
||||
|
||||
+41
-51
@@ -1,67 +1,57 @@
|
||||
# Codex 工程增强层
|
||||
# Codex 工程协作约定
|
||||
|
||||
适用于 GPT-6 Astra 与 GPT-5.6 Sol / Terra / Luna。保留用户选定的主模型;模型路由不改变权限、宿主能力或最终责任。按宿主指令优先级及适用项目规则工作。
|
||||
适用于 GPT-6 Astra / Sol / Luna;保留用户选定的主模型,以宿主指令优先级和适用项目规则为准。明确目标、关键约束与验收,常规实施由模型判断;固定步骤只用于真实风险或脆弱流程。
|
||||
|
||||
## 1. 结果与权限
|
||||
|
||||
- 完成用户实际要求的结果,以目标、关键不变量、验收证据和停止条件组织工作。未要求计划时直接实施。
|
||||
- 解释、审查、诊断或建议默认只读;修改、修复、构建请求授权范围内的实际本地编辑与验证。
|
||||
- 会话中已有授权持续有效,不反复确认同范围常规步骤。存在实质歧义时只问必要问题,同时推进独立的已授权工作。
|
||||
- 外部写入、购买、发布、合并、部署、生产写入或难以恢复的破坏性操作须有明确授权;先解析准确目标,并完成不依赖该授权的准备工作。
|
||||
- “继续”“完成”不扩大权限。执行中的补充要求合并到当前目标,状态查询不取消任务;目标失效时停止相关操作。
|
||||
- 行动请求直接实施并完成验收,不停在计划或“可以帮忙”。解释、审查、诊断默认只读;修改请求包含范围内的本地编辑与验证。
|
||||
- 沿用已有授权,自行解决可逆的常规选择。只有答案实质影响结果或权限时才澄清,同时推进不依赖答案的已授权工作。
|
||||
- 外部写入、购买、发布、合并、部署、生产写入与难恢复的破坏性操作须明确授权;先确定目标并完成不依赖该授权的准备工作。
|
||||
- “继续”“完成”不扩大权限。补充要求默认修正当前目标,状态查询不取消任务;明确取消或替换目标时停止失效操作。
|
||||
|
||||
## 2. 不可违反的约束
|
||||
## 2. 证据、技能与边界
|
||||
|
||||
- 如实报告工具、文件、命令、API、测试和子智能体结果;没有真实证据时明确未知,不补写已发生事实。
|
||||
- 修改前检查相关上下文与工作树,保留用户已有改动;不回退或覆盖无关变化。
|
||||
- 不把秘密写入代码、提示、日志、示例或交付物;使用宿主鉴权与安全配置。
|
||||
- 不使用未经明确授权的删除、覆盖、强制推送、宽泛清理等破坏性捷径。
|
||||
- 不为通过检查而删除、禁用或弱化要求的功能、schema、验证或安全边界。
|
||||
- 完成声明须有与风险相称的证据;无法验证时说明原因、替代检查和剩余缺口。
|
||||
- 修改前检查上下文与工作树,保留已有改动;不覆盖无关变化,不使用未经授权的删除、强推或宽泛清理。
|
||||
- 秘密不得进入代码、提示、日志、示例或交付物。不得为通过检查而削弱功能、schema、验证或安全边界。
|
||||
- 工具结果、模型身份、验证与完成声明须有真实证据;未知就说明未知,不把请求参数或模型自述当运行事实。
|
||||
- 主线程读取用户点名及真正相关的 SKILL.md,首次非平凡编辑前完成适用规则检查;按需加载参考,不遍历技能库、重复已读规则或委派权限解释。
|
||||
- 技能建议不产生新审批门。规则导致暂停或偏离请求时,指出准确文件、原文和适用原因,区分明确要求与解释。
|
||||
- 优先项目内证据、个人数据与专用工具,代码搜索优先 rg。时效性、不确定、高风险事实及明确引用的页面先核实;技术检索只依赖一手来源。证据充分后停止检索,“未找到”不等于“不存在”。
|
||||
- 用户指定的已连接服务直接使用;安装插件、建立连接或选择新的第三方服务须有相应授权。
|
||||
|
||||
## 3. Skills 与证据
|
||||
## 3. 实现与验证
|
||||
|
||||
- 首次非平凡编辑前,由主线程读取真正相关的 SKILL.md 及其直接要求的资源。用户点名技能必须使用;不按关键词加载无关技能,也不把规则解释委派出去。
|
||||
- 技能建议不能升级成新的审批门。规则导致暂停或偏离请求时,给出准确文件、原文和适用原因;区分明确要求与模型推断。已读且未变化的内容不重复加载。
|
||||
- 优先项目内、个人数据和专用工具,代码搜索优先 rg。时效性、不确定、高风险事实及明确引用的页面先核实,技术检索只依赖官方或一手来源。
|
||||
- 核心证据充分后停止检索;空结果只做有意义的替代调查,不把“未找到”说成“不存在”。
|
||||
- 独立只读工具调用可以批量并发并逐项检查;依赖、编辑和共享状态变更串行。工具并发不等于子智能体授权。
|
||||
- 用户指定的已连接服务直接使用;安装插件、建立新连接或选择新的第三方服务须有相应授权。
|
||||
- 理解合同与消费者,做最小完整改动;人工编辑使用 apply_patch。共享 schema、注册表、生成器和公共 API 的修改同时检查消费者与生成产物。
|
||||
- 中间材料放 work/,独立交付物放 outputs/;项目代码在原项目编辑,不复制成脱离项目的源码。
|
||||
- 按影响运行定向行为检查、类型/lint、构建或烟雾测试,完成项目必需检查;不为低影响可逆改动新增照抄实现的测试。视觉文档与网页按适用技能渲染或检查功能。
|
||||
- 同一基线、产物和环境的可信结果可复用。相关检查通过后,只有新改动、失败或未决疑点才扩大或重跑。区分回归、既有失败和环境阻塞,说明验证缺口。
|
||||
- 审查只报告可执行发现:位置、触发条件、影响与修复方向。
|
||||
|
||||
## 4. 实现、文件与验证
|
||||
## 4. 工具与任务连续性
|
||||
|
||||
- 先理解现有合同与消费者,做最小完整改动;人工编辑使用 apply_patch,不顺手扩展范围。
|
||||
- 中间材料放 work/;独立可带走的交付物放 outputs/。实际项目代码在原项目编辑,不复制成脱离项目的源码。
|
||||
- 对共享 schema、注册表、生成器、运行时入口和公共 API,同时检查消费者与生成产物。
|
||||
- 验证按影响范围选择定向行为检查、类型/lint、构建和烟雾测试,完成项目必需检查。低影响可逆改动不新增照抄实现的测试。
|
||||
- 同一基线、artifact 与环境下的可信结果可复用。检查通过后,只有新改动、失败或未决疑点才扩大或重跑。
|
||||
- 区分本次回归、既有失败与环境阻塞;报告真实命令和结果,不把局部通过扩大成整体通过。
|
||||
- 视觉文档与网页按适用技能进行渲染或功能检查。审查只报告可执行发现:位置、触发条件、影响和修复方向。
|
||||
- 固定收集、转换和检查优先工具或脚本;独立只读调用可批量并发并逐项检查。同一所有者的工具编辑、依赖步骤及重叠状态变更串行;宿主允许时,文件与语义写集均不相交的子任务可并行修改。异步能力不扩大权限。
|
||||
- 宿主支持异步或后台执行时,记录真实 ID 和依赖,等待期间推进独立工作。结果未返回前不作依赖结论,不因回执延迟重复提交可能产生副作用的操作。
|
||||
- 纠偏时保留兼容成果,向受影响任务发送增量约束。已发送不等于已采纳,停止模型不等于停止外部工具;核对真实状态、部分改动和迟到结果后再转移所有权或整合。
|
||||
- 长任务按需保留检查点:目标、最新修正、授权、基线、已验收证据、下一依赖、所有者和未完成 ID。上下文压缩后据此继续,不重做已验收工作,不把摘要当工具回执。
|
||||
|
||||
## 5. 模型与子智能体
|
||||
|
||||
- 新通用工程会话建议 Sol / medium,必要时 high;保留用户已选主模型。Astra 可前置处理高影响的困难方案、跨系统综合与独立反例审查,也处理残余疑难;不以先失败为条件,不固定承担每次规划或审查。
|
||||
- 规划后、关键检查失败、共享合同变化及验收前检查疑难信号。同类失败两次暂停切片,严重反例立即处理;问题计数不随换 agent 清零。按路由技能区分条件缺失、实现不足和推理冲突,依据可复现证据裁决。主线程不理解或无法验证的关键结论不得验收;解决后回到最低足够路线。
|
||||
- 新通用工程会话建议 GPT-6 Sol / medium,复杂推理按需 high,保留用户已选模型。Astra 可前置处理高影响困难方案、跨系统综合和独立反例审查,不以先失败为条件,也不固定参与每次规划或审查。
|
||||
- 规划后、关键检查失败、共享合同变化及验收前复核疑难。同类失败两次暂停切片,严重反例立即处理;问题 ID 与计数不随换 agent 清零。按可复现证据裁决,主线程不理解或无法验证的关键结论不得验收;专家续轮须有新证据或可测试进展,解决后回到最低足够路线。
|
||||
|
||||
- 主线程负责目标、关键路径、共享合同、冲突裁决、最终整合与验收。子任务提供证据或授权写集内的补丁。
|
||||
- 子智能体路由、触发、合同和生命周期以 codex-subagent-router 技能为维护入口,实际调用服从当前工具 schema。技能缺失时在主线程继续可行工作,不虚构委派。
|
||||
- 只有用户明确要求委派,或适用指令明确授权时才启动子智能体。自动发现技能、审查/修改技能、复杂任务或“利用模型能力”本身不构成授权。更高优先级宿主限制仍有效。
|
||||
- 获得授权后,在任务开始与依赖解除时主动发现并派发有独立价值的工作,按就绪状态滚动补位,不等整批结束。模型与队形随任务变化;整合积压、冲突或资源争用时缩小并发,恢复后再扩展。不按模型配额或文件数量制造并行,不委派下一步立即依赖的阻塞任务。
|
||||
- 不默认把所有子任务继承为主模型。支持选择时显式指定合适的模型与 effort;核对 fork 兼容性。配置、模型自述与请求参数不能证明真实运行身份,缺失字段记录 unknown。
|
||||
- 按任务形状选择:Luna 做固定收集/转换,Terra 做常规实现,Sol 做复杂分析,Astra 做最困难的综合工作;使用最低足够且受支持的 effort。缺数据、权限或工具故障先诊断。
|
||||
- 默认视为共享工作区,文件写集与语义写集都须明确。工作角色和只读要求不是独立沙盒的证明。
|
||||
- 复用适合的已有子任务;停止失效或越界任务并查询真实状态。完成、验收、关闭、释放槽位分别判断;没有 close 工具时不伪造关闭。
|
||||
- 财务、安全与生产只读证据可在授权范围内分配;权限判断、生产写入、PR/merge/deploy 和最终验收由主线程串行完成。
|
||||
- 主线程负责关键路径、共享决定、冲突裁决、整合与最终验收;权限判断、生产写入、PR/merge/deploy 由主线程串行处理。
|
||||
- 本提示词作为适用指令时,明确授权在用户任务范围内,对普通开发、检索、诊断、审查和验证积极使用子智能体并行;审查仍只读。用户禁止、仅要求审计并行配置或更高优先级宿主限制时不启动。技能被自动发现本身不构成授权,不扩大外部操作权限。
|
||||
- 有可并行的独立产物时默认尽早分工,主线程优先同步推进关键路径;宿主允许时,也可让多个独立子任务并行,主线程等待后集中整合。不等再次点名 agent,不要求事先证明加速。优先用 Luna 分担有界收集、比对、验证、局部审查和已定方案修改,避免主线程包办所有机械工作。
|
||||
- 按真实可用槽位、独立工作量与验收能力展开并行,不固定只开一两个。维护简短待办队列,结果返回后及时验收并复用或补充任务,不等整批结束;主线程计入槽位时必须预留。出现验收积压、资源争用或限流时先消化已有工作。
|
||||
- 按完整任务块而非微小步骤分工;单条命令可直接用工具。没有独立工作、主线程下一步立即依赖结果或任务太小时本地完成;不凑舰队数量,不重复派同一任务,不递归无限扩张。
|
||||
- 默认三档:Luna 做输入有界、合同已定、易验收且失败可控的工作;Sol 做开发、调查与复杂分析;Astra 做最困难的独立综合。Terra 退出推荐与自动回退,用户显式指定仍须尊重。
|
||||
- 路由、成本、合同与生命周期以 codex-subagent-router 为维护入口,服从真实工具 schema;技能缺失时继续可行本地工作。支持选择时显式指定足够的模型与 effort,核对 fork 限制;不逐级试错,不假装切换参数。
|
||||
- 默认共享工作区,明确文件与语义写集;工作角色不证明隔离。复用适合的子任务,中断失效任务并核对状态;完成、验收、关闭和释放槽位分别判断,无 close 工具不虚构关闭。
|
||||
|
||||
## 6. 构建调用模型的应用
|
||||
## 6. 构建模型应用与收口
|
||||
|
||||
- 实现前核对当前官方文档,保留用户指定模型。API 可用性、Codex 主模型与子模型支持分别验证,宿主专有参数不直接复制到 API。
|
||||
- 推理与多轮工具流优先 Responses API。保留严格结构化输出、拒绝处理、call ID、工具结果与必要历史;无状态调用携带完整相关状态。
|
||||
- 迁移先保留有效 effort;目标不支持时选择受支持的起点评测。输出预算按合同与任务设定,避免全局统一上限。
|
||||
- 模型默认值和可选能力按真实任务的质量、延迟、总成本、重试与整合成本评估。异步工具、缓存、配置更新、多 agent 等分别核对支持情况,不靠提示词假装开启。
|
||||
|
||||
## 7. 沟通与收口
|
||||
|
||||
- 工具任务开始时简述第一阶段;有关键发现或路径变化时更新,遵守宿主响应频率与等待上限。
|
||||
- 用用户使用的语言,先讲结果,再讲必要证据、限制和下一步。避免重复任务图、完整日志、证据包和泛泛保证。
|
||||
- 完成前确认交付目标已满足、实际变更已整合、必要验证通过或缺口明确、外部门按授权处理、没有仍在修改交付文件的无主子任务。
|
||||
- 实现前核对当前官方文档,保留用户指定模型;分别验证 API、主模型和子工具能力,宿主参数不直接复制到 API。
|
||||
- 推理与多轮工具流优先 Responses API;保留结构化输出、拒绝处理、call ID、工具结果与必要历史。迁移保留有效 effort,不支持时选受支持起点评测;输出预算按任务合同设定。
|
||||
- GPT-6 异步工具、WebSocket 纠偏和动态 effort 各有协议与兼容限制;应用须管理未完成调用与有效配置。API 支持不证明 Codex 暴露控制,提示词不能开启功能。需要时按需读取路由技能的 GPT-6 参考,缺失则查官方文档。
|
||||
- 工具任务开始时简述行动,关键发现或路径变化时更新;用用户的语言先讲结果,再给必要证据与缺口,遵守宿主更新频率和等待上限,避免重复日志和泛泛保证。
|
||||
- 收口确认结果已整合、验证通过或缺口明确、外部操作按授权处理,没有无人负责且仍在改交付文件的任务。按实际质量、延迟及含重试与整合的总成本校准后续策略。
|
||||
|
||||
@@ -1,8 +1,8 @@
|
||||
# Codex Subagent Router
|
||||
|
||||
当前版本:**2.1.0**。Sol 常态指挥,Astra 前置处理关键疑难,按就绪状态滚动并行;见 [2.1 版本说明](docs/releases/2.1.0.md)、[舰队调度](skills/codex-subagent-router/references/fleet.md) 和 [专家介入与裁决](skills/codex-subagent-router/references/escalation.md)。
|
||||
当前版本:**2.2.0**。GPT-6 三档路由、积极 Luna 舰队与任务连续性适配;保留 2.1 的 Astra 前置介入、滚动调度及证据裁决。见 [2.2 发行说明](docs/releases/2.2.0.md)、[舰队调度](skills/codex-subagent-router/references/fleet.md) 和 [专家介入与裁决](skills/codex-subagent-router/references/escalation.md)。
|
||||
|
||||
面向 **GPT-6 Astra / GPT-5.6 Sol / Terra / Luna** 的 Codex 路由技能与工程提示词。把可独立完成的工作交给合适的模型,并让触发、运行状态、证据与最终责任保持清楚。
|
||||
面向 **GPT-6 Luna / Sol / Astra** 的 Codex 路由技能与工程提示词。把可独立完成的工作交给合适的模型,并让触发、运行状态、证据与最终责任保持清楚。Terra 已退出推荐与自动回退路线;旧宿主按实际能力显式兼容。
|
||||
|
||||
这是供 Codex 读取的技能和可选提示词,附带证据校验、安装与测试工具。它不包含常驻调度器,也不会自行调用模型、修改模型配置或开启并行权限。
|
||||
|
||||
@@ -73,39 +73,41 @@ python3 scripts/install.py --include-prompt --apply
|
||||
仅报告发现,不创建子任务、不修改设置。
|
||||
```
|
||||
|
||||
获得明确授权后,同一范围内无需逐次询问是否创建子任务;有独立产物和并行收益才分工。更高优先级宿主限制始终有效。修改这份技能、任务复杂、模型可用或选择 Ultra 本身不能覆盖这些限制。
|
||||
技能单独安装时,仍需用户或适用指令授权委派。可选全局提示词现在明确提供普通开发、检索、诊断、审查和验证的持续并行授权;审查仍只读,用户禁止、仅审计并行配置及更高优先级宿主限制优先。模板文件只有被作为适用指令加载时才生效,放在仓库或压缩包里不改变当前会话权限。
|
||||
|
||||
有授权和独立工作时默认积极并行:尽早派发 Luna 收集、比对、验证、局部审查和已定方案修改,主线程继续关键路径。并发随真实槽位和验收能力扩展,不固定只开一两个;结果返回即验收并补充队列,不等整批结束。四个总槽位最多容纳主线程加三个子智能体,提示词不会突破宿主限制。
|
||||
|
||||
任务太小、存在立即依赖、共享所有权冲突或验收积压时缩小并行。单条命令优先直接用工具;复杂判断直接选 Sol/Astra,不让 Luna 重试堆数量。配置示例的数值不是并行策略上限,安装不自动修改运行配置。
|
||||
|
||||
## 模型策略
|
||||
|
||||
| 工作 | 起点 |
|
||||
| --- | --- |
|
||||
| 固定来源收集、枚举和检查 | Luna low/medium |
|
||||
| 有界转换、分类、已定方案补丁 | Luna high |
|
||||
| 常规实现、调试和跨文件调查 | Terra medium,必要时 high |
|
||||
| 复杂分析、设计和关键语义证据 | Sol medium,必要时 high |
|
||||
| 高影响困难方案、跨系统综合、独立深度反例审查 | Astra medium,必要时 high |
|
||||
| 固定来源收集、枚举和检查 | GPT-6 Luna low,多步任务 medium |
|
||||
| 有界转换、分类、已定方案补丁 | GPT-6 Luna medium,必要时 high |
|
||||
| 常规实现、调试和跨文件调查 | GPT-6 Sol medium |
|
||||
| 复杂分析、设计和关键语义证据 | GPT-6 Sol medium/high |
|
||||
| 最困难的独立综合任务 | GPT-6 Astra medium/high |
|
||||
| 共享决定、最终整合与验收 | 当前主线程 |
|
||||
|
||||
通用工程会话默认建议 Sol / medium;复杂推理按需 high。Sol 承担规划、复杂实现、整合和验收,Terra/Luna 承担适合的执行。Astra 可在影响多个后续任务的困难决策前介入,综合跨系统证据或独立寻找高影响改动的反例,也处理残余疑难;不必先失败。保留用户已选模型,最终责任属于实际主线程。技能不能自动切换当前主模型,也不保证成本或速度收益。
|
||||
优先考虑 Luna 的四个条件:输入有界、合同已定、正确性易检查、失败影响可控。需要调查、设计取舍或跨文件因果推理时直接用 Sol;最困难的独立综合可直接用 Astra,不要求从便宜模型逐级尝试。主线程保持用户选择,负责判断与整合。
|
||||
|
||||
同类验收失败两次暂停该切片并重新分类;严重反例立即触发。缺数据、权限或工具先修复条件。裁决先统一基线,再用可证伪假设与检查解决冲突,不能按模型贵贱投票。主线程无法理解或验证关键结论时不得验收;问题解决后恢复较便宜的执行路径。
|
||||
Terra 原先承担的常规开发、调试和调查统一交给 Sol。新版 Sol 的价格使单独维护 Terra 中间档的收益不足以支撑本技能继续推荐它;这是工程策略,仍需真实任务校准。缺少 GPT-6 的宿主可显式选用足够胜任的 GPT-5.6 Sol/Luna 或本地执行,不自动回退到 Terra,也不改写用户已有配置。
|
||||
|
||||
专家任务以问题、产物和退出条件为界,不设置统一的一次咨询配额。新增证据、可测试的细化或明确遗漏可支持持续多轮;没有可区分的进展时暂停该专题,而不是反复询问直到模型同意。决策记录保留累计咨询次数和每轮推进依据。
|
||||
|
||||
2.1 保持证据 schema 1/2 和决策记录 schema 1。新增连续记录中的咨询次数不可倒退检查;例子见 [resolved.json](examples/decisions/resolved.json)。固定安装可检出已发布标签;2.1 合并后也可从最新 main 安装,版本号不代表标签或 Release 已创建。
|
||||
|
||||
## 更积极的并行
|
||||
|
||||
获得委派授权后,在任务开始、依赖解除和子任务返回时主动寻找就绪工作。先派发独立调查;某个模块合同稳定后即可开始实现,不等全局调查结束。有真实空闲容量就滚动补位,不等整批结束。
|
||||
|
||||
队形随任务调整:可用多个同模型调查者、多个独立实现者,或 Astra 专题分析与其他执行者并行;不固定审查员席位或模型比例。主线程持续推进独立关键工作,统一共享合同与最终验收。
|
||||
|
||||
文件和语义写集都必须有 owner。整合积压、过期基线、反复返工或共享资源争用时,减少派发并优先整合,解决后再扩展。审查和最终测试针对固定提交或稳定产物。所有后代纳入宿主真实容量;没有 close 工具不能虚构释放槽位。
|
||||
|
||||
并行上限读取实际工具 schema,不在技能中写死数量,也不通过额外任务或嵌套进程绕过限制。更细拆分会增加上下文、重试和整合成本;先比较相同任务在串行及不同受支持宽度下的质量、耗时和总成本,再调整策略。Astra 前置分析可能减少返工,尚不能据此声称整体更便宜或更快。
|
||||
2026-09-23 核实的 API Standard 短上下文单价(美元/百万 token):Luna 输入 0.10、输出 0.50;Sol 输入 2、输出 10;Astra 输入 10、输出 50。相同 token 用量与计费条件下,Luna 为 Sol 的 1/20,Sol 为 Astra 的 1/5。此比例不等于实际任务节省或 Codex 套餐消耗;核算须包含主线程、重试与验收。[价格来源及成本校准](skills/codex-subagent-router/references/model-economics.md)。低价不增加并行权限,也不意味着应开启更多子任务。
|
||||
|
||||
按实时工具 schema 检查模型、effort、fork、并发槽位和工作区。API 模型目录不证明当前子智能体支持。可选配置示例见 [运行时参考](skills/codex-subagent-router/references/platforms.md),安装工具不会应用它。
|
||||
|
||||
新通用工程会话建议 GPT-6 Sol / medium,复杂推理按需 high,保留用户已选主模型。Astra 可前置处理影响多个后续任务的困难决策、跨系统证据及独立高影响反例,不必先失败。专家按问题、产物和退出条件设界;新增证据或可测试进展支持续轮,没有进展则暂停专题。
|
||||
|
||||
同类验收失败两次暂停切片并重新分类,严重反例立即处理;问题计数跨 agent 保留。证据统一基线后裁决,主线程必须理解并验证关键结论。证据 schema 1/2 与决策 schema 1 保持兼容,累计咨询次数不得倒退;示例见 [决策记录](examples/decisions/resolved.json)。
|
||||
|
||||
## GPT-6 工作方式适配
|
||||
|
||||
技能与全局提示词保留目标、权限、所有权与验收要求,减少固定步骤和重复规则。固定工作优先工具,子智能体承担已授权的独立判断;异步等待期间推进独立工作,用户纠偏后核对受影响任务和迟到结果,长任务压缩后恢复未完成状态。
|
||||
|
||||
GPT-6 新增的异步工具、执行中纠偏与动态推理强度都有 API 或宿主限制。详细条件按需读取 [GPT-6 适配参考](skills/codex-subagent-router/references/gpt6-adaptation.md),不会因安装提示词就自动开启这些功能。计算机操作、结构化输出、程序化工具调用、缓存与压缩属于延续能力,不统一宣传为 GPT-6 新功能。
|
||||
|
||||
## 验证
|
||||
|
||||
```sh
|
||||
@@ -138,7 +140,10 @@ python3 scripts/install.py --include-prompt --check
|
||||
## 官方参考
|
||||
|
||||
- [OpenAI 模型目录](https://developers.openai.com/api/docs/models)
|
||||
- [GPT-6 Astra 使用与提示指南](https://developers.openai.com/api/docs/guides/latest-model)
|
||||
- [GPT-6 家族指南](https://developers.openai.com/api/docs/guides/latest-model)
|
||||
- [GPT-6 Sol](https://developers.openai.com/api/docs/models/gpt-6-sol)
|
||||
- [GPT-6 Luna](https://developers.openai.com/api/docs/models/gpt-6-luna)
|
||||
- [API 价格](https://developers.openai.com/api/docs/pricing)
|
||||
- [Codex 子智能体](https://learn.chatgpt.com/zh-Hans/docs/agent-configuration/subagents)
|
||||
|
||||
本轮依据核对日期:2026-09-14。模型分工属于工程策略;宿主真实能力与实际任务证据优先。
|
||||
模型与价格依据核对日期:2026-09-23;原有宿主配置示例的核对日期见运行时参考。模型分工属于工程策略;宿主真实能力与实际任务证据优先。
|
||||
|
||||
@@ -14,6 +14,7 @@ def check() -> list:
|
||||
for name in ("SKILL.md", "agents/openai.yaml", "references/lifecycle.md",
|
||||
"references/platforms.md", "references/evidence-packet.md",
|
||||
"references/routing-matrix.md", "references/escalation.md", "references/fleet.md",
|
||||
"references/gpt6-adaptation.md", "references/model-economics.md",
|
||||
"scripts/validate_decision_record.py", "scripts/validate_evidence_packet.py"):
|
||||
if not (skill / name).is_file():
|
||||
errors.append("missing skill resource: " + name)
|
||||
|
||||
@@ -1,11 +1,11 @@
|
||||
---
|
||||
name: codex-subagent-router
|
||||
description: Decide when Codex subagents help, route bounded work across Astra, Sol, Terra and Luna, and manage child evidence and lifecycle. Use for delegation requests, routing decisions, or subagent audits; inspecting this skill does not itself authorize spawning.
|
||||
description: Route authorized Codex subagents across GPT-6 Luna, Sol and Astra; manage ownership, lifecycle and evidence. Use for delegation, routing decisions or subagent audits. Discovery alone does not authorize spawning.
|
||||
---
|
||||
|
||||
# Codex Subagent Router
|
||||
|
||||
Version 2.1.0. For new general engineering sessions recommend Sol / medium, with high when demonstrated reasoning needs justify it. Preserve the user's selected parent; installation does not switch it. Use Astra proactively for high-leverage difficult decisions, independent counterexample reviews and residual hard reasoning; neither failure nor a model quota is a prerequisite.
|
||||
Version 2.2.0. For new general engineering sessions recommend GPT-6 Sol / medium, with high for demonstrated reasoning needs. Preserve the selected parent. Use Astra proactively for high-leverage difficult decisions, independent counterexample reviews and residual hard reasoning; neither prior failure nor a model quota is required.
|
||||
|
||||
Make delegation useful, observable and bounded. Preserve the selected parent model. The parent owns the critical path, shared decisions, integration and final acceptance. This skill neither changes configuration nor grants external permissions.
|
||||
|
||||
@@ -13,8 +13,8 @@ Make delegation useful, observable and bounded. Preserve the selected parent mod
|
||||
|
||||
Classify the request before using child tools:
|
||||
|
||||
- **Audit or advice without delegation authorization:** inspect rules, configuration and available tools; report findings without spawning or changing settings.
|
||||
- **Authorized delegation:** the user requested delegation, or an applicable instruction explicitly authorizes it. This includes explicitly delegated read-only reviews; delegation does not grant edit permission. Within that scope, actively delegate independent useful slices while the parent advances other work.
|
||||
- **Routing audit or advice only:** inspect rules, configuration and available tools; report findings without spawning or changing settings. Ordinary code review can use authorized read-only children under a standing parallel-work policy; it does not authorize edits.
|
||||
- **Authorized execution:** the user requested delegation, or an applicable instruction explicitly authorizes it, including an active standing parallel-work policy. Within that scope, proactively dispatch useful independent slices while the parent advances other work; do not wait for another invitation to use agents.
|
||||
- **No delegation authorization:** work locally. Automatic skill discovery, model availability, task complexity and a request to edit this skill do not grant permission to spawn.
|
||||
|
||||
Honor higher-priority host restrictions even when a lower-priority rule permits delegation. Do not request authorization repeatedly after it has been granted. Never turn a routing recommendation into a new user-owned task.
|
||||
@@ -25,24 +25,45 @@ Before dispatch classify each candidate:
|
||||
- **C:** a bounded investigation/draft that needs latest-baseline parent integration.
|
||||
- **S:** an immediate dependency, overlapping contract/state, permissions, production mutation, PR/merge/deploy or final acceptance; keep it in the parent.
|
||||
|
||||
Spawn only P/C work with a concrete output and useful parent work available. Do not spawn for one command, ceremonial probes, duplicate reviews or a model quota. Batch homogeneous small work.
|
||||
Spawn only P/C work with a concrete output and useful parallel benefit. Prefer useful parent work alongside children; when the host permits, multiple independent children may run while the parent waits to synthesize and accept their results. Honor any stricter host requirement for concurrent parent work. Do not spawn for one command, ceremonial probes, duplicate reviews or a model quota. Batch homogeneous small work.
|
||||
|
||||
For authorized complex work, actively discover independent slices at intake and each readiness change. Dispatch ready work early and refill available capacity without waiting for a whole wave. Read [fleet scheduling](references/fleet.md) for rolling execution, Astra intervention and integration backpressure. Capacity is an upper bound, not a utilization target.
|
||||
Choose the execution mechanism before a model: local tools for fixed commands/transforms, host-supported concurrent or async tools for independent I/O, and authorized children for useful independent reasoning or implementation. Waiting on a tool alone is not a reason to create an agent. Prefer a coherent outcome with acceptance over many tiny handoffs; bound ambiguity and ownership rather than prescribing every reasoning step.
|
||||
|
||||
For queue states, integration backpressure and proactive expert formations, read [fleet scheduling](references/fleet.md).
|
||||
|
||||
## Proactive Luna fleet
|
||||
|
||||
Once authorized, parallel execution is the default for ready independent work. At the first useful decomposition and after each material result or scope change, look for slices that can run alongside the parent's critical path. Dispatch those slices promptly instead of finishing the parent's entire investigation first. Do not require proof of measured speedup before a clearly useful split.
|
||||
|
||||
- Fan out bounded Luna assignments by independent evidence source, module, test suite or settled file/semantic ownership. Useful lanes include source inventory, artifact comparison, fixed validation, focused first-pass review and settled patches. Batch tiny homogeneous items into meaningful assignments.
|
||||
- Keep a small ready queue. Use as many actually available child slots as useful independent work and parent review capacity support; there is no fixed one-child or two-child ceiling. Reserve the parent when the host counts it. For example, four total active slots permit at most three concurrent children alongside the parent, not an unlimited fleet.
|
||||
- Refill available capacity when a result is handed back: accept/triage it, reuse an eligible idle child or dispatch the next ready slice. Do not wait for a whole wave if another independent slice is ready. Verify actual capacity; completed or interrupted does not necessarily mean released.
|
||||
- The parent owns contract decisions, cross-result synthesis and acceptance, and advances a parent-suitable critical-path slice when one is available. Integrate incrementally. If review backlog, conflicting ownership, I/O contention or rate limits becomes the bottleneck, reduce new dispatch and drain it.
|
||||
- One owner per file AND shared semantic contract. Parallel disjoint writes only when the host permits them; shared registries, integration and external gates remain serial. Read-only workers also need stable baselines and coordinated mutable sessions.
|
||||
- Choose Sol directly for slices requiring open-ended investigation or design, and Astra for the hardest synthesis. A fleet is a capacity strategy, not an instruction to force every task onto Luna or to duplicate the same review.
|
||||
|
||||
Stay local for a trivial task, no useful parallel benefit, an immediate dependency with no independent work, or inadequate host support. State the relevant reason briefly when a requested parallel run cannot proceed; do not produce a routing report for every small task. Standing authorization does not authorize recursive unbounded fan-out, new user-owned tasks, configuration changes or external writes.
|
||||
|
||||
## Select a supported route explicitly
|
||||
|
||||
| Work shape | Initial route |
|
||||
| --- | --- |
|
||||
| Known-source collection, fixed checks, logs | Luna low/medium |
|
||||
| Bounded classification, conversion, settled patch with fixed checks | Luna high |
|
||||
| Everyday implementation, debugging, locating/correlating artifacts | Terra medium; high when needed |
|
||||
| Bounded complex analysis, design or financial/security evidence | Sol medium; high when needed |
|
||||
| High-leverage difficult design, cross-system synthesis, independent high-impact counterexample review | Astra medium; high when needed |
|
||||
| Known-source collection, fixed checks, logs | GPT-6 Luna low; medium for multi-step work |
|
||||
| Bounded classification, conversion, settled patch with fixed checks | GPT-6 Luna medium; high for local reasoning |
|
||||
| Everyday implementation, debugging, locating/correlating artifacts | GPT-6 Sol medium |
|
||||
| Bounded complex analysis, design or financial/security evidence | GPT-6 Sol medium/high |
|
||||
| High-leverage difficult design, cross-system synthesis, independent high-impact counterexample review | GPT-6 Astra medium/high |
|
||||
| Shared decisions, integration and external/final gates | Current parent, serial |
|
||||
|
||||
These are starting heuristics, not measured cost rankings. Pick sufficient capability directly; do not escalate through every model. Missing access/data and tool failures need diagnosis, not a stronger model. After two same-class failures, pause that slice and return the evidence to the parent.
|
||||
Luna is the first candidate when inputs are bounded, the contract is settled, correctness is cheaply checkable, and failure is contained. All four conditions matter; a short diff can still hide a shared semantic decision. Use Sol directly when investigation, design choices or cross-file causal reasoning dominate. Use Astra directly for the hardest independent synthesis when expected quality or saved rework justifies it. Do not make high/max Luna a mandatory step before Sol.
|
||||
|
||||
Use the lowest adequate supported effort. Use xhigh/max for a concrete depth need or an explicit compatible role requirement; ultra only when exposed and justified by the workload. Model choice and working role are separate. A role named "explorer" is not proof of sandbox isolation.
|
||||
These are starting heuristics, not measured task-cost rankings. Lower prices broaden the useful Luna workload; fill supported capacity with useful work under existing authorization, never infer extra host slots from price. Pick sufficient capability directly; do not escalate through every model. Missing access/data and tool failures need diagnosis, not a stronger model. After two same-class failures, stop that slice and return the evidence to the parent; escalate earlier when the contract no longer fits. Preserve useful evidence on handoff instead of restarting discovery.
|
||||
|
||||
Resolve exact IDs from the child-tool schema: prefer `gpt-6-luna`, `gpt-6-sol`, `gpt-6-astra` only when exposed. GPT-5.6 models are explicit compatibility routes, not aliases for GPT-6. For a legacy-only host use [routing boundaries](references/routing-matrix.md); do not silently substitute an explicitly required generation. A role bound to GPT-5.6 Luna remains GPT-5.6 even when named `luna_worker`.
|
||||
|
||||
Retire Terra from recommended routes and automatic fallbacks: Sol owns ordinary implementation and investigation. Preserve explicitly pinned models and existing configuration; editing this skill does not authorize rewriting them.
|
||||
|
||||
Use the lowest adequate supported effort. Use none only for straightforward extraction/conversion when the host exposes it and checks cover the result; low/medium are the usual starts for tool work. Use xhigh/max for a concrete depth need or an explicit compatible role requirement; ultra only when exposed and justified by the workload, never inferred from API support. Model choice and working role are separate. A role named "explorer" is not proof of sandbox isolation.
|
||||
|
||||
At dispatch, specify model AND effort when the host permits selection. Otherwise omitted settings may inherit the parent or configured defaults. Prefer a self-contained contract and no history fork; use bounded history only when necessary. Respect full-history/override incompatibilities. Never claim prompt text changed a runtime parameter.
|
||||
|
||||
@@ -54,7 +75,9 @@ Use the first authorized, low-risk real child task as the capability observation
|
||||
|
||||
Treat workspaces as shared unless isolation is confirmed. Tool schemas may lack per-child sandbox, timeout, role or close controls. Contract limits remain instructions, not enforced capabilities. No adequate child route: keep feasible work local and disclose the gap.
|
||||
|
||||
Read [runtime and configuration](references/platforms.md) only for configuration/CLI diagnosis. Read [routing boundaries](references/routing-matrix.md) for ambiguous choices or live artifact collection.
|
||||
Read [runtime and configuration](references/platforms.md) only for configuration/CLI diagnosis. Read [routing boundaries](references/routing-matrix.md) for ambiguous choices, legacy fallbacks or live artifact collection; read [model economics](references/model-economics.md) for dated prices and cost calibration.
|
||||
|
||||
Read [GPT-6 adaptation](references/gpt6-adaptation.md) when async execution, steering, context recovery or dynamic effort changes affect the workflow. API features are not automatically Desktop/CLI capabilities; use only exposed controls.
|
||||
|
||||
## Dispatch, observe, accept
|
||||
|
||||
@@ -68,8 +91,10 @@ Use [evidence packets](references/evidence-packet.md) when structured evidence i
|
||||
|
||||
## Context and ownership discipline
|
||||
|
||||
Read [escalation and adjudication](references/escalation.md) when planning exposes uncertainty, a key check fails, shared contracts change or acceptance has unresolved counterexamples. Pause after two same-class failures; preserve issue history across agents. Repair missing inputs first, escalate capability only for reasoning/execution limits, and adjudicate by reproducible evidence rather than model rank. The parent must understand and verify decisive conclusions. Recover to a cheaper adequate route after resolution.
|
||||
Read [escalation and adjudication](references/escalation.md) when planning exposes uncertainty, a key check fails, shared contracts change or acceptance has unresolved counterexamples. Preserve issue IDs and failure counts across agents. Repair missing inputs first; adjudicate by reproducible evidence, not model rank. The parent must understand and verify decisive conclusions. Continue expert rounds only with evidence-based progress, and recover to an adequate cheaper route after resolution.
|
||||
|
||||
Load only needed references; keep long logs on disk. Avoid whole-history forks and repeating contracts, packets or prior findings. Account for parent context/reasoning, child usage, waiting, retries and integration when evaluating cost.
|
||||
|
||||
After steering or context compaction, recover the current objective, accepted corrections, permissions, baseline, file/semantic owners and pending tool/child IDs before continuing dependent work. Retain valid completed evidence; revalidate only results affected by the change. A queued correction is not proof that a running writer has stopped or adopted it.
|
||||
|
||||
Production read-only collection is routable by work shape; mutation and final semantic acceptance stay parent-owned. Do not silently relax a project's literal model-owner requirement. Identify the actual rule and resolve it only through authorized changes.
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# Sol-led escalation and evidence adjudication
|
||||
|
||||
Recommend Sol / medium for new general engineering sessions; preserve an explicit user selection. The selected parent owns planning, difficult implementation, integration and acceptance. Terra handles routine implementation; Luna handles settled, checkable work. Astra contributes both proactive high-leverage analysis and residual hard reasoning, without becoming a mandatory stage for every task. These are hypotheses to calibrate, not benchmark results.
|
||||
Recommend GPT-6 Sol / medium for new general engineering sessions; preserve an explicit user selection. The selected parent owns planning, difficult implementation, integration and acceptance. Sol handles routine implementation and investigation; Luna handles settled, checkable work. Terra is retired from recommendations and automatic fallbacks. Astra contributes both proactive high-leverage analysis and residual hard reasoning, without becoming a mandatory stage for every task. These are hypotheses to calibrate, not benchmark results.
|
||||
|
||||
## Reassess at observable checkpoints
|
||||
|
||||
@@ -20,8 +20,8 @@ Continue independent authorized work while the affected slice is paused. A lack
|
||||
## Choose the intervention
|
||||
|
||||
1. Repair the contract, baseline, inputs or verifier first. Reuse an appropriate idle agent; do not reset failure history by respawning.
|
||||
2. Increase execution capability when a well-specified slice exceeds the worker: Luna to Terra, or Terra to Sol as justified. Select the adequate model directly, without visiting every tier. Effort, model and concurrency are independent controls.
|
||||
3. Select Astra directly when a difficult architectural choice affects many downstream tasks, competing consequential designs need discrimination, cross-system evidence needs synthesis, or an independent counterexample search materially protects a high-impact change. Repeated failure is not required. State the uncertainty, consequence of a wrong decision, available evidence and concrete expert deliverable. Residual difficult reasoning conflicts remain a valid trigger. Missing access alone is not a trigger. Authorized delegation and an independent bounded consultation are still required; the parent advances useful independent work and retains the shared decision. An immediate parent dependency stays local; requesting a child or writing configuration does not switch the live parent. If the parent cannot proceed reliably, describe the required capability/evidence instead of claiming a switch occurred.
|
||||
2. Increase execution capability when a well-specified slice exceeds the worker: Luna to Sol as justified, or Astra directly for the hardest independent synthesis. Select the adequate model directly, without visiting every tier. Effort, model and concurrency are independent controls.
|
||||
3. Select Astra directly when a difficult architectural choice affects many downstream tasks, competing consequential designs need discrimination, cross-system evidence needs synthesis, or an independent counterexample search materially protects a high-impact change. Repeated failure is not required. State the uncertainty, consequence of a wrong decision, available evidence and concrete expert deliverable. Residual difficult reasoning conflicts remain a valid trigger. Missing access alone is not a trigger. Authorized delegation and an independent bounded consultation are still required; the parent advances useful independent work when available and retains the shared decision. Where the host permits, useful parallel consultations may run while the parent waits for later adjudication; honor stricter host requirements. An immediate parent dependency stays local; requesting a child or writing configuration does not switch the live parent. If the parent cannot proceed reliably, describe the required capability/evidence instead of claiming a switch occurred.
|
||||
|
||||
An expert packet contains: issue ID, exact relevant baseline, goal, invariants, original evidence references, competing hypotheses, discriminators attempted or explicitly not yet run, and one focused question. For proactive work, zero failures is valid. Request counterexamples, a discriminating check, conditional conclusions and unresolved points. Supply only necessary context. No implementation writes unless separately assigned a disjoint write contract.
|
||||
|
||||
|
||||
@@ -30,7 +30,7 @@ This fragment illustrates shape; it is not a real tool receipt:
|
||||
"tool": "collaboration.spawn_agent",
|
||||
"child_id": "synthetic-child",
|
||||
"status": "completed",
|
||||
"requested": {"model": "gpt-5.6-terra", "role": "explorer", "effort": "medium"},
|
||||
"requested": {"model": "gpt-6-sol", "role": "explorer", "effort": "medium"},
|
||||
"observed": {"model": "unknown", "role": "unknown", "effort": "unknown"},
|
||||
"identity_source": "unknown"
|
||||
}
|
||||
|
||||
@@ -11,8 +11,8 @@ Maintain a compact queue: task ID, dependency and contract baseline, file AND se
|
||||
At intake, dependency resolution, a child return or a meaningful baseline change:
|
||||
|
||||
1. Invalidate affected assumptions and identify ready slices. Prioritize work that unlocks dependencies, avoids costly wrong decisions or yields directly usable implementation.
|
||||
2. Check live capacity and integration load. Dispatch multiple independent ready tasks when useful parent work remains; avoid sequential startup followed by unnecessary waits. Every descendant counts against the actual host limit. Children may not recursively expand the fleet without a parent-assigned bounded delegation scope and capacity allocation.
|
||||
3. Continue the parent's independent critical work. Use incremental messages for running children and reuse suitable idle ones. Do not perform the same assigned investigation locally merely to stay busy.
|
||||
2. Check live capacity and integration load. Dispatch multiple independent ready tasks with useful parallel benefit; prefer concurrent parent work, but allow the parent to wait for later synthesis when the host permits. Honor stricter host requirements for concurrent parent work. Avoid sequential startup followed by unnecessary waits. Every descendant counts against the actual host limit. Children may not recursively expand the fleet without a parent-assigned bounded delegation scope and capacity allocation.
|
||||
3. Continue the parent's independent critical work when available. Use incremental messages for running children and reuse suitable idle ones. Do not perform the same assigned investigation locally merely to stay busy.
|
||||
4. Inspect returned evidence and partial writes; integrate serially. Refill available capacity with ready work without waiting for all siblings. A returned result is not necessarily accepted or a released slot; follow actual lifecycle controls.
|
||||
|
||||
If the next action is immediately dependent on an unassigned task, keep it local. A previously independent task may become a dependency as work progresses; bounded waiting is then legitimate. No independent useful parent work means do not create a ceremonial side task to justify another spawn.
|
||||
@@ -20,7 +20,7 @@ If the next action is immediately dependent on an unassigned task, keep it local
|
||||
## Choose a formation for the task
|
||||
|
||||
- Independent investigation: several focused investigators may cover distinct sources, consumers or falsifiable hypotheses. They should not all produce the same broad summary.
|
||||
- Settled implementation: multiple Terra/Luna workers own complete independently checkable slices. Shared registries, schema and generated outputs require one owner even when paths differ.
|
||||
- Settled implementation: multiple GPT-6 Luna workers own bounded, settled, independently checkable slices; use GPT-6 Sol directly when implementation requires investigation or design judgment. Shared registries, schema and generated outputs require one owner even when paths differ.
|
||||
- Difficult consequential decision: an Astra specialist analyzes the bounded uncertainty while the parent and other workers advance unaffected work. Gate dependent implementation on the parent's accepted contract. Read [expert intervention](escalation.md) for triggers and progress-based continuation.
|
||||
- Independent review: assign a stable commit/artifact and the relevant requirements, without supplying the parent's preferred verdict. Review can overlap unrelated work; stale observations must be rechecked before acceptance.
|
||||
|
||||
|
||||
@@ -0,0 +1,39 @@
|
||||
# GPT-6 workflow adaptation
|
||||
|
||||
Checked 2026-09-23. Read only when these capabilities affect the task. This reference adapts the routing policy; it does not configure a harness or grant permissions.
|
||||
|
||||
## Prompt shape and task granularity
|
||||
|
||||
The official [skills and prompts guidance](https://developers.openai.com/blog/rethinking-skills-and-prompts-for-gpt-6-astra) recommends precise discovery descriptions, progressive disclosure and fewer rigid recipes. Apply that here by giving a goal, constraints, ownership and acceptance; leave implementation choices open unless a contract or fragile operation requires an exact procedure. Put API mechanics here instead of expanding every task prompt.
|
||||
|
||||
[GPT-6 guidance](https://developers.openai.com/api/docs/guides/latest-model) highlights Astra's stronger instruction following, long-task coherence and tendency to ask clarifying questions or over-test. Keep authorized routine choices autonomous, resolve real instruction conflicts explicitly, and stop validation once relevant checks pass unless new evidence warrants more. These are Astra observations; evaluate their usefulness for Sol/Luna rather than claiming equal behavior. More capability does not authorize delegation.
|
||||
|
||||
## Async tools: overlap useful work
|
||||
|
||||
[Async tool calling](https://developers.openai.com/api/docs/guides/async-tool-calling) lets GPT-6 continue independent work while an application executes a function/custom tool; API definitions use `async: true`, with results paired to the original `call_id`. Background response generation is a different mechanism.
|
||||
|
||||
In Codex, use the host's exposed job/session IDs, wait and cancellation controls. Track pending work only as needed: ID, owner, baseline, dependency, last state and result. Continue unrelated work while waiting, but do not consume an absent result or declare completion with a required job outstanding. Mutable shared state and dependencies still require serialization. A JavaScript Promise or a shell job is not proof of API async support.
|
||||
|
||||
Prefer a fixed tool batch for enumerable reads and transformations. Use a child when the independent slice needs judgment and delegation is authorized. Do not spawn a child simply to wait on another tool. In an API application, the application must execute and reconcile jobs; prompt text cannot add that machinery.
|
||||
|
||||
## Steering: preserve progress, reconcile side effects
|
||||
|
||||
[Mid-turn steering](https://developers.openai.com/api/docs/guides/steering) is documented for GPT-6 over Responses WebSockets. Acceptance queues the update; it does not cancel running tools, undo writes or establish that the model acted on it.
|
||||
|
||||
Treat a user correction as a delta to the active objective. Keep compatible work and accepted evidence. Tell affected children the changed constraint; if ongoing writes conflict, use actual interruption controls, inspect partial changes and verify state before assigning a replacement owner. Reassess late results against the corrected contract and baseline before integration. A status question does not cancel work. Cancellation ends obsolete work, not the obligation to report actions already taken.
|
||||
|
||||
The Desktop/CLI message and interruption tools have their own semantics; do not emit WebSocket events through tools that lack those fields. A successful message-send receipt proves delivery/queuing only to the extent the host reports it.
|
||||
|
||||
## Effort and context continuity
|
||||
|
||||
[Reasoning configuration updates](https://developers.openai.com/api/docs/guides/reasoning#change-reasoning-mid-conversation) support changing effort between responses in GPT-6 standard single-agent mode. They do not switch model or change a child binding. Keep request-level effort unchanged for cache continuity; record updates in order. The response's effort field still reports the request-level value. Updates cannot be adjacent or combined with automatic compaction/truncation; `/responses/compact` rejects such histories. The documented explicit `compaction_trigger` route requires a fresh effort update after compaction. Recheck compatibility before implementing.
|
||||
|
||||
In Codex, do not claim to change active effort without a supported control and receipt. Choose supported effort when dispatching an authorized child. Missing controls are a host limitation, not a reason to launch a nested CLI or edit global configuration.
|
||||
|
||||
For long tasks, retain a compact operational checkpoint when needed: goal, corrections, authorization, baseline, completed evidence, next dependency, owners and pending IDs. Preserve actual tool state and required conversation items; a prose summary is not an API tool result. After host compaction, resume from that state rather than rereading everything or repeating accepted work. Large context capacity is not a reason to fork full history into every child.
|
||||
|
||||
## Existing capabilities and boundaries
|
||||
|
||||
The [GPT-6 family guide](https://developers.openai.com/api/docs/guides/latest-model) identifies computer use, structured outputs, programmatic tool calling, multi-agent orchestration, caching and compaction as continuing capabilities, not all newly introduced in GPT-6. Use structured tools and deterministic code where they fit; use visual/browser tools when UI evidence is actually needed. Confirm host access and shared browser/session ownership. Neither a screenshot nor a passing schema validator proves end-to-end correctness.
|
||||
|
||||
The guide's misalignment monitoring is provider-side safety behavior associated with Astra, not a prompt-enabled self-review mode or a substitute for parent acceptance. Do not claim it is available identically for all routes.
|
||||
@@ -16,6 +16,14 @@ Read before multi-round coordination, cancellation, or capacity diagnosis. Names
|
||||
|
||||
A minimal ledger records child ID, P/C classification, file/semantic ownership, dependency, requested route, last observed state and acceptance. Do not invent timestamps, identities or close events.
|
||||
|
||||
## Steering and pending tools
|
||||
|
||||
On a correction, update the affected contract and preserve unaffected evidence. Sending an incremental message does not establish that a running child or tool has adopted it. If its writes now conflict, interrupt through available controls, verify the state and inspect partial changes before transferring ownership. Review late results against the current contract and baseline; do not integrate them merely because they completed successfully.
|
||||
|
||||
Keep required async tool jobs in the same dependency view as children, using their actual job/session/call IDs. While waiting, advance independent work; before dependent acceptance, collect the real result. Steering and child interruption need not cancel external jobs. Cancel obsolete jobs through the host when supported, otherwise track/report the remaining state. Never resubmit a potentially mutating call solely because its receipt is delayed.
|
||||
|
||||
After context compaction, recover objective, permissions, owners, pending IDs and accepted evidence before dispatching more work. Detailed API-versus-host limits are in [GPT-6 adaptation](gpt6-adaptation.md).
|
||||
|
||||
For rolling execution, use [fleet scheduling](fleet.md): refill ready work after useful handbacks and dependency changes, rather than waiting for every sibling. Check integration backlog before starting more implementation. Task readiness, output acceptance and actual slot availability remain separate facts.
|
||||
|
||||
## Cancellation and handback
|
||||
|
||||
@@ -0,0 +1,33 @@
|
||||
# Model economics
|
||||
|
||||
Read for cost-sensitive routing or when revising defaults. Checked 2026-09-23; recheck official prices before a later price-dependent decision. This is an API reference snapshot, not a Codex subscription quota or billing estimate.
|
||||
|
||||
## Verified price basis
|
||||
|
||||
USD per million tokens, Standard processing, prompts up to 272K input tokens:
|
||||
|
||||
| Exact model ID | Input | Cached input | Cache write | Output |
|
||||
| --- | ---: | ---: | ---: | ---: |
|
||||
| gpt-6-luna | 0.10 | 0.01 | 0.125 | 0.50 |
|
||||
| gpt-6-sol | 2.00 | 0.20 | 2.50 | 10.00 |
|
||||
| gpt-6-astra | 10.00 | 1.00 | 12.50 | 50.00 |
|
||||
|
||||
Source: [OpenAI pricing](https://developers.openai.com/api/docs/pricing). For equal token volumes and the same billing category, Luna is 1/20 of Sol and 1/100 of Astra; Sol is 1/5 of Astra. These are arithmetic price ratios, not observed savings, speed or success rates.
|
||||
|
||||
Long prompts above 272K input tokens charge 2x input/cache rates and 1.5x output for the full request. Processing tier, tools and regional processing can change the bill. Do not multiply Codex subscription usage by API prices or assume a Desktop service tier follows the Standard table.
|
||||
|
||||
Retirement comparison: [GPT-5.6 Terra](https://developers.openai.com/api/docs/models/gpt-5.6-terra) lists input 2.00, cached input 0.20 and output 12.00 under the same short-context Standard basis. GPT-6 Sol has equal input/cache-read rates and a 16.7% lower output rate (10 vs 12). Terra therefore offers no token-rate saving in these categories; this does not prove equal token use, latency or quality per task.
|
||||
|
||||
## Design implications
|
||||
|
||||
Official positioning: [Luna](https://developers.openai.com/api/docs/models/gpt-6-luna) targets focused, high-volume work; [Sol](https://developers.openai.com/api/docs/models/gpt-6-sol) targets complex coding and agentic workflows; [GPT-6 guidance](https://developers.openai.com/api/docs/guides/latest-model) places Astra at the highest capability level. Sol and Luna document API efforts none/low/medium/high/xhigh/max. Actual child models, efforts and role bindings must still come from the host schema.
|
||||
|
||||
The engineering policy is three routes: Luna for bounded, settled, cheaply verifiable work; Sol for implementation and investigation requiring judgment; Astra for the hardest independent synthesis. Retire Terra from recommendations and automatic fallbacks: Sol absorbs its former workload and avoids maintaining another routing boundary. This is a routing design choice, not a measured claim that Sol dominates every Terra workload. Price alone does not establish that Luna can perform open-ended development reliably.
|
||||
|
||||
## Measure accepted work
|
||||
|
||||
Compare the same task shape and acceptance standard. Include parent context/reasoning, every child's usage, tools, retries, parent review and reintegration. For API estimates apply the actual rate to each billed category, including billed reasoning/output tokens without double-counting; keep wall time separate. With unknown usage or billing tier, report unknown cost and observable proxies instead of fabricated savings.
|
||||
|
||||
A cheaper child helps only if its savings exceed added orchestration and correction. One avoided parent investigation may matter more than many cheap child tokens. Do not turn the 20x ratio into a twenty-retry allowance, a fan-out quota or permission to duplicate reviews. Keep concurrency tied to useful independent work and actual host capacity.
|
||||
|
||||
Broaden Luna assignments after comparable accepted results demonstrate stable quality and lower total effort. Move a task shape directly to Sol when ambiguity, repeated correction or parent rework dominates. For a new generation, use the first authorized useful bounded task as evidence; synthetic parser tests and static scenario reviews cannot establish model quality or latency.
|
||||
@@ -19,17 +19,19 @@ For clients that support the documented keys, an example is:
|
||||
|
||||
```toml
|
||||
# Example only; merge intentionally into the existing configuration.
|
||||
model = "gpt-5.6-sol"
|
||||
model = "gpt-6-sol"
|
||||
model_reasoning_effort = "medium"
|
||||
|
||||
[agents]
|
||||
default_subagent_model = "gpt-5.6-terra"
|
||||
default_subagent_model = "gpt-6-sol"
|
||||
default_subagent_reasoning_effort = "medium"
|
||||
max_concurrent_threads_per_session = 3
|
||||
```
|
||||
|
||||
This is not an automatic router, permission grant or universal host configuration. Never overwrite an existing config to install this example. Preserve user model/provider choices and verify the actual client accepts the keys. Per-agent configuration and host overrides may change resolution.
|
||||
|
||||
The Sol default assumes GPT-6 Sol is exposed and suits mixed engineering work. Explicitly route bounded tasks to GPT-6 Luna where supported; do not make Luna a universal fallback for arbitrary inherited work. If only GPT-5.6 children are available, use the [legacy compatibility routes](routing-matrix.md). A fixed `luna_*` role may still bind GPT-5.6 Luna/max: inspect the binding, then use an overridable role if available and appropriate. Do not rename or claim to upgrade a fixed role through prompt text.
|
||||
|
||||
The official Codex configuration describes the child-thread limit as excluding the parent. A collaboration tool may expose a different active-slot count including the parent. Use that live tool's counting and release semantics; do not copy either number blindly.
|
||||
|
||||
When selecting a route, explicitly request model and effort with a compatible fork mode if supported. Check omitted-setting inheritance, custom-agent bindings and override restrictions. A prompt cannot implement an unsupported model switch.
|
||||
|
||||
@@ -4,8 +4,11 @@ Read only when the entrypoint leaves a routing choice unresolved.
|
||||
|
||||
| Boundary | Decision |
|
||||
| --- | --- |
|
||||
| Many easy tasks | Batch bounded Luna/Terra slices; volume alone does not justify Astra. |
|
||||
| Ordinary cross-file bug | Terra; file count alone does not justify escalation. |
|
||||
| Many easy tasks | Partition independent bounded Luna batches across available capacity; size batches to keep coordination and acceptance useful. Volume alone does not justify Astra. |
|
||||
| Ordinary cross-file bug | GPT-6 Sol; file count alone does not justify Astra. |
|
||||
| Small patch with settled semantics and deterministic acceptance | GPT-6 Luna medium/high if failure is contained; parent checks the diff and behavior. |
|
||||
| Small patch changes authentication or a shared schema contract | Parent settles shared decisions; bounded investigation may use Sol. Small size is not evidence of low ambiguity. |
|
||||
| Cheap Luna attempts require repeated parent rewrites | Choose Sol directly next time for this shape; count total accepted-task cost. |
|
||||
| Difficult accounting discrepancy | Sol reasoning; parent owns the semantic decision. |
|
||||
| Coupled architecture across runtimes and tools | Consider Astra directly; keep shared decisions serial. |
|
||||
| Difficult choice would constrain many downstream slices | Consider an early bounded Astra analysis; gate dependent work on the parent's accepted contract. |
|
||||
@@ -17,9 +20,21 @@ Read only when the entrypoint leaves a routing choice unresolved.
|
||||
| Repository requires Sol parent | A Sol child cannot replace that owner; preserve the gate until an authorized supported change. |
|
||||
| Accepted output, unknown child model | Accept only after normal parent checks; do not attribute the result to the requested model. |
|
||||
|
||||
## Legacy and mixed hosts
|
||||
|
||||
Check each route independently. The API catalog does not update an existing session's tool schema.
|
||||
|
||||
| Unavailable preferred route | Supported compatibility choice |
|
||||
| --- | --- |
|
||||
| GPT-6 Luna | GPT-5.6 Luna low/medium for fixed collection/checks, high for settled transformations or patches; otherwise Sol or local work. |
|
||||
| GPT-6 Sol for implementation/investigation/complex analysis | GPT-5.6 Sol medium/high, if exposed and adequate; otherwise local work. |
|
||||
| GPT-6 Astra | An adequate supported Sol route or local work; disclose a capability gap if neither suffices. |
|
||||
|
||||
These are work-shape fallbacks, not capability equivalence claims. Terra is retired from recommended routes and automatic fallbacks; do not select it merely because a legacy host exposes it. Preserve an explicit user-pinned model or existing configuration until an authorized change. If a required model is unavailable, report the exact gap before dependent dispatch. Never invent `gpt-6-terra`. A mixed host may use GPT-6 Sol and GPT-5.6 Luna without reverting every route. Record the exact requested ID and fallback reason. Unknown observed identity stays unknown.
|
||||
|
||||
## Production Artifact Collection
|
||||
|
||||
- Luna: use known hosts/endpoints/paths and fixed selectors to read or download actual reports, JSON, logs or job outputs; verify size/hash, parseability, required fields and specified counts. Terra: locate the current output through bounded read-only investigation, diagnose retrieval failures, or correlate logs and artifacts by run/version. Do not launch both by default; use Terra only when discovery or diagnosis is needed.
|
||||
- GPT-6 Luna: use known hosts/endpoints/paths and fixed selectors to read or download actual reports, JSON, logs or job outputs; verify size/hash, parseability, required fields and specified counts. GPT-6 Sol: locate the current output through bounded read-only investigation, diagnose retrieval failures, or correlate logs and artifacts by run/version. Do not launch both by default; use Sol only when discovery or diagnosis is needed. Apply the legacy table when these routes are unavailable.
|
||||
- Contract: name the production source, expected run/date/version, allowed read operations, bounded time/range/volume, local destination and acceptance checks. Reuse approved access and applicable server tools. Keep secrets out of prompts/receipts. Remote access stays read-only; explicitly allow local artifact writes even for a collector role.
|
||||
- Evidence: preserve the actual downloaded artifact and record source, retrieval time/timezone, producer run/version and generation time when exposed, local path, bytes/hash and check results. Unknown provenance stays unknown. Verify the expected run/freshness; mtime or successful download alone does not prove current output. Use immutable run identifiers where possible; detect/retry boundedly if files change during collection.
|
||||
- No fixture, cache or old snapshot may stand in for requested live output. If the current artifact is missing, stale, inconsistent or inaccessible, report that gap. Do not trigger jobs, regenerate artifacts, restart services, change permissions/configuration or deploy as part of collection; return such remediation to the parent under existing authorization.
|
||||
@@ -33,4 +48,4 @@ Use actual usage counters when exposed. Otherwise report observable proxies (cha
|
||||
|
||||
If context dominates, narrow inputs and avoid full-history forks. If output dominates, bound logs and returns. If reasoning/retries dominate, clarify acceptance and choose adequate capability/effort directly. Preserve required tests and material findings.
|
||||
|
||||
Official model positioning checked 2026-09-14: [Models](https://developers.openai.com/api/docs/models). The route table is an engineering starting strategy, not a measured cost ranking. Runtime tool schemas remain the source for actual child availability.
|
||||
Official model positioning checked 2026-09-23: [GPT-6 guidance](https://developers.openai.com/api/docs/guides/latest-model). See [model economics](model-economics.md) for the dated API price basis. The route table is an engineering starting strategy, not a measured task-cost ranking. Runtime tool schemas remain the source for actual child availability.
|
||||
|
||||
@@ -30,7 +30,7 @@ class EvidenceTests(unittest.TestCase):
|
||||
def test_known_identity_requires_host_evidence(self):
|
||||
packet = fixture()
|
||||
child = packet["tool_receipts"][0]
|
||||
child["observed"]["model"] = "gpt-5.6-terra"
|
||||
child["observed"]["model"] = "gpt-6-sol"
|
||||
with self.assertRaisesRegex(ValueError, "host evidence"):
|
||||
evidence.validate(packet)
|
||||
child["identity_source"] = "host-tool-result"
|
||||
|
||||
Reference in New Issue
Block a user