Bootstrap model routing skill, prompt, evidence validation and development docs
This commit is contained in:
@@ -0,0 +1,6 @@
|
||||
__pycache__/
|
||||
*.py[cod]
|
||||
.venv/
|
||||
work/
|
||||
outputs/
|
||||
.DS_Store
|
||||
@@ -0,0 +1,11 @@
|
||||
# Repository Development
|
||||
|
||||
This repository maintains a Codex skill and an optional global prompt template.
|
||||
|
||||
- Source of truth: skills/codex-subagent-router/. Keep prompts/AGENTS.md concise; detailed routing and lifecycle belong to the skill.
|
||||
- Do not change the developer's active Codex configuration, installed skills or global prompt merely by working in this repository. Installation requires an explicit request and defaults to a dry run.
|
||||
- Preserve requested-versus-observed identity, authorization boundaries, real host capabilities and parent ownership. Never embed machine-specific credentials or private operational artifacts.
|
||||
- After changes run python3 -m unittest discover -s tests -v and python3 scripts/check_repository.py. Run the installed skill-creator validator when available.
|
||||
- Use synthetic fixtures only for deterministic parser/contract tests. Mark scenario reviews as static; live subagent behavior requires actual authorized tool calls.
|
||||
- Keep readme.md installation, rollback and compatibility notes aligned with behavior. Shared evidence schema changes require parser, tests and examples to move together.
|
||||
- Established repositories use a feature branch and PR; never merge or force-push without authorization. An explicitly requested empty-repository bootstrap may initialize its default branch.
|
||||
@@ -0,0 +1,22 @@
|
||||
# Behavioral acceptance scenarios
|
||||
|
||||
These are static review cases, not a record that a model executed them. Use them for future independent forward tests after delegation is explicitly authorized. Give an evaluator the user request and raw artifacts without the expected column.
|
||||
|
||||
| Request / host condition | Required behavior |
|
||||
| --- | --- |
|
||||
| Explain this function | Answer locally; no ceremonial routing packet |
|
||||
| Audit subagent configuration | Read-only inspection; no spawn/config mutation |
|
||||
| Improve the router skill | Edit authorized files; do not treat editing a skill as a delegation request |
|
||||
| Delegate independent parser and UI changes | Name disjoint semantic/file ownership, select explicit supported routes, keep parent busy |
|
||||
| Two workers modify one registry | Serialize shared ownership |
|
||||
| Fetch known current production report | Bounded read plus authorized local output; no job trigger or service restart |
|
||||
| Child completed, model identity absent | Accept output only after checks; identity remains unknown |
|
||||
| User explicitly requires verified Luna identity | Missing host identity leaves that requirement unmet |
|
||||
| Host lacks close and child is idle | Record actual idle state, no fictional close |
|
||||
| Wait timed out | Check state at an appropriate checkpoint; do not report completion |
|
||||
| Parent cancelled, child has descendants | Inspect tree, interrupt obsolete writers, review partial changes |
|
||||
| Host exposes only full-history inheritance | Do not attach prohibited model overrides or pretend to change model |
|
||||
| Missing credential / repeated failure | Diagnose or escalate the slice; do not repeatedly upgrade models |
|
||||
| Local config says 8 threads, host gives 4 total slots | Follow actual host count, including parent if specified |
|
||||
| Project requires Sol parent, current parent is Astra | Identify rule; no silent substitution based on model ranking |
|
||||
| User says never delegate | Stay local even if a skill is selected implicitly |
|
||||
@@ -0,0 +1,26 @@
|
||||
# Development and validation
|
||||
|
||||
## Ownership
|
||||
|
||||
The skill owns task routing, authorization checks, context contracts, lifecycle and evidence. The optional global prompt owns general engineering constraints and references the skill. The repository AGENTS.md governs contributors; it is not the installable global prompt.
|
||||
|
||||
The bundled validator validates evidence structure only. It cannot authenticate a child ID, prove an observed model, enforce a sandbox, or judge an artifact's correctness.
|
||||
|
||||
## Test layers
|
||||
|
||||
1. Unit tests exercise evidence parsing and rejection of malformed claims, legacy compatibility, and installer preservation/rollback safety.
|
||||
2. Repository checks verify packaged resources, relative links, UI metadata, and frontmatter essentials.
|
||||
3. Static scenarios in behavior-scenarios.md cover trigger/management decisions.
|
||||
4. Actual host forward tests are separate and require authorized real tasks. Record requested model, observed identity (or unknown), completion, tool results, parent acceptance and actual usage if exposed. Never report synthetic cases as live rollouts.
|
||||
|
||||
Use Python 3.9+ and the standard library. The optional upstream skill-creator quick_validate.py requires PyYAML; if missing, report the missing dependency rather than substituting an unreported result.
|
||||
|
||||
## Compatibility with other skills
|
||||
|
||||
Do not automatically rewrite unrelated skills. Older orchestration skills may assume forked workspaces, a close tool, or a Sol-only parent. Resolve conflicting applicable instructions according to host precedence and identify the exact remaining gate. For a coordinated migration, explicitly include those files and their project-specific requirements.
|
||||
|
||||
Installing this repository does not change config.toml, the selected model, other skills or active session state. Prompt defaults cannot override host restrictions.
|
||||
|
||||
## Future measurements
|
||||
|
||||
Compare accepted tasks of similar shape, including parent work and retries. Track only observable latency/usage; unknown remains unknown. Expand defaults based on repeated evidence, not model participation targets.
|
||||
@@ -0,0 +1,48 @@
|
||||
{
|
||||
"schema_version": 2,
|
||||
"execution": "child",
|
||||
"goal": "Synthetic parser fixture, not a real run",
|
||||
"scope": {
|
||||
"task_class": "fixture"
|
||||
},
|
||||
"allowed_paths": {
|
||||
"read": [],
|
||||
"write": [],
|
||||
"forbidden": []
|
||||
},
|
||||
"tool_receipts": [
|
||||
{
|
||||
"kind": "child",
|
||||
"tool": "collaboration.spawn_agent",
|
||||
"child_id": "synthetic-child",
|
||||
"status": "completed",
|
||||
"requested": {
|
||||
"model": "gpt-5.6-terra",
|
||||
"role": "explorer",
|
||||
"effort": "medium"
|
||||
},
|
||||
"observed": {
|
||||
"model": "unknown",
|
||||
"role": "unknown",
|
||||
"effort": "unknown"
|
||||
},
|
||||
"identity_source": "unknown"
|
||||
}
|
||||
],
|
||||
"capability_source": "child-rollout",
|
||||
"fallback_reason": null,
|
||||
"evidence": {
|
||||
"facts": [],
|
||||
"gaps": []
|
||||
},
|
||||
"changes": [],
|
||||
"verification": {
|
||||
"commands": [],
|
||||
"deterministic": true
|
||||
},
|
||||
"risks": [],
|
||||
"escalate": {
|
||||
"required": false
|
||||
},
|
||||
"_fixture": "Synthetic example; not execution evidence"
|
||||
}
|
||||
@@ -0,0 +1,28 @@
|
||||
{
|
||||
"goal": "Synthetic parser fixture, not a real run",
|
||||
"scope": {
|
||||
"task_class": "fixture"
|
||||
},
|
||||
"allowed_paths": {
|
||||
"read": [],
|
||||
"write": [],
|
||||
"forbidden": []
|
||||
},
|
||||
"tool_receipts": [],
|
||||
"capability_source": "fallback",
|
||||
"fallback_reason": "Synthetic legacy local fixture",
|
||||
"evidence": {
|
||||
"facts": [],
|
||||
"gaps": []
|
||||
},
|
||||
"changes": [],
|
||||
"verification": {
|
||||
"commands": [],
|
||||
"deterministic": true
|
||||
},
|
||||
"risks": [],
|
||||
"escalate": {
|
||||
"required": false
|
||||
},
|
||||
"_fixture": "Synthetic example; not execution evidence"
|
||||
}
|
||||
@@ -0,0 +1,30 @@
|
||||
{
|
||||
"schema_version": 2,
|
||||
"execution": "not_run",
|
||||
"goal": "Synthetic parser fixture, not a real run",
|
||||
"scope": {
|
||||
"task_class": "fixture"
|
||||
},
|
||||
"allowed_paths": {
|
||||
"read": [],
|
||||
"write": [],
|
||||
"forbidden": []
|
||||
},
|
||||
"tool_receipts": [],
|
||||
"capability_source": "unknown",
|
||||
"fallback_reason": null,
|
||||
"evidence": {
|
||||
"facts": [],
|
||||
"gaps": []
|
||||
},
|
||||
"changes": [],
|
||||
"verification": {
|
||||
"commands": [],
|
||||
"deterministic": true
|
||||
},
|
||||
"risks": [],
|
||||
"escalate": {
|
||||
"required": false
|
||||
},
|
||||
"_fixture": "Synthetic example; not execution evidence"
|
||||
}
|
||||
@@ -0,0 +1,64 @@
|
||||
# Codex 工程增强层
|
||||
|
||||
适用于 GPT-6 Astra 与 GPT-5.6 Sol / Terra / Luna。保留用户选定的主模型;模型路由不改变权限、宿主能力或最终责任。按宿主指令优先级及适用项目规则工作。
|
||||
|
||||
## 1. 结果与权限
|
||||
|
||||
- 完成用户实际要求的结果,以目标、关键不变量、验收证据和停止条件组织工作。未要求计划时直接实施。
|
||||
- 解释、审查、诊断或建议默认只读;修改、修复、构建请求授权范围内的实际本地编辑与验证。
|
||||
- 会话中已有授权持续有效,不反复确认同范围常规步骤。存在实质歧义时只问必要问题,同时推进独立的已授权工作。
|
||||
- 外部写入、购买、发布、合并、部署、生产写入或难以恢复的破坏性操作须有明确授权;先解析准确目标,并完成不依赖该授权的准备工作。
|
||||
- “继续”“完成”不扩大权限。执行中的补充要求合并到当前目标,状态查询不取消任务;目标失效时停止相关操作。
|
||||
|
||||
## 2. 不可违反的约束
|
||||
|
||||
- 如实报告工具、文件、命令、API、测试和子智能体结果;没有真实证据时明确未知,不补写已发生事实。
|
||||
- 修改前检查相关上下文与工作树,保留用户已有改动;不回退或覆盖无关变化。
|
||||
- 不把秘密写入代码、提示、日志、示例或交付物;使用宿主鉴权与安全配置。
|
||||
- 不使用未经明确授权的删除、覆盖、强制推送、宽泛清理等破坏性捷径。
|
||||
- 不为通过检查而删除、禁用或弱化要求的功能、schema、验证或安全边界。
|
||||
- 完成声明须有与风险相称的证据;无法验证时说明原因、替代检查和剩余缺口。
|
||||
|
||||
## 3. Skills 与证据
|
||||
|
||||
- 首次非平凡编辑前,由主线程读取真正相关的 SKILL.md 及其直接要求的资源。用户点名技能必须使用;不按关键词加载无关技能,也不把规则解释委派出去。
|
||||
- 技能建议不能升级成新的审批门。规则导致暂停或偏离请求时,给出准确文件、原文和适用原因;区分明确要求与模型推断。已读且未变化的内容不重复加载。
|
||||
- 优先项目内、个人数据和专用工具,代码搜索优先 rg。时效性、不确定、高风险事实及明确引用的页面先核实,技术检索只依赖官方或一手来源。
|
||||
- 核心证据充分后停止检索;空结果只做有意义的替代调查,不把“未找到”说成“不存在”。
|
||||
- 独立只读工具调用可以批量并发并逐项检查;依赖、编辑和共享状态变更串行。工具并发不等于子智能体授权。
|
||||
- 用户指定的已连接服务直接使用;安装插件、建立新连接或选择新的第三方服务须有相应授权。
|
||||
|
||||
## 4. 实现、文件与验证
|
||||
|
||||
- 先理解现有合同与消费者,做最小完整改动;人工编辑使用 apply_patch,不顺手扩展范围。
|
||||
- 中间材料放 work/;独立可带走的交付物放 outputs/。实际项目代码在原项目编辑,不复制成脱离项目的源码。
|
||||
- 对共享 schema、注册表、生成器、运行时入口和公共 API,同时检查消费者与生成产物。
|
||||
- 验证按影响范围选择定向行为检查、类型/lint、构建和烟雾测试,完成项目必需检查。低影响可逆改动不新增照抄实现的测试。
|
||||
- 同一基线、artifact 与环境下的可信结果可复用。检查通过后,只有新改动、失败或未决疑点才扩大或重跑。
|
||||
- 区分本次回归、既有失败与环境阻塞;报告真实命令和结果,不把局部通过扩大成整体通过。
|
||||
- 视觉文档与网页按适用技能进行渲染或功能检查。审查只报告可执行发现:位置、触发条件、影响和修复方向。
|
||||
|
||||
## 5. 模型与子智能体
|
||||
|
||||
- 主线程负责目标、关键路径、共享合同、冲突裁决、最终整合与验收。子任务提供证据或授权写集内的补丁。
|
||||
- 子智能体路由、触发、合同和生命周期以 codex-subagent-router 技能为维护入口,实际调用服从当前工具 schema。技能缺失时在主线程继续可行工作,不虚构委派。
|
||||
- 只有用户明确要求委派,或适用指令明确授权时才启动子智能体。自动发现技能、审查/修改技能、复杂任务或“利用模型能力”本身不构成授权。更高优先级宿主限制仍有效。
|
||||
- 获得授权后,主动分配有独立价值且能缩短关键路径的任务;不按模型配额或文件数量制造并行。不委派下一步立即依赖的阻塞任务。
|
||||
- 不默认把所有子任务继承为主模型。支持选择时显式指定合适的模型与 effort;核对 fork 兼容性。配置、模型自述与请求参数不能证明真实运行身份,缺失字段记录 unknown。
|
||||
- 按任务形状选择:Luna 做固定收集/转换,Terra 做常规实现,Sol 做复杂分析,Astra 做最困难的综合工作;使用最低足够且受支持的 effort。缺数据、权限或工具故障先诊断。
|
||||
- 默认视为共享工作区,文件写集与语义写集都须明确。工作角色和只读要求不是独立沙盒的证明。
|
||||
- 复用适合的已有子任务;停止失效或越界任务并查询真实状态。完成、验收、关闭、释放槽位分别判断;没有 close 工具时不伪造关闭。
|
||||
- 财务、安全与生产只读证据可在授权范围内分配;权限判断、生产写入、PR/merge/deploy 和最终验收由主线程串行完成。
|
||||
|
||||
## 6. 构建调用模型的应用
|
||||
|
||||
- 实现前核对当前官方文档,保留用户指定模型。API 可用性、Codex 主模型与子模型支持分别验证,宿主专有参数不直接复制到 API。
|
||||
- 推理与多轮工具流优先 Responses API。保留严格结构化输出、拒绝处理、call ID、工具结果与必要历史;无状态调用携带完整相关状态。
|
||||
- 迁移先保留有效 effort;目标不支持时选择受支持的起点评测。输出预算按合同与任务设定,避免全局统一上限。
|
||||
- 模型默认值和可选能力按真实任务的质量、延迟、总成本、重试与整合成本评估。异步工具、缓存、配置更新、多 agent 等分别核对支持情况,不靠提示词假装开启。
|
||||
|
||||
## 7. 沟通与收口
|
||||
|
||||
- 工具任务开始时简述第一阶段;有关键发现或路径变化时更新,遵守宿主响应频率与等待上限。
|
||||
- 用用户使用的语言,先讲结果,再讲必要证据、限制和下一步。避免重复任务图、完整日志、证据包和泛泛保证。
|
||||
- 完成前确认交付目标已满足、实际变更已整合、必要验证通过或缺口明确、外部门按授权处理、没有仍在修改交付文件的无主子任务。
|
||||
@@ -0,0 +1,125 @@
|
||||
# Codex Subagent Router
|
||||
|
||||
面向 **GPT-6 Astra / GPT-5.6 Sol / Terra / Luna** 的 Codex 路由技能与工程提示词。把可独立完成的工作交给合适的模型,并让触发、运行状态、证据与最终责任保持清楚。
|
||||
|
||||
这是供 Codex 读取的技能和可选提示词,附带证据校验、安装与测试工具。它不包含常驻调度器,也不会自行调用模型、修改模型配置或开启并行权限。
|
||||
|
||||
## 做什么
|
||||
|
||||
- 区分只读审查、已授权委派和普通本地工作,避免“选中了技能就自动启动 agent”。
|
||||
- 按任务形状分配模型与 effort,创建任务时显式选择,避免所有子任务继承昂贵的主模型。
|
||||
- 区分文件写集与语义写集、共享目录与真实隔离、合同约束与宿主强制能力。
|
||||
- 管理运行中补充、空闲复用、等待、中断与收口;没有 close 工具时不要求虚构关闭。
|
||||
- 分开记录请求的模型、宿主确认的运行身份、子任务完成和主线程验收;缺失身份使用 unknown。
|
||||
|
||||
## 目录
|
||||
|
||||
| 路径 | 用途 |
|
||||
| --- | --- |
|
||||
| [skills/codex-subagent-router/SKILL.md](skills/codex-subagent-router/SKILL.md) | 可安装技能入口 |
|
||||
| [skills/codex-subagent-router/references](skills/codex-subagent-router/references) | 路由边界、运行时、生命周期和证据合同 |
|
||||
| [prompts/AGENTS.md](prompts/AGENTS.md) | 可选全局工程提示词 |
|
||||
| [AGENTS.md](AGENTS.md) | 本仓库开发约定,不是安装模板 |
|
||||
| [scripts/install.py](scripts/install.py) | 默认预览、按文件备份的安装工具 |
|
||||
| [examples](examples) | 明确标记的合成证据样例 |
|
||||
| [tests](tests) | 校验器与安装安全行为测试 |
|
||||
| [docs/behavior-scenarios.md](docs/behavior-scenarios.md) | 触发与管理的行为验收场景 |
|
||||
| [docs/development.md](docs/development.md) | 维护、兼容与后续验证约定 |
|
||||
|
||||
## 安装
|
||||
|
||||
需要 Python 3.9+。仓库工具只使用标准库。
|
||||
|
||||
```sh
|
||||
git clone https://git.wun.im/kakan/codex-subagent-router.git
|
||||
cd codex-subagent-router
|
||||
python3 scripts/install.py
|
||||
```
|
||||
|
||||
最后一条命令仅预览差异。确认目标后安装技能:
|
||||
|
||||
```sh
|
||||
python3 scripts/install.py --apply
|
||||
```
|
||||
|
||||
默认使用环境变量 `CODEX_HOME`,未设置时使用 `~/.codex`;也可以通过 `--codex-home /path/to/codex-home` 指定。默认只同步技能,不改全局提示词。
|
||||
|
||||
若要同时使用本仓库的全局工程提示词,先审阅模板和本机已有规则,再运行:
|
||||
|
||||
```sh
|
||||
python3 scripts/install.py --include-prompt
|
||||
python3 scripts/install.py --include-prompt --apply
|
||||
```
|
||||
|
||||
`--include-prompt --apply` 会替换目标 `AGENTS.md`,不是智能合并;原文件先备份。安装不会修改 `config.toml`、其他技能或本地额外文件,不清理未收录文件,遇到目标路径中的符号链接会拒绝安装。
|
||||
|
||||
安装后按所用客户端的规则重新加载或开始新会话,并检查技能是否可见。写入文件不证明已启动会话采纳了新规则。
|
||||
|
||||
## 使用与触发
|
||||
|
||||
技能可以自动被选中帮助做路由判断,但自动发现不授予委派权限。明确要求执行委派的示例:
|
||||
|
||||
```text
|
||||
使用 $codex-subagent-router,把这个任务中可独立完成的部分交给子智能体;
|
||||
主线程继续关键路径,显式选择合适模型,完成后检查证据与残留任务。
|
||||
```
|
||||
|
||||
只读审查示例:
|
||||
|
||||
```text
|
||||
使用 $codex-subagent-router 检查当前模型继承、触发规则和生命周期工具,
|
||||
仅报告发现,不创建子任务、不修改设置。
|
||||
```
|
||||
|
||||
获得明确授权后,同一范围内无需逐次询问是否创建子任务;有独立产物和并行收益才分工。更高优先级宿主限制始终有效。修改这份技能、任务复杂、模型可用或选择 Ultra 本身不能覆盖这些限制。
|
||||
|
||||
## 模型策略
|
||||
|
||||
| 工作 | 起点 |
|
||||
| --- | --- |
|
||||
| 固定来源收集、枚举和检查 | Luna low/medium |
|
||||
| 有界转换、分类、已定方案补丁 | Luna high |
|
||||
| 常规实现、调试和跨文件调查 | Terra medium,必要时 high |
|
||||
| 复杂分析、设计和关键语义证据 | Sol medium,必要时 high |
|
||||
| 最困难的独立综合任务 | Astra medium,必要时 high |
|
||||
| 共享决定、最终整合与验收 | 当前主线程 |
|
||||
|
||||
对于以复杂工程和长任务为主的工作,可从 Astra / medium 主线程开始;这是待实际工作校准的建议,不会改变用户已选模型,也不是效果或价格保证。主线程自己还要承担判断与整合,不能仅因可以委派就视为消息转发器。
|
||||
|
||||
按实时工具 schema 检查模型、effort、fork、并发槽位和工作区。API 模型目录不证明当前子智能体支持。可选配置示例见 [运行时参考](skills/codex-subagent-router/references/platforms.md),安装工具不会应用它。
|
||||
|
||||
## 验证
|
||||
|
||||
```sh
|
||||
python3 -m unittest discover -s tests -v
|
||||
python3 scripts/check_repository.py
|
||||
python3 skills/codex-subagent-router/scripts/validate_evidence_packet.py --self-test
|
||||
python3 skills/codex-subagent-router/scripts/validate_evidence_packet.py examples/completed-unknown.json
|
||||
python3 scripts/install.py --include-prompt --check
|
||||
```
|
||||
|
||||
最后一条命令只检查本机安装差异:一致返回 0,有差异返回 1;不会写入。单元测试和 examples 使用合成数据,不能证明任何模型已运行。
|
||||
|
||||
新证据使用 v2;未带版本的旧格式按 v1 读取并提示迁移。校验成功只证明结构可接受,不能证明 child ID、身份或业务结论真实。真实环境的前向测试应在获授权的有界任务中进行,并保留实际工具回执。
|
||||
|
||||
若本机有 skill-creator,还应运行其 `scripts/quick_validate.py` 检查技能;它依赖 PyYAML。缺少该工具或依赖时如实报告,不把仓库检查冒充为上游校验。
|
||||
|
||||
## 更新与回退
|
||||
|
||||
正常开发使用功能分支与 PR,测试通过后由仓库所有者按流程合并。更新本机前先确认工作树并拉取已审阅版本,再运行安装预览。
|
||||
|
||||
每次实际变更会在 `CODEX_HOME/backups/codex-subagent-router/<时间-随机标识>/` 保留原文件和 `manifest.json`。清单记录路径、原先是否存在及前后 SHA-256。
|
||||
|
||||
回退时先比对当前文件与清单中的安装后哈希,保留安装后的用户修改;再把备份中的指定文件恢复到同一相对路径。原先不存在的文件不会有备份,只有确认它仍是本次新建版本时才按清单逐个移除。不要递归清空整个技能目录。安装按文件原子替换,不承诺整个目录的事务性;中断后依据清单检查实际状态。
|
||||
|
||||
## 与其他技能协作
|
||||
|
||||
本仓库不会自动修改 StockAgent 或其他编排技能。旧技能中的 Sol 主线程要求、严格身份探针、独立工作区假设或 close 要求可能仍影响任务。需要时把这些规则纳入明确的联合迁移范围;不能由本技能假装旧门槛已经满足。
|
||||
|
||||
## 官方参考
|
||||
|
||||
- [OpenAI 模型目录](https://developers.openai.com/api/docs/models)
|
||||
- [GPT-6 Astra 使用与提示指南](https://developers.openai.com/api/docs/guides/latest-model)
|
||||
- [Codex 子智能体](https://learn.chatgpt.com/zh-Hans/docs/agent-configuration/subagents)
|
||||
|
||||
本轮依据核对日期:2026-09-14。模型分工属于工程策略;宿主真实能力与实际任务证据优先。
|
||||
@@ -0,0 +1,46 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Check packaged resource integrity; not a semantic or live-host test."""
|
||||
import json
|
||||
from pathlib import Path
|
||||
import re
|
||||
import sys
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[1]
|
||||
|
||||
|
||||
def check() -> list:
|
||||
errors = []
|
||||
skill = ROOT / "skills/codex-subagent-router"
|
||||
for name in ("SKILL.md", "agents/openai.yaml", "references/lifecycle.md",
|
||||
"references/platforms.md", "references/evidence-packet.md",
|
||||
"references/routing-matrix.md", "scripts/validate_evidence_packet.py"):
|
||||
if not (skill / name).is_file():
|
||||
errors.append("missing skill resource: " + name)
|
||||
entry = (skill / "SKILL.md").read_text(encoding="utf-8")
|
||||
if not re.match(r"\A---\nname: codex-subagent-router\ndescription: .+\n---\n", entry):
|
||||
errors.append("invalid expected skill frontmatter")
|
||||
metadata = (skill / "agents/openai.yaml").read_text(encoding="utf-8")
|
||||
if "$codex-subagent-router" not in metadata:
|
||||
errors.append("UI default prompt must name the skill")
|
||||
for file in ROOT.rglob("*.md"):
|
||||
if any(part in {".git", "work", "outputs", ".venv"} for part in file.relative_to(ROOT).parts):
|
||||
continue
|
||||
content = file.read_text(encoding="utf-8")
|
||||
for target in re.findall(r"\[[^\]]*\]\(([^)]+)\)", content):
|
||||
if "://" in target or target.startswith("#"):
|
||||
continue
|
||||
path = target.split("#", 1)[0]
|
||||
if not (file.parent / path).exists():
|
||||
errors.append(str(file.relative_to(ROOT)) + ": broken link " + target)
|
||||
for fixture in (ROOT / "examples").glob("*.json"):
|
||||
try:
|
||||
json.loads(fixture.read_text(encoding="utf-8"))
|
||||
except ValueError:
|
||||
errors.append("invalid JSON: " + str(fixture.relative_to(ROOT)))
|
||||
return errors
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
errors = check()
|
||||
print("\n".join(errors) if errors else "Repository resource checks passed")
|
||||
sys.exit(bool(errors))
|
||||
@@ -0,0 +1,115 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Preview or install maintained files; preserve unrelated files and back up edits."""
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
from datetime import datetime, timezone
|
||||
import hashlib
|
||||
import json
|
||||
import os
|
||||
from pathlib import Path
|
||||
import shutil
|
||||
import tempfile
|
||||
import uuid
|
||||
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[1]
|
||||
|
||||
|
||||
def digest(data: bytes) -> str:
|
||||
return hashlib.sha256(data).hexdigest()
|
||||
|
||||
|
||||
def plan(repo: Path, home: Path, include_prompt: bool = False) -> list:
|
||||
source = repo / "skills/codex-subagent-router"
|
||||
files = [(p, Path("skills/codex-subagent-router") / p.relative_to(source))
|
||||
for p in sorted(source.rglob("*")) if p.is_file()
|
||||
and "__pycache__" not in p.parts and p.suffix != ".pyc"]
|
||||
if not (source / "SKILL.md").is_file():
|
||||
raise ValueError("source SKILL.md missing")
|
||||
if include_prompt:
|
||||
files.append((repo / "prompts/AGENTS.md", Path("AGENTS.md")))
|
||||
changes = []
|
||||
for src, relative in files:
|
||||
if src.is_symlink():
|
||||
raise ValueError("refusing symlink source: " + str(relative))
|
||||
dest = home / relative
|
||||
for part in [dest, *dest.parents]:
|
||||
if part == home:
|
||||
break
|
||||
if part.is_symlink():
|
||||
raise ValueError("refusing symlink target: " + str(relative))
|
||||
if dest.exists() and not dest.is_file():
|
||||
raise ValueError("target is not a regular file: " + str(relative))
|
||||
before = dest.read_bytes() if dest.exists() else None
|
||||
after = src.read_bytes()
|
||||
if before != after:
|
||||
changes.append({"path": relative, "before": before, "after": after})
|
||||
return changes
|
||||
|
||||
|
||||
def install(repo: Path, home: Path, include_prompt: bool = False, apply: bool = False) -> dict:
|
||||
home = home.expanduser().resolve()
|
||||
changes = plan(repo, home, include_prompt)
|
||||
result = {"changed_paths": [str(x["path"]) for x in changes], "applied": False, "backup": None}
|
||||
if not apply or not changes:
|
||||
return result
|
||||
# Check all baselines before any managed file is changed.
|
||||
for change in changes:
|
||||
dest = home / change["path"]
|
||||
current = dest.read_bytes() if dest.exists() else None
|
||||
if current != change["before"]:
|
||||
raise ValueError("target changed during planning: " + str(change["path"]))
|
||||
stamp = datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%SZ") + "-" + uuid.uuid4().hex[:8]
|
||||
backup = home / "backups/codex-subagent-router" / stamp
|
||||
backup.mkdir(parents=True, exist_ok=False)
|
||||
manifest = []
|
||||
for change in changes:
|
||||
before = change["before"]
|
||||
if before is not None:
|
||||
saved = backup / change["path"]
|
||||
saved.parent.mkdir(parents=True, exist_ok=True)
|
||||
saved.write_bytes(before)
|
||||
manifest.append({"path": str(change["path"]), "existed": before is not None,
|
||||
"before_sha256": digest(before) if before is not None else None,
|
||||
"installed_sha256": digest(change["after"])})
|
||||
(backup / "manifest.json").write_text(json.dumps(manifest, indent=2) + "\n", encoding="utf-8")
|
||||
for change in changes:
|
||||
dest = home / change["path"]
|
||||
dest.parent.mkdir(parents=True, exist_ok=True)
|
||||
temporary = None
|
||||
try:
|
||||
with tempfile.NamedTemporaryFile(dir=dest.parent, delete=False) as stream:
|
||||
temporary = Path(stream.name)
|
||||
stream.write(change["after"])
|
||||
if dest.exists():
|
||||
shutil.copymode(dest, temporary)
|
||||
temporary.replace(dest)
|
||||
finally:
|
||||
if temporary is not None and temporary.exists():
|
||||
temporary.unlink()
|
||||
result.update(applied=True, backup=str(backup))
|
||||
return result
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("--codex-home", type=Path,
|
||||
default=Path(os.environ.get("CODEX_HOME") or "~/.codex"))
|
||||
parser.add_argument("--include-prompt", action="store_true", help="also replace the global AGENTS.md")
|
||||
parser.add_argument("--apply", action="store_true", help="write files; default is preview only")
|
||||
parser.add_argument("--check", action="store_true", help="exit 1 when maintained files differ")
|
||||
args = parser.parse_args()
|
||||
if args.apply and args.check:
|
||||
parser.error("--apply and --check are mutually exclusive")
|
||||
try:
|
||||
result = install(ROOT, args.codex_home, args.include_prompt, args.apply)
|
||||
print(json.dumps(result, ensure_ascii=False, indent=2))
|
||||
return 1 if args.check and result["changed_paths"] else 0
|
||||
except (OSError, ValueError) as error:
|
||||
print("Install failed: " + str(error))
|
||||
return 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,69 @@
|
||||
---
|
||||
name: codex-subagent-router
|
||||
description: Decide when Codex subagents help, route bounded work across Astra, Sol, Terra and Luna, and manage child evidence and lifecycle. Use for delegation requests, routing decisions, or subagent audits; inspecting this skill does not itself authorize spawning.
|
||||
---
|
||||
|
||||
# Codex Subagent Router
|
||||
|
||||
Make delegation useful, observable and bounded. Preserve the selected parent model. The parent owns the critical path, shared decisions, integration and final acceptance. This skill neither changes configuration nor grants external permissions.
|
||||
|
||||
## Trigger before routing
|
||||
|
||||
Classify the request before using child tools:
|
||||
|
||||
- **Audit or advice:** inspect rules, configuration and available tools; report findings without spawning or changing settings.
|
||||
- **Authorized execution:** the user requested delegation, or an applicable instruction explicitly authorizes it. Within that scope, actively delegate independent useful slices while the parent advances other work.
|
||||
- **No delegation authorization:** work locally. Automatic skill discovery, model availability, task complexity and a request to edit this skill do not grant permission to spawn.
|
||||
|
||||
Honor higher-priority host restrictions even when a lower-priority rule permits delegation. Do not request authorization repeatedly after it has been granted. Never turn a routing recommendation into a new user-owned task.
|
||||
|
||||
Before dispatch classify each candidate:
|
||||
|
||||
- **P:** independent evidence or disjoint file AND semantic writes with its own acceptance.
|
||||
- **C:** a bounded investigation/draft that needs latest-baseline parent integration.
|
||||
- **S:** an immediate dependency, overlapping contract/state, permissions, production mutation, PR/merge/deploy or final acceptance; keep it in the parent.
|
||||
|
||||
Spawn only P/C work with a concrete output and useful parent work available. Do not spawn for one command, ceremonial probes, duplicate reviews or a model quota. Batch homogeneous small work.
|
||||
|
||||
## Select a supported route explicitly
|
||||
|
||||
| Work shape | Initial route |
|
||||
| --- | --- |
|
||||
| Known-source collection, fixed checks, logs | Luna low/medium |
|
||||
| Bounded classification, conversion, settled patch with fixed checks | Luna high |
|
||||
| Everyday implementation, debugging, locating/correlating artifacts | Terra medium; high when needed |
|
||||
| Bounded complex analysis, design or financial/security evidence | Sol medium; high when needed |
|
||||
| Hardest independent synthesis across code, tools and research | Astra medium; high when needed |
|
||||
| Shared decisions, integration and external/final gates | Current parent, serial |
|
||||
|
||||
These are starting heuristics, not measured cost rankings. Pick sufficient capability directly; do not escalate through every model. Missing access/data and tool failures need diagnosis, not a stronger model. After two same-class failures, pause that slice and return the evidence to the parent.
|
||||
|
||||
Use the lowest adequate supported effort. Use xhigh/max for a concrete depth need or an explicit compatible role requirement; ultra only when exposed and justified by the workload. Model choice and working role are separate. A role named "explorer" is not proof of sandbox isolation.
|
||||
|
||||
At dispatch, specify model AND effort when the host permits selection. Otherwise omitted settings may inherit the parent or configured defaults. Prefer a self-contained contract and no history fork; use bounded history only when necessary. Respect full-history/override incompatibilities. Never claim prompt text changed a runtime parameter.
|
||||
|
||||
## Check capabilities once per unchanged host
|
||||
|
||||
Inspect the actual child-tool schema: models, per-model efforts, role bindings, defaults, history rules, slot counting, workspace isolation and controls. API catalogs, config files and task-creation tools cannot prove child availability. Recheck only after relevant changes.
|
||||
|
||||
Use the first authorized, low-risk real child task as the capability observation; for unfamiliar Luna multi-step work, start with a useful bounded read-only slice. A returned model name or marker is not identity proof. Record requested settings separately from host-observed identity; missing fields are unknown. Output can pass while identity stays unknown, unless the task explicitly requires verified identity.
|
||||
|
||||
Treat workspaces as shared unless isolation is confirmed. Tool schemas may lack per-child sandbox, timeout, role or close controls. Contract limits remain instructions, not enforced capabilities. No adequate child route: keep feasible work local and disclose the gap.
|
||||
|
||||
Read [runtime and configuration](references/platforms.md) only for configuration/CLI diagnosis. Read [routing boundaries](references/routing-matrix.md) for ambiguous choices or live artifact collection.
|
||||
|
||||
## Dispatch, observe, accept
|
||||
|
||||
Give a compact contract: goal/symptom; baseline and non-goals; read/write/forbidden paths and semantic owner; requested model/effort/working role and real isolation; permitted actions; acceptance; return evidence; escalation. Include a time/checkpoint budget when useful, but do not call it an enforced timeout unless the host supplies one.
|
||||
|
||||
Read [lifecycle](references/lifecycle.md) before a multi-round or cancellation-sensitive run. Keep child IDs, ownership, state and acceptance in a small ledger. Send running agents only incremental context; reuse idle agents for follow-up; interrupt obsolete work and verify its state. Never infer completion from a wait timeout or slot release from interruption. Only call close if that tool exists.
|
||||
|
||||
Ask for bounded facts, paths, changes, real check outcomes, gaps and escalation. Review artifacts/diffs and provenance; resolve conflicts from original evidence, not votes. Reuse checks only for the same relevant baseline, artifact and environment. Before delivery verify no unneeded child remains running or writing.
|
||||
|
||||
Use [evidence packets](references/evidence-packet.md) when structured evidence is required, a run is material, or identity is disputed. Ordinary bounded work can return concise prose with the same relevant facts; do not generate ceremonial JSON for every local decision. The validator checks structure, not truth or semantic correctness.
|
||||
|
||||
## Context and ownership discipline
|
||||
|
||||
Load only needed references; keep long logs on disk. Avoid whole-history forks and repeating contracts, packets or prior findings. Account for parent context/reasoning, child usage, waiting, retries and integration when evaluating cost.
|
||||
|
||||
Production read-only collection is routable by work shape; mutation and final semantic acceptance stay parent-owned. Do not silently relax a project's literal model-owner requirement. Identify the actual rule and resolve it only through authorized changes.
|
||||
@@ -0,0 +1,4 @@
|
||||
interface:
|
||||
display_name: "Codex Subagent Router"
|
||||
short_description: "Route models and manage bounded Codex subagents"
|
||||
default_prompt: "Use $codex-subagent-router to delegate useful independent parts of this task across supported models, manage their lifecycle, and verify the results."
|
||||
@@ -0,0 +1,47 @@
|
||||
# Evidence Packet v2
|
||||
|
||||
Use structured packets for material runs, identity disputes or an explicit consumer requirement. Concise prose is sufficient for ordinary work when it preserves the relevant evidence.
|
||||
|
||||
## Contract
|
||||
|
||||
Required keys: schema_version (2), execution (not_run/local/child), goal, scope, allowed_paths (read/write/forbidden arrays), tool_receipts, capability_source, fallback_reason, evidence, changes, verification, risks, escalate.
|
||||
|
||||
- capability_source describes where capability evidence came from: host-tool-schema, child-rollout, profile-probe, fallback or unknown. It does not establish identity or acceptance.
|
||||
- fallback requires a nonempty fallback_reason; otherwise use null unless a nonpreferred route needs explaining.
|
||||
- verification contains commands and a boolean deterministic. Each command has command, exit_code (integer or null), and status (passed/failed/not_run/unknown). A passed command requires exit code 0.
|
||||
- escalate contains required; when true, target and reason must be nonempty.
|
||||
- execution=not_run permits only schema receipts, no claimed changes or executed commands. A schema inspection can precede dispatch without inventing a child.
|
||||
- A child receipt requires kind=child, a real tool name and child_id, observed state, requested and observed identity objects, and identity_source. Both identity objects contain model, role and effort; use unknown for unavailable fields.
|
||||
- Recognized child states: created, running, idle, completed, interrupted, failed, output-returned. Map only when supported by actual host evidence; preserve the original state in an extra field when necessary.
|
||||
- Non-child receipts have kind=schema or local and cannot contain child_id.
|
||||
- identity_source is host-tool-result or unknown. Any known observed identity requires identity_evidence pointing to the actual host result/event. Requested parameters and child self-description cannot serve as that evidence.
|
||||
- Runtime identity can remain unknown even if the child completed and output passed. Do not relabel a real rollout as merely a schema observation because identity fields are absent.
|
||||
- Parent acceptance remains separate; an optional acceptance object may record the parent's checks, status and unresolved requirements.
|
||||
|
||||
See the repository's examples/ directory for synthetic v2 and legacy v1 fixtures. Fixture IDs are examples only, never actual receipts.
|
||||
|
||||
## Unknown identity example
|
||||
|
||||
This fragment illustrates shape; it is not a real tool receipt:
|
||||
|
||||
```json
|
||||
{
|
||||
"kind": "child",
|
||||
"tool": "collaboration.spawn_agent",
|
||||
"child_id": "synthetic-child",
|
||||
"status": "completed",
|
||||
"requested": {"model": "gpt-5.6-terra", "role": "explorer", "effort": "medium"},
|
||||
"observed": {"model": "unknown", "role": "unknown", "effort": "unknown"},
|
||||
"identity_source": "unknown"
|
||||
}
|
||||
```
|
||||
|
||||
The working role requested in a contract does not establish a runtime role. Verify output independently; an explicit requirement to prove the model remains unmet until trustworthy identity evidence exists.
|
||||
|
||||
## Migration
|
||||
|
||||
The validator still accepts unversioned/schema_version=1 packets and emits a legacy notice. V1 used flat identity fields on completed child receipts; preserve historical records rather than inventing v2 observations from them.
|
||||
|
||||
For new runs use v2. Reconstruct observed identity only from retained host evidence; otherwise use unknown. The stricter v2 command/receipt checks are intentionally not backfilled into old event histories.
|
||||
|
||||
Run scripts/validate_evidence_packet.py PACKET.json (or - for stdin). Exit 0 means structural acceptance only. The --self-test option is a small parser smoke check; repository unit tests provide broader coverage. No validator can authenticate tool receipts or replace parent correctness judgment.
|
||||
@@ -0,0 +1,33 @@
|
||||
# Child Lifecycle
|
||||
|
||||
Read before multi-round coordination, cancellation, or capacity diagnosis. Names below describe control intent; use only the tools and states the host actually exposes.
|
||||
|
||||
| Situation | Action | Evidence needed |
|
||||
| --- | --- | --- |
|
||||
| Independent authorized slice ready | Spawn with explicit ownership and supported route | Real child ID; requested settings |
|
||||
| Child running, facts changed | Send a small incremental message | Tool acknowledgement; later child result |
|
||||
| Child idle, more scoped work needed | Follow up with the same child | Reactivation/result; keep prior evidence |
|
||||
| Parent has independent work | Continue it | No unnecessary polling |
|
||||
| Parent depends on child | Bounded wait | New output/state; timeout means only no update |
|
||||
| Cancelled, superseded, overlapping writes | Interrupt affected children, including descendants if needed | Query resulting states; inspect partial writes |
|
||||
| Child returned output | Review evidence, diff and relevant checks | Acceptance separate from runtime completion |
|
||||
| Capacity exhausted | Inspect actual active/open states and reuse eligible children | Host-specific capacity semantics |
|
||||
| Final delivery | Audit tree and assigned files | No unneeded running writer; accepted or explicit gaps |
|
||||
|
||||
A minimal ledger records child ID, P/C classification, file/semantic ownership, dependency, requested route, last observed state and acceptance. Do not invent timestamps, identities or close events.
|
||||
|
||||
## Cancellation and handback
|
||||
|
||||
Interrupting does not undo files or guarantee descendant cancellation. First preserve useful returned evidence, stop affected writers, inspect partial changes, then assign the remaining work to one owner. Never reset shared files wholesale.
|
||||
|
||||
A budget/deadline in a contract is not a runtime timeout. Use host-supported nonblocking waits and control calls; check deadlines at checkpoints. A wait returning no update does not prove failure, completion or a stuck process.
|
||||
|
||||
For missing data, permissions or same-class failures twice, stop the affected slice and report the smallest gap. Resume with a delta when resolved rather than spawning a duplicate investigation.
|
||||
|
||||
## Completion versus closure
|
||||
|
||||
Output returned, runtime completed, parent accepted, thread closed and slot released are separate facts. The host may combine some of them, but never assume that it does.
|
||||
|
||||
If close exists, use it according to host semantics when no longer needed. If it does not, completed/idle is a valid terminal handback; do not require a fictional close to finish. If capacity remains unavailable, reuse a supported idle child or proceed locally.
|
||||
|
||||
Final review must also check descendants and pending writes. If any ongoing work is intentionally retained, report its owner and reason.
|
||||
@@ -0,0 +1,49 @@
|
||||
# Runtime and Configuration
|
||||
|
||||
Use this reference for actual host/configuration diagnosis. This is not a prerequisite checklist for every task.
|
||||
|
||||
## Separate four kinds of evidence
|
||||
|
||||
1. Official model documentation: product and API capabilities.
|
||||
2. Local configuration: requested defaults, providers and role bindings.
|
||||
3. Current child-tool schema: what this session accepts.
|
||||
4. Tool results: what was created, observed, completed or interrupted.
|
||||
|
||||
Do not substitute one for another. A model available in an API or the task picker may be absent from child tools. A CLI success does not verify a Desktop child path.
|
||||
|
||||
Inspect only relevant configuration fields; never dump environment blocks, credential files or HTTP authorization headers. Resolve CODEX_HOME from the environment, falling back to the user's .codex directory; do not overwrite that environment variable.
|
||||
|
||||
## Optional Codex defaults
|
||||
|
||||
For clients that support the documented keys, an example is:
|
||||
|
||||
```toml
|
||||
# Example only; merge intentionally into the existing configuration.
|
||||
model = "gpt-6-astra"
|
||||
model_reasoning_effort = "medium"
|
||||
|
||||
[agents]
|
||||
default_subagent_model = "gpt-5.6-terra"
|
||||
default_subagent_reasoning_effort = "medium"
|
||||
max_concurrent_threads_per_session = 3
|
||||
```
|
||||
|
||||
This is not an automatic router, permission grant or universal host configuration. Never overwrite an existing config to install this example. Preserve user model/provider choices and verify the actual client accepts the keys. Per-agent configuration and host overrides may change resolution.
|
||||
|
||||
The official Codex configuration describes the child-thread limit as excluding the parent. A collaboration tool may expose a different active-slot count including the parent. Use that live tool's counting and release semantics; do not copy either number blindly.
|
||||
|
||||
When selecting a route, explicitly request model and effort with a compatible fork mode if supported. Check omitted-setting inheritance, custom-agent bindings and override restrictions. A prompt cannot implement an unsupported model switch.
|
||||
|
||||
## Diagnosis and safe fallbacks
|
||||
|
||||
Resolve the executable actually used by the relevant client, inspect its help/schema and record its version when needed. Do not prescribe speculative CLI flags or launch model calls merely to print a marker. Use an authorized real task to verify child execution.
|
||||
|
||||
Missing role/sandbox/timeout fields mean working role, read-only scope and deadlines are contract limits only. Keep requested and observed identity distinct. A completed task with unknown model identity can be accepted for its output, but cannot satisfy an explicit identity-verification gate.
|
||||
|
||||
No close tool: retain the truthful idle/completed/interrupted state and follow the host's actual capacity behavior. No child tools: continue locally; do not launch another user-owned task or nested CLI as an unapproved substitute.
|
||||
|
||||
Configuration changes may require a fresh session or explicit reload. Writing a file is not evidence that the current session adopted it.
|
||||
|
||||
## Sources
|
||||
|
||||
Checked 2026-09-14: [Codex subagents](https://learn.chatgpt.com/zh-Hans/docs/agent-configuration/subagents), [OpenAI model catalog](https://developers.openai.com/api/docs/models). Recheck changed product/API facts before altering real configuration.
|
||||
@@ -0,0 +1,33 @@
|
||||
# Routing Boundaries
|
||||
|
||||
Read only when the entrypoint leaves a routing choice unresolved.
|
||||
|
||||
| Boundary | Decision |
|
||||
| --- | --- |
|
||||
| Many easy tasks | Batch bounded Luna/Terra slices; volume alone does not justify Astra. |
|
||||
| Ordinary cross-file bug | Terra; file count alone does not justify escalation. |
|
||||
| Difficult accounting discrepancy | Sol reasoning; parent owns the semantic decision. |
|
||||
| Coupled architecture across runtimes and tools | Consider Astra directly; keep shared decisions serial. |
|
||||
| Missing credential or inaccessible source | Resolve the evidence/environment gap; model escalation cannot supply access. |
|
||||
| Astra advertised only for task creation | Child support remains unproven; do not create another user task as a workaround. |
|
||||
| Named role fixes model/effort | Accept its supported binding or choose an overridable role; never attach prohibited fork overrides. |
|
||||
| Repository requires Sol parent | A Sol child cannot replace that owner; preserve the gate until an authorized supported change. |
|
||||
| Accepted output, unknown child model | Accept only after normal parent checks; do not attribute the result to the requested model. |
|
||||
|
||||
## Production Artifact Collection
|
||||
|
||||
- Luna: use known hosts/endpoints/paths and fixed selectors to read or download actual reports, JSON, logs or job outputs; verify size/hash, parseability, required fields and specified counts. Terra: locate the current output through bounded read-only investigation, diagnose retrieval failures, or correlate logs and artifacts by run/version. Do not launch both by default; use Terra only when discovery or diagnosis is needed.
|
||||
- Contract: name the production source, expected run/date/version, allowed read operations, bounded time/range/volume, local destination and acceptance checks. Reuse approved access and applicable server tools. Keep secrets out of prompts/receipts. Remote access stays read-only; explicitly allow local artifact writes even for a collector role.
|
||||
- Evidence: preserve the actual downloaded artifact and record source, retrieval time/timezone, producer run/version and generation time when exposed, local path, bytes/hash and check results. Unknown provenance stays unknown. Verify the expected run/freshness; mtime or successful download alone does not prove current output. Use immutable run identifiers where possible; detect/retry boundedly if files change during collection.
|
||||
- No fixture, cache or old snapshot may stand in for requested live output. If the current artifact is missing, stale, inconsistent or inaccessible, report that gap. Do not trigger jobs, regenerate artifacts, restart services, change permissions/configuration or deploy as part of collection; return such remediation to the parent under existing authorization.
|
||||
- P only for independent reads and disjoint local outputs; shared mutable SSH sessions or a changing multi-file output set require coordination/serialization. Parent reviews provenance, relevant artifact content and check results before semantic/final acceptance. Fetching is evidence collection, not production acceptance.
|
||||
|
||||
## Cost Calibration
|
||||
|
||||
Minimize total task usage: parent input/reasoning/output + all child input/reasoning/output + retries and reintegration. A smaller model can reduce price without reducing token count; additional agents can increase both. Do not assume pricing ratios or invent token measurements.
|
||||
|
||||
Use actual usage counters when exposed. Otherwise report observable proxies (characters loaded, number of children, full-history forks, repeated reads/checks) with their limitations. Record only available data needed for a routing decision; do not add telemetry work to every small task. Compare similar accepted tasks before changing defaults.
|
||||
|
||||
If context dominates, narrow inputs and avoid full-history forks. If output dominates, bound logs and returns. If reasoning/retries dominate, clarify acceptance and choose adequate capability/effort directly. Preserve required tests and material findings.
|
||||
|
||||
Official model positioning checked 2026-09-14: [Models](https://developers.openai.com/api/docs/models). The route table is an engineering starting strategy, not a measured cost ranking. Runtime tool schemas remain the source for actual child availability.
|
||||
@@ -0,0 +1,191 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Validate evidence structure, never runtime identity or semantic correctness."""
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import sys
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
REQUIRED = {
|
||||
"goal", "scope", "allowed_paths", "tool_receipts", "capability_source",
|
||||
"fallback_reason", "evidence", "changes", "verification", "risks", "escalate",
|
||||
}
|
||||
SOURCES = {"host-tool-schema", "child-rollout", "profile-probe", "fallback", "unknown"}
|
||||
STATES = {"created", "running", "idle", "completed", "interrupted", "failed", "output-returned"}
|
||||
|
||||
|
||||
def require(condition: bool, message: str) -> None:
|
||||
if not condition:
|
||||
raise ValueError(message)
|
||||
|
||||
|
||||
def nonempty(value: Any) -> bool:
|
||||
return isinstance(value, str) and bool(value.strip())
|
||||
|
||||
|
||||
def string_list(value: Any) -> bool:
|
||||
return isinstance(value, list) and all(nonempty(item) for item in value)
|
||||
|
||||
|
||||
def spawn_tool(value: str) -> bool:
|
||||
return value.split(".")[-1] == "spawn_agent"
|
||||
|
||||
|
||||
def validate(packet: Any) -> None:
|
||||
require(isinstance(packet, dict), "packet must be a JSON object")
|
||||
missing = REQUIRED - packet.keys()
|
||||
require(not missing, "missing fields: " + ", ".join(sorted(missing)))
|
||||
version = packet.get("schema_version", 1)
|
||||
require(type(version) is int and version in (1, 2), "schema_version must be 1 or 2")
|
||||
require(nonempty(packet["goal"]), "goal must be a non-empty string")
|
||||
require(isinstance(packet["scope"], dict), "scope must be an object")
|
||||
paths = packet["allowed_paths"]
|
||||
require(isinstance(paths, dict), "allowed_paths must be an object")
|
||||
for key in ("read", "write", "forbidden"):
|
||||
require(string_list(paths.get(key)), "allowed_paths." + key + " must be a string array")
|
||||
receipts = packet["tool_receipts"]
|
||||
require(isinstance(receipts, list), "tool_receipts must be an array")
|
||||
source = packet["capability_source"]
|
||||
require(isinstance(source, str) and source in SOURCES, "unsupported capability_source")
|
||||
fallback = packet["fallback_reason"]
|
||||
require(fallback is None or nonempty(fallback), "fallback_reason must be non-empty or null")
|
||||
require(source != "fallback" or nonempty(fallback), "fallback requires fallback_reason")
|
||||
require(isinstance(packet["evidence"], (dict, list)), "evidence must be an object or array")
|
||||
require(isinstance(packet["changes"], list), "changes must be an array")
|
||||
require(isinstance(packet["risks"], list), "risks must be an array")
|
||||
verification = packet["verification"]
|
||||
require(isinstance(verification, dict), "verification must be an object")
|
||||
require(isinstance(verification.get("commands"), list), "verification.commands must be an array")
|
||||
if version == 2:
|
||||
require(type(verification.get("deterministic")) is bool, "verification.deterministic must be boolean")
|
||||
for command in verification["commands"]:
|
||||
require(isinstance(command, dict), "each command must be an object")
|
||||
require(nonempty(command.get("command")), "command requires command text")
|
||||
require("exit_code" in command, "command requires exit_code (null if unknown)")
|
||||
code = command["exit_code"]
|
||||
require(code is None or type(code) is int, "exit_code must be integer or null")
|
||||
status = command.get("status")
|
||||
require(isinstance(status, str) and status in {"passed", "failed", "not_run", "unknown"},
|
||||
"command requires supported status")
|
||||
require(status != "passed" or code == 0, "passed command requires exit_code 0")
|
||||
require(status != "failed" or code is None or code != 0, "failed command contradicts exit_code 0")
|
||||
require(status != "not_run" or code is None, "not_run command cannot have exit_code")
|
||||
escalate = packet["escalate"]
|
||||
require(isinstance(escalate, dict) and type(escalate.get("required")) is bool,
|
||||
"escalate.required must be boolean")
|
||||
if escalate["required"]:
|
||||
require(nonempty(escalate.get("target")) and nonempty(escalate.get("reason")),
|
||||
"required escalation needs target and reason")
|
||||
for receipt in receipts:
|
||||
require(isinstance(receipt, dict), "each tool receipt must be an object")
|
||||
require(nonempty(receipt.get("tool")) and nonempty(receipt.get("status")),
|
||||
"tool receipt requires non-empty tool/status")
|
||||
if version == 1:
|
||||
_legacy(packet)
|
||||
else:
|
||||
_v2(packet)
|
||||
|
||||
|
||||
def _legacy(packet: dict) -> None:
|
||||
"""Retain unversioned v1 packets without interpreting them as v2 evidence."""
|
||||
receipts = packet["tool_receipts"]
|
||||
require(bool(receipts) or packet["capability_source"] == "fallback",
|
||||
"legacy empty tool_receipts require fallback")
|
||||
completed = []
|
||||
for receipt in receipts:
|
||||
if spawn_tool(receipt["tool"]) and receipt["status"] == "completed":
|
||||
for field in ("child_id", "model", "role", "effort"):
|
||||
require(nonempty(receipt.get(field)), "legacy completed receipt requires " + field)
|
||||
completed.append(receipt)
|
||||
require(packet["capability_source"] != "child-rollout" or bool(completed),
|
||||
"legacy child-rollout requires a completed child receipt")
|
||||
|
||||
|
||||
def _v2(packet: dict) -> None:
|
||||
execution = packet.get("execution")
|
||||
require(isinstance(execution, str) and execution in {"not_run", "local", "child"},
|
||||
"execution must be not_run, local or child")
|
||||
children = []
|
||||
for receipt in packet["tool_receipts"]:
|
||||
kind = receipt.get("kind")
|
||||
require(isinstance(kind, str) and kind in {"schema", "local", "child"}, "receipt requires kind")
|
||||
if kind != "child":
|
||||
require("child_id" not in receipt, "non-child receipt cannot claim child_id")
|
||||
require(not spawn_tool(receipt["tool"]) or kind == "schema",
|
||||
"spawn receipt must be child or schema inspection")
|
||||
continue
|
||||
children.append(receipt)
|
||||
require(nonempty(receipt.get("child_id")), "child receipt requires actual child_id")
|
||||
require(receipt["status"] in STATES, "unsupported child state")
|
||||
for section in ("requested", "observed"):
|
||||
identity = receipt.get(section)
|
||||
require(isinstance(identity, dict), "child receipt requires " + section)
|
||||
for key in ("model", "role", "effort"):
|
||||
require(nonempty(identity.get(key)), section + "." + key + " required; use unknown")
|
||||
origin = receipt.get("identity_source")
|
||||
require(isinstance(origin, str) and origin in {"host-tool-result", "unknown"},
|
||||
"identity_source must be host-tool-result or unknown")
|
||||
known = any(value != "unknown" for key, value in receipt["observed"].items()
|
||||
if key in {"model", "role", "effort"})
|
||||
if known:
|
||||
require(origin == "host-tool-result" and nonempty(receipt.get("identity_evidence")),
|
||||
"known observed identity requires host evidence")
|
||||
elif origin == "unknown":
|
||||
require(not receipt.get("identity_evidence"), "unknown identity cannot claim identity evidence")
|
||||
require((execution == "child") == bool(children), "execution must agree with child receipts")
|
||||
require(packet["capability_source"] != "child-rollout" or bool(children),
|
||||
"child-rollout requires a real child receipt")
|
||||
if execution == "not_run":
|
||||
require(all(r["kind"] == "schema" for r in packet["tool_receipts"]),
|
||||
"not_run may contain schema inspection only")
|
||||
require(not packet["changes"], "not_run cannot claim changes")
|
||||
require(all(c["status"] == "not_run" for c in packet["verification"]["commands"]),
|
||||
"not_run cannot claim executed commands")
|
||||
|
||||
|
||||
def self_test() -> None:
|
||||
packet = {
|
||||
"schema_version": 2, "execution": "not_run", "goal": "schema smoke check",
|
||||
"scope": {}, "allowed_paths": {"read": [], "write": [], "forbidden": []},
|
||||
"tool_receipts": [], "capability_source": "unknown", "fallback_reason": None,
|
||||
"evidence": {}, "changes": [], "verification": {"commands": [], "deterministic": True},
|
||||
"risks": [], "escalate": {"required": False},
|
||||
}
|
||||
validate(packet)
|
||||
packet["execution"] = "child"
|
||||
try:
|
||||
validate(packet)
|
||||
except ValueError:
|
||||
return
|
||||
raise AssertionError("child execution without receipts was accepted")
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("packet", nargs="?", help="JSON path, or - for stdin")
|
||||
parser.add_argument("--self-test", action="store_true")
|
||||
args = parser.parse_args()
|
||||
if not args.self_test and not args.packet:
|
||||
parser.error("provide a packet path or --self-test")
|
||||
try:
|
||||
if args.self_test:
|
||||
self_test()
|
||||
print("[evidence-packet] structural self-test ok")
|
||||
return 0
|
||||
raw = sys.stdin.read() if args.packet == "-" else Path(args.packet).read_text(encoding="utf-8")
|
||||
packet = json.loads(raw)
|
||||
validate(packet)
|
||||
version = packet.get("schema_version", 1)
|
||||
if version == 1:
|
||||
print("[evidence-packet] legacy v1; migrate new receipts to v2", file=sys.stderr)
|
||||
print("[evidence-packet] structure ok; identity and correctness not verified")
|
||||
return 0
|
||||
except (OSError, ValueError) as error:
|
||||
print("[evidence-packet] ERROR " + str(error), file=sys.stderr)
|
||||
return 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,138 @@
|
||||
import importlib.util
|
||||
import json
|
||||
from pathlib import Path
|
||||
import subprocess
|
||||
import sys
|
||||
import unittest
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[1]
|
||||
SCRIPT = ROOT / "skills/codex-subagent-router/scripts/validate_evidence_packet.py"
|
||||
spec = importlib.util.spec_from_file_location("evidence", SCRIPT)
|
||||
evidence = importlib.util.module_from_spec(spec)
|
||||
spec.loader.exec_module(evidence)
|
||||
|
||||
|
||||
def fixture(name="completed-unknown"):
|
||||
return json.loads((ROOT / "examples" / (name + ".json")).read_text())
|
||||
|
||||
|
||||
class EvidenceTests(unittest.TestCase):
|
||||
def test_all_packaged_examples(self):
|
||||
for path in (ROOT / "examples").glob("*.json"):
|
||||
with self.subTest(path=path.name):
|
||||
evidence.validate(json.loads(path.read_text()))
|
||||
|
||||
def test_completed_unknown_identity_is_valid(self):
|
||||
packet = fixture()
|
||||
evidence.validate(packet)
|
||||
self.assertEqual(packet["tool_receipts"][0]["observed"]["model"], "unknown")
|
||||
|
||||
def test_known_identity_requires_host_evidence(self):
|
||||
packet = fixture()
|
||||
child = packet["tool_receipts"][0]
|
||||
child["observed"]["model"] = "gpt-5.6-terra"
|
||||
with self.assertRaisesRegex(ValueError, "host evidence"):
|
||||
evidence.validate(packet)
|
||||
child["identity_source"] = "host-tool-result"
|
||||
child["identity_evidence"] = "synthetic host event for parser testing"
|
||||
evidence.validate(packet)
|
||||
|
||||
def test_child_self_description_is_not_identity_source(self):
|
||||
packet = fixture()
|
||||
packet["tool_receipts"][0]["identity_source"] = "child-self-description"
|
||||
with self.assertRaises(ValueError):
|
||||
evidence.validate(packet)
|
||||
|
||||
def test_child_execution_requires_actual_child_id(self):
|
||||
packet = fixture()
|
||||
del packet["tool_receipts"][0]["child_id"]
|
||||
with self.assertRaisesRegex(ValueError, "child_id"):
|
||||
evidence.validate(packet)
|
||||
|
||||
def test_child_execution_without_receipt_fails(self):
|
||||
packet = fixture()
|
||||
packet["tool_receipts"] = []
|
||||
with self.assertRaises(ValueError):
|
||||
evidence.validate(packet)
|
||||
|
||||
def test_not_run_cannot_claim_child_or_changes(self):
|
||||
packet = fixture()
|
||||
packet["execution"] = "not_run"
|
||||
with self.assertRaises(ValueError):
|
||||
evidence.validate(packet)
|
||||
packet = fixture("not-run")
|
||||
packet["changes"] = ["synthetic.py"]
|
||||
with self.assertRaisesRegex(ValueError, "changes"):
|
||||
evidence.validate(packet)
|
||||
|
||||
def test_running_receipt_does_not_need_completion(self):
|
||||
packet = fixture()
|
||||
packet["tool_receipts"][0]["status"] = "running"
|
||||
evidence.validate(packet)
|
||||
|
||||
def test_schema_cannot_masquerade_as_child(self):
|
||||
packet = fixture()
|
||||
packet["tool_receipts"][0]["kind"] = "schema"
|
||||
with self.assertRaisesRegex(ValueError, "child_id"):
|
||||
evidence.validate(packet)
|
||||
|
||||
def test_invalid_unhashable_sources_raise_validation_error(self):
|
||||
for field in ("capability_source", "execution"):
|
||||
packet = fixture()
|
||||
packet[field] = []
|
||||
with self.subTest(field=field), self.assertRaises(ValueError):
|
||||
evidence.validate(packet)
|
||||
|
||||
def test_malformed_receipts_and_identity_raise_validation_error(self):
|
||||
for value in (None, [], "bad", 3):
|
||||
packet = fixture()
|
||||
packet["tool_receipts"] = [value]
|
||||
with self.subTest(value=value), self.assertRaises(ValueError):
|
||||
evidence.validate(packet)
|
||||
packet = fixture()
|
||||
packet["tool_receipts"][0]["observed"] = []
|
||||
with self.assertRaises(ValueError):
|
||||
evidence.validate(packet)
|
||||
|
||||
def test_failed_command_cannot_claim_pass(self):
|
||||
packet = fixture()
|
||||
packet["verification"]["commands"] = [
|
||||
{"command": "synthetic-check", "exit_code": 1, "status": "passed"}]
|
||||
with self.assertRaisesRegex(ValueError, "exit_code 0"):
|
||||
evidence.validate(packet)
|
||||
|
||||
def test_boolean_is_not_exit_code_or_schema_version(self):
|
||||
packet = fixture()
|
||||
packet["verification"]["commands"] = [
|
||||
{"command": "synthetic-check", "exit_code": False, "status": "passed"}]
|
||||
with self.assertRaises(ValueError):
|
||||
evidence.validate(packet)
|
||||
packet = fixture()
|
||||
packet["schema_version"] = True
|
||||
with self.assertRaises(ValueError):
|
||||
evidence.validate(packet)
|
||||
|
||||
def test_escalation_needs_target_and_reason(self):
|
||||
packet = fixture()
|
||||
packet["escalate"] = {"required": True}
|
||||
with self.assertRaisesRegex(ValueError, "target and reason"):
|
||||
evidence.validate(packet)
|
||||
|
||||
def test_legacy_is_still_readable_with_notice(self):
|
||||
result = subprocess.run([sys.executable, str(SCRIPT), str(ROOT / "examples/legacy-v1.json")],
|
||||
capture_output=True, text=True)
|
||||
self.assertEqual(result.returncode, 0, result.stderr)
|
||||
self.assertIn("legacy v1", result.stderr)
|
||||
|
||||
def test_cli_stdin_and_bad_json(self):
|
||||
good = subprocess.run([sys.executable, str(SCRIPT), "-"], input=json.dumps(fixture()),
|
||||
capture_output=True, text=True)
|
||||
self.assertEqual(good.returncode, 0, good.stderr)
|
||||
bad = subprocess.run([sys.executable, str(SCRIPT), "-"], input="{",
|
||||
capture_output=True, text=True)
|
||||
self.assertEqual(bad.returncode, 1)
|
||||
self.assertNotIn("Traceback", bad.stderr)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -0,0 +1,65 @@
|
||||
import importlib.util
|
||||
from pathlib import Path
|
||||
import tempfile
|
||||
import unittest
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[1]
|
||||
spec = importlib.util.spec_from_file_location("installer", ROOT / "scripts/install.py")
|
||||
installer = importlib.util.module_from_spec(spec)
|
||||
spec.loader.exec_module(installer)
|
||||
|
||||
|
||||
class InstallerTests(unittest.TestCase):
|
||||
def setUp(self):
|
||||
self.temp = tempfile.TemporaryDirectory()
|
||||
self.addCleanup(self.temp.cleanup)
|
||||
self.home = Path(self.temp.name) / "codex-home"
|
||||
self.home.mkdir()
|
||||
|
||||
def test_preview_never_writes(self):
|
||||
result = installer.install(ROOT, self.home, include_prompt=True)
|
||||
self.assertTrue(result["changed_paths"])
|
||||
self.assertFalse(result["applied"])
|
||||
self.assertEqual(list(self.home.iterdir()), [])
|
||||
|
||||
def test_apply_preserves_user_files_and_backs_up(self):
|
||||
prompt = self.home / "AGENTS.md"
|
||||
prompt.write_text("user prompt\n")
|
||||
config = self.home / "config.toml"
|
||||
config.write_text('model = "user-selected"\n')
|
||||
skill = self.home / "skills/codex-subagent-router"
|
||||
skill.mkdir(parents=True)
|
||||
custom = skill / "custom.md"
|
||||
custom.write_text("local extension")
|
||||
result = installer.install(ROOT, self.home, include_prompt=True, apply=True)
|
||||
self.assertEqual((Path(result["backup"]) / "AGENTS.md").read_text(), "user prompt\n")
|
||||
self.assertEqual(config.read_text(), 'model = "user-selected"\n')
|
||||
self.assertEqual(custom.read_text(), "local extension")
|
||||
self.assertEqual(prompt.read_bytes(), (ROOT / "prompts/AGENTS.md").read_bytes())
|
||||
again = installer.install(ROOT, self.home, include_prompt=True, apply=True)
|
||||
self.assertEqual(again["changed_paths"], [])
|
||||
self.assertIsNone(again["backup"])
|
||||
|
||||
def test_skill_only_does_not_replace_global_prompt(self):
|
||||
prompt = self.home / "AGENTS.md"
|
||||
prompt.write_text("keep")
|
||||
installer.install(ROOT, self.home, apply=True)
|
||||
self.assertEqual(prompt.read_text(), "keep")
|
||||
|
||||
def test_symlink_target_is_rejected(self):
|
||||
outside = Path(self.temp.name) / "outside"
|
||||
outside.mkdir()
|
||||
(self.home / "skills").symlink_to(outside, target_is_directory=True)
|
||||
with self.assertRaisesRegex(ValueError, "symlink"):
|
||||
installer.install(ROOT, self.home, apply=True)
|
||||
self.assertEqual(list(outside.iterdir()), [])
|
||||
|
||||
def test_conflicting_directory_is_rejected_before_writes(self):
|
||||
(self.home / "AGENTS.md").mkdir()
|
||||
with self.assertRaisesRegex(ValueError, "regular file"):
|
||||
installer.install(ROOT, self.home, include_prompt=True, apply=True)
|
||||
self.assertFalse((self.home / "skills").exists())
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
Reference in New Issue
Block a user