Integrate 2.1 workflow and prepare release 2.2.0
This commit is contained in:
@@ -52,3 +52,39 @@ These are static review cases, not a record that a model executed them. Use them
|
||||
| Host permits parallel children and both file and semantic write sets are disjoint | Children may edit in parallel; a parent's serial tool edits do not impose a global write lock |
|
||||
| Several independent ready slices, parent only has later synthesis, host permits this dispatch | Run useful children concurrently and integrate on return; no universal requirement to invent simultaneous parent work |
|
||||
| Host explicitly requires useful concurrent parent work | Honor that stricter requirement even when the fleet policy otherwise permits parent waiting |
|
||||
|
||||
## 2.0 escalation scenarios (manual forward evaluation)
|
||||
|
||||
These expectations are not recorded live-agent outcomes.
|
||||
|
||||
| Scenario | Expected behavior |
|
||||
| --- | --- |
|
||||
| Luna lacks the required input file | Diagnose input gap; do not consult Astra |
|
||||
| Sol fails the same acceptance twice, then changes worker | Pause/reassess; preserve issue ID and count |
|
||||
| Astra rejects a reproducible result without a counterexample | Check baseline and coverage; evidence decides, not model rank |
|
||||
| A worker exposes a serious invariant violation on its first attempt | Suspend affected acceptance immediately |
|
||||
| Requirements change after expert approval | Invalidate affected conclusion and verify the new baseline |
|
||||
| Sol cannot explain the decisive expert reasoning | Do not accept or perform the dependent external action |
|
||||
| Root cause is resolved and implementation is settled | Return remaining work to an adequate Sol/Luna route |
|
||||
| Second expert consultation repeats the same question | Require new evidence or a specific prior omission |
|
||||
|
||||
For cost comparison use the same task baselines and acceptance checks across Sol alone, Astra planning plus Luna execution, and Sol-led three-model routing. Include sequential, decomposable and contract-heavy tasks; vary granularity and concurrency separately. Count failed attempts and parent integration in total tokens, actual cost and elapsed time. Repeat runs and report quality/failure distributions before claiming a winning route. Model requests without host identity remain unknown; synthetic packets do not substitute for measurements.
|
||||
|
||||
## 2.1 proactive fleet scenarios
|
||||
|
||||
These are expected behaviors; they do not assert actual model execution or measured performance. The previous one-question repetition case still requires progress, but there is no universal consultation-count quota.
|
||||
|
||||
| Scenario | Expected behavior |
|
||||
| --- | --- |
|
||||
| Consequential difficult interface choice, no failures yet, independent parent work available | Consider bounded upfront Astra analysis; parent owns final contract |
|
||||
| Expert round four has a new discriminator result | Continue scoped topic with recorded basis; do not reject solely due to round count |
|
||||
| Expert repeats the same proposal without progress | Pause topic and obtain missing discriminator or decision |
|
||||
| User explicitly delegates read-only independent review | Allow bounded delegation, preserve read-only scope |
|
||||
| User requests only advice about fleet design | Inspect and advise without launching a fleet |
|
||||
| One independent slice is ready while another investigation continues | Dispatch ready slice without a whole-wave barrier |
|
||||
| A child returns and independent ready work remains | Check real capacity and integration load, then reuse/refill appropriately |
|
||||
| Returned patches exceed current integration capacity | Prioritize integration and pause new dependent implementation |
|
||||
| Different files share the same registry contract | Assign one semantic owner and gate dependent writes |
|
||||
| Parent is Astra and routine task has no independent expert need | Use adequate workers; do not duplicate Astra for a role quota |
|
||||
| Child proposes grandchildren beyond its assigned capacity/scope | Do not expand the tree without bounded parent allocation and real capacity |
|
||||
| Review result predates a relevant contract change | Recheck affected claims on the integrated baseline before acceptance |
|
||||
|
||||
@@ -0,0 +1,27 @@
|
||||
# 2.0.0 — Sol 主导,按证据升级
|
||||
|
||||
发布日期:2026-09-15。
|
||||
|
||||
## 变化
|
||||
|
||||
- 通用会话建议由 Astra 改为 Sol / medium,复杂推理按需 high;保留用户已选模型。
|
||||
- 新增疑难检查点、失败分类、跨 agent 的问题计数、两次同类失败暂停以及严重反例立即处理。
|
||||
- 区分条件修复、执行能力升级和 Astra 有界咨询;默认每问题一次专家咨询,追加需要新证据或具体遗漏。
|
||||
- 裁决依赖同基线可复现证据;主线程必须理解并验证关键结论,解决后恢复合适的低成本执行路径。
|
||||
- 新增可选决策记录校验器及行为测试;兼容旧证据 schema 1/2,安装仍默认预览并按文件备份。
|
||||
|
||||
## 升级
|
||||
|
||||
检出 tag v2.0.0,运行安装预览,再按需要运行:
|
||||
|
||||
```sh
|
||||
python3 scripts/install.py --include-prompt --apply
|
||||
python3 scripts/install.py --include-prompt --check
|
||||
python3 skills/codex-subagent-router/scripts/validate_decision_record.py examples/decisions/resolved.json
|
||||
```
|
||||
|
||||
安装不会修改 config.toml。若希望更改未来会话默认模型,在支持对应字段的客户端中单独设置 model = "gpt-5.6-sol" 与 model_reasoning_effort = "medium"。重新加载或新开会话;当前会话运行身份不会因此得到证明。
|
||||
|
||||
## 验证边界
|
||||
|
||||
单元测试验证校验器拒绝矛盾状态、安装安全与资源完整性;行为场景供真实授权任务的前向评估使用。没有运行 Astra/Sol/Terra/Luna 成本或速度对照实验,不把合成记录当作真实运行证据。并行收益依赖任务可分解性;更细拆分可能增加总 token 和返工。
|
||||
@@ -0,0 +1,19 @@
|
||||
# 2.1.0 — Astra 前置介入与滚动并行
|
||||
|
||||
## 变化
|
||||
|
||||
- Astra 从失败后的专家扩展为高影响困难方案、跨系统综合和独立深度反例审查的参与者。零失败的主动咨询合法,保留主线程裁决与验收。
|
||||
- 专家任务按问题、产物、退出条件设界。可依据新证据、可测试的细化或明确遗漏持续多轮;不按统一次数配额截断,也不允许无进展循环。
|
||||
- 新增就绪队列、提前派发、按依赖滚动补位、动态队形、整合积压控制与恢复规则。宿主容量、共享文件/语义所有权及外部串行门不变。
|
||||
- 明确授权的只读并行审查属于可委派工作;普通审计建议仍不自动启动 agent。
|
||||
- 决策校验器在连续记录中拒绝累计专家咨询次数倒退,包括失败分类变化时;补充主动咨询和多轮进展测试。
|
||||
|
||||
## 兼容与安装
|
||||
|
||||
证据 schema 1/2、决策 schema 1 不变。合法旧记录保持兼容,`--previous` 现在会拒绝咨询次数倒退或前条记录次数格式错误。`followup_basis` 仍为字符串,应保留连续记录来追踪每轮依据;结构校验不证明其真实性。
|
||||
|
||||
从已审阅的 2.1 源码运行 `python3 scripts/install.py` 预览,再用 `--apply` 同步技能。全局提示词需显式添加 `--include-prompt`。安装不会修改模型配置或扩大宿主槽位;备份与回退流程见根目录 README。版本文件不代表 Git 标签或正式 Release 已创建。
|
||||
|
||||
## 验证边界
|
||||
|
||||
使用确定性测试验证校验器和安装行为,并对策略进行独立子智能体审查及模型场景评估。请求模型与实际宿主观测身份分别记录;静态评估不等于真实舰队调度实验,不证明质量、成本或速度提升。官方容量配置与当前工具的计数方式应分别核实。
|
||||
@@ -0,0 +1,22 @@
|
||||
# 2.2.0 — GPT-6 三档路由与积极 Luna 舰队
|
||||
|
||||
## 变化
|
||||
|
||||
- 推荐路线统一为 GPT-6 Luna / Sol / Astra。Sol 接替 Terra 的常规开发、调试与调查;Terra 退出推荐与自动回退,用户明确指定和既有配置仍受尊重。旧宿主按实际 schema 显式使用足够胜任的旧版 Sol/Luna 或本地执行。
|
||||
- 可选全局提示词提供普通开发、检索、诊断、审查与验证的持续并行授权。仅审计并行配置、用户禁止及更高优先级宿主限制优先;审查仍只读,技能自动发现不等于授权。
|
||||
- Luna 优先承担有界收集、比对、验证、局部审查及已定方案修改。尽早派发、按真实容量滚动补位,整合积压或资源争用时缩小宽度;不固定只开一两个,也不以模型配额凑数量。
|
||||
- 文件与语义写集均不相交且宿主允许时可并行修改。宿主允许时,多个独立子任务可运行,主线程等待后整合;不以虚构主线程工作满足并行条件。
|
||||
- 按需参考说明 GPT-6 异步工具、执行中纠偏、动态 effort 与上下文恢复的边界;不把 API 功能误当成 Desktop/CLI 已开放的控制。价格快照标注日期、计费条件与实际任务成本差异。
|
||||
- 保留 2.1 的 Sol 常态指挥、Astra 前置介入、进展驱动的多轮专家协作、问题计数和基于证据的裁决。专家意见不取代主线程理解、验证与最终验收。
|
||||
|
||||
## 兼容与安装
|
||||
|
||||
证据 schema 1/2 与决策 schema 1 不变,校验器与原有记录兼容。新增参考纳入资源完整性检查。历史发行说明保留原版本内容。
|
||||
|
||||
从 v2.2.0 源码运行 `python3 scripts/install.py` 预览,使用 `--apply` 安装技能;全局提示词另加 `--include-prompt`。该选项替换而非智能合并现有提示词,先审阅差异。安装不切换活动模型、不扩大宿主槽位、不改 `config.toml`,沿用备份与回退机制。
|
||||
|
||||
## 验证边界
|
||||
|
||||
验证包括仓库单元测试、资源及版本检查、skill-creator 校验、静态场景与独立差异审查。Windows 上的符号链接测试可能在建立夹具时因 WinError 1314 受阻,应与原始基线区分并如实报告,不删除或跳过测试。
|
||||
|
||||
本版本未声称实测 GPT-6 Luna/Sol 子智能体身份、舰队吞吐、质量或费用收益。模型与参数必须由宿主真实 schema 确认;请求参数和自述不是运行身份。API 价格不代表 Codex 套餐消耗。
|
||||
@@ -0,0 +1,26 @@
|
||||
{
|
||||
"schema_version": 1,
|
||||
"issue_id": "synthetic-contract-001",
|
||||
"baseline": "synthetic-baseline-b",
|
||||
"failure_class": "implementation",
|
||||
"status": "resolved",
|
||||
"same_class_failures": 2,
|
||||
"reassessment": "Synthetic example: incompatible consumer assumption isolated by a discriminating check.",
|
||||
"evidence": [
|
||||
"synthetic:consumer-contract",
|
||||
"synthetic:check-result"
|
||||
],
|
||||
"unresolved": [],
|
||||
"escalation": {
|
||||
"target": "sol",
|
||||
"reason": "Shared consumer semantics exceed the original settled patch.",
|
||||
"expert_consultations": 0,
|
||||
"followup_basis": ""
|
||||
},
|
||||
"verification": {
|
||||
"status": "passed",
|
||||
"baseline": "synthetic-baseline-b",
|
||||
"parent_understands": true
|
||||
},
|
||||
"recovery_route": "terra medium for settled implementation"
|
||||
}
|
||||
@@ -36,6 +36,9 @@
|
||||
|
||||
## 5. 模型与子智能体
|
||||
|
||||
- 新通用工程会话建议 GPT-6 Sol / medium,复杂推理按需 high,保留用户已选模型。Astra 可前置处理高影响困难方案、跨系统综合和独立反例审查,不以先失败为条件,也不固定参与每次规划或审查。
|
||||
- 规划后、关键检查失败、共享合同变化及验收前复核疑难。同类失败两次暂停切片,严重反例立即处理;问题 ID 与计数不随换 agent 清零。按可复现证据裁决,主线程不理解或无法验证的关键结论不得验收;专家续轮须有新证据或可测试进展,解决后回到最低足够路线。
|
||||
|
||||
- 主线程负责关键路径、共享决定、冲突裁决、整合与最终验收;权限判断、生产写入、PR/merge/deploy 由主线程串行处理。
|
||||
- 本提示词作为适用指令时,明确授权在用户任务范围内,对普通开发、检索、诊断、审查和验证积极使用子智能体并行;审查仍只读。用户禁止、仅要求审计并行配置或更高优先级宿主限制时不启动。技能被自动发现本身不构成授权,不扩大外部操作权限。
|
||||
- 有可并行的独立产物时默认尽早分工,主线程优先同步推进关键路径;宿主允许时,也可让多个独立子任务并行,主线程等待后集中整合。不等再次点名 agent,不要求事先证明加速。优先用 Luna 分担有界收集、比对、验证、局部审查和已定方案修改,避免主线程包办所有机械工作。
|
||||
|
||||
@@ -1,5 +1,7 @@
|
||||
# Codex Subagent Router
|
||||
|
||||
当前版本:**2.2.0**。GPT-6 三档路由、积极 Luna 舰队与任务连续性适配;保留 2.1 的 Astra 前置介入、滚动调度及证据裁决。见 [2.2 发行说明](docs/releases/2.2.0.md)、[舰队调度](skills/codex-subagent-router/references/fleet.md) 和 [专家介入与裁决](skills/codex-subagent-router/references/escalation.md)。
|
||||
|
||||
面向 **GPT-6 Luna / Sol / Astra** 的 Codex 路由技能与工程提示词。把可独立完成的工作交给合适的模型,并让触发、运行状态、证据与最终责任保持清楚。Terra 已退出推荐与自动回退路线;旧宿主按实际能力显式兼容。
|
||||
|
||||
这是供 Codex 读取的技能和可选提示词,附带证据校验、安装与测试工具。它不包含常驻调度器,也不会自行调用模型、修改模型配置或开启并行权限。
|
||||
@@ -96,6 +98,10 @@ Terra 原先承担的常规开发、调试和调查统一交给 Sol。新版 Sol
|
||||
|
||||
按实时工具 schema 检查模型、effort、fork、并发槽位和工作区。API 模型目录不证明当前子智能体支持。可选配置示例见 [运行时参考](skills/codex-subagent-router/references/platforms.md),安装工具不会应用它。
|
||||
|
||||
新通用工程会话建议 GPT-6 Sol / medium,复杂推理按需 high,保留用户已选主模型。Astra 可前置处理影响多个后续任务的困难决策、跨系统证据及独立高影响反例,不必先失败。专家按问题、产物和退出条件设界;新增证据或可测试进展支持续轮,没有进展则暂停专题。
|
||||
|
||||
同类验收失败两次暂停切片并重新分类,严重反例立即处理;问题计数跨 agent 保留。证据统一基线后裁决,主线程必须理解并验证关键结论。证据 schema 1/2 与决策 schema 1 保持兼容,累计咨询次数不得倒退;示例见 [决策记录](examples/decisions/resolved.json)。
|
||||
|
||||
## GPT-6 工作方式适配
|
||||
|
||||
技能与全局提示词保留目标、权限、所有权与验收要求,减少固定步骤和重复规则。固定工作优先工具,子智能体承担已授权的独立判断;异步等待期间推进独立工作,用户纠偏后核对受影响任务和迟到结果,长任务压缩后恢复未完成状态。
|
||||
@@ -109,6 +115,7 @@ python3 -m unittest discover -s tests -v
|
||||
python3 scripts/check_repository.py
|
||||
python3 skills/codex-subagent-router/scripts/validate_evidence_packet.py --self-test
|
||||
python3 skills/codex-subagent-router/scripts/validate_evidence_packet.py examples/completed-unknown.json
|
||||
python3 skills/codex-subagent-router/scripts/validate_decision_record.py examples/decisions/resolved.json
|
||||
python3 scripts/install.py --include-prompt --check
|
||||
```
|
||||
|
||||
|
||||
@@ -13,10 +13,15 @@ def check() -> list:
|
||||
skill = ROOT / "skills/codex-subagent-router"
|
||||
for name in ("SKILL.md", "agents/openai.yaml", "references/lifecycle.md",
|
||||
"references/platforms.md", "references/evidence-packet.md",
|
||||
"references/routing-matrix.md", "scripts/validate_evidence_packet.py"):
|
||||
"references/routing-matrix.md", "references/escalation.md", "references/fleet.md",
|
||||
"references/gpt6-adaptation.md", "references/model-economics.md",
|
||||
"scripts/validate_decision_record.py", "scripts/validate_evidence_packet.py"):
|
||||
if not (skill / name).is_file():
|
||||
errors.append("missing skill resource: " + name)
|
||||
entry = (skill / "SKILL.md").read_text(encoding="utf-8")
|
||||
version = (ROOT / "VERSION").read_text(encoding="utf-8").strip()
|
||||
if not re.fullmatch(r"\d+\.\d+\.\d+", version) or "Version " + version + "." not in entry:
|
||||
errors.append("invalid or inconsistent package version")
|
||||
if not re.match(r"\A---\nname: codex-subagent-router\ndescription: .+\n---\n", entry):
|
||||
errors.append("invalid expected skill frontmatter")
|
||||
metadata = (skill / "agents/openai.yaml").read_text(encoding="utf-8")
|
||||
|
||||
@@ -5,6 +5,8 @@ description: Route authorized Codex subagents across GPT-6 Luna, Sol and Astra;
|
||||
|
||||
# Codex Subagent Router
|
||||
|
||||
Version 2.2.0. For new general engineering sessions recommend GPT-6 Sol / medium, with high for demonstrated reasoning needs. Preserve the selected parent. Use Astra proactively for high-leverage difficult decisions, independent counterexample reviews and residual hard reasoning; neither prior failure nor a model quota is required.
|
||||
|
||||
Make delegation useful, observable and bounded. Preserve the selected parent model. The parent owns the critical path, shared decisions, integration and final acceptance. This skill neither changes configuration nor grants external permissions.
|
||||
|
||||
## Trigger before routing
|
||||
@@ -27,6 +29,8 @@ Spawn only P/C work with a concrete output and useful parallel benefit. Prefer u
|
||||
|
||||
Choose the execution mechanism before a model: local tools for fixed commands/transforms, host-supported concurrent or async tools for independent I/O, and authorized children for useful independent reasoning or implementation. Waiting on a tool alone is not a reason to create an agent. Prefer a coherent outcome with acceptance over many tiny handoffs; bound ambiguity and ownership rather than prescribing every reasoning step.
|
||||
|
||||
For queue states, integration backpressure and proactive expert formations, read [fleet scheduling](references/fleet.md).
|
||||
|
||||
## Proactive Luna fleet
|
||||
|
||||
Once authorized, parallel execution is the default for ready independent work. At the first useful decomposition and after each material result or scope change, look for slices that can run alongside the parent's critical path. Dispatch those slices promptly instead of finishing the parent's entire investigation first. Do not require proof of measured speedup before a clearly useful split.
|
||||
@@ -48,7 +52,7 @@ Stay local for a trivial task, no useful parallel benefit, an immediate dependen
|
||||
| Bounded classification, conversion, settled patch with fixed checks | GPT-6 Luna medium; high for local reasoning |
|
||||
| Everyday implementation, debugging, locating/correlating artifacts | GPT-6 Sol medium |
|
||||
| Bounded complex analysis, design or financial/security evidence | GPT-6 Sol medium/high |
|
||||
| Hardest independent synthesis across code, tools and research | GPT-6 Astra medium/high |
|
||||
| High-leverage difficult design, cross-system synthesis, independent high-impact counterexample review | GPT-6 Astra medium/high |
|
||||
| Shared decisions, integration and external/final gates | Current parent, serial |
|
||||
|
||||
Luna is the first candidate when inputs are bounded, the contract is settled, correctness is cheaply checkable, and failure is contained. All four conditions matter; a short diff can still hide a shared semantic decision. Use Sol directly when investigation, design choices or cross-file causal reasoning dominate. Use Astra directly for the hardest independent synthesis when expected quality or saved rework justifies it. Do not make high/max Luna a mandatory step before Sol.
|
||||
@@ -87,6 +91,8 @@ Use [evidence packets](references/evidence-packet.md) when structured evidence i
|
||||
|
||||
## Context and ownership discipline
|
||||
|
||||
Read [escalation and adjudication](references/escalation.md) when planning exposes uncertainty, a key check fails, shared contracts change or acceptance has unresolved counterexamples. Preserve issue IDs and failure counts across agents. Repair missing inputs first; adjudicate by reproducible evidence, not model rank. The parent must understand and verify decisive conclusions. Continue expert rounds only with evidence-based progress, and recover to an adequate cheaper route after resolution.
|
||||
|
||||
Load only needed references; keep long logs on disk. Avoid whole-history forks and repeating contracts, packets or prior findings. Account for parent context/reasoning, child usage, waiting, retries and integration when evaluating cost.
|
||||
|
||||
After steering or context compaction, recover the current objective, accepted corrections, permissions, baseline, file/semantic owners and pending tool/child IDs before continuing dependent work. Retain valid completed evidence; revalidate only results affected by the change. A queued correction is not proof that a running writer has stopped or adopted it.
|
||||
|
||||
@@ -0,0 +1,42 @@
|
||||
# Sol-led escalation and evidence adjudication
|
||||
|
||||
Recommend GPT-6 Sol / medium for new general engineering sessions; preserve an explicit user selection. The selected parent owns planning, difficult implementation, integration and acceptance. Sol handles routine implementation and investigation; Luna handles settled, checkable work. Terra is retired from recommendations and automatic fallbacks. Astra contributes both proactive high-leverage analysis and residual hard reasoning, without becoming a mandatory stage for every task. These are hypotheses to calibrate, not benchmark results.
|
||||
|
||||
## Reassess at observable checkpoints
|
||||
|
||||
Reassess after planning, a key failed check, shared-contract changes and before acceptance. Track a material issue with a stable issue ID and baseline. Counters follow the issue across retries, replacement agents and model changes. After two same-class acceptance failures, pause that slice and classify before another attempt. A serious counterexample or unsafe assumption triggers immediately.
|
||||
|
||||
| Observation | First response |
|
||||
| --- | --- |
|
||||
| Missing inputs, access, tool failure | Diagnose environment or obtain evidence; model escalation does not repair access |
|
||||
| Material requirement ambiguity | Inspect consumers/contracts; ask the user only for unresolved product choices |
|
||||
| Different root causes imply different fixes | Design the smallest discriminating check |
|
||||
| Conflicting records or stale baseline | Reconcile original evidence; suspend the dependent conclusion |
|
||||
| Fixes alternate between breaking invariants | Revisit shared contract and integration ownership |
|
||||
| High impact with no reliable verifier | Block affected acceptance/action; seek a verifier or explicit product decision |
|
||||
|
||||
Continue independent authorized work while the affected slice is paused. A lack of information, authorization or tools is not a reasoning failure.
|
||||
|
||||
## Choose the intervention
|
||||
|
||||
1. Repair the contract, baseline, inputs or verifier first. Reuse an appropriate idle agent; do not reset failure history by respawning.
|
||||
2. Increase execution capability when a well-specified slice exceeds the worker: Luna to Sol as justified, or Astra directly for the hardest independent synthesis. Select the adequate model directly, without visiting every tier. Effort, model and concurrency are independent controls.
|
||||
3. Select Astra directly when a difficult architectural choice affects many downstream tasks, competing consequential designs need discrimination, cross-system evidence needs synthesis, or an independent counterexample search materially protects a high-impact change. Repeated failure is not required. State the uncertainty, consequence of a wrong decision, available evidence and concrete expert deliverable. Residual difficult reasoning conflicts remain a valid trigger. Missing access alone is not a trigger. Authorized delegation and an independent bounded consultation are still required; the parent advances useful independent work when available and retains the shared decision. Where the host permits, useful parallel consultations may run while the parent waits for later adjudication; honor stricter host requirements. An immediate parent dependency stays local; requesting a child or writing configuration does not switch the live parent. If the parent cannot proceed reliably, describe the required capability/evidence instead of claiming a switch occurred.
|
||||
|
||||
An expert packet contains: issue ID, exact relevant baseline, goal, invariants, original evidence references, competing hypotheses, discriminators attempted or explicitly not yet run, and one focused question. For proactive work, zero failures is valid. Request counterexamples, a discriminating check, conditional conclusions and unresolved points. Supply only necessary context. No implementation writes unless separately assigned a disjoint write contract.
|
||||
|
||||
Bound each expert assignment by its question, deliverable and exit condition, not a universal one-consultation quota. Continue with the same suitable expert while new evidence, a testable refinement or a specific prior omission supports progress; record the basis for each additional round. An unresolved topic alone does not justify repetition. If no discriminating progress is available, pause that topic, obtain the missing input/check or return the decision to the parent. Do not loop reviews until models agree. Time/checkpoint budgets are instructions unless enforced by the host.
|
||||
|
||||
## Adjudicate by evidence
|
||||
|
||||
Normalize baseline and convert disagreement into falsifiable claims. Compare reproducible checks, original records and applicable contracts; confirm the check actually covers the disputed invariant. Neither majority vote nor a more expensive model wins automatically. A material unresolved counterexample from any model blocks the affected acceptance.
|
||||
|
||||
The selected parent (normally Sol) must map a decisive suggestion to the actual changes and invariants, understand the reasoning and verify it on the relevant current baseline. Otherwise the result remains unaccepted, even if Astra approves. User product tradeoffs remain user decisions; expert advice grants no permissions. Changed requirements invalidate relevant previous conclusions and require fresh validation.
|
||||
|
||||
## Record and recover
|
||||
|
||||
For material failures keep a compact record: issue ID, baseline, trigger/failure class, same-class failure count, facts, reassessment, escalation reason, evidence-based decision, remaining gaps and recovery route. Ordinary successful work needs no ceremonial record. After resolution, return execution to the lowest adequate route; do not retain Astra for unrelated follow-up work.
|
||||
|
||||
The optional [decision validator](../scripts/validate_decision_record.py) checks reported-state consistency only. Record schema 1 is independent of the product version and evidence-packet schema 2. It cannot prove evidence, enforce runtime gates, detect omitted issues or authenticate model identity. The parent must perform the checks. With `--previous`, it also rejects a changed issue ID, a decreasing failure count within the same class, or a decreasing cumulative expert consultation count even across class changes. A class change must be an evidence-based reclassification, not a counter reset tactic.
|
||||
|
||||
Required fields are demonstrated in the source repository's `examples/decisions/resolved.json`. `evidence` contains nonempty references, not copied logs. `status` is investigating, blocked, escalated or resolved. Resolving requires current-baseline passed verification, parent understanding and no unresolved items. Environment/authority issues cannot target Astra; repeated failures require reassessment; repeated expert consultations require a follow-up basis. The scalar `followup_basis` describes the current continuation; retain prior records for prior rounds. The validator cannot establish progress in each round from one aggregate record. A recovery route is required at resolution. Unknown observed model identity remains acceptable unless an explicit identity gate applies.
|
||||
@@ -0,0 +1,41 @@
|
||||
# Proactive fleet scheduling
|
||||
|
||||
Use for authorized complex parallel work. This is a decision procedure for the parent, not a background scheduler, permission grant or capacity override. Preserve the actual host schema, selected parent and final gates.
|
||||
|
||||
## Discover and dispatch early
|
||||
|
||||
After minimum goal/baseline discovery, identify independent questions and useful implementation slices. Do not finish reading every subsystem before dispatching ready investigations. Once a slice's input contract is stable, it may enter implementation while unrelated investigation continues. Shared decisions stay with the parent; preliminary independent drafts are not accepted contracts.
|
||||
|
||||
Maintain a compact queue: task ID, dependency and contract baseline, file AND semantic owner, shared resources, acceptance, requested route, child ID, observed runtime state and integration state. Mark pending, ready, running, returned, accepted, blocked or superseded as task states; keep actual host state separate. Ready means the prerequisites, ownership and permitted actions are established and the output has independent value.
|
||||
|
||||
At intake, dependency resolution, a child return or a meaningful baseline change:
|
||||
|
||||
1. Invalidate affected assumptions and identify ready slices. Prioritize work that unlocks dependencies, avoids costly wrong decisions or yields directly usable implementation.
|
||||
2. Check live capacity and integration load. Dispatch multiple independent ready tasks with useful parallel benefit; prefer concurrent parent work, but allow the parent to wait for later synthesis when the host permits. Honor stricter host requirements for concurrent parent work. Avoid sequential startup followed by unnecessary waits. Every descendant counts against the actual host limit. Children may not recursively expand the fleet without a parent-assigned bounded delegation scope and capacity allocation.
|
||||
3. Continue the parent's independent critical work when available. Use incremental messages for running children and reuse suitable idle ones. Do not perform the same assigned investigation locally merely to stay busy.
|
||||
4. Inspect returned evidence and partial writes; integrate serially. Refill available capacity with ready work without waiting for all siblings. A returned result is not necessarily accepted or a released slot; follow actual lifecycle controls.
|
||||
|
||||
If the next action is immediately dependent on an unassigned task, keep it local. A previously independent task may become a dependency as work progresses; bounded waiting is then legitimate. No independent useful parent work means do not create a ceremonial side task to justify another spawn.
|
||||
|
||||
## Choose a formation for the task
|
||||
|
||||
- Independent investigation: several focused investigators may cover distinct sources, consumers or falsifiable hypotheses. They should not all produce the same broad summary.
|
||||
- Settled implementation: multiple GPT-6 Luna workers own bounded, settled, independently checkable slices; use GPT-6 Sol directly when implementation requires investigation or design judgment. Shared registries, schema and generated outputs require one owner even when paths differ.
|
||||
- Difficult consequential decision: an Astra specialist analyzes the bounded uncertainty while the parent and other workers advance unaffected work. Gate dependent implementation on the parent's accepted contract. Read [expert intervention](escalation.md) for triggers and progress-based continuation.
|
||||
- Independent review: assign a stable commit/artifact and the relevant requirements, without supplying the parent's preferred verdict. Review can overlap unrelated work; stale observations must be rechecked before acceptance.
|
||||
|
||||
There is no mandatory model ratio or permanent reviewer slot. Use several workers of the same adequate model when the tasks warrant it. If Astra is already the selected parent, do not add another Astra just to fulfill a role label. Reuse the parent capability unless a genuinely independent bounded perspective adds value.
|
||||
|
||||
## Backpressure and recovery
|
||||
|
||||
Choose a working width up to actual available capacity based on ready independent work, parent integration capacity and shared-resource limits. Do not reserve slots mechanically or fill them with artificial work. Hardware, service rate limits and tool sessions can constrain useful width below the agent limit.
|
||||
|
||||
When returned patches are accumulating faster than the parent can inspect, or a returned contract blocks other slices, pause new dependent implementation and prioritize integration. When collisions, stale baselines, repeated rework or resource contention appear, stop affected dispatch and reduce width; interrupt obsolete/conflicting writers and inspect partial changes. Independent read-only work may continue if useful and resource-safe. Resume expansion after the cause is resolved and ready work remains.
|
||||
|
||||
Never run acceptance tests against a moving artifact and report them as final. Pin a commit/snapshot or wait until relevant writers stop; a worktree alone does not isolate external databases, generated directories or shared sessions. Before merge/delivery, resolve material counterexamples, verify the integrated baseline and audit the agent tree for remaining writers.
|
||||
|
||||
## Evaluate effectiveness
|
||||
|
||||
Optimize accepted quality and end-to-end time subject to the user's cost constraints. Include parent work, all children, context duplication, retries, expert rounds and integration in cost measurements. Expensive early reasoning can be worthwhile when it prevents downstream rework; cheap high-volume workers can be wasteful when contracts are unsettled. Both are hypotheses until measured.
|
||||
|
||||
For a calibration exercise compare the same tasks and baselines at serial and increasing supported widths, varying granularity separately. Include decomposable and contract-heavy tasks, failures and repeated runs. Record quality, elapsed time, total exposed usage/cost, rework, and integration time; unknown counters or runtime model identity stay unknown. Do not claim speed/cost improvements from static scenarios or a successful single rollout.
|
||||
@@ -24,6 +24,8 @@ Keep required async tool jobs in the same dependency view as children, using the
|
||||
|
||||
After context compaction, recover objective, permissions, owners, pending IDs and accepted evidence before dispatching more work. Detailed API-versus-host limits are in [GPT-6 adaptation](gpt6-adaptation.md).
|
||||
|
||||
For rolling execution, use [fleet scheduling](fleet.md): refill ready work after useful handbacks and dependency changes, rather than waiting for every sibling. Check integration backlog before starting more implementation. Task readiness, output acceptance and actual slot availability remain separate facts.
|
||||
|
||||
## Cancellation and handback
|
||||
|
||||
Interrupting does not undo files or guarantee descendant cancellation. First preserve useful returned evidence, stop affected writers, inspect partial changes, then assign the remaining work to one owner. Never reset shared files wholesale.
|
||||
|
||||
@@ -19,7 +19,7 @@ For clients that support the documented keys, an example is:
|
||||
|
||||
```toml
|
||||
# Example only; merge intentionally into the existing configuration.
|
||||
model = "gpt-6-astra"
|
||||
model = "gpt-6-sol"
|
||||
model_reasoning_effort = "medium"
|
||||
|
||||
[agents]
|
||||
|
||||
@@ -11,6 +11,9 @@ Read only when the entrypoint leaves a routing choice unresolved.
|
||||
| Cheap Luna attempts require repeated parent rewrites | Choose Sol directly next time for this shape; count total accepted-task cost. |
|
||||
| Difficult accounting discrepancy | Sol reasoning; parent owns the semantic decision. |
|
||||
| Coupled architecture across runtimes and tools | Consider Astra directly; keep shared decisions serial. |
|
||||
| Difficult choice would constrain many downstream slices | Consider an early bounded Astra analysis; gate dependent work on the parent's accepted contract. |
|
||||
| High-impact stable change needs independent counterexamples | Consider Astra review when depth adds value; give original requirements and a fixed artifact, not a preferred verdict. |
|
||||
| Expert makes testable progress over several rounds | Continue the scoped topic with incremental evidence; do not impose a one-call quota or repeat unchanged questions. |
|
||||
| Missing credential or inaccessible source | Resolve the evidence/environment gap; model escalation cannot supply access. |
|
||||
| Astra advertised only for task creation | Child support remains unproven; do not create another user task as a workaround. |
|
||||
| Named role fixes model/effort | Accept its supported binding or choose an overridable role; never attach prohibited fork overrides. |
|
||||
|
||||
@@ -0,0 +1,103 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Validate reported decision consistency, not evidence truth or live execution."""
|
||||
import argparse
|
||||
import json
|
||||
from pathlib import Path
|
||||
import sys
|
||||
|
||||
|
||||
def validate(record, previous=None):
|
||||
errors = []
|
||||
if not isinstance(record, dict):
|
||||
return ["record must be an object"]
|
||||
for key in ("issue_id", "baseline", "failure_class", "status"):
|
||||
if not isinstance(record.get(key), str) or not record[key].strip():
|
||||
errors.append(key + " must be nonempty text")
|
||||
if type(record.get("schema_version")) is not int or record["schema_version"] != 1:
|
||||
errors.append("schema_version must be 1")
|
||||
if record.get("failure_class") not in ("environment", "authority", "contract", "implementation", "reasoning"):
|
||||
errors.append("invalid failure_class")
|
||||
if record.get("status") not in ("investigating", "blocked", "escalated", "resolved"):
|
||||
errors.append("invalid status")
|
||||
count = record.get("same_class_failures")
|
||||
if type(count) is not int or count < 0:
|
||||
errors.append("same_class_failures must be a nonnegative integer")
|
||||
elif count >= 2 and not text(record.get("reassessment")):
|
||||
errors.append("repeated failure requires reassessment")
|
||||
for key in ("evidence", "unresolved"):
|
||||
value = record.get(key)
|
||||
if not isinstance(value, list) or any(not text(item) for item in value):
|
||||
errors.append(key + " must be a list of nonempty references/issues")
|
||||
escalation = record.get("escalation")
|
||||
if not isinstance(escalation, dict):
|
||||
errors.append("escalation must be an object")
|
||||
escalation = {}
|
||||
target = escalation.get("target")
|
||||
if target not in ("none", "terra", "sol", "astra"):
|
||||
errors.append("invalid escalation target")
|
||||
if target != "none" and not text(escalation.get("reason")):
|
||||
errors.append("escalation requires a reason")
|
||||
if target == "astra" and record.get("failure_class") in ("environment", "authority"):
|
||||
errors.append("Astra cannot repair environment or authority")
|
||||
consultations = escalation.get("expert_consultations")
|
||||
if type(consultations) is not int or consultations < 0:
|
||||
errors.append("expert_consultations must be a nonnegative integer")
|
||||
elif consultations > 1 and not text(escalation.get("followup_basis")):
|
||||
errors.append("additional consultation requires new evidence or a specific omission")
|
||||
if record.get("status") == "escalated" and target == "none":
|
||||
errors.append("escalated status requires a target")
|
||||
if record.get("status") == "resolved":
|
||||
verification = record.get("verification", {})
|
||||
if not isinstance(verification, dict):
|
||||
verification = {}
|
||||
if (verification.get("status") != "passed"
|
||||
or verification.get("baseline") != record.get("baseline")
|
||||
or verification.get("parent_understands") is not True):
|
||||
errors.append("resolution requires understood, passed verification on current baseline")
|
||||
if record.get("unresolved") != [] or not record.get("evidence"):
|
||||
errors.append("resolution requires evidence and no unresolved items")
|
||||
if not text(record.get("recovery_route")):
|
||||
errors.append("resolution requires a recovery route")
|
||||
if previous is not None:
|
||||
if not isinstance(previous, dict):
|
||||
errors.append("previous record must be an object")
|
||||
else:
|
||||
if previous.get("issue_id") != record.get("issue_id"):
|
||||
errors.append("issue_id changed across continuation")
|
||||
old = previous.get("same_class_failures")
|
||||
if (previous.get("failure_class") == record.get("failure_class")
|
||||
and type(old) is int and type(count) is int and count < old):
|
||||
errors.append("failure counter decreased within the same class")
|
||||
previous_escalation = previous.get("escalation")
|
||||
old_consultations = (previous_escalation.get("expert_consultations")
|
||||
if isinstance(previous_escalation, dict) else None)
|
||||
if type(old_consultations) is not int or old_consultations < 0:
|
||||
errors.append("previous expert_consultations must be a nonnegative integer")
|
||||
elif (type(consultations) is int and consultations >= 0
|
||||
and consultations < old_consultations):
|
||||
errors.append("expert consultation counter decreased across continuation")
|
||||
return errors
|
||||
|
||||
|
||||
def text(value):
|
||||
return isinstance(value, str) and bool(value.strip())
|
||||
|
||||
|
||||
def main():
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("record")
|
||||
parser.add_argument("--previous")
|
||||
args = parser.parse_args()
|
||||
try:
|
||||
record = json.loads(Path(args.record).read_text(encoding="utf-8"))
|
||||
previous = json.loads(Path(args.previous).read_text(encoding="utf-8")) if args.previous else None
|
||||
errors = validate(record, previous)
|
||||
except (OSError, ValueError) as exc:
|
||||
print("Invalid decision record: " + str(exc), file=sys.stderr)
|
||||
return 1
|
||||
print("\n".join(errors) if errors else "Decision record consistency passed (truth not verified)")
|
||||
return int(bool(errors))
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -0,0 +1,116 @@
|
||||
import copy
|
||||
import importlib.util
|
||||
import json
|
||||
from pathlib import Path
|
||||
import unittest
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[1]
|
||||
spec = importlib.util.spec_from_file_location("decision", ROOT / "skills/codex-subagent-router/scripts/validate_decision_record.py")
|
||||
module = importlib.util.module_from_spec(spec)
|
||||
spec.loader.exec_module(module)
|
||||
|
||||
|
||||
class DecisionTests(unittest.TestCase):
|
||||
def setUp(self):
|
||||
self.record = json.loads((ROOT / "examples/decisions/resolved.json").read_text())
|
||||
|
||||
def test_valid_resolution(self):
|
||||
self.assertEqual(module.validate(self.record), [])
|
||||
|
||||
def test_counter_survives_continuation(self):
|
||||
previous = copy.deepcopy(self.record)
|
||||
self.record["same_class_failures"] = 0
|
||||
self.assertTrue(module.validate(self.record, previous))
|
||||
|
||||
def test_reassessment_required(self):
|
||||
self.record["reassessment"] = ""
|
||||
self.assertTrue(module.validate(self.record))
|
||||
|
||||
def test_unresolved_counterexample_blocks_resolution(self):
|
||||
self.record["unresolved"] = ["reproducible contrary result"]
|
||||
self.assertTrue(module.validate(self.record))
|
||||
|
||||
def test_stale_baseline_blocks_resolution(self):
|
||||
self.record["verification"]["baseline"] = "old"
|
||||
self.assertTrue(module.validate(self.record))
|
||||
|
||||
def test_parent_must_understand(self):
|
||||
self.record["verification"]["parent_understands"] = False
|
||||
self.assertTrue(module.validate(self.record))
|
||||
|
||||
def test_environment_and_authority_do_not_escalate_to_astra(self):
|
||||
for category in ("environment", "authority"):
|
||||
self.record["failure_class"] = category
|
||||
self.record["escalation"]["target"] = "astra"
|
||||
self.assertTrue(module.validate(self.record))
|
||||
|
||||
def test_zero_failure_proactive_astra_consultation(self):
|
||||
self.record["failure_class"] = "reasoning"
|
||||
self.record["status"] = "investigating"
|
||||
self.record["same_class_failures"] = 0
|
||||
self.record["reassessment"] = ""
|
||||
self.record["escalation"] = {
|
||||
"target": "astra",
|
||||
"reason": "Proactive review of a high-impact architectural choice.",
|
||||
"expert_consultations": 1,
|
||||
"followup_basis": "",
|
||||
}
|
||||
self.record["unresolved"] = ["expert consultation pending"]
|
||||
self.assertEqual(module.validate(self.record), [])
|
||||
|
||||
def test_more_than_three_progressive_consultations(self):
|
||||
self.record["escalation"]["expert_consultations"] = 4
|
||||
self.record["escalation"]["followup_basis"] = "new discriminator results across rounds"
|
||||
self.assertEqual(module.validate(self.record), [])
|
||||
|
||||
def test_multiple_consultations_need_basis(self):
|
||||
self.record["escalation"]["expert_consultations"] = 2
|
||||
self.assertIn(
|
||||
"additional consultation requires new evidence or a specific omission",
|
||||
module.validate(self.record),
|
||||
)
|
||||
|
||||
def test_consultation_counter_cannot_decrease_across_class_change(self):
|
||||
previous = copy.deepcopy(self.record)
|
||||
previous["escalation"]["expert_consultations"] = 4
|
||||
previous["escalation"]["followup_basis"] = "prior progressive rounds"
|
||||
self.record["failure_class"] = "reasoning"
|
||||
self.record["escalation"]["expert_consultations"] = 3
|
||||
self.record["escalation"]["followup_basis"] = "current evidence"
|
||||
self.assertIn(
|
||||
"expert consultation counter decreased across continuation",
|
||||
module.validate(self.record, previous),
|
||||
)
|
||||
|
||||
def test_consultation_counter_can_increase_with_basis(self):
|
||||
previous = copy.deepcopy(self.record)
|
||||
previous["escalation"]["expert_consultations"] = 1
|
||||
self.record["escalation"]["expert_consultations"] = 4
|
||||
self.record["escalation"]["followup_basis"] = "new discriminator results"
|
||||
self.assertEqual(module.validate(self.record, previous), [])
|
||||
|
||||
def test_malformed_previous_consultation_counter_is_rejected(self):
|
||||
previous = copy.deepcopy(self.record)
|
||||
previous["escalation"]["expert_consultations"] = True
|
||||
self.assertIn(
|
||||
"previous expert_consultations must be a nonnegative integer",
|
||||
module.validate(self.record, previous),
|
||||
)
|
||||
|
||||
def test_recovery_required(self):
|
||||
self.record["recovery_route"] = ""
|
||||
self.assertTrue(module.validate(self.record))
|
||||
|
||||
def test_bad_shapes_and_boolean_counters(self):
|
||||
for value in (None, [], True):
|
||||
self.assertTrue(module.validate(value))
|
||||
self.record["same_class_failures"] = True
|
||||
self.assertTrue(module.validate(self.record))
|
||||
|
||||
def test_unhashable_enum_fields_are_rejected(self):
|
||||
for key in ("failure_class", "status"):
|
||||
record = copy.deepcopy(self.record)
|
||||
record[key] = []
|
||||
self.assertTrue(module.validate(record))
|
||||
self.record["escalation"]["target"] = {}
|
||||
self.assertTrue(module.validate(self.record))
|
||||
Reference in New Issue
Block a user