Develop 2.1.0 proactive Astra intervention and rolling fleet scheduling

This commit is contained in:
Codex
2026-09-15 23:51:54 +08:00
parent e2fbb211b6
commit e66dcc1810
13 changed files with 175 additions and 24 deletions
+1 -1
View File
@@ -1 +1 @@
2.0.0
2.1.0
+19
View File
@@ -37,3 +37,22 @@ These expectations are not recorded live-agent outcomes.
| Second expert consultation repeats the same question | Require new evidence or a specific prior omission |
For cost comparison use the same task baselines and acceptance checks across Sol alone, Astra planning plus Luna execution, and Sol-led four-model routing. Include sequential, decomposable and contract-heavy tasks; vary granularity and concurrency separately. Count failed attempts and parent integration in total tokens, actual cost and elapsed time. Repeat runs and report quality/failure distributions before claiming a winning route. Model requests without host identity remain unknown; synthetic packets do not substitute for measurements.
## 2.1 proactive fleet scenarios
These are expected behaviors; they do not assert actual model execution or measured performance. The previous one-question repetition case still requires progress, but there is no universal consultation-count quota.
| Scenario | Expected behavior |
| --- | --- |
| Consequential difficult interface choice, no failures yet, independent parent work available | Consider bounded upfront Astra analysis; parent owns final contract |
| Expert round four has a new discriminator result | Continue scoped topic with recorded basis; do not reject solely due to round count |
| Expert repeats the same proposal without progress | Pause topic and obtain missing discriminator or decision |
| User explicitly delegates read-only independent review | Allow bounded delegation, preserve read-only scope |
| User requests only advice about fleet design | Inspect and advise without launching a fleet |
| One independent slice is ready while another investigation continues | Dispatch ready slice without a whole-wave barrier |
| A child returns and independent ready work remains | Check real capacity and integration load, then reuse/refill appropriately |
| Returned patches exceed current integration capacity | Prioritize integration and pause new dependent implementation |
| Different files share the same registry contract | Assign one semantic owner and gate dependent writes |
| Parent is Astra and routine task has no independent expert need | Use adequate workers; do not duplicate Astra for a role quota |
| Child proposes grandchildren beyond its assigned capacity/scope | Do not expand the tree without bounded parent allocation and real capacity |
| Review result predates a relevant contract change | Recheck affected claims on the integrated baseline before acceptance |
+19
View File
@@ -0,0 +1,19 @@
# 2.1.0 — Astra 前置介入与滚动并行
## 变化
- Astra 从失败后的专家扩展为高影响困难方案、跨系统综合和独立深度反例审查的参与者。零失败的主动咨询合法,保留主线程裁决与验收。
- 专家任务按问题、产物、退出条件设界。可依据新证据、可测试的细化或明确遗漏持续多轮;不按统一次数配额截断,也不允许无进展循环。
- 新增就绪队列、提前派发、按依赖滚动补位、动态队形、整合积压控制与恢复规则。宿主容量、共享文件/语义所有权及外部串行门不变。
- 明确授权的只读并行审查属于可委派工作;普通审计建议仍不自动启动 agent。
- 决策校验器在连续记录中拒绝累计专家咨询次数倒退,包括失败分类变化时;补充主动咨询和多轮进展测试。
## 兼容与安装
证据 schema 1/2、决策 schema 1 不变。合法旧记录保持兼容,`--previous` 现在会拒绝咨询次数倒退或前条记录次数格式错误。`followup_basis` 仍为字符串,应保留连续记录来追踪每轮依据;结构校验不证明其真实性。
从已审阅的 2.1 源码运行 `python3 scripts/install.py` 预览,再用 `--apply` 同步技能。全局提示词需显式添加 `--include-prompt`。安装不会修改模型配置或扩大宿主槽位;备份与回退流程见根目录 README。版本文件不代表 Git 标签或正式 Release 已创建。
## 验证边界
使用确定性测试验证校验器和安装行为,并对策略进行独立子智能体审查及模型场景评估。请求模型与实际宿主观测身份分别记录;静态评估不等于真实舰队调度实验,不证明质量、成本或速度提升。官方容量配置与当前工具的计数方式应分别核实。
+2 -2
View File
@@ -40,13 +40,13 @@
## 5. 模型与子智能体
- 新通用工程会话建议 Sol / medium,必要时 high;保留用户已选主模型。Astra 用于证据充分后仍未解决的困难综合问题,不固定承担每次规划或审查。
- 新通用工程会话建议 Sol / medium,必要时 high;保留用户已选主模型。Astra 可前置处理高影响的困难方案、跨系统综合与独立反例审查,也处理残余疑难;不以先失败为条件,不固定承担每次规划或审查。
- 规划后、关键检查失败、共享合同变化及验收前检查疑难信号。同类失败两次暂停切片,严重反例立即处理;问题计数不随换 agent 清零。按路由技能区分条件缺失、实现不足和推理冲突,依据可复现证据裁决。主线程不理解或无法验证的关键结论不得验收;解决后回到最低足够路线。
- 主线程负责目标、关键路径、共享合同、冲突裁决、最终整合与验收。子任务提供证据或授权写集内的补丁。
- 子智能体路由、触发、合同和生命周期以 codex-subagent-router 技能为维护入口,实际调用服从当前工具 schema。技能缺失时在主线程继续可行工作,不虚构委派。
- 只有用户明确要求委派,或适用指令明确授权时才启动子智能体。自动发现技能、审查/修改技能、复杂任务或“利用模型能力”本身不构成授权。更高优先级宿主限制仍有效。
- 获得授权后,主动分配有独立价值且能缩短关键路径的任务;不按模型配额或文件数量制造并行。不委派下一步立即依赖的阻塞任务。
- 获得授权后,在任务开始与依赖解除时主动发现并派发有独立价值的工作,按就绪状态滚动补位,不等整批结束。模型与队形随任务变化;整合积压、冲突或资源争用时缩小并发,恢复后再扩展。不按模型配额或文件数量制造并行,不委派下一步立即依赖的阻塞任务。
- 不默认把所有子任务继承为主模型。支持选择时显式指定合适的模型与 effort;核对 fork 兼容性。配置、模型自述与请求参数不能证明真实运行身份,缺失字段记录 unknown。
- 按任务形状选择:Luna 做固定收集/转换,Terra 做常规实现,Sol 做复杂分析,Astra 做最困难的综合工作;使用最低足够且受支持的 effort。缺数据、权限或工具故障先诊断。
- 默认视为共享工作区,文件写集与语义写集都须明确。工作角色和只读要求不是独立沙盒的证明。
+15 -5
View File
@@ -1,6 +1,6 @@
# Codex Subagent Router
当前版本:**2.0.0**。推荐 Sol 主会话,按证据升级;见 [2.0 发布说明](docs/releases/2.0.0.md) 和 [升级与裁决机制](skills/codex-subagent-router/references/escalation.md)。
当前版本:**2.1.0**。Sol 常态指挥,Astra 前置处理关键疑难,按就绪状态滚动并行;见 [2.1 版本说明](docs/releases/2.1.0.md)、[舰队调度](skills/codex-subagent-router/references/fleet.md) 和 [专家介入与裁决](skills/codex-subagent-router/references/escalation.md)。
面向 **GPT-6 Astra / GPT-5.6 Sol / Terra / Luna** 的 Codex 路由技能与工程提示词。把可独立完成的工作交给合适的模型,并让触发、运行状态、证据与最终责任保持清楚。
@@ -83,16 +83,26 @@ python3 scripts/install.py --include-prompt --apply
| 有界转换、分类、已定方案补丁 | Luna high |
| 常规实现、调试和跨文件调查 | Terra medium,必要时 high |
| 复杂分析、设计和关键语义证据 | Sol medium,必要时 high |
| 最困难的独立综合任务 | Astra medium,必要时 high |
| 高影响困难方案、跨系统综合、独立深度反例审查 | Astra medium,必要时 high |
| 共享决定、最终整合与验收 | 当前主线程 |
通用工程会话默认建议 Sol / medium;复杂推理按需 high。Sol 承担规划、复杂实现、整合和验收,Terra/Luna 承担适合的执行,Astra 只处理证据充分后仍未解决的困难综合问题。保留用户已选模型。技能不能自动切换当前主模型,也不保证成本或速度收益。
通用工程会话默认建议 Sol / medium;复杂推理按需 high。Sol 承担规划、复杂实现、整合和验收,Terra/Luna 承担适合的执行。Astra 可在影响多个后续任务的困难决策前介入,综合跨系统证据或独立寻找高影响改动的反例,也处理残余疑难;不必先失败。保留用户已选模型,最终责任属于实际主线程。技能不能自动切换当前主模型,也不保证成本或速度收益。
同类验收失败两次暂停该切片并重新分类;严重反例立即触发。缺数据、权限或工具先修复条件。裁决先统一基线,再用可证伪假设与检查解决冲突,不能按模型贵贱投票。主线程无法理解或验证关键结论时不得验收;问题解决后恢复较便宜的执行路径。
默认每个问题一次有界专家咨询,追加必须有新证据或明确遗漏。并发仅用于具有独立产物、写集不冲突且能缩短关键路径的工作;细拆和更多 agent 会增加上下文复制、重试与整合成本。Astra 规划加 Luna 执行适合方案稳定且验收明确的任务,尚无本仓库真实对照数据证明其普遍优于四层路由。后续比较须同时记录质量、总 token、实际费用、耗时、返工及主线程整合开销。
专家任务以问题、产物和退出条件为界,不设置统一的一次咨询配额。新增证据、可测试的细化或明确遗漏可支持持续多轮;没有可区分的进展时暂停该专题,而不是反复询问直到模型同意。决策记录保留累计咨询次数和每轮推进依据。
升级到固定版本可执行 `git checkout v2.0.0`,然后预览并应用安装。产品 2.0.0 保持原证据 schema 1/2 兼容,新增可选决策记录 schema 1;例子见 [resolved.json](examples/decisions/resolved.json)。
2.1 保持证据 schema 1/2 和决策记录 schema 1。新增连续记录中的咨询次数不可倒退检查;例子见 [resolved.json](examples/decisions/resolved.json)。固定安装可检出已发布标签;2.1 合并后也可从最新 main 安装,版本号不代表标签或 Release 已创建。
## 更积极的并行
获得委派授权后,在任务开始、依赖解除和子任务返回时主动寻找就绪工作。先派发独立调查;某个模块合同稳定后即可开始实现,不等全局调查结束。有真实空闲容量就滚动补位,不等整批结束。
队形随任务调整:可用多个同模型调查者、多个独立实现者,或 Astra 专题分析与其他执行者并行;不固定审查员席位或模型比例。主线程持续推进独立关键工作,统一共享合同与最终验收。
文件和语义写集都必须有 owner。整合积压、过期基线、反复返工或共享资源争用时,减少派发并优先整合,解决后再扩展。审查和最终测试针对固定提交或稳定产物。所有后代纳入宿主真实容量;没有 close 工具不能虚构释放槽位。
并行上限读取实际工具 schema,不在技能中写死数量,也不通过额外任务或嵌套进程绕过限制。更细拆分会增加上下文、重试和整合成本;先比较相同任务在串行及不同受支持宽度下的质量、耗时和总成本,再调整策略。Astra 前置分析可能减少返工,尚不能据此声称整体更便宜或更快。
按实时工具 schema 检查模型、effort、fork、并发槽位和工作区。API 模型目录不证明当前子智能体支持。可选配置示例见 [运行时参考](skills/codex-subagent-router/references/platforms.md),安装工具不会应用它。
+1 -1
View File
@@ -13,7 +13,7 @@ def check() -> list:
skill = ROOT / "skills/codex-subagent-router"
for name in ("SKILL.md", "agents/openai.yaml", "references/lifecycle.md",
"references/platforms.md", "references/evidence-packet.md",
"references/routing-matrix.md", "references/escalation.md",
"references/routing-matrix.md", "references/escalation.md", "references/fleet.md",
"scripts/validate_decision_record.py", "scripts/validate_evidence_packet.py"):
if not (skill / name).is_file():
errors.append("missing skill resource: " + name)
+6 -4
View File
@@ -5,7 +5,7 @@ description: Decide when Codex subagents help, route bounded work across Astra,
# Codex Subagent Router
Version 2.0.0. For new general engineering sessions recommend Sol / medium, with high when demonstrated reasoning needs justify it. Preserve the user's selected parent; installation does not switch it. Use Astra for bounded residual hard reasoning, rather than mandatory planning of every task.
Version 2.1.0. For new general engineering sessions recommend Sol / medium, with high when demonstrated reasoning needs justify it. Preserve the user's selected parent; installation does not switch it. Use Astra proactively for high-leverage difficult decisions, independent counterexample reviews and residual hard reasoning; neither failure nor a model quota is a prerequisite.
Make delegation useful, observable and bounded. Preserve the selected parent model. The parent owns the critical path, shared decisions, integration and final acceptance. This skill neither changes configuration nor grants external permissions.
@@ -13,8 +13,8 @@ Make delegation useful, observable and bounded. Preserve the selected parent mod
Classify the request before using child tools:
- **Audit or advice:** inspect rules, configuration and available tools; report findings without spawning or changing settings.
- **Authorized execution:** the user requested delegation, or an applicable instruction explicitly authorizes it. Within that scope, actively delegate independent useful slices while the parent advances other work.
- **Audit or advice without delegation authorization:** inspect rules, configuration and available tools; report findings without spawning or changing settings.
- **Authorized delegation:** the user requested delegation, or an applicable instruction explicitly authorizes it. This includes explicitly delegated read-only reviews; delegation does not grant edit permission. Within that scope, actively delegate independent useful slices while the parent advances other work.
- **No delegation authorization:** work locally. Automatic skill discovery, model availability, task complexity and a request to edit this skill do not grant permission to spawn.
Honor higher-priority host restrictions even when a lower-priority rule permits delegation. Do not request authorization repeatedly after it has been granted. Never turn a routing recommendation into a new user-owned task.
@@ -27,6 +27,8 @@ Before dispatch classify each candidate:
Spawn only P/C work with a concrete output and useful parent work available. Do not spawn for one command, ceremonial probes, duplicate reviews or a model quota. Batch homogeneous small work.
For authorized complex work, actively discover independent slices at intake and each readiness change. Dispatch ready work early and refill available capacity without waiting for a whole wave. Read [fleet scheduling](references/fleet.md) for rolling execution, Astra intervention and integration backpressure. Capacity is an upper bound, not a utilization target.
## Select a supported route explicitly
| Work shape | Initial route |
@@ -35,7 +37,7 @@ Spawn only P/C work with a concrete output and useful parent work available. Do
| Bounded classification, conversion, settled patch with fixed checks | Luna high |
| Everyday implementation, debugging, locating/correlating artifacts | Terra medium; high when needed |
| Bounded complex analysis, design or financial/security evidence | Sol medium; high when needed |
| Hardest independent synthesis across code, tools and research | Astra medium; high when needed |
| High-leverage difficult design, cross-system synthesis, independent high-impact counterexample review | Astra medium; high when needed |
| Shared decisions, integration and external/final gates | Current parent, serial |
These are starting heuristics, not measured cost rankings. Pick sufficient capability directly; do not escalate through every model. Missing access/data and tool failures need diagnosis, not a stronger model. After two same-class failures, pause that slice and return the evidence to the parent.
@@ -1,6 +1,6 @@
# Sol-led escalation and evidence adjudication
Recommend Sol / medium for new general engineering sessions; preserve an explicit user selection. Sol owns planning, difficult implementation, integration and acceptance. Terra handles routine implementation; Luna handles settled, checkable work. Astra is a bounded expert for residual hard reasoning, not a mandatory planning or review stage. These are hypotheses to calibrate, not benchmark results.
Recommend Sol / medium for new general engineering sessions; preserve an explicit user selection. The selected parent owns planning, difficult implementation, integration and acceptance. Terra handles routine implementation; Luna handles settled, checkable work. Astra contributes both proactive high-leverage analysis and residual hard reasoning, without becoming a mandatory stage for every task. These are hypotheses to calibrate, not benchmark results.
## Reassess at observable checkpoints
@@ -21,22 +21,22 @@ Continue independent authorized work while the affected slice is paused. A lack
1. Repair the contract, baseline, inputs or verifier first. Reuse an appropriate idle agent; do not reset failure history by respawning.
2. Increase execution capability when a well-specified slice exceeds the worker: Luna to Terra, or Terra to Sol as justified. Select the adequate model directly, without visiting every tier. Effort, model and concurrency are independent controls.
3. Use Astra only for a remaining difficult reasoning conflict with adequate evidence. Authorized delegation and an independent bounded consultation are still required. An immediate parent dependency stays local; requesting a child or writing configuration does not switch the live parent. If the parent cannot proceed reliably, describe the required capability/evidence instead of claiming a switch occurred.
3. Select Astra directly when a difficult architectural choice affects many downstream tasks, competing consequential designs need discrimination, cross-system evidence needs synthesis, or an independent counterexample search materially protects a high-impact change. Repeated failure is not required. State the uncertainty, consequence of a wrong decision, available evidence and concrete expert deliverable. Residual difficult reasoning conflicts remain a valid trigger. Missing access alone is not a trigger. Authorized delegation and an independent bounded consultation are still required; the parent advances useful independent work and retains the shared decision. An immediate parent dependency stays local; requesting a child or writing configuration does not switch the live parent. If the parent cannot proceed reliably, describe the required capability/evidence instead of claiming a switch occurred.
An expert packet contains: issue ID, exact relevant baseline, goal, invariants, original evidence references, competing hypotheses, attempted discriminators and one focused question. Request counterexamples, a discriminating check, conditional conclusions and unresolved points. Supply only necessary context. No implementation writes unless separately assigned a disjoint write contract.
An expert packet contains: issue ID, exact relevant baseline, goal, invariants, original evidence references, competing hypotheses, discriminators attempted or explicitly not yet run, and one focused question. For proactive work, zero failures is valid. Request counterexamples, a discriminating check, conditional conclusions and unresolved points. Supply only necessary context. No implementation writes unless separately assigned a disjoint write contract.
Default to one bounded expert consultation per issue. A follow-up requires new evidence or a specific omission in the first answer; record that basis. Do not loop reviews until models agree. Time/checkpoint budgets are instructions unless enforced by the host.
Bound each expert assignment by its question, deliverable and exit condition, not a universal one-consultation quota. Continue with the same suitable expert while new evidence, a testable refinement or a specific prior omission supports progress; record the basis for each additional round. An unresolved topic alone does not justify repetition. If no discriminating progress is available, pause that topic, obtain the missing input/check or return the decision to the parent. Do not loop reviews until models agree. Time/checkpoint budgets are instructions unless enforced by the host.
## Adjudicate by evidence
Normalize baseline and convert disagreement into falsifiable claims. Compare reproducible checks, original records and applicable contracts; confirm the check actually covers the disputed invariant. Neither majority vote nor a more expensive model wins automatically. A material unresolved counterexample from any model blocks the affected acceptance.
Sol must map a decisive suggestion to the actual changes and invariants, understand the reasoning and verify it on the relevant current baseline. Otherwise the result remains unaccepted, even if Astra approves. User product tradeoffs remain user decisions; expert advice grants no permissions. Changed requirements invalidate relevant previous conclusions and require fresh validation.
The selected parent (normally Sol) must map a decisive suggestion to the actual changes and invariants, understand the reasoning and verify it on the relevant current baseline. Otherwise the result remains unaccepted, even if Astra approves. User product tradeoffs remain user decisions; expert advice grants no permissions. Changed requirements invalidate relevant previous conclusions and require fresh validation.
## Record and recover
For material failures keep a compact record: issue ID, baseline, trigger/failure class, same-class failure count, facts, reassessment, escalation reason, evidence-based decision, remaining gaps and recovery route. Ordinary successful work needs no ceremonial record. After resolution, return execution to the lowest adequate route; do not retain Astra for unrelated follow-up work.
The optional [decision validator](../scripts/validate_decision_record.py) checks reported-state consistency only. Record schema 1 is independent of product version 2.0 and evidence-packet schema 2. It cannot prove evidence, enforce runtime gates, detect omitted issues or authenticate model identity. The parent must perform the checks. With `--previous`, it also rejects a changed issue ID or a decreasing count within the same failure class. A class change must be an evidence-based reclassification, not a counter reset tactic.
The optional [decision validator](../scripts/validate_decision_record.py) checks reported-state consistency only. Record schema 1 is independent of the product version and evidence-packet schema 2. It cannot prove evidence, enforce runtime gates, detect omitted issues or authenticate model identity. The parent must perform the checks. With `--previous`, it also rejects a changed issue ID, a decreasing failure count within the same class, or a decreasing cumulative expert consultation count even across class changes. A class change must be an evidence-based reclassification, not a counter reset tactic.
Required fields are demonstrated in the source repository's `examples/decisions/resolved.json`. `evidence` contains nonempty references, not copied logs. `status` is investigating, blocked, escalated or resolved. Resolving requires current-baseline passed verification, parent understanding and no unresolved items. Environment/authority issues cannot target Astra; repeated failures require reassessment; repeated expert consultations require a follow-up basis. A recovery route is required at resolution. Unknown observed model identity remains acceptable unless an explicit identity gate applies.
Required fields are demonstrated in the source repository's `examples/decisions/resolved.json`. `evidence` contains nonempty references, not copied logs. `status` is investigating, blocked, escalated or resolved. Resolving requires current-baseline passed verification, parent understanding and no unresolved items. Environment/authority issues cannot target Astra; repeated failures require reassessment; repeated expert consultations require a follow-up basis. The scalar `followup_basis` describes the current continuation; retain prior records for prior rounds. The validator cannot establish progress in each round from one aggregate record. A recovery route is required at resolution. Unknown observed model identity remains acceptable unless an explicit identity gate applies.
@@ -0,0 +1,41 @@
# Proactive fleet scheduling
Use for authorized complex parallel work. This is a decision procedure for the parent, not a background scheduler, permission grant or capacity override. Preserve the actual host schema, selected parent and final gates.
## Discover and dispatch early
After minimum goal/baseline discovery, identify independent questions and useful implementation slices. Do not finish reading every subsystem before dispatching ready investigations. Once a slice's input contract is stable, it may enter implementation while unrelated investigation continues. Shared decisions stay with the parent; preliminary independent drafts are not accepted contracts.
Maintain a compact queue: task ID, dependency and contract baseline, file AND semantic owner, shared resources, acceptance, requested route, child ID, observed runtime state and integration state. Mark pending, ready, running, returned, accepted, blocked or superseded as task states; keep actual host state separate. Ready means the prerequisites, ownership and permitted actions are established and the output has independent value.
At intake, dependency resolution, a child return or a meaningful baseline change:
1. Invalidate affected assumptions and identify ready slices. Prioritize work that unlocks dependencies, avoids costly wrong decisions or yields directly usable implementation.
2. Check live capacity and integration load. Dispatch multiple independent ready tasks when useful parent work remains; avoid sequential startup followed by unnecessary waits. Every descendant counts against the actual host limit. Children may not recursively expand the fleet without a parent-assigned bounded delegation scope and capacity allocation.
3. Continue the parent's independent critical work. Use incremental messages for running children and reuse suitable idle ones. Do not perform the same assigned investigation locally merely to stay busy.
4. Inspect returned evidence and partial writes; integrate serially. Refill available capacity with ready work without waiting for all siblings. A returned result is not necessarily accepted or a released slot; follow actual lifecycle controls.
If the next action is immediately dependent on an unassigned task, keep it local. A previously independent task may become a dependency as work progresses; bounded waiting is then legitimate. No independent useful parent work means do not create a ceremonial side task to justify another spawn.
## Choose a formation for the task
- Independent investigation: several focused investigators may cover distinct sources, consumers or falsifiable hypotheses. They should not all produce the same broad summary.
- Settled implementation: multiple Terra/Luna workers own complete independently checkable slices. Shared registries, schema and generated outputs require one owner even when paths differ.
- Difficult consequential decision: an Astra specialist analyzes the bounded uncertainty while the parent and other workers advance unaffected work. Gate dependent implementation on the parent's accepted contract. Read [expert intervention](escalation.md) for triggers and progress-based continuation.
- Independent review: assign a stable commit/artifact and the relevant requirements, without supplying the parent's preferred verdict. Review can overlap unrelated work; stale observations must be rechecked before acceptance.
There is no mandatory model ratio or permanent reviewer slot. Use several workers of the same adequate model when the tasks warrant it. If Astra is already the selected parent, do not add another Astra just to fulfill a role label. Reuse the parent capability unless a genuinely independent bounded perspective adds value.
## Backpressure and recovery
Choose a working width up to actual available capacity based on ready independent work, parent integration capacity and shared-resource limits. Do not reserve slots mechanically or fill them with artificial work. Hardware, service rate limits and tool sessions can constrain useful width below the agent limit.
When returned patches are accumulating faster than the parent can inspect, or a returned contract blocks other slices, pause new dependent implementation and prioritize integration. When collisions, stale baselines, repeated rework or resource contention appear, stop affected dispatch and reduce width; interrupt obsolete/conflicting writers and inspect partial changes. Independent read-only work may continue if useful and resource-safe. Resume expansion after the cause is resolved and ready work remains.
Never run acceptance tests against a moving artifact and report them as final. Pin a commit/snapshot or wait until relevant writers stop; a worktree alone does not isolate external databases, generated directories or shared sessions. Before merge/delivery, resolve material counterexamples, verify the integrated baseline and audit the agent tree for remaining writers.
## Evaluate effectiveness
Optimize accepted quality and end-to-end time subject to the user's cost constraints. Include parent work, all children, context duplication, retries, expert rounds and integration in cost measurements. Expensive early reasoning can be worthwhile when it prevents downstream rework; cheap high-volume workers can be wasteful when contracts are unsettled. Both are hypotheses until measured.
For a calibration exercise compare the same tasks and baselines at serial and increasing supported widths, varying granularity separately. Include decomposable and contract-heavy tasks, failures and repeated runs. Record quality, elapsed time, total exposed usage/cost, rework, and integration time; unknown counters or runtime model identity stay unknown. Do not claim speed/cost improvements from static scenarios or a successful single rollout.
@@ -16,6 +16,8 @@ Read before multi-round coordination, cancellation, or capacity diagnosis. Names
A minimal ledger records child ID, P/C classification, file/semantic ownership, dependency, requested route, last observed state and acceptance. Do not invent timestamps, identities or close events.
For rolling execution, use [fleet scheduling](fleet.md): refill ready work after useful handbacks and dependency changes, rather than waiting for every sibling. Check integration backlog before starting more implementation. Task readiness, output acceptance and actual slot availability remain separate facts.
## Cancellation and handback
Interrupting does not undo files or guarantee descendant cancellation. First preserve useful returned evidence, stop affected writers, inspect partial changes, then assign the remaining work to one owner. Never reset shared files wholesale.
@@ -8,6 +8,9 @@ Read only when the entrypoint leaves a routing choice unresolved.
| Ordinary cross-file bug | Terra; file count alone does not justify escalation. |
| Difficult accounting discrepancy | Sol reasoning; parent owns the semantic decision. |
| Coupled architecture across runtimes and tools | Consider Astra directly; keep shared decisions serial. |
| Difficult choice would constrain many downstream slices | Consider an early bounded Astra analysis; gate dependent work on the parent's accepted contract. |
| High-impact stable change needs independent counterexamples | Consider Astra review when depth adds value; give original requirements and a fixed artifact, not a preferred verdict. |
| Expert makes testable progress over several rounds | Continue the scoped topic with incremental evidence; do not impose a one-call quota or repeat unchanged questions. |
| Missing credential or inaccessible source | Resolve the evidence/environment gap; model escalation cannot supply access. |
| Astra advertised only for task creation | Child support remains unproven; do not create another user task as a workaround. |
| Named role fixes model/effort | Accept its supported binding or choose an overridable role; never attach prohibited fork overrides. |
@@ -68,6 +68,14 @@ def validate(record, previous=None):
if (previous.get("failure_class") == record.get("failure_class")
and type(old) is int and type(count) is int and count < old):
errors.append("failure counter decreased within the same class")
previous_escalation = previous.get("escalation")
old_consultations = (previous_escalation.get("expert_consultations")
if isinstance(previous_escalation, dict) else None)
if type(old_consultations) is not int or old_consultations < 0:
errors.append("previous expert_consultations must be a nonnegative integer")
elif (type(consultations) is int and consultations >= 0
and consultations < old_consultations):
errors.append("expert consultation counter decreased across continuation")
return errors
+51 -4
View File
@@ -44,12 +44,59 @@ class DecisionTests(unittest.TestCase):
self.record["escalation"]["target"] = "astra"
self.assertTrue(module.validate(self.record))
def test_extra_consultation_needs_basis(self):
self.record["escalation"]["expert_consultations"] = 2
self.assertTrue(module.validate(self.record))
self.record["escalation"]["followup_basis"] = "new discriminator result"
def test_zero_failure_proactive_astra_consultation(self):
self.record["failure_class"] = "reasoning"
self.record["status"] = "investigating"
self.record["same_class_failures"] = 0
self.record["reassessment"] = ""
self.record["escalation"] = {
"target": "astra",
"reason": "Proactive review of a high-impact architectural choice.",
"expert_consultations": 1,
"followup_basis": "",
}
self.record["unresolved"] = ["expert consultation pending"]
self.assertEqual(module.validate(self.record), [])
def test_more_than_three_progressive_consultations(self):
self.record["escalation"]["expert_consultations"] = 4
self.record["escalation"]["followup_basis"] = "new discriminator results across rounds"
self.assertEqual(module.validate(self.record), [])
def test_multiple_consultations_need_basis(self):
self.record["escalation"]["expert_consultations"] = 2
self.assertIn(
"additional consultation requires new evidence or a specific omission",
module.validate(self.record),
)
def test_consultation_counter_cannot_decrease_across_class_change(self):
previous = copy.deepcopy(self.record)
previous["escalation"]["expert_consultations"] = 4
previous["escalation"]["followup_basis"] = "prior progressive rounds"
self.record["failure_class"] = "reasoning"
self.record["escalation"]["expert_consultations"] = 3
self.record["escalation"]["followup_basis"] = "current evidence"
self.assertIn(
"expert consultation counter decreased across continuation",
module.validate(self.record, previous),
)
def test_consultation_counter_can_increase_with_basis(self):
previous = copy.deepcopy(self.record)
previous["escalation"]["expert_consultations"] = 1
self.record["escalation"]["expert_consultations"] = 4
self.record["escalation"]["followup_basis"] = "new discriminator results"
self.assertEqual(module.validate(self.record, previous), [])
def test_malformed_previous_consultation_counter_is_rejected(self):
previous = copy.deepcopy(self.record)
previous["escalation"]["expert_consultations"] = True
self.assertIn(
"previous expert_consultations must be a nonnegative integer",
module.validate(self.record, previous),
)
def test_recovery_required(self):
self.record["recovery_route"] = ""
self.assertTrue(module.validate(self.record))