Integrate 2.1 workflow and prepare release 2.2.0

This commit is contained in:
Codex Agent
2026-09-23 10:44:41 +08:00
17 changed files with 462 additions and 3 deletions
+36
View File
@@ -52,3 +52,39 @@ These are static review cases, not a record that a model executed them. Use them
| Host permits parallel children and both file and semantic write sets are disjoint | Children may edit in parallel; a parent's serial tool edits do not impose a global write lock |
| Several independent ready slices, parent only has later synthesis, host permits this dispatch | Run useful children concurrently and integrate on return; no universal requirement to invent simultaneous parent work |
| Host explicitly requires useful concurrent parent work | Honor that stricter requirement even when the fleet policy otherwise permits parent waiting |
## 2.0 escalation scenarios (manual forward evaluation)
These expectations are not recorded live-agent outcomes.
| Scenario | Expected behavior |
| --- | --- |
| Luna lacks the required input file | Diagnose input gap; do not consult Astra |
| Sol fails the same acceptance twice, then changes worker | Pause/reassess; preserve issue ID and count |
| Astra rejects a reproducible result without a counterexample | Check baseline and coverage; evidence decides, not model rank |
| A worker exposes a serious invariant violation on its first attempt | Suspend affected acceptance immediately |
| Requirements change after expert approval | Invalidate affected conclusion and verify the new baseline |
| Sol cannot explain the decisive expert reasoning | Do not accept or perform the dependent external action |
| Root cause is resolved and implementation is settled | Return remaining work to an adequate Sol/Luna route |
| Second expert consultation repeats the same question | Require new evidence or a specific prior omission |
For cost comparison use the same task baselines and acceptance checks across Sol alone, Astra planning plus Luna execution, and Sol-led three-model routing. Include sequential, decomposable and contract-heavy tasks; vary granularity and concurrency separately. Count failed attempts and parent integration in total tokens, actual cost and elapsed time. Repeat runs and report quality/failure distributions before claiming a winning route. Model requests without host identity remain unknown; synthetic packets do not substitute for measurements.
## 2.1 proactive fleet scenarios
These are expected behaviors; they do not assert actual model execution or measured performance. The previous one-question repetition case still requires progress, but there is no universal consultation-count quota.
| Scenario | Expected behavior |
| --- | --- |
| Consequential difficult interface choice, no failures yet, independent parent work available | Consider bounded upfront Astra analysis; parent owns final contract |
| Expert round four has a new discriminator result | Continue scoped topic with recorded basis; do not reject solely due to round count |
| Expert repeats the same proposal without progress | Pause topic and obtain missing discriminator or decision |
| User explicitly delegates read-only independent review | Allow bounded delegation, preserve read-only scope |
| User requests only advice about fleet design | Inspect and advise without launching a fleet |
| One independent slice is ready while another investigation continues | Dispatch ready slice without a whole-wave barrier |
| A child returns and independent ready work remains | Check real capacity and integration load, then reuse/refill appropriately |
| Returned patches exceed current integration capacity | Prioritize integration and pause new dependent implementation |
| Different files share the same registry contract | Assign one semantic owner and gate dependent writes |
| Parent is Astra and routine task has no independent expert need | Use adequate workers; do not duplicate Astra for a role quota |
| Child proposes grandchildren beyond its assigned capacity/scope | Do not expand the tree without bounded parent allocation and real capacity |
| Review result predates a relevant contract change | Recheck affected claims on the integrated baseline before acceptance |
+27
View File
@@ -0,0 +1,27 @@
# 2.0.0 — Sol 主导,按证据升级
发布日期:2026-09-15。
## 变化
- 通用会话建议由 Astra 改为 Sol / medium,复杂推理按需 high;保留用户已选模型。
- 新增疑难检查点、失败分类、跨 agent 的问题计数、两次同类失败暂停以及严重反例立即处理。
- 区分条件修复、执行能力升级和 Astra 有界咨询;默认每问题一次专家咨询,追加需要新证据或具体遗漏。
- 裁决依赖同基线可复现证据;主线程必须理解并验证关键结论,解决后恢复合适的低成本执行路径。
- 新增可选决策记录校验器及行为测试;兼容旧证据 schema 1/2,安装仍默认预览并按文件备份。
## 升级
检出 tag v2.0.0,运行安装预览,再按需要运行:
```sh
python3 scripts/install.py --include-prompt --apply
python3 scripts/install.py --include-prompt --check
python3 skills/codex-subagent-router/scripts/validate_decision_record.py examples/decisions/resolved.json
```
安装不会修改 config.toml。若希望更改未来会话默认模型,在支持对应字段的客户端中单独设置 model = "gpt-5.6-sol" 与 model_reasoning_effort = "medium"。重新加载或新开会话;当前会话运行身份不会因此得到证明。
## 验证边界
单元测试验证校验器拒绝矛盾状态、安装安全与资源完整性;行为场景供真实授权任务的前向评估使用。没有运行 Astra/Sol/Terra/Luna 成本或速度对照实验,不把合成记录当作真实运行证据。并行收益依赖任务可分解性;更细拆分可能增加总 token 和返工。
+19
View File
@@ -0,0 +1,19 @@
# 2.1.0 — Astra 前置介入与滚动并行
## 变化
- Astra 从失败后的专家扩展为高影响困难方案、跨系统综合和独立深度反例审查的参与者。零失败的主动咨询合法,保留主线程裁决与验收。
- 专家任务按问题、产物、退出条件设界。可依据新证据、可测试的细化或明确遗漏持续多轮;不按统一次数配额截断,也不允许无进展循环。
- 新增就绪队列、提前派发、按依赖滚动补位、动态队形、整合积压控制与恢复规则。宿主容量、共享文件/语义所有权及外部串行门不变。
- 明确授权的只读并行审查属于可委派工作;普通审计建议仍不自动启动 agent。
- 决策校验器在连续记录中拒绝累计专家咨询次数倒退,包括失败分类变化时;补充主动咨询和多轮进展测试。
## 兼容与安装
证据 schema 1/2、决策 schema 1 不变。合法旧记录保持兼容,`--previous` 现在会拒绝咨询次数倒退或前条记录次数格式错误。`followup_basis` 仍为字符串,应保留连续记录来追踪每轮依据;结构校验不证明其真实性。
从已审阅的 2.1 源码运行 `python3 scripts/install.py` 预览,再用 `--apply` 同步技能。全局提示词需显式添加 `--include-prompt`。安装不会修改模型配置或扩大宿主槽位;备份与回退流程见根目录 README。版本文件不代表 Git 标签或正式 Release 已创建。
## 验证边界
使用确定性测试验证校验器和安装行为,并对策略进行独立子智能体审查及模型场景评估。请求模型与实际宿主观测身份分别记录;静态评估不等于真实舰队调度实验,不证明质量、成本或速度提升。官方容量配置与当前工具的计数方式应分别核实。
+22
View File
@@ -0,0 +1,22 @@
# 2.2.0 — GPT-6 三档路由与积极 Luna 舰队
## 变化
- 推荐路线统一为 GPT-6 Luna / Sol / Astra。Sol 接替 Terra 的常规开发、调试与调查;Terra 退出推荐与自动回退,用户明确指定和既有配置仍受尊重。旧宿主按实际 schema 显式使用足够胜任的旧版 Sol/Luna 或本地执行。
- 可选全局提示词提供普通开发、检索、诊断、审查与验证的持续并行授权。仅审计并行配置、用户禁止及更高优先级宿主限制优先;审查仍只读,技能自动发现不等于授权。
- Luna 优先承担有界收集、比对、验证、局部审查及已定方案修改。尽早派发、按真实容量滚动补位,整合积压或资源争用时缩小宽度;不固定只开一两个,也不以模型配额凑数量。
- 文件与语义写集均不相交且宿主允许时可并行修改。宿主允许时,多个独立子任务可运行,主线程等待后整合;不以虚构主线程工作满足并行条件。
- 按需参考说明 GPT-6 异步工具、执行中纠偏、动态 effort 与上下文恢复的边界;不把 API 功能误当成 Desktop/CLI 已开放的控制。价格快照标注日期、计费条件与实际任务成本差异。
- 保留 2.1 的 Sol 常态指挥、Astra 前置介入、进展驱动的多轮专家协作、问题计数和基于证据的裁决。专家意见不取代主线程理解、验证与最终验收。
## 兼容与安装
证据 schema 1/2 与决策 schema 1 不变,校验器与原有记录兼容。新增参考纳入资源完整性检查。历史发行说明保留原版本内容。
从 v2.2.0 源码运行 `python3 scripts/install.py` 预览,使用 `--apply` 安装技能;全局提示词另加 `--include-prompt`。该选项替换而非智能合并现有提示词,先审阅差异。安装不切换活动模型、不扩大宿主槽位、不改 `config.toml`,沿用备份与回退机制。
## 验证边界
验证包括仓库单元测试、资源及版本检查、skill-creator 校验、静态场景与独立差异审查。Windows 上的符号链接测试可能在建立夹具时因 WinError 1314 受阻,应与原始基线区分并如实报告,不删除或跳过测试。
本版本未声称实测 GPT-6 Luna/Sol 子智能体身份、舰队吞吐、质量或费用收益。模型与参数必须由宿主真实 schema 确认;请求参数和自述不是运行身份。API 价格不代表 Codex 套餐消耗。