Release 2.0.0 with Sol-led escalation and evidence adjudication
This commit is contained in:
@@ -20,3 +20,20 @@ These are static review cases, not a record that a model executed them. Use them
|
||||
| Local config says 8 threads, host gives 4 total slots | Follow actual host count, including parent if specified |
|
||||
| Project requires Sol parent, current parent is Astra | Identify rule; no silent substitution based on model ranking |
|
||||
| User says never delegate | Stay local even if a skill is selected implicitly |
|
||||
|
||||
## 2.0 escalation scenarios (manual forward evaluation)
|
||||
|
||||
These expectations are not recorded live-agent outcomes.
|
||||
|
||||
| Scenario | Expected behavior |
|
||||
| --- | --- |
|
||||
| Luna lacks the required input file | Diagnose input gap; do not consult Astra |
|
||||
| Sol fails the same acceptance twice, then changes worker | Pause/reassess; preserve issue ID and count |
|
||||
| Astra rejects a reproducible result without a counterexample | Check baseline and coverage; evidence decides, not model rank |
|
||||
| A worker exposes a serious invariant violation on its first attempt | Suspend affected acceptance immediately |
|
||||
| Requirements change after expert approval | Invalidate affected conclusion and verify the new baseline |
|
||||
| Sol cannot explain the decisive expert reasoning | Do not accept or perform the dependent external action |
|
||||
| Root cause is resolved and implementation is settled | Return remaining work to an adequate Terra/Luna route |
|
||||
| Second expert consultation repeats the same question | Require new evidence or a specific prior omission |
|
||||
|
||||
For cost comparison use the same task baselines and acceptance checks across Sol alone, Astra planning plus Luna execution, and Sol-led four-model routing. Include sequential, decomposable and contract-heavy tasks; vary granularity and concurrency separately. Count failed attempts and parent integration in total tokens, actual cost and elapsed time. Repeat runs and report quality/failure distributions before claiming a winning route. Model requests without host identity remain unknown; synthetic packets do not substitute for measurements.
|
||||
|
||||
@@ -0,0 +1,27 @@
|
||||
# 2.0.0 — Sol 主导,按证据升级
|
||||
|
||||
发布日期:2026-09-15。
|
||||
|
||||
## 变化
|
||||
|
||||
- 通用会话建议由 Astra 改为 Sol / medium,复杂推理按需 high;保留用户已选模型。
|
||||
- 新增疑难检查点、失败分类、跨 agent 的问题计数、两次同类失败暂停以及严重反例立即处理。
|
||||
- 区分条件修复、执行能力升级和 Astra 有界咨询;默认每问题一次专家咨询,追加需要新证据或具体遗漏。
|
||||
- 裁决依赖同基线可复现证据;主线程必须理解并验证关键结论,解决后恢复合适的低成本执行路径。
|
||||
- 新增可选决策记录校验器及行为测试;兼容旧证据 schema 1/2,安装仍默认预览并按文件备份。
|
||||
|
||||
## 升级
|
||||
|
||||
检出 tag v2.0.0,运行安装预览,再按需要运行:
|
||||
|
||||
```sh
|
||||
python3 scripts/install.py --include-prompt --apply
|
||||
python3 scripts/install.py --include-prompt --check
|
||||
python3 skills/codex-subagent-router/scripts/validate_decision_record.py examples/decisions/resolved.json
|
||||
```
|
||||
|
||||
安装不会修改 config.toml。若希望更改未来会话默认模型,在支持对应字段的客户端中单独设置 model = "gpt-5.6-sol" 与 model_reasoning_effort = "medium"。重新加载或新开会话;当前会话运行身份不会因此得到证明。
|
||||
|
||||
## 验证边界
|
||||
|
||||
单元测试验证校验器拒绝矛盾状态、安装安全与资源完整性;行为场景供真实授权任务的前向评估使用。没有运行 Astra/Sol/Terra/Luna 成本或速度对照实验,不把合成记录当作真实运行证据。并行收益依赖任务可分解性;更细拆分可能增加总 token 和返工。
|
||||
Reference in New Issue
Block a user