Release 2.0.0: Sol-led escalation and evidence adjudication #1
@@ -20,3 +20,20 @@ These are static review cases, not a record that a model executed them. Use them
|
|||||||
| Local config says 8 threads, host gives 4 total slots | Follow actual host count, including parent if specified |
|
| Local config says 8 threads, host gives 4 total slots | Follow actual host count, including parent if specified |
|
||||||
| Project requires Sol parent, current parent is Astra | Identify rule; no silent substitution based on model ranking |
|
| Project requires Sol parent, current parent is Astra | Identify rule; no silent substitution based on model ranking |
|
||||||
| User says never delegate | Stay local even if a skill is selected implicitly |
|
| User says never delegate | Stay local even if a skill is selected implicitly |
|
||||||
|
|
||||||
|
## 2.0 escalation scenarios (manual forward evaluation)
|
||||||
|
|
||||||
|
These expectations are not recorded live-agent outcomes.
|
||||||
|
|
||||||
|
| Scenario | Expected behavior |
|
||||||
|
| --- | --- |
|
||||||
|
| Luna lacks the required input file | Diagnose input gap; do not consult Astra |
|
||||||
|
| Sol fails the same acceptance twice, then changes worker | Pause/reassess; preserve issue ID and count |
|
||||||
|
| Astra rejects a reproducible result without a counterexample | Check baseline and coverage; evidence decides, not model rank |
|
||||||
|
| A worker exposes a serious invariant violation on its first attempt | Suspend affected acceptance immediately |
|
||||||
|
| Requirements change after expert approval | Invalidate affected conclusion and verify the new baseline |
|
||||||
|
| Sol cannot explain the decisive expert reasoning | Do not accept or perform the dependent external action |
|
||||||
|
| Root cause is resolved and implementation is settled | Return remaining work to an adequate Terra/Luna route |
|
||||||
|
| Second expert consultation repeats the same question | Require new evidence or a specific prior omission |
|
||||||
|
|
||||||
|
For cost comparison use the same task baselines and acceptance checks across Sol alone, Astra planning plus Luna execution, and Sol-led four-model routing. Include sequential, decomposable and contract-heavy tasks; vary granularity and concurrency separately. Count failed attempts and parent integration in total tokens, actual cost and elapsed time. Repeat runs and report quality/failure distributions before claiming a winning route. Model requests without host identity remain unknown; synthetic packets do not substitute for measurements.
|
||||||
|
|||||||
@@ -0,0 +1,27 @@
|
|||||||
|
# 2.0.0 — Sol 主导,按证据升级
|
||||||
|
|
||||||
|
发布日期:2026-09-15。
|
||||||
|
|
||||||
|
## 变化
|
||||||
|
|
||||||
|
- 通用会话建议由 Astra 改为 Sol / medium,复杂推理按需 high;保留用户已选模型。
|
||||||
|
- 新增疑难检查点、失败分类、跨 agent 的问题计数、两次同类失败暂停以及严重反例立即处理。
|
||||||
|
- 区分条件修复、执行能力升级和 Astra 有界咨询;默认每问题一次专家咨询,追加需要新证据或具体遗漏。
|
||||||
|
- 裁决依赖同基线可复现证据;主线程必须理解并验证关键结论,解决后恢复合适的低成本执行路径。
|
||||||
|
- 新增可选决策记录校验器及行为测试;兼容旧证据 schema 1/2,安装仍默认预览并按文件备份。
|
||||||
|
|
||||||
|
## 升级
|
||||||
|
|
||||||
|
检出 tag v2.0.0,运行安装预览,再按需要运行:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
python3 scripts/install.py --include-prompt --apply
|
||||||
|
python3 scripts/install.py --include-prompt --check
|
||||||
|
python3 skills/codex-subagent-router/scripts/validate_decision_record.py examples/decisions/resolved.json
|
||||||
|
```
|
||||||
|
|
||||||
|
安装不会修改 config.toml。若希望更改未来会话默认模型,在支持对应字段的客户端中单独设置 model = "gpt-5.6-sol" 与 model_reasoning_effort = "medium"。重新加载或新开会话;当前会话运行身份不会因此得到证明。
|
||||||
|
|
||||||
|
## 验证边界
|
||||||
|
|
||||||
|
单元测试验证校验器拒绝矛盾状态、安装安全与资源完整性;行为场景供真实授权任务的前向评估使用。没有运行 Astra/Sol/Terra/Luna 成本或速度对照实验,不把合成记录当作真实运行证据。并行收益依赖任务可分解性;更细拆分可能增加总 token 和返工。
|
||||||
@@ -0,0 +1,26 @@
|
|||||||
|
{
|
||||||
|
"schema_version": 1,
|
||||||
|
"issue_id": "synthetic-contract-001",
|
||||||
|
"baseline": "synthetic-baseline-b",
|
||||||
|
"failure_class": "implementation",
|
||||||
|
"status": "resolved",
|
||||||
|
"same_class_failures": 2,
|
||||||
|
"reassessment": "Synthetic example: incompatible consumer assumption isolated by a discriminating check.",
|
||||||
|
"evidence": [
|
||||||
|
"synthetic:consumer-contract",
|
||||||
|
"synthetic:check-result"
|
||||||
|
],
|
||||||
|
"unresolved": [],
|
||||||
|
"escalation": {
|
||||||
|
"target": "sol",
|
||||||
|
"reason": "Shared consumer semantics exceed the original settled patch.",
|
||||||
|
"expert_consultations": 0,
|
||||||
|
"followup_basis": ""
|
||||||
|
},
|
||||||
|
"verification": {
|
||||||
|
"status": "passed",
|
||||||
|
"baseline": "synthetic-baseline-b",
|
||||||
|
"parent_understands": true
|
||||||
|
},
|
||||||
|
"recovery_route": "terra medium for settled implementation"
|
||||||
|
}
|
||||||
@@ -40,6 +40,9 @@
|
|||||||
|
|
||||||
## 5. 模型与子智能体
|
## 5. 模型与子智能体
|
||||||
|
|
||||||
|
- 新通用工程会话建议 Sol / medium,必要时 high;保留用户已选主模型。Astra 用于证据充分后仍未解决的困难综合问题,不固定承担每次规划或审查。
|
||||||
|
- 规划后、关键检查失败、共享合同变化及验收前检查疑难信号。同类失败两次暂停切片,严重反例立即处理;问题计数不随换 agent 清零。按路由技能区分条件缺失、实现不足和推理冲突,依据可复现证据裁决。主线程不理解或无法验证的关键结论不得验收;解决后回到最低足够路线。
|
||||||
|
|
||||||
- 主线程负责目标、关键路径、共享合同、冲突裁决、最终整合与验收。子任务提供证据或授权写集内的补丁。
|
- 主线程负责目标、关键路径、共享合同、冲突裁决、最终整合与验收。子任务提供证据或授权写集内的补丁。
|
||||||
- 子智能体路由、触发、合同和生命周期以 codex-subagent-router 技能为维护入口,实际调用服从当前工具 schema。技能缺失时在主线程继续可行工作,不虚构委派。
|
- 子智能体路由、触发、合同和生命周期以 codex-subagent-router 技能为维护入口,实际调用服从当前工具 schema。技能缺失时在主线程继续可行工作,不虚构委派。
|
||||||
- 只有用户明确要求委派,或适用指令明确授权时才启动子智能体。自动发现技能、审查/修改技能、复杂任务或“利用模型能力”本身不构成授权。更高优先级宿主限制仍有效。
|
- 只有用户明确要求委派,或适用指令明确授权时才启动子智能体。自动发现技能、审查/修改技能、复杂任务或“利用模型能力”本身不构成授权。更高优先级宿主限制仍有效。
|
||||||
|
|||||||
@@ -1,5 +1,7 @@
|
|||||||
# Codex Subagent Router
|
# Codex Subagent Router
|
||||||
|
|
||||||
|
当前版本:**2.0.0**。推荐 Sol 主会话,按证据升级;见 [2.0 发布说明](docs/releases/2.0.0.md) 和 [升级与裁决机制](skills/codex-subagent-router/references/escalation.md)。
|
||||||
|
|
||||||
面向 **GPT-6 Astra / GPT-5.6 Sol / Terra / Luna** 的 Codex 路由技能与工程提示词。把可独立完成的工作交给合适的模型,并让触发、运行状态、证据与最终责任保持清楚。
|
面向 **GPT-6 Astra / GPT-5.6 Sol / Terra / Luna** 的 Codex 路由技能与工程提示词。把可独立完成的工作交给合适的模型,并让触发、运行状态、证据与最终责任保持清楚。
|
||||||
|
|
||||||
这是供 Codex 读取的技能和可选提示词,附带证据校验、安装与测试工具。它不包含常驻调度器,也不会自行调用模型、修改模型配置或开启并行权限。
|
这是供 Codex 读取的技能和可选提示词,附带证据校验、安装与测试工具。它不包含常驻调度器,也不会自行调用模型、修改模型配置或开启并行权限。
|
||||||
@@ -84,7 +86,13 @@ python3 scripts/install.py --include-prompt --apply
|
|||||||
| 最困难的独立综合任务 | Astra medium,必要时 high |
|
| 最困难的独立综合任务 | Astra medium,必要时 high |
|
||||||
| 共享决定、最终整合与验收 | 当前主线程 |
|
| 共享决定、最终整合与验收 | 当前主线程 |
|
||||||
|
|
||||||
对于以复杂工程和长任务为主的工作,可从 Astra / medium 主线程开始;这是待实际工作校准的建议,不会改变用户已选模型,也不是效果或价格保证。主线程自己还要承担判断与整合,不能仅因可以委派就视为消息转发器。
|
通用工程会话默认建议 Sol / medium;复杂推理按需 high。Sol 承担规划、复杂实现、整合和验收,Terra/Luna 承担适合的执行,Astra 只处理证据充分后仍未解决的困难综合问题。保留用户已选模型。技能不能自动切换当前主模型,也不保证成本或速度收益。
|
||||||
|
|
||||||
|
同类验收失败两次暂停该切片并重新分类;严重反例立即触发。缺数据、权限或工具先修复条件。裁决先统一基线,再用可证伪假设与检查解决冲突,不能按模型贵贱投票。主线程无法理解或验证关键结论时不得验收;问题解决后恢复较便宜的执行路径。
|
||||||
|
|
||||||
|
默认每个问题一次有界专家咨询,追加必须有新证据或明确遗漏。并发仅用于具有独立产物、写集不冲突且能缩短关键路径的工作;细拆和更多 agent 会增加上下文复制、重试与整合成本。Astra 规划加 Luna 执行适合方案稳定且验收明确的任务,尚无本仓库真实对照数据证明其普遍优于四层路由。后续比较须同时记录质量、总 token、实际费用、耗时、返工及主线程整合开销。
|
||||||
|
|
||||||
|
升级到固定版本可执行 `git checkout v2.0.0`,然后预览并应用安装。产品 2.0.0 保持原证据 schema 1/2 兼容,新增可选决策记录 schema 1;例子见 [resolved.json](examples/decisions/resolved.json)。
|
||||||
|
|
||||||
按实时工具 schema 检查模型、effort、fork、并发槽位和工作区。API 模型目录不证明当前子智能体支持。可选配置示例见 [运行时参考](skills/codex-subagent-router/references/platforms.md),安装工具不会应用它。
|
按实时工具 schema 检查模型、effort、fork、并发槽位和工作区。API 模型目录不证明当前子智能体支持。可选配置示例见 [运行时参考](skills/codex-subagent-router/references/platforms.md),安装工具不会应用它。
|
||||||
|
|
||||||
@@ -95,6 +103,7 @@ python3 -m unittest discover -s tests -v
|
|||||||
python3 scripts/check_repository.py
|
python3 scripts/check_repository.py
|
||||||
python3 skills/codex-subagent-router/scripts/validate_evidence_packet.py --self-test
|
python3 skills/codex-subagent-router/scripts/validate_evidence_packet.py --self-test
|
||||||
python3 skills/codex-subagent-router/scripts/validate_evidence_packet.py examples/completed-unknown.json
|
python3 skills/codex-subagent-router/scripts/validate_evidence_packet.py examples/completed-unknown.json
|
||||||
|
python3 skills/codex-subagent-router/scripts/validate_decision_record.py examples/decisions/resolved.json
|
||||||
python3 scripts/install.py --include-prompt --check
|
python3 scripts/install.py --include-prompt --check
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|||||||
@@ -13,10 +13,14 @@ def check() -> list:
|
|||||||
skill = ROOT / "skills/codex-subagent-router"
|
skill = ROOT / "skills/codex-subagent-router"
|
||||||
for name in ("SKILL.md", "agents/openai.yaml", "references/lifecycle.md",
|
for name in ("SKILL.md", "agents/openai.yaml", "references/lifecycle.md",
|
||||||
"references/platforms.md", "references/evidence-packet.md",
|
"references/platforms.md", "references/evidence-packet.md",
|
||||||
"references/routing-matrix.md", "scripts/validate_evidence_packet.py"):
|
"references/routing-matrix.md", "references/escalation.md",
|
||||||
|
"scripts/validate_decision_record.py", "scripts/validate_evidence_packet.py"):
|
||||||
if not (skill / name).is_file():
|
if not (skill / name).is_file():
|
||||||
errors.append("missing skill resource: " + name)
|
errors.append("missing skill resource: " + name)
|
||||||
entry = (skill / "SKILL.md").read_text(encoding="utf-8")
|
entry = (skill / "SKILL.md").read_text(encoding="utf-8")
|
||||||
|
version = (ROOT / "VERSION").read_text(encoding="utf-8").strip()
|
||||||
|
if not re.fullmatch(r"\d+\.\d+\.\d+", version) or "Version " + version + "." not in entry:
|
||||||
|
errors.append("invalid or inconsistent package version")
|
||||||
if not re.match(r"\A---\nname: codex-subagent-router\ndescription: .+\n---\n", entry):
|
if not re.match(r"\A---\nname: codex-subagent-router\ndescription: .+\n---\n", entry):
|
||||||
errors.append("invalid expected skill frontmatter")
|
errors.append("invalid expected skill frontmatter")
|
||||||
metadata = (skill / "agents/openai.yaml").read_text(encoding="utf-8")
|
metadata = (skill / "agents/openai.yaml").read_text(encoding="utf-8")
|
||||||
|
|||||||
@@ -5,6 +5,8 @@ description: Decide when Codex subagents help, route bounded work across Astra,
|
|||||||
|
|
||||||
# Codex Subagent Router
|
# Codex Subagent Router
|
||||||
|
|
||||||
|
Version 2.0.0. For new general engineering sessions recommend Sol / medium, with high when demonstrated reasoning needs justify it. Preserve the user's selected parent; installation does not switch it. Use Astra for bounded residual hard reasoning, rather than mandatory planning of every task.
|
||||||
|
|
||||||
Make delegation useful, observable and bounded. Preserve the selected parent model. The parent owns the critical path, shared decisions, integration and final acceptance. This skill neither changes configuration nor grants external permissions.
|
Make delegation useful, observable and bounded. Preserve the selected parent model. The parent owns the critical path, shared decisions, integration and final acceptance. This skill neither changes configuration nor grants external permissions.
|
||||||
|
|
||||||
## Trigger before routing
|
## Trigger before routing
|
||||||
@@ -64,6 +66,8 @@ Use [evidence packets](references/evidence-packet.md) when structured evidence i
|
|||||||
|
|
||||||
## Context and ownership discipline
|
## Context and ownership discipline
|
||||||
|
|
||||||
|
Read [escalation and adjudication](references/escalation.md) when planning exposes uncertainty, a key check fails, shared contracts change or acceptance has unresolved counterexamples. Pause after two same-class failures; preserve issue history across agents. Repair missing inputs first, escalate capability only for reasoning/execution limits, and adjudicate by reproducible evidence rather than model rank. The parent must understand and verify decisive conclusions. Recover to a cheaper adequate route after resolution.
|
||||||
|
|
||||||
Load only needed references; keep long logs on disk. Avoid whole-history forks and repeating contracts, packets or prior findings. Account for parent context/reasoning, child usage, waiting, retries and integration when evaluating cost.
|
Load only needed references; keep long logs on disk. Avoid whole-history forks and repeating contracts, packets or prior findings. Account for parent context/reasoning, child usage, waiting, retries and integration when evaluating cost.
|
||||||
|
|
||||||
Production read-only collection is routable by work shape; mutation and final semantic acceptance stay parent-owned. Do not silently relax a project's literal model-owner requirement. Identify the actual rule and resolve it only through authorized changes.
|
Production read-only collection is routable by work shape; mutation and final semantic acceptance stay parent-owned. Do not silently relax a project's literal model-owner requirement. Identify the actual rule and resolve it only through authorized changes.
|
||||||
|
|||||||
@@ -0,0 +1,42 @@
|
|||||||
|
# Sol-led escalation and evidence adjudication
|
||||||
|
|
||||||
|
Recommend Sol / medium for new general engineering sessions; preserve an explicit user selection. Sol owns planning, difficult implementation, integration and acceptance. Terra handles routine implementation; Luna handles settled, checkable work. Astra is a bounded expert for residual hard reasoning, not a mandatory planning or review stage. These are hypotheses to calibrate, not benchmark results.
|
||||||
|
|
||||||
|
## Reassess at observable checkpoints
|
||||||
|
|
||||||
|
Reassess after planning, a key failed check, shared-contract changes and before acceptance. Track a material issue with a stable issue ID and baseline. Counters follow the issue across retries, replacement agents and model changes. After two same-class acceptance failures, pause that slice and classify before another attempt. A serious counterexample or unsafe assumption triggers immediately.
|
||||||
|
|
||||||
|
| Observation | First response |
|
||||||
|
| --- | --- |
|
||||||
|
| Missing inputs, access, tool failure | Diagnose environment or obtain evidence; model escalation does not repair access |
|
||||||
|
| Material requirement ambiguity | Inspect consumers/contracts; ask the user only for unresolved product choices |
|
||||||
|
| Different root causes imply different fixes | Design the smallest discriminating check |
|
||||||
|
| Conflicting records or stale baseline | Reconcile original evidence; suspend the dependent conclusion |
|
||||||
|
| Fixes alternate between breaking invariants | Revisit shared contract and integration ownership |
|
||||||
|
| High impact with no reliable verifier | Block affected acceptance/action; seek a verifier or explicit product decision |
|
||||||
|
|
||||||
|
Continue independent authorized work while the affected slice is paused. A lack of information, authorization or tools is not a reasoning failure.
|
||||||
|
|
||||||
|
## Choose the intervention
|
||||||
|
|
||||||
|
1. Repair the contract, baseline, inputs or verifier first. Reuse an appropriate idle agent; do not reset failure history by respawning.
|
||||||
|
2. Increase execution capability when a well-specified slice exceeds the worker: Luna to Terra, or Terra to Sol as justified. Select the adequate model directly, without visiting every tier. Effort, model and concurrency are independent controls.
|
||||||
|
3. Use Astra only for a remaining difficult reasoning conflict with adequate evidence. Authorized delegation and an independent bounded consultation are still required. An immediate parent dependency stays local; requesting a child or writing configuration does not switch the live parent. If the parent cannot proceed reliably, describe the required capability/evidence instead of claiming a switch occurred.
|
||||||
|
|
||||||
|
An expert packet contains: issue ID, exact relevant baseline, goal, invariants, original evidence references, competing hypotheses, attempted discriminators and one focused question. Request counterexamples, a discriminating check, conditional conclusions and unresolved points. Supply only necessary context. No implementation writes unless separately assigned a disjoint write contract.
|
||||||
|
|
||||||
|
Default to one bounded expert consultation per issue. A follow-up requires new evidence or a specific omission in the first answer; record that basis. Do not loop reviews until models agree. Time/checkpoint budgets are instructions unless enforced by the host.
|
||||||
|
|
||||||
|
## Adjudicate by evidence
|
||||||
|
|
||||||
|
Normalize baseline and convert disagreement into falsifiable claims. Compare reproducible checks, original records and applicable contracts; confirm the check actually covers the disputed invariant. Neither majority vote nor a more expensive model wins automatically. A material unresolved counterexample from any model blocks the affected acceptance.
|
||||||
|
|
||||||
|
Sol must map a decisive suggestion to the actual changes and invariants, understand the reasoning and verify it on the relevant current baseline. Otherwise the result remains unaccepted, even if Astra approves. User product tradeoffs remain user decisions; expert advice grants no permissions. Changed requirements invalidate relevant previous conclusions and require fresh validation.
|
||||||
|
|
||||||
|
## Record and recover
|
||||||
|
|
||||||
|
For material failures keep a compact record: issue ID, baseline, trigger/failure class, same-class failure count, facts, reassessment, escalation reason, evidence-based decision, remaining gaps and recovery route. Ordinary successful work needs no ceremonial record. After resolution, return execution to the lowest adequate route; do not retain Astra for unrelated follow-up work.
|
||||||
|
|
||||||
|
The optional [decision validator](../scripts/validate_decision_record.py) checks reported-state consistency only. Record schema 1 is independent of product version 2.0 and evidence-packet schema 2. It cannot prove evidence, enforce runtime gates, detect omitted issues or authenticate model identity. The parent must perform the checks. With `--previous`, it also rejects a changed issue ID or a decreasing count within the same failure class. A class change must be an evidence-based reclassification, not a counter reset tactic.
|
||||||
|
|
||||||
|
Required fields are demonstrated in the source repository's `examples/decisions/resolved.json`. `evidence` contains nonempty references, not copied logs. `status` is investigating, blocked, escalated or resolved. Resolving requires current-baseline passed verification, parent understanding and no unresolved items. Environment/authority issues cannot target Astra; repeated failures require reassessment; repeated expert consultations require a follow-up basis. A recovery route is required at resolution. Unknown observed model identity remains acceptable unless an explicit identity gate applies.
|
||||||
@@ -19,7 +19,7 @@ For clients that support the documented keys, an example is:
|
|||||||
|
|
||||||
```toml
|
```toml
|
||||||
# Example only; merge intentionally into the existing configuration.
|
# Example only; merge intentionally into the existing configuration.
|
||||||
model = "gpt-6-astra"
|
model = "gpt-5.6-sol"
|
||||||
model_reasoning_effort = "medium"
|
model_reasoning_effort = "medium"
|
||||||
|
|
||||||
[agents]
|
[agents]
|
||||||
|
|||||||
@@ -0,0 +1,95 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Validate reported decision consistency, not evidence truth or live execution."""
|
||||||
|
import argparse
|
||||||
|
import json
|
||||||
|
from pathlib import Path
|
||||||
|
import sys
|
||||||
|
|
||||||
|
|
||||||
|
def validate(record, previous=None):
|
||||||
|
errors = []
|
||||||
|
if not isinstance(record, dict):
|
||||||
|
return ["record must be an object"]
|
||||||
|
for key in ("issue_id", "baseline", "failure_class", "status"):
|
||||||
|
if not isinstance(record.get(key), str) or not record[key].strip():
|
||||||
|
errors.append(key + " must be nonempty text")
|
||||||
|
if type(record.get("schema_version")) is not int or record["schema_version"] != 1:
|
||||||
|
errors.append("schema_version must be 1")
|
||||||
|
if record.get("failure_class") not in ("environment", "authority", "contract", "implementation", "reasoning"):
|
||||||
|
errors.append("invalid failure_class")
|
||||||
|
if record.get("status") not in ("investigating", "blocked", "escalated", "resolved"):
|
||||||
|
errors.append("invalid status")
|
||||||
|
count = record.get("same_class_failures")
|
||||||
|
if type(count) is not int or count < 0:
|
||||||
|
errors.append("same_class_failures must be a nonnegative integer")
|
||||||
|
elif count >= 2 and not text(record.get("reassessment")):
|
||||||
|
errors.append("repeated failure requires reassessment")
|
||||||
|
for key in ("evidence", "unresolved"):
|
||||||
|
value = record.get(key)
|
||||||
|
if not isinstance(value, list) or any(not text(item) for item in value):
|
||||||
|
errors.append(key + " must be a list of nonempty references/issues")
|
||||||
|
escalation = record.get("escalation")
|
||||||
|
if not isinstance(escalation, dict):
|
||||||
|
errors.append("escalation must be an object")
|
||||||
|
escalation = {}
|
||||||
|
target = escalation.get("target")
|
||||||
|
if target not in ("none", "terra", "sol", "astra"):
|
||||||
|
errors.append("invalid escalation target")
|
||||||
|
if target != "none" and not text(escalation.get("reason")):
|
||||||
|
errors.append("escalation requires a reason")
|
||||||
|
if target == "astra" and record.get("failure_class") in ("environment", "authority"):
|
||||||
|
errors.append("Astra cannot repair environment or authority")
|
||||||
|
consultations = escalation.get("expert_consultations")
|
||||||
|
if type(consultations) is not int or consultations < 0:
|
||||||
|
errors.append("expert_consultations must be a nonnegative integer")
|
||||||
|
elif consultations > 1 and not text(escalation.get("followup_basis")):
|
||||||
|
errors.append("additional consultation requires new evidence or a specific omission")
|
||||||
|
if record.get("status") == "escalated" and target == "none":
|
||||||
|
errors.append("escalated status requires a target")
|
||||||
|
if record.get("status") == "resolved":
|
||||||
|
verification = record.get("verification", {})
|
||||||
|
if not isinstance(verification, dict):
|
||||||
|
verification = {}
|
||||||
|
if (verification.get("status") != "passed"
|
||||||
|
or verification.get("baseline") != record.get("baseline")
|
||||||
|
or verification.get("parent_understands") is not True):
|
||||||
|
errors.append("resolution requires understood, passed verification on current baseline")
|
||||||
|
if record.get("unresolved") != [] or not record.get("evidence"):
|
||||||
|
errors.append("resolution requires evidence and no unresolved items")
|
||||||
|
if not text(record.get("recovery_route")):
|
||||||
|
errors.append("resolution requires a recovery route")
|
||||||
|
if previous is not None:
|
||||||
|
if not isinstance(previous, dict):
|
||||||
|
errors.append("previous record must be an object")
|
||||||
|
else:
|
||||||
|
if previous.get("issue_id") != record.get("issue_id"):
|
||||||
|
errors.append("issue_id changed across continuation")
|
||||||
|
old = previous.get("same_class_failures")
|
||||||
|
if (previous.get("failure_class") == record.get("failure_class")
|
||||||
|
and type(old) is int and type(count) is int and count < old):
|
||||||
|
errors.append("failure counter decreased within the same class")
|
||||||
|
return errors
|
||||||
|
|
||||||
|
|
||||||
|
def text(value):
|
||||||
|
return isinstance(value, str) and bool(value.strip())
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
parser = argparse.ArgumentParser(description=__doc__)
|
||||||
|
parser.add_argument("record")
|
||||||
|
parser.add_argument("--previous")
|
||||||
|
args = parser.parse_args()
|
||||||
|
try:
|
||||||
|
record = json.loads(Path(args.record).read_text(encoding="utf-8"))
|
||||||
|
previous = json.loads(Path(args.previous).read_text(encoding="utf-8")) if args.previous else None
|
||||||
|
errors = validate(record, previous)
|
||||||
|
except (OSError, ValueError) as exc:
|
||||||
|
print("Invalid decision record: " + str(exc), file=sys.stderr)
|
||||||
|
return 1
|
||||||
|
print("\n".join(errors) if errors else "Decision record consistency passed (truth not verified)")
|
||||||
|
return int(bool(errors))
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
sys.exit(main())
|
||||||
@@ -0,0 +1,69 @@
|
|||||||
|
import copy
|
||||||
|
import importlib.util
|
||||||
|
import json
|
||||||
|
from pathlib import Path
|
||||||
|
import unittest
|
||||||
|
|
||||||
|
ROOT = Path(__file__).resolve().parents[1]
|
||||||
|
spec = importlib.util.spec_from_file_location("decision", ROOT / "skills/codex-subagent-router/scripts/validate_decision_record.py")
|
||||||
|
module = importlib.util.module_from_spec(spec)
|
||||||
|
spec.loader.exec_module(module)
|
||||||
|
|
||||||
|
|
||||||
|
class DecisionTests(unittest.TestCase):
|
||||||
|
def setUp(self):
|
||||||
|
self.record = json.loads((ROOT / "examples/decisions/resolved.json").read_text())
|
||||||
|
|
||||||
|
def test_valid_resolution(self):
|
||||||
|
self.assertEqual(module.validate(self.record), [])
|
||||||
|
|
||||||
|
def test_counter_survives_continuation(self):
|
||||||
|
previous = copy.deepcopy(self.record)
|
||||||
|
self.record["same_class_failures"] = 0
|
||||||
|
self.assertTrue(module.validate(self.record, previous))
|
||||||
|
|
||||||
|
def test_reassessment_required(self):
|
||||||
|
self.record["reassessment"] = ""
|
||||||
|
self.assertTrue(module.validate(self.record))
|
||||||
|
|
||||||
|
def test_unresolved_counterexample_blocks_resolution(self):
|
||||||
|
self.record["unresolved"] = ["reproducible contrary result"]
|
||||||
|
self.assertTrue(module.validate(self.record))
|
||||||
|
|
||||||
|
def test_stale_baseline_blocks_resolution(self):
|
||||||
|
self.record["verification"]["baseline"] = "old"
|
||||||
|
self.assertTrue(module.validate(self.record))
|
||||||
|
|
||||||
|
def test_parent_must_understand(self):
|
||||||
|
self.record["verification"]["parent_understands"] = False
|
||||||
|
self.assertTrue(module.validate(self.record))
|
||||||
|
|
||||||
|
def test_environment_and_authority_do_not_escalate_to_astra(self):
|
||||||
|
for category in ("environment", "authority"):
|
||||||
|
self.record["failure_class"] = category
|
||||||
|
self.record["escalation"]["target"] = "astra"
|
||||||
|
self.assertTrue(module.validate(self.record))
|
||||||
|
|
||||||
|
def test_extra_consultation_needs_basis(self):
|
||||||
|
self.record["escalation"]["expert_consultations"] = 2
|
||||||
|
self.assertTrue(module.validate(self.record))
|
||||||
|
self.record["escalation"]["followup_basis"] = "new discriminator result"
|
||||||
|
self.assertEqual(module.validate(self.record), [])
|
||||||
|
|
||||||
|
def test_recovery_required(self):
|
||||||
|
self.record["recovery_route"] = ""
|
||||||
|
self.assertTrue(module.validate(self.record))
|
||||||
|
|
||||||
|
def test_bad_shapes_and_boolean_counters(self):
|
||||||
|
for value in (None, [], True):
|
||||||
|
self.assertTrue(module.validate(value))
|
||||||
|
self.record["same_class_failures"] = True
|
||||||
|
self.assertTrue(module.validate(self.record))
|
||||||
|
|
||||||
|
def test_unhashable_enum_fields_are_rejected(self):
|
||||||
|
for key in ("failure_class", "status"):
|
||||||
|
record = copy.deepcopy(self.record)
|
||||||
|
record[key] = []
|
||||||
|
self.assertTrue(module.validate(record))
|
||||||
|
self.record["escalation"]["target"] = {}
|
||||||
|
self.assertTrue(module.validate(self.record))
|
||||||
Reference in New Issue
Block a user