Release 2.0.0 with Sol-led escalation and evidence adjudication

This commit is contained in:
Codex
2026-09-15 21:17:41 +08:00
parent 7969489b7e
commit dc0360cefb
12 changed files with 300 additions and 3 deletions
+4
View File
@@ -5,6 +5,8 @@ description: Decide when Codex subagents help, route bounded work across Astra,
# Codex Subagent Router
Version 2.0.0. For new general engineering sessions recommend Sol / medium, with high when demonstrated reasoning needs justify it. Preserve the user's selected parent; installation does not switch it. Use Astra for bounded residual hard reasoning, rather than mandatory planning of every task.
Make delegation useful, observable and bounded. Preserve the selected parent model. The parent owns the critical path, shared decisions, integration and final acceptance. This skill neither changes configuration nor grants external permissions.
## Trigger before routing
@@ -64,6 +66,8 @@ Use [evidence packets](references/evidence-packet.md) when structured evidence i
## Context and ownership discipline
Read [escalation and adjudication](references/escalation.md) when planning exposes uncertainty, a key check fails, shared contracts change or acceptance has unresolved counterexamples. Pause after two same-class failures; preserve issue history across agents. Repair missing inputs first, escalate capability only for reasoning/execution limits, and adjudicate by reproducible evidence rather than model rank. The parent must understand and verify decisive conclusions. Recover to a cheaper adequate route after resolution.
Load only needed references; keep long logs on disk. Avoid whole-history forks and repeating contracts, packets or prior findings. Account for parent context/reasoning, child usage, waiting, retries and integration when evaluating cost.
Production read-only collection is routable by work shape; mutation and final semantic acceptance stay parent-owned. Do not silently relax a project's literal model-owner requirement. Identify the actual rule and resolve it only through authorized changes.
@@ -0,0 +1,42 @@
# Sol-led escalation and evidence adjudication
Recommend Sol / medium for new general engineering sessions; preserve an explicit user selection. Sol owns planning, difficult implementation, integration and acceptance. Terra handles routine implementation; Luna handles settled, checkable work. Astra is a bounded expert for residual hard reasoning, not a mandatory planning or review stage. These are hypotheses to calibrate, not benchmark results.
## Reassess at observable checkpoints
Reassess after planning, a key failed check, shared-contract changes and before acceptance. Track a material issue with a stable issue ID and baseline. Counters follow the issue across retries, replacement agents and model changes. After two same-class acceptance failures, pause that slice and classify before another attempt. A serious counterexample or unsafe assumption triggers immediately.
| Observation | First response |
| --- | --- |
| Missing inputs, access, tool failure | Diagnose environment or obtain evidence; model escalation does not repair access |
| Material requirement ambiguity | Inspect consumers/contracts; ask the user only for unresolved product choices |
| Different root causes imply different fixes | Design the smallest discriminating check |
| Conflicting records or stale baseline | Reconcile original evidence; suspend the dependent conclusion |
| Fixes alternate between breaking invariants | Revisit shared contract and integration ownership |
| High impact with no reliable verifier | Block affected acceptance/action; seek a verifier or explicit product decision |
Continue independent authorized work while the affected slice is paused. A lack of information, authorization or tools is not a reasoning failure.
## Choose the intervention
1. Repair the contract, baseline, inputs or verifier first. Reuse an appropriate idle agent; do not reset failure history by respawning.
2. Increase execution capability when a well-specified slice exceeds the worker: Luna to Terra, or Terra to Sol as justified. Select the adequate model directly, without visiting every tier. Effort, model and concurrency are independent controls.
3. Use Astra only for a remaining difficult reasoning conflict with adequate evidence. Authorized delegation and an independent bounded consultation are still required. An immediate parent dependency stays local; requesting a child or writing configuration does not switch the live parent. If the parent cannot proceed reliably, describe the required capability/evidence instead of claiming a switch occurred.
An expert packet contains: issue ID, exact relevant baseline, goal, invariants, original evidence references, competing hypotheses, attempted discriminators and one focused question. Request counterexamples, a discriminating check, conditional conclusions and unresolved points. Supply only necessary context. No implementation writes unless separately assigned a disjoint write contract.
Default to one bounded expert consultation per issue. A follow-up requires new evidence or a specific omission in the first answer; record that basis. Do not loop reviews until models agree. Time/checkpoint budgets are instructions unless enforced by the host.
## Adjudicate by evidence
Normalize baseline and convert disagreement into falsifiable claims. Compare reproducible checks, original records and applicable contracts; confirm the check actually covers the disputed invariant. Neither majority vote nor a more expensive model wins automatically. A material unresolved counterexample from any model blocks the affected acceptance.
Sol must map a decisive suggestion to the actual changes and invariants, understand the reasoning and verify it on the relevant current baseline. Otherwise the result remains unaccepted, even if Astra approves. User product tradeoffs remain user decisions; expert advice grants no permissions. Changed requirements invalidate relevant previous conclusions and require fresh validation.
## Record and recover
For material failures keep a compact record: issue ID, baseline, trigger/failure class, same-class failure count, facts, reassessment, escalation reason, evidence-based decision, remaining gaps and recovery route. Ordinary successful work needs no ceremonial record. After resolution, return execution to the lowest adequate route; do not retain Astra for unrelated follow-up work.
The optional [decision validator](../scripts/validate_decision_record.py) checks reported-state consistency only. Record schema 1 is independent of product version 2.0 and evidence-packet schema 2. It cannot prove evidence, enforce runtime gates, detect omitted issues or authenticate model identity. The parent must perform the checks. With `--previous`, it also rejects a changed issue ID or a decreasing count within the same failure class. A class change must be an evidence-based reclassification, not a counter reset tactic.
Required fields are demonstrated in the source repository's `examples/decisions/resolved.json`. `evidence` contains nonempty references, not copied logs. `status` is investigating, blocked, escalated or resolved. Resolving requires current-baseline passed verification, parent understanding and no unresolved items. Environment/authority issues cannot target Astra; repeated failures require reassessment; repeated expert consultations require a follow-up basis. A recovery route is required at resolution. Unknown observed model identity remains acceptable unless an explicit identity gate applies.
@@ -19,7 +19,7 @@ For clients that support the documented keys, an example is:
```toml
# Example only; merge intentionally into the existing configuration.
model = "gpt-6-astra"
model = "gpt-5.6-sol"
model_reasoning_effort = "medium"
[agents]
@@ -0,0 +1,95 @@
#!/usr/bin/env python3
"""Validate reported decision consistency, not evidence truth or live execution."""
import argparse
import json
from pathlib import Path
import sys
def validate(record, previous=None):
errors = []
if not isinstance(record, dict):
return ["record must be an object"]
for key in ("issue_id", "baseline", "failure_class", "status"):
if not isinstance(record.get(key), str) or not record[key].strip():
errors.append(key + " must be nonempty text")
if type(record.get("schema_version")) is not int or record["schema_version"] != 1:
errors.append("schema_version must be 1")
if record.get("failure_class") not in ("environment", "authority", "contract", "implementation", "reasoning"):
errors.append("invalid failure_class")
if record.get("status") not in ("investigating", "blocked", "escalated", "resolved"):
errors.append("invalid status")
count = record.get("same_class_failures")
if type(count) is not int or count < 0:
errors.append("same_class_failures must be a nonnegative integer")
elif count >= 2 and not text(record.get("reassessment")):
errors.append("repeated failure requires reassessment")
for key in ("evidence", "unresolved"):
value = record.get(key)
if not isinstance(value, list) or any(not text(item) for item in value):
errors.append(key + " must be a list of nonempty references/issues")
escalation = record.get("escalation")
if not isinstance(escalation, dict):
errors.append("escalation must be an object")
escalation = {}
target = escalation.get("target")
if target not in ("none", "terra", "sol", "astra"):
errors.append("invalid escalation target")
if target != "none" and not text(escalation.get("reason")):
errors.append("escalation requires a reason")
if target == "astra" and record.get("failure_class") in ("environment", "authority"):
errors.append("Astra cannot repair environment or authority")
consultations = escalation.get("expert_consultations")
if type(consultations) is not int or consultations < 0:
errors.append("expert_consultations must be a nonnegative integer")
elif consultations > 1 and not text(escalation.get("followup_basis")):
errors.append("additional consultation requires new evidence or a specific omission")
if record.get("status") == "escalated" and target == "none":
errors.append("escalated status requires a target")
if record.get("status") == "resolved":
verification = record.get("verification", {})
if not isinstance(verification, dict):
verification = {}
if (verification.get("status") != "passed"
or verification.get("baseline") != record.get("baseline")
or verification.get("parent_understands") is not True):
errors.append("resolution requires understood, passed verification on current baseline")
if record.get("unresolved") != [] or not record.get("evidence"):
errors.append("resolution requires evidence and no unresolved items")
if not text(record.get("recovery_route")):
errors.append("resolution requires a recovery route")
if previous is not None:
if not isinstance(previous, dict):
errors.append("previous record must be an object")
else:
if previous.get("issue_id") != record.get("issue_id"):
errors.append("issue_id changed across continuation")
old = previous.get("same_class_failures")
if (previous.get("failure_class") == record.get("failure_class")
and type(old) is int and type(count) is int and count < old):
errors.append("failure counter decreased within the same class")
return errors
def text(value):
return isinstance(value, str) and bool(value.strip())
def main():
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("record")
parser.add_argument("--previous")
args = parser.parse_args()
try:
record = json.loads(Path(args.record).read_text(encoding="utf-8"))
previous = json.loads(Path(args.previous).read_text(encoding="utf-8")) if args.previous else None
errors = validate(record, previous)
except (OSError, ValueError) as exc:
print("Invalid decision record: " + str(exc), file=sys.stderr)
return 1
print("\n".join(errors) if errors else "Decision record consistency passed (truth not verified)")
return int(bool(errors))
if __name__ == "__main__":
sys.exit(main())