Files
codex-subagent-router/skills/codex-subagent-router/references/escalation.md
T
2026-09-23 10:44:41 +08:00

6.9 KiB

Sol-led escalation and evidence adjudication

Recommend GPT-6 Sol / medium for new general engineering sessions; preserve an explicit user selection. The selected parent owns planning, difficult implementation, integration and acceptance. Sol handles routine implementation and investigation; Luna handles settled, checkable work. Terra is retired from recommendations and automatic fallbacks. Astra contributes both proactive high-leverage analysis and residual hard reasoning, without becoming a mandatory stage for every task. These are hypotheses to calibrate, not benchmark results.

Reassess at observable checkpoints

Reassess after planning, a key failed check, shared-contract changes and before acceptance. Track a material issue with a stable issue ID and baseline. Counters follow the issue across retries, replacement agents and model changes. After two same-class acceptance failures, pause that slice and classify before another attempt. A serious counterexample or unsafe assumption triggers immediately.

Observation First response
Missing inputs, access, tool failure Diagnose environment or obtain evidence; model escalation does not repair access
Material requirement ambiguity Inspect consumers/contracts; ask the user only for unresolved product choices
Different root causes imply different fixes Design the smallest discriminating check
Conflicting records or stale baseline Reconcile original evidence; suspend the dependent conclusion
Fixes alternate between breaking invariants Revisit shared contract and integration ownership
High impact with no reliable verifier Block affected acceptance/action; seek a verifier or explicit product decision

Continue independent authorized work while the affected slice is paused. A lack of information, authorization or tools is not a reasoning failure.

Choose the intervention

  1. Repair the contract, baseline, inputs or verifier first. Reuse an appropriate idle agent; do not reset failure history by respawning.
  2. Increase execution capability when a well-specified slice exceeds the worker: Luna to Sol as justified, or Astra directly for the hardest independent synthesis. Select the adequate model directly, without visiting every tier. Effort, model and concurrency are independent controls.
  3. Select Astra directly when a difficult architectural choice affects many downstream tasks, competing consequential designs need discrimination, cross-system evidence needs synthesis, or an independent counterexample search materially protects a high-impact change. Repeated failure is not required. State the uncertainty, consequence of a wrong decision, available evidence and concrete expert deliverable. Residual difficult reasoning conflicts remain a valid trigger. Missing access alone is not a trigger. Authorized delegation and an independent bounded consultation are still required; the parent advances useful independent work when available and retains the shared decision. Where the host permits, useful parallel consultations may run while the parent waits for later adjudication; honor stricter host requirements. An immediate parent dependency stays local; requesting a child or writing configuration does not switch the live parent. If the parent cannot proceed reliably, describe the required capability/evidence instead of claiming a switch occurred.

An expert packet contains: issue ID, exact relevant baseline, goal, invariants, original evidence references, competing hypotheses, discriminators attempted or explicitly not yet run, and one focused question. For proactive work, zero failures is valid. Request counterexamples, a discriminating check, conditional conclusions and unresolved points. Supply only necessary context. No implementation writes unless separately assigned a disjoint write contract.

Bound each expert assignment by its question, deliverable and exit condition, not a universal one-consultation quota. Continue with the same suitable expert while new evidence, a testable refinement or a specific prior omission supports progress; record the basis for each additional round. An unresolved topic alone does not justify repetition. If no discriminating progress is available, pause that topic, obtain the missing input/check or return the decision to the parent. Do not loop reviews until models agree. Time/checkpoint budgets are instructions unless enforced by the host.

Adjudicate by evidence

Normalize baseline and convert disagreement into falsifiable claims. Compare reproducible checks, original records and applicable contracts; confirm the check actually covers the disputed invariant. Neither majority vote nor a more expensive model wins automatically. A material unresolved counterexample from any model blocks the affected acceptance.

The selected parent (normally Sol) must map a decisive suggestion to the actual changes and invariants, understand the reasoning and verify it on the relevant current baseline. Otherwise the result remains unaccepted, even if Astra approves. User product tradeoffs remain user decisions; expert advice grants no permissions. Changed requirements invalidate relevant previous conclusions and require fresh validation.

Record and recover

For material failures keep a compact record: issue ID, baseline, trigger/failure class, same-class failure count, facts, reassessment, escalation reason, evidence-based decision, remaining gaps and recovery route. Ordinary successful work needs no ceremonial record. After resolution, return execution to the lowest adequate route; do not retain Astra for unrelated follow-up work.

The optional decision validator checks reported-state consistency only. Record schema 1 is independent of the product version and evidence-packet schema 2. It cannot prove evidence, enforce runtime gates, detect omitted issues or authenticate model identity. The parent must perform the checks. With --previous, it also rejects a changed issue ID, a decreasing failure count within the same class, or a decreasing cumulative expert consultation count even across class changes. A class change must be an evidence-based reclassification, not a counter reset tactic.

Required fields are demonstrated in the source repository's examples/decisions/resolved.json. evidence contains nonempty references, not copied logs. status is investigating, blocked, escalated or resolved. Resolving requires current-baseline passed verification, parent understanding and no unresolved items. Environment/authority issues cannot target Astra; repeated failures require reassessment; repeated expert consultations require a follow-up basis. The scalar followup_basis describes the current continuation; retain prior records for prior rounds. The validator cannot establish progress in each round from one aggregate record. A recovery route is required at resolution. Unknown observed model identity remains acceptable unless an explicit identity gate applies.