Integrate 2.1 workflow and prepare release 2.2.0
This commit is contained in:
@@ -5,6 +5,8 @@ description: Route authorized Codex subagents across GPT-6 Luna, Sol and Astra;
|
||||
|
||||
# Codex Subagent Router
|
||||
|
||||
Version 2.2.0. For new general engineering sessions recommend GPT-6 Sol / medium, with high for demonstrated reasoning needs. Preserve the selected parent. Use Astra proactively for high-leverage difficult decisions, independent counterexample reviews and residual hard reasoning; neither prior failure nor a model quota is required.
|
||||
|
||||
Make delegation useful, observable and bounded. Preserve the selected parent model. The parent owns the critical path, shared decisions, integration and final acceptance. This skill neither changes configuration nor grants external permissions.
|
||||
|
||||
## Trigger before routing
|
||||
@@ -27,6 +29,8 @@ Spawn only P/C work with a concrete output and useful parallel benefit. Prefer u
|
||||
|
||||
Choose the execution mechanism before a model: local tools for fixed commands/transforms, host-supported concurrent or async tools for independent I/O, and authorized children for useful independent reasoning or implementation. Waiting on a tool alone is not a reason to create an agent. Prefer a coherent outcome with acceptance over many tiny handoffs; bound ambiguity and ownership rather than prescribing every reasoning step.
|
||||
|
||||
For queue states, integration backpressure and proactive expert formations, read [fleet scheduling](references/fleet.md).
|
||||
|
||||
## Proactive Luna fleet
|
||||
|
||||
Once authorized, parallel execution is the default for ready independent work. At the first useful decomposition and after each material result or scope change, look for slices that can run alongside the parent's critical path. Dispatch those slices promptly instead of finishing the parent's entire investigation first. Do not require proof of measured speedup before a clearly useful split.
|
||||
@@ -48,7 +52,7 @@ Stay local for a trivial task, no useful parallel benefit, an immediate dependen
|
||||
| Bounded classification, conversion, settled patch with fixed checks | GPT-6 Luna medium; high for local reasoning |
|
||||
| Everyday implementation, debugging, locating/correlating artifacts | GPT-6 Sol medium |
|
||||
| Bounded complex analysis, design or financial/security evidence | GPT-6 Sol medium/high |
|
||||
| Hardest independent synthesis across code, tools and research | GPT-6 Astra medium/high |
|
||||
| High-leverage difficult design, cross-system synthesis, independent high-impact counterexample review | GPT-6 Astra medium/high |
|
||||
| Shared decisions, integration and external/final gates | Current parent, serial |
|
||||
|
||||
Luna is the first candidate when inputs are bounded, the contract is settled, correctness is cheaply checkable, and failure is contained. All four conditions matter; a short diff can still hide a shared semantic decision. Use Sol directly when investigation, design choices or cross-file causal reasoning dominate. Use Astra directly for the hardest independent synthesis when expected quality or saved rework justifies it. Do not make high/max Luna a mandatory step before Sol.
|
||||
@@ -87,6 +91,8 @@ Use [evidence packets](references/evidence-packet.md) when structured evidence i
|
||||
|
||||
## Context and ownership discipline
|
||||
|
||||
Read [escalation and adjudication](references/escalation.md) when planning exposes uncertainty, a key check fails, shared contracts change or acceptance has unresolved counterexamples. Preserve issue IDs and failure counts across agents. Repair missing inputs first; adjudicate by reproducible evidence, not model rank. The parent must understand and verify decisive conclusions. Continue expert rounds only with evidence-based progress, and recover to an adequate cheaper route after resolution.
|
||||
|
||||
Load only needed references; keep long logs on disk. Avoid whole-history forks and repeating contracts, packets or prior findings. Account for parent context/reasoning, child usage, waiting, retries and integration when evaluating cost.
|
||||
|
||||
After steering or context compaction, recover the current objective, accepted corrections, permissions, baseline, file/semantic owners and pending tool/child IDs before continuing dependent work. Retain valid completed evidence; revalidate only results affected by the change. A queued correction is not proof that a running writer has stopped or adopted it.
|
||||
|
||||
@@ -0,0 +1,42 @@
|
||||
# Sol-led escalation and evidence adjudication
|
||||
|
||||
Recommend GPT-6 Sol / medium for new general engineering sessions; preserve an explicit user selection. The selected parent owns planning, difficult implementation, integration and acceptance. Sol handles routine implementation and investigation; Luna handles settled, checkable work. Terra is retired from recommendations and automatic fallbacks. Astra contributes both proactive high-leverage analysis and residual hard reasoning, without becoming a mandatory stage for every task. These are hypotheses to calibrate, not benchmark results.
|
||||
|
||||
## Reassess at observable checkpoints
|
||||
|
||||
Reassess after planning, a key failed check, shared-contract changes and before acceptance. Track a material issue with a stable issue ID and baseline. Counters follow the issue across retries, replacement agents and model changes. After two same-class acceptance failures, pause that slice and classify before another attempt. A serious counterexample or unsafe assumption triggers immediately.
|
||||
|
||||
| Observation | First response |
|
||||
| --- | --- |
|
||||
| Missing inputs, access, tool failure | Diagnose environment or obtain evidence; model escalation does not repair access |
|
||||
| Material requirement ambiguity | Inspect consumers/contracts; ask the user only for unresolved product choices |
|
||||
| Different root causes imply different fixes | Design the smallest discriminating check |
|
||||
| Conflicting records or stale baseline | Reconcile original evidence; suspend the dependent conclusion |
|
||||
| Fixes alternate between breaking invariants | Revisit shared contract and integration ownership |
|
||||
| High impact with no reliable verifier | Block affected acceptance/action; seek a verifier or explicit product decision |
|
||||
|
||||
Continue independent authorized work while the affected slice is paused. A lack of information, authorization or tools is not a reasoning failure.
|
||||
|
||||
## Choose the intervention
|
||||
|
||||
1. Repair the contract, baseline, inputs or verifier first. Reuse an appropriate idle agent; do not reset failure history by respawning.
|
||||
2. Increase execution capability when a well-specified slice exceeds the worker: Luna to Sol as justified, or Astra directly for the hardest independent synthesis. Select the adequate model directly, without visiting every tier. Effort, model and concurrency are independent controls.
|
||||
3. Select Astra directly when a difficult architectural choice affects many downstream tasks, competing consequential designs need discrimination, cross-system evidence needs synthesis, or an independent counterexample search materially protects a high-impact change. Repeated failure is not required. State the uncertainty, consequence of a wrong decision, available evidence and concrete expert deliverable. Residual difficult reasoning conflicts remain a valid trigger. Missing access alone is not a trigger. Authorized delegation and an independent bounded consultation are still required; the parent advances useful independent work when available and retains the shared decision. Where the host permits, useful parallel consultations may run while the parent waits for later adjudication; honor stricter host requirements. An immediate parent dependency stays local; requesting a child or writing configuration does not switch the live parent. If the parent cannot proceed reliably, describe the required capability/evidence instead of claiming a switch occurred.
|
||||
|
||||
An expert packet contains: issue ID, exact relevant baseline, goal, invariants, original evidence references, competing hypotheses, discriminators attempted or explicitly not yet run, and one focused question. For proactive work, zero failures is valid. Request counterexamples, a discriminating check, conditional conclusions and unresolved points. Supply only necessary context. No implementation writes unless separately assigned a disjoint write contract.
|
||||
|
||||
Bound each expert assignment by its question, deliverable and exit condition, not a universal one-consultation quota. Continue with the same suitable expert while new evidence, a testable refinement or a specific prior omission supports progress; record the basis for each additional round. An unresolved topic alone does not justify repetition. If no discriminating progress is available, pause that topic, obtain the missing input/check or return the decision to the parent. Do not loop reviews until models agree. Time/checkpoint budgets are instructions unless enforced by the host.
|
||||
|
||||
## Adjudicate by evidence
|
||||
|
||||
Normalize baseline and convert disagreement into falsifiable claims. Compare reproducible checks, original records and applicable contracts; confirm the check actually covers the disputed invariant. Neither majority vote nor a more expensive model wins automatically. A material unresolved counterexample from any model blocks the affected acceptance.
|
||||
|
||||
The selected parent (normally Sol) must map a decisive suggestion to the actual changes and invariants, understand the reasoning and verify it on the relevant current baseline. Otherwise the result remains unaccepted, even if Astra approves. User product tradeoffs remain user decisions; expert advice grants no permissions. Changed requirements invalidate relevant previous conclusions and require fresh validation.
|
||||
|
||||
## Record and recover
|
||||
|
||||
For material failures keep a compact record: issue ID, baseline, trigger/failure class, same-class failure count, facts, reassessment, escalation reason, evidence-based decision, remaining gaps and recovery route. Ordinary successful work needs no ceremonial record. After resolution, return execution to the lowest adequate route; do not retain Astra for unrelated follow-up work.
|
||||
|
||||
The optional [decision validator](../scripts/validate_decision_record.py) checks reported-state consistency only. Record schema 1 is independent of the product version and evidence-packet schema 2. It cannot prove evidence, enforce runtime gates, detect omitted issues or authenticate model identity. The parent must perform the checks. With `--previous`, it also rejects a changed issue ID, a decreasing failure count within the same class, or a decreasing cumulative expert consultation count even across class changes. A class change must be an evidence-based reclassification, not a counter reset tactic.
|
||||
|
||||
Required fields are demonstrated in the source repository's `examples/decisions/resolved.json`. `evidence` contains nonempty references, not copied logs. `status` is investigating, blocked, escalated or resolved. Resolving requires current-baseline passed verification, parent understanding and no unresolved items. Environment/authority issues cannot target Astra; repeated failures require reassessment; repeated expert consultations require a follow-up basis. The scalar `followup_basis` describes the current continuation; retain prior records for prior rounds. The validator cannot establish progress in each round from one aggregate record. A recovery route is required at resolution. Unknown observed model identity remains acceptable unless an explicit identity gate applies.
|
||||
@@ -0,0 +1,41 @@
|
||||
# Proactive fleet scheduling
|
||||
|
||||
Use for authorized complex parallel work. This is a decision procedure for the parent, not a background scheduler, permission grant or capacity override. Preserve the actual host schema, selected parent and final gates.
|
||||
|
||||
## Discover and dispatch early
|
||||
|
||||
After minimum goal/baseline discovery, identify independent questions and useful implementation slices. Do not finish reading every subsystem before dispatching ready investigations. Once a slice's input contract is stable, it may enter implementation while unrelated investigation continues. Shared decisions stay with the parent; preliminary independent drafts are not accepted contracts.
|
||||
|
||||
Maintain a compact queue: task ID, dependency and contract baseline, file AND semantic owner, shared resources, acceptance, requested route, child ID, observed runtime state and integration state. Mark pending, ready, running, returned, accepted, blocked or superseded as task states; keep actual host state separate. Ready means the prerequisites, ownership and permitted actions are established and the output has independent value.
|
||||
|
||||
At intake, dependency resolution, a child return or a meaningful baseline change:
|
||||
|
||||
1. Invalidate affected assumptions and identify ready slices. Prioritize work that unlocks dependencies, avoids costly wrong decisions or yields directly usable implementation.
|
||||
2. Check live capacity and integration load. Dispatch multiple independent ready tasks with useful parallel benefit; prefer concurrent parent work, but allow the parent to wait for later synthesis when the host permits. Honor stricter host requirements for concurrent parent work. Avoid sequential startup followed by unnecessary waits. Every descendant counts against the actual host limit. Children may not recursively expand the fleet without a parent-assigned bounded delegation scope and capacity allocation.
|
||||
3. Continue the parent's independent critical work when available. Use incremental messages for running children and reuse suitable idle ones. Do not perform the same assigned investigation locally merely to stay busy.
|
||||
4. Inspect returned evidence and partial writes; integrate serially. Refill available capacity with ready work without waiting for all siblings. A returned result is not necessarily accepted or a released slot; follow actual lifecycle controls.
|
||||
|
||||
If the next action is immediately dependent on an unassigned task, keep it local. A previously independent task may become a dependency as work progresses; bounded waiting is then legitimate. No independent useful parent work means do not create a ceremonial side task to justify another spawn.
|
||||
|
||||
## Choose a formation for the task
|
||||
|
||||
- Independent investigation: several focused investigators may cover distinct sources, consumers or falsifiable hypotheses. They should not all produce the same broad summary.
|
||||
- Settled implementation: multiple GPT-6 Luna workers own bounded, settled, independently checkable slices; use GPT-6 Sol directly when implementation requires investigation or design judgment. Shared registries, schema and generated outputs require one owner even when paths differ.
|
||||
- Difficult consequential decision: an Astra specialist analyzes the bounded uncertainty while the parent and other workers advance unaffected work. Gate dependent implementation on the parent's accepted contract. Read [expert intervention](escalation.md) for triggers and progress-based continuation.
|
||||
- Independent review: assign a stable commit/artifact and the relevant requirements, without supplying the parent's preferred verdict. Review can overlap unrelated work; stale observations must be rechecked before acceptance.
|
||||
|
||||
There is no mandatory model ratio or permanent reviewer slot. Use several workers of the same adequate model when the tasks warrant it. If Astra is already the selected parent, do not add another Astra just to fulfill a role label. Reuse the parent capability unless a genuinely independent bounded perspective adds value.
|
||||
|
||||
## Backpressure and recovery
|
||||
|
||||
Choose a working width up to actual available capacity based on ready independent work, parent integration capacity and shared-resource limits. Do not reserve slots mechanically or fill them with artificial work. Hardware, service rate limits and tool sessions can constrain useful width below the agent limit.
|
||||
|
||||
When returned patches are accumulating faster than the parent can inspect, or a returned contract blocks other slices, pause new dependent implementation and prioritize integration. When collisions, stale baselines, repeated rework or resource contention appear, stop affected dispatch and reduce width; interrupt obsolete/conflicting writers and inspect partial changes. Independent read-only work may continue if useful and resource-safe. Resume expansion after the cause is resolved and ready work remains.
|
||||
|
||||
Never run acceptance tests against a moving artifact and report them as final. Pin a commit/snapshot or wait until relevant writers stop; a worktree alone does not isolate external databases, generated directories or shared sessions. Before merge/delivery, resolve material counterexamples, verify the integrated baseline and audit the agent tree for remaining writers.
|
||||
|
||||
## Evaluate effectiveness
|
||||
|
||||
Optimize accepted quality and end-to-end time subject to the user's cost constraints. Include parent work, all children, context duplication, retries, expert rounds and integration in cost measurements. Expensive early reasoning can be worthwhile when it prevents downstream rework; cheap high-volume workers can be wasteful when contracts are unsettled. Both are hypotheses until measured.
|
||||
|
||||
For a calibration exercise compare the same tasks and baselines at serial and increasing supported widths, varying granularity separately. Include decomposable and contract-heavy tasks, failures and repeated runs. Record quality, elapsed time, total exposed usage/cost, rework, and integration time; unknown counters or runtime model identity stay unknown. Do not claim speed/cost improvements from static scenarios or a successful single rollout.
|
||||
@@ -24,6 +24,8 @@ Keep required async tool jobs in the same dependency view as children, using the
|
||||
|
||||
After context compaction, recover objective, permissions, owners, pending IDs and accepted evidence before dispatching more work. Detailed API-versus-host limits are in [GPT-6 adaptation](gpt6-adaptation.md).
|
||||
|
||||
For rolling execution, use [fleet scheduling](fleet.md): refill ready work after useful handbacks and dependency changes, rather than waiting for every sibling. Check integration backlog before starting more implementation. Task readiness, output acceptance and actual slot availability remain separate facts.
|
||||
|
||||
## Cancellation and handback
|
||||
|
||||
Interrupting does not undo files or guarantee descendant cancellation. First preserve useful returned evidence, stop affected writers, inspect partial changes, then assign the remaining work to one owner. Never reset shared files wholesale.
|
||||
|
||||
@@ -19,7 +19,7 @@ For clients that support the documented keys, an example is:
|
||||
|
||||
```toml
|
||||
# Example only; merge intentionally into the existing configuration.
|
||||
model = "gpt-6-astra"
|
||||
model = "gpt-6-sol"
|
||||
model_reasoning_effort = "medium"
|
||||
|
||||
[agents]
|
||||
|
||||
@@ -11,6 +11,9 @@ Read only when the entrypoint leaves a routing choice unresolved.
|
||||
| Cheap Luna attempts require repeated parent rewrites | Choose Sol directly next time for this shape; count total accepted-task cost. |
|
||||
| Difficult accounting discrepancy | Sol reasoning; parent owns the semantic decision. |
|
||||
| Coupled architecture across runtimes and tools | Consider Astra directly; keep shared decisions serial. |
|
||||
| Difficult choice would constrain many downstream slices | Consider an early bounded Astra analysis; gate dependent work on the parent's accepted contract. |
|
||||
| High-impact stable change needs independent counterexamples | Consider Astra review when depth adds value; give original requirements and a fixed artifact, not a preferred verdict. |
|
||||
| Expert makes testable progress over several rounds | Continue the scoped topic with incremental evidence; do not impose a one-call quota or repeat unchanged questions. |
|
||||
| Missing credential or inaccessible source | Resolve the evidence/environment gap; model escalation cannot supply access. |
|
||||
| Astra advertised only for task creation | Child support remains unproven; do not create another user task as a workaround. |
|
||||
| Named role fixes model/effort | Accept its supported binding or choose an overridable role; never attach prohibited fork overrides. |
|
||||
|
||||
@@ -0,0 +1,103 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Validate reported decision consistency, not evidence truth or live execution."""
|
||||
import argparse
|
||||
import json
|
||||
from pathlib import Path
|
||||
import sys
|
||||
|
||||
|
||||
def validate(record, previous=None):
|
||||
errors = []
|
||||
if not isinstance(record, dict):
|
||||
return ["record must be an object"]
|
||||
for key in ("issue_id", "baseline", "failure_class", "status"):
|
||||
if not isinstance(record.get(key), str) or not record[key].strip():
|
||||
errors.append(key + " must be nonempty text")
|
||||
if type(record.get("schema_version")) is not int or record["schema_version"] != 1:
|
||||
errors.append("schema_version must be 1")
|
||||
if record.get("failure_class") not in ("environment", "authority", "contract", "implementation", "reasoning"):
|
||||
errors.append("invalid failure_class")
|
||||
if record.get("status") not in ("investigating", "blocked", "escalated", "resolved"):
|
||||
errors.append("invalid status")
|
||||
count = record.get("same_class_failures")
|
||||
if type(count) is not int or count < 0:
|
||||
errors.append("same_class_failures must be a nonnegative integer")
|
||||
elif count >= 2 and not text(record.get("reassessment")):
|
||||
errors.append("repeated failure requires reassessment")
|
||||
for key in ("evidence", "unresolved"):
|
||||
value = record.get(key)
|
||||
if not isinstance(value, list) or any(not text(item) for item in value):
|
||||
errors.append(key + " must be a list of nonempty references/issues")
|
||||
escalation = record.get("escalation")
|
||||
if not isinstance(escalation, dict):
|
||||
errors.append("escalation must be an object")
|
||||
escalation = {}
|
||||
target = escalation.get("target")
|
||||
if target not in ("none", "terra", "sol", "astra"):
|
||||
errors.append("invalid escalation target")
|
||||
if target != "none" and not text(escalation.get("reason")):
|
||||
errors.append("escalation requires a reason")
|
||||
if target == "astra" and record.get("failure_class") in ("environment", "authority"):
|
||||
errors.append("Astra cannot repair environment or authority")
|
||||
consultations = escalation.get("expert_consultations")
|
||||
if type(consultations) is not int or consultations < 0:
|
||||
errors.append("expert_consultations must be a nonnegative integer")
|
||||
elif consultations > 1 and not text(escalation.get("followup_basis")):
|
||||
errors.append("additional consultation requires new evidence or a specific omission")
|
||||
if record.get("status") == "escalated" and target == "none":
|
||||
errors.append("escalated status requires a target")
|
||||
if record.get("status") == "resolved":
|
||||
verification = record.get("verification", {})
|
||||
if not isinstance(verification, dict):
|
||||
verification = {}
|
||||
if (verification.get("status") != "passed"
|
||||
or verification.get("baseline") != record.get("baseline")
|
||||
or verification.get("parent_understands") is not True):
|
||||
errors.append("resolution requires understood, passed verification on current baseline")
|
||||
if record.get("unresolved") != [] or not record.get("evidence"):
|
||||
errors.append("resolution requires evidence and no unresolved items")
|
||||
if not text(record.get("recovery_route")):
|
||||
errors.append("resolution requires a recovery route")
|
||||
if previous is not None:
|
||||
if not isinstance(previous, dict):
|
||||
errors.append("previous record must be an object")
|
||||
else:
|
||||
if previous.get("issue_id") != record.get("issue_id"):
|
||||
errors.append("issue_id changed across continuation")
|
||||
old = previous.get("same_class_failures")
|
||||
if (previous.get("failure_class") == record.get("failure_class")
|
||||
and type(old) is int and type(count) is int and count < old):
|
||||
errors.append("failure counter decreased within the same class")
|
||||
previous_escalation = previous.get("escalation")
|
||||
old_consultations = (previous_escalation.get("expert_consultations")
|
||||
if isinstance(previous_escalation, dict) else None)
|
||||
if type(old_consultations) is not int or old_consultations < 0:
|
||||
errors.append("previous expert_consultations must be a nonnegative integer")
|
||||
elif (type(consultations) is int and consultations >= 0
|
||||
and consultations < old_consultations):
|
||||
errors.append("expert consultation counter decreased across continuation")
|
||||
return errors
|
||||
|
||||
|
||||
def text(value):
|
||||
return isinstance(value, str) and bool(value.strip())
|
||||
|
||||
|
||||
def main():
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("record")
|
||||
parser.add_argument("--previous")
|
||||
args = parser.parse_args()
|
||||
try:
|
||||
record = json.loads(Path(args.record).read_text(encoding="utf-8"))
|
||||
previous = json.loads(Path(args.previous).read_text(encoding="utf-8")) if args.previous else None
|
||||
errors = validate(record, previous)
|
||||
except (OSError, ValueError) as exc:
|
||||
print("Invalid decision record: " + str(exc), file=sys.stderr)
|
||||
return 1
|
||||
print("\n".join(errors) if errors else "Decision record consistency passed (truth not verified)")
|
||||
return int(bool(errors))
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
Reference in New Issue
Block a user