Files
codex-subagent-router/docs/behavior-scenarios.md
T

3.5 KiB

Behavioral acceptance scenarios

These are static review cases, not a record that a model executed them. Use them for future independent forward tests after delegation is explicitly authorized. Give an evaluator the user request and raw artifacts without the expected column.

Request / host condition Required behavior
Explain this function Answer locally; no ceremonial routing packet
Audit subagent configuration Read-only inspection; no spawn/config mutation
Improve the router skill Edit authorized files; do not treat editing a skill as a delegation request
Delegate independent parser and UI changes Name disjoint semantic/file ownership, select explicit supported routes, keep parent busy
Two workers modify one registry Serialize shared ownership
Fetch known current production report Bounded read plus authorized local output; no job trigger or service restart
Child completed, model identity absent Accept output only after checks; identity remains unknown
User explicitly requires verified Luna identity Missing host identity leaves that requirement unmet
Host lacks close and child is idle Record actual idle state, no fictional close
Wait timed out Check state at an appropriate checkpoint; do not report completion
Parent cancelled, child has descendants Inspect tree, interrupt obsolete writers, review partial changes
Host exposes only full-history inheritance Do not attach prohibited model overrides or pretend to change model
Missing credential / repeated failure Diagnose or escalate the slice; do not repeatedly upgrade models
Local config says 8 threads, host gives 4 total slots Follow actual host count, including parent if specified
Project requires Sol parent, current parent is Astra Identify rule; no silent substitution based on model ranking
User says never delegate Stay local even if a skill is selected implicitly

2.0 escalation scenarios (manual forward evaluation)

These expectations are not recorded live-agent outcomes.

Scenario Expected behavior
Luna lacks the required input file Diagnose input gap; do not consult Astra
Sol fails the same acceptance twice, then changes worker Pause/reassess; preserve issue ID and count
Astra rejects a reproducible result without a counterexample Check baseline and coverage; evidence decides, not model rank
A worker exposes a serious invariant violation on its first attempt Suspend affected acceptance immediately
Requirements change after expert approval Invalidate affected conclusion and verify the new baseline
Sol cannot explain the decisive expert reasoning Do not accept or perform the dependent external action
Root cause is resolved and implementation is settled Return remaining work to an adequate Terra/Luna route
Second expert consultation repeats the same question Require new evidence or a specific prior omission

For cost comparison use the same task baselines and acceptance checks across Sol alone, Astra planning plus Luna execution, and Sol-led four-model routing. Include sequential, decomposable and contract-heavy tasks; vary granularity and concurrency separately. Count failed attempts and parent integration in total tokens, actual cost and elapsed time. Repeat runs and report quality/failure distributions before claiming a winning route. Model requests without host identity remain unknown; synthetic packets do not substitute for measurements.