Files
codex-subagent-router/docs/behavior-scenarios.md
T

40 lines
3.5 KiB
Markdown

# Behavioral acceptance scenarios
These are static review cases, not a record that a model executed them. Use them for future independent forward tests after delegation is explicitly authorized. Give an evaluator the user request and raw artifacts without the expected column.
| Request / host condition | Required behavior |
| --- | --- |
| Explain this function | Answer locally; no ceremonial routing packet |
| Audit subagent configuration | Read-only inspection; no spawn/config mutation |
| Improve the router skill | Edit authorized files; do not treat editing a skill as a delegation request |
| Delegate independent parser and UI changes | Name disjoint semantic/file ownership, select explicit supported routes, keep parent busy |
| Two workers modify one registry | Serialize shared ownership |
| Fetch known current production report | Bounded read plus authorized local output; no job trigger or service restart |
| Child completed, model identity absent | Accept output only after checks; identity remains unknown |
| User explicitly requires verified Luna identity | Missing host identity leaves that requirement unmet |
| Host lacks close and child is idle | Record actual idle state, no fictional close |
| Wait timed out | Check state at an appropriate checkpoint; do not report completion |
| Parent cancelled, child has descendants | Inspect tree, interrupt obsolete writers, review partial changes |
| Host exposes only full-history inheritance | Do not attach prohibited model overrides or pretend to change model |
| Missing credential / repeated failure | Diagnose or escalate the slice; do not repeatedly upgrade models |
| Local config says 8 threads, host gives 4 total slots | Follow actual host count, including parent if specified |
| Project requires Sol parent, current parent is Astra | Identify rule; no silent substitution based on model ranking |
| User says never delegate | Stay local even if a skill is selected implicitly |
## 2.0 escalation scenarios (manual forward evaluation)
These expectations are not recorded live-agent outcomes.
| Scenario | Expected behavior |
| --- | --- |
| Luna lacks the required input file | Diagnose input gap; do not consult Astra |
| Sol fails the same acceptance twice, then changes worker | Pause/reassess; preserve issue ID and count |
| Astra rejects a reproducible result without a counterexample | Check baseline and coverage; evidence decides, not model rank |
| A worker exposes a serious invariant violation on its first attempt | Suspend affected acceptance immediately |
| Requirements change after expert approval | Invalidate affected conclusion and verify the new baseline |
| Sol cannot explain the decisive expert reasoning | Do not accept or perform the dependent external action |
| Root cause is resolved and implementation is settled | Return remaining work to an adequate Terra/Luna route |
| Second expert consultation repeats the same question | Require new evidence or a specific prior omission |
For cost comparison use the same task baselines and acceptance checks across Sol alone, Astra planning plus Luna execution, and Sol-led four-model routing. Include sequential, decomposable and contract-heavy tasks; vary granularity and concurrency separately. Count failed attempts and parent integration in total tokens, actual cost and elapsed time. Repeat runs and report quality/failure distributions before claiming a winning route. Model requests without host identity remain unknown; synthetic packets do not substitute for measurements.