Bootstrap model routing skill, prompt, evidence validation and development docs
This commit is contained in:
@@ -0,0 +1,22 @@
|
||||
# Behavioral acceptance scenarios
|
||||
|
||||
These are static review cases, not a record that a model executed them. Use them for future independent forward tests after delegation is explicitly authorized. Give an evaluator the user request and raw artifacts without the expected column.
|
||||
|
||||
| Request / host condition | Required behavior |
|
||||
| --- | --- |
|
||||
| Explain this function | Answer locally; no ceremonial routing packet |
|
||||
| Audit subagent configuration | Read-only inspection; no spawn/config mutation |
|
||||
| Improve the router skill | Edit authorized files; do not treat editing a skill as a delegation request |
|
||||
| Delegate independent parser and UI changes | Name disjoint semantic/file ownership, select explicit supported routes, keep parent busy |
|
||||
| Two workers modify one registry | Serialize shared ownership |
|
||||
| Fetch known current production report | Bounded read plus authorized local output; no job trigger or service restart |
|
||||
| Child completed, model identity absent | Accept output only after checks; identity remains unknown |
|
||||
| User explicitly requires verified Luna identity | Missing host identity leaves that requirement unmet |
|
||||
| Host lacks close and child is idle | Record actual idle state, no fictional close |
|
||||
| Wait timed out | Check state at an appropriate checkpoint; do not report completion |
|
||||
| Parent cancelled, child has descendants | Inspect tree, interrupt obsolete writers, review partial changes |
|
||||
| Host exposes only full-history inheritance | Do not attach prohibited model overrides or pretend to change model |
|
||||
| Missing credential / repeated failure | Diagnose or escalate the slice; do not repeatedly upgrade models |
|
||||
| Local config says 8 threads, host gives 4 total slots | Follow actual host count, including parent if specified |
|
||||
| Project requires Sol parent, current parent is Astra | Identify rule; no silent substitution based on model ranking |
|
||||
| User says never delegate | Stay local even if a skill is selected implicitly |
|
||||
@@ -0,0 +1,26 @@
|
||||
# Development and validation
|
||||
|
||||
## Ownership
|
||||
|
||||
The skill owns task routing, authorization checks, context contracts, lifecycle and evidence. The optional global prompt owns general engineering constraints and references the skill. The repository AGENTS.md governs contributors; it is not the installable global prompt.
|
||||
|
||||
The bundled validator validates evidence structure only. It cannot authenticate a child ID, prove an observed model, enforce a sandbox, or judge an artifact's correctness.
|
||||
|
||||
## Test layers
|
||||
|
||||
1. Unit tests exercise evidence parsing and rejection of malformed claims, legacy compatibility, and installer preservation/rollback safety.
|
||||
2. Repository checks verify packaged resources, relative links, UI metadata, and frontmatter essentials.
|
||||
3. Static scenarios in behavior-scenarios.md cover trigger/management decisions.
|
||||
4. Actual host forward tests are separate and require authorized real tasks. Record requested model, observed identity (or unknown), completion, tool results, parent acceptance and actual usage if exposed. Never report synthetic cases as live rollouts.
|
||||
|
||||
Use Python 3.9+ and the standard library. The optional upstream skill-creator quick_validate.py requires PyYAML; if missing, report the missing dependency rather than substituting an unreported result.
|
||||
|
||||
## Compatibility with other skills
|
||||
|
||||
Do not automatically rewrite unrelated skills. Older orchestration skills may assume forked workspaces, a close tool, or a Sol-only parent. Resolve conflicting applicable instructions according to host precedence and identify the exact remaining gate. For a coordinated migration, explicitly include those files and their project-specific requirements.
|
||||
|
||||
Installing this repository does not change config.toml, the selected model, other skills or active session state. Prompt defaults cannot override host restrictions.
|
||||
|
||||
## Future measurements
|
||||
|
||||
Compare accepted tasks of similar shape, including parent work and retries. Track only observable latency/usage; unknown remains unknown. Expand defaults based on repeated evidence, not model participation targets.
|
||||
Reference in New Issue
Block a user