Bootstrap model routing skill, prompt, evidence validation and development docs
This commit is contained in:
@@ -0,0 +1,69 @@
|
||||
---
|
||||
name: codex-subagent-router
|
||||
description: Decide when Codex subagents help, route bounded work across Astra, Sol, Terra and Luna, and manage child evidence and lifecycle. Use for delegation requests, routing decisions, or subagent audits; inspecting this skill does not itself authorize spawning.
|
||||
---
|
||||
|
||||
# Codex Subagent Router
|
||||
|
||||
Make delegation useful, observable and bounded. Preserve the selected parent model. The parent owns the critical path, shared decisions, integration and final acceptance. This skill neither changes configuration nor grants external permissions.
|
||||
|
||||
## Trigger before routing
|
||||
|
||||
Classify the request before using child tools:
|
||||
|
||||
- **Audit or advice:** inspect rules, configuration and available tools; report findings without spawning or changing settings.
|
||||
- **Authorized execution:** the user requested delegation, or an applicable instruction explicitly authorizes it. Within that scope, actively delegate independent useful slices while the parent advances other work.
|
||||
- **No delegation authorization:** work locally. Automatic skill discovery, model availability, task complexity and a request to edit this skill do not grant permission to spawn.
|
||||
|
||||
Honor higher-priority host restrictions even when a lower-priority rule permits delegation. Do not request authorization repeatedly after it has been granted. Never turn a routing recommendation into a new user-owned task.
|
||||
|
||||
Before dispatch classify each candidate:
|
||||
|
||||
- **P:** independent evidence or disjoint file AND semantic writes with its own acceptance.
|
||||
- **C:** a bounded investigation/draft that needs latest-baseline parent integration.
|
||||
- **S:** an immediate dependency, overlapping contract/state, permissions, production mutation, PR/merge/deploy or final acceptance; keep it in the parent.
|
||||
|
||||
Spawn only P/C work with a concrete output and useful parent work available. Do not spawn for one command, ceremonial probes, duplicate reviews or a model quota. Batch homogeneous small work.
|
||||
|
||||
## Select a supported route explicitly
|
||||
|
||||
| Work shape | Initial route |
|
||||
| --- | --- |
|
||||
| Known-source collection, fixed checks, logs | Luna low/medium |
|
||||
| Bounded classification, conversion, settled patch with fixed checks | Luna high |
|
||||
| Everyday implementation, debugging, locating/correlating artifacts | Terra medium; high when needed |
|
||||
| Bounded complex analysis, design or financial/security evidence | Sol medium; high when needed |
|
||||
| Hardest independent synthesis across code, tools and research | Astra medium; high when needed |
|
||||
| Shared decisions, integration and external/final gates | Current parent, serial |
|
||||
|
||||
These are starting heuristics, not measured cost rankings. Pick sufficient capability directly; do not escalate through every model. Missing access/data and tool failures need diagnosis, not a stronger model. After two same-class failures, pause that slice and return the evidence to the parent.
|
||||
|
||||
Use the lowest adequate supported effort. Use xhigh/max for a concrete depth need or an explicit compatible role requirement; ultra only when exposed and justified by the workload. Model choice and working role are separate. A role named "explorer" is not proof of sandbox isolation.
|
||||
|
||||
At dispatch, specify model AND effort when the host permits selection. Otherwise omitted settings may inherit the parent or configured defaults. Prefer a self-contained contract and no history fork; use bounded history only when necessary. Respect full-history/override incompatibilities. Never claim prompt text changed a runtime parameter.
|
||||
|
||||
## Check capabilities once per unchanged host
|
||||
|
||||
Inspect the actual child-tool schema: models, per-model efforts, role bindings, defaults, history rules, slot counting, workspace isolation and controls. API catalogs, config files and task-creation tools cannot prove child availability. Recheck only after relevant changes.
|
||||
|
||||
Use the first authorized, low-risk real child task as the capability observation; for unfamiliar Luna multi-step work, start with a useful bounded read-only slice. A returned model name or marker is not identity proof. Record requested settings separately from host-observed identity; missing fields are unknown. Output can pass while identity stays unknown, unless the task explicitly requires verified identity.
|
||||
|
||||
Treat workspaces as shared unless isolation is confirmed. Tool schemas may lack per-child sandbox, timeout, role or close controls. Contract limits remain instructions, not enforced capabilities. No adequate child route: keep feasible work local and disclose the gap.
|
||||
|
||||
Read [runtime and configuration](references/platforms.md) only for configuration/CLI diagnosis. Read [routing boundaries](references/routing-matrix.md) for ambiguous choices or live artifact collection.
|
||||
|
||||
## Dispatch, observe, accept
|
||||
|
||||
Give a compact contract: goal/symptom; baseline and non-goals; read/write/forbidden paths and semantic owner; requested model/effort/working role and real isolation; permitted actions; acceptance; return evidence; escalation. Include a time/checkpoint budget when useful, but do not call it an enforced timeout unless the host supplies one.
|
||||
|
||||
Read [lifecycle](references/lifecycle.md) before a multi-round or cancellation-sensitive run. Keep child IDs, ownership, state and acceptance in a small ledger. Send running agents only incremental context; reuse idle agents for follow-up; interrupt obsolete work and verify its state. Never infer completion from a wait timeout or slot release from interruption. Only call close if that tool exists.
|
||||
|
||||
Ask for bounded facts, paths, changes, real check outcomes, gaps and escalation. Review artifacts/diffs and provenance; resolve conflicts from original evidence, not votes. Reuse checks only for the same relevant baseline, artifact and environment. Before delivery verify no unneeded child remains running or writing.
|
||||
|
||||
Use [evidence packets](references/evidence-packet.md) when structured evidence is required, a run is material, or identity is disputed. Ordinary bounded work can return concise prose with the same relevant facts; do not generate ceremonial JSON for every local decision. The validator checks structure, not truth or semantic correctness.
|
||||
|
||||
## Context and ownership discipline
|
||||
|
||||
Load only needed references; keep long logs on disk. Avoid whole-history forks and repeating contracts, packets or prior findings. Account for parent context/reasoning, child usage, waiting, retries and integration when evaluating cost.
|
||||
|
||||
Production read-only collection is routable by work shape; mutation and final semantic acceptance stay parent-owned. Do not silently relax a project's literal model-owner requirement. Identify the actual rule and resolve it only through authorized changes.
|
||||
@@ -0,0 +1,4 @@
|
||||
interface:
|
||||
display_name: "Codex Subagent Router"
|
||||
short_description: "Route models and manage bounded Codex subagents"
|
||||
default_prompt: "Use $codex-subagent-router to delegate useful independent parts of this task across supported models, manage their lifecycle, and verify the results."
|
||||
@@ -0,0 +1,47 @@
|
||||
# Evidence Packet v2
|
||||
|
||||
Use structured packets for material runs, identity disputes or an explicit consumer requirement. Concise prose is sufficient for ordinary work when it preserves the relevant evidence.
|
||||
|
||||
## Contract
|
||||
|
||||
Required keys: schema_version (2), execution (not_run/local/child), goal, scope, allowed_paths (read/write/forbidden arrays), tool_receipts, capability_source, fallback_reason, evidence, changes, verification, risks, escalate.
|
||||
|
||||
- capability_source describes where capability evidence came from: host-tool-schema, child-rollout, profile-probe, fallback or unknown. It does not establish identity or acceptance.
|
||||
- fallback requires a nonempty fallback_reason; otherwise use null unless a nonpreferred route needs explaining.
|
||||
- verification contains commands and a boolean deterministic. Each command has command, exit_code (integer or null), and status (passed/failed/not_run/unknown). A passed command requires exit code 0.
|
||||
- escalate contains required; when true, target and reason must be nonempty.
|
||||
- execution=not_run permits only schema receipts, no claimed changes or executed commands. A schema inspection can precede dispatch without inventing a child.
|
||||
- A child receipt requires kind=child, a real tool name and child_id, observed state, requested and observed identity objects, and identity_source. Both identity objects contain model, role and effort; use unknown for unavailable fields.
|
||||
- Recognized child states: created, running, idle, completed, interrupted, failed, output-returned. Map only when supported by actual host evidence; preserve the original state in an extra field when necessary.
|
||||
- Non-child receipts have kind=schema or local and cannot contain child_id.
|
||||
- identity_source is host-tool-result or unknown. Any known observed identity requires identity_evidence pointing to the actual host result/event. Requested parameters and child self-description cannot serve as that evidence.
|
||||
- Runtime identity can remain unknown even if the child completed and output passed. Do not relabel a real rollout as merely a schema observation because identity fields are absent.
|
||||
- Parent acceptance remains separate; an optional acceptance object may record the parent's checks, status and unresolved requirements.
|
||||
|
||||
See the repository's examples/ directory for synthetic v2 and legacy v1 fixtures. Fixture IDs are examples only, never actual receipts.
|
||||
|
||||
## Unknown identity example
|
||||
|
||||
This fragment illustrates shape; it is not a real tool receipt:
|
||||
|
||||
```json
|
||||
{
|
||||
"kind": "child",
|
||||
"tool": "collaboration.spawn_agent",
|
||||
"child_id": "synthetic-child",
|
||||
"status": "completed",
|
||||
"requested": {"model": "gpt-5.6-terra", "role": "explorer", "effort": "medium"},
|
||||
"observed": {"model": "unknown", "role": "unknown", "effort": "unknown"},
|
||||
"identity_source": "unknown"
|
||||
}
|
||||
```
|
||||
|
||||
The working role requested in a contract does not establish a runtime role. Verify output independently; an explicit requirement to prove the model remains unmet until trustworthy identity evidence exists.
|
||||
|
||||
## Migration
|
||||
|
||||
The validator still accepts unversioned/schema_version=1 packets and emits a legacy notice. V1 used flat identity fields on completed child receipts; preserve historical records rather than inventing v2 observations from them.
|
||||
|
||||
For new runs use v2. Reconstruct observed identity only from retained host evidence; otherwise use unknown. The stricter v2 command/receipt checks are intentionally not backfilled into old event histories.
|
||||
|
||||
Run scripts/validate_evidence_packet.py PACKET.json (or - for stdin). Exit 0 means structural acceptance only. The --self-test option is a small parser smoke check; repository unit tests provide broader coverage. No validator can authenticate tool receipts or replace parent correctness judgment.
|
||||
@@ -0,0 +1,33 @@
|
||||
# Child Lifecycle
|
||||
|
||||
Read before multi-round coordination, cancellation, or capacity diagnosis. Names below describe control intent; use only the tools and states the host actually exposes.
|
||||
|
||||
| Situation | Action | Evidence needed |
|
||||
| --- | --- | --- |
|
||||
| Independent authorized slice ready | Spawn with explicit ownership and supported route | Real child ID; requested settings |
|
||||
| Child running, facts changed | Send a small incremental message | Tool acknowledgement; later child result |
|
||||
| Child idle, more scoped work needed | Follow up with the same child | Reactivation/result; keep prior evidence |
|
||||
| Parent has independent work | Continue it | No unnecessary polling |
|
||||
| Parent depends on child | Bounded wait | New output/state; timeout means only no update |
|
||||
| Cancelled, superseded, overlapping writes | Interrupt affected children, including descendants if needed | Query resulting states; inspect partial writes |
|
||||
| Child returned output | Review evidence, diff and relevant checks | Acceptance separate from runtime completion |
|
||||
| Capacity exhausted | Inspect actual active/open states and reuse eligible children | Host-specific capacity semantics |
|
||||
| Final delivery | Audit tree and assigned files | No unneeded running writer; accepted or explicit gaps |
|
||||
|
||||
A minimal ledger records child ID, P/C classification, file/semantic ownership, dependency, requested route, last observed state and acceptance. Do not invent timestamps, identities or close events.
|
||||
|
||||
## Cancellation and handback
|
||||
|
||||
Interrupting does not undo files or guarantee descendant cancellation. First preserve useful returned evidence, stop affected writers, inspect partial changes, then assign the remaining work to one owner. Never reset shared files wholesale.
|
||||
|
||||
A budget/deadline in a contract is not a runtime timeout. Use host-supported nonblocking waits and control calls; check deadlines at checkpoints. A wait returning no update does not prove failure, completion or a stuck process.
|
||||
|
||||
For missing data, permissions or same-class failures twice, stop the affected slice and report the smallest gap. Resume with a delta when resolved rather than spawning a duplicate investigation.
|
||||
|
||||
## Completion versus closure
|
||||
|
||||
Output returned, runtime completed, parent accepted, thread closed and slot released are separate facts. The host may combine some of them, but never assume that it does.
|
||||
|
||||
If close exists, use it according to host semantics when no longer needed. If it does not, completed/idle is a valid terminal handback; do not require a fictional close to finish. If capacity remains unavailable, reuse a supported idle child or proceed locally.
|
||||
|
||||
Final review must also check descendants and pending writes. If any ongoing work is intentionally retained, report its owner and reason.
|
||||
@@ -0,0 +1,49 @@
|
||||
# Runtime and Configuration
|
||||
|
||||
Use this reference for actual host/configuration diagnosis. This is not a prerequisite checklist for every task.
|
||||
|
||||
## Separate four kinds of evidence
|
||||
|
||||
1. Official model documentation: product and API capabilities.
|
||||
2. Local configuration: requested defaults, providers and role bindings.
|
||||
3. Current child-tool schema: what this session accepts.
|
||||
4. Tool results: what was created, observed, completed or interrupted.
|
||||
|
||||
Do not substitute one for another. A model available in an API or the task picker may be absent from child tools. A CLI success does not verify a Desktop child path.
|
||||
|
||||
Inspect only relevant configuration fields; never dump environment blocks, credential files or HTTP authorization headers. Resolve CODEX_HOME from the environment, falling back to the user's .codex directory; do not overwrite that environment variable.
|
||||
|
||||
## Optional Codex defaults
|
||||
|
||||
For clients that support the documented keys, an example is:
|
||||
|
||||
```toml
|
||||
# Example only; merge intentionally into the existing configuration.
|
||||
model = "gpt-6-astra"
|
||||
model_reasoning_effort = "medium"
|
||||
|
||||
[agents]
|
||||
default_subagent_model = "gpt-5.6-terra"
|
||||
default_subagent_reasoning_effort = "medium"
|
||||
max_concurrent_threads_per_session = 3
|
||||
```
|
||||
|
||||
This is not an automatic router, permission grant or universal host configuration. Never overwrite an existing config to install this example. Preserve user model/provider choices and verify the actual client accepts the keys. Per-agent configuration and host overrides may change resolution.
|
||||
|
||||
The official Codex configuration describes the child-thread limit as excluding the parent. A collaboration tool may expose a different active-slot count including the parent. Use that live tool's counting and release semantics; do not copy either number blindly.
|
||||
|
||||
When selecting a route, explicitly request model and effort with a compatible fork mode if supported. Check omitted-setting inheritance, custom-agent bindings and override restrictions. A prompt cannot implement an unsupported model switch.
|
||||
|
||||
## Diagnosis and safe fallbacks
|
||||
|
||||
Resolve the executable actually used by the relevant client, inspect its help/schema and record its version when needed. Do not prescribe speculative CLI flags or launch model calls merely to print a marker. Use an authorized real task to verify child execution.
|
||||
|
||||
Missing role/sandbox/timeout fields mean working role, read-only scope and deadlines are contract limits only. Keep requested and observed identity distinct. A completed task with unknown model identity can be accepted for its output, but cannot satisfy an explicit identity-verification gate.
|
||||
|
||||
No close tool: retain the truthful idle/completed/interrupted state and follow the host's actual capacity behavior. No child tools: continue locally; do not launch another user-owned task or nested CLI as an unapproved substitute.
|
||||
|
||||
Configuration changes may require a fresh session or explicit reload. Writing a file is not evidence that the current session adopted it.
|
||||
|
||||
## Sources
|
||||
|
||||
Checked 2026-09-14: [Codex subagents](https://learn.chatgpt.com/zh-Hans/docs/agent-configuration/subagents), [OpenAI model catalog](https://developers.openai.com/api/docs/models). Recheck changed product/API facts before altering real configuration.
|
||||
@@ -0,0 +1,33 @@
|
||||
# Routing Boundaries
|
||||
|
||||
Read only when the entrypoint leaves a routing choice unresolved.
|
||||
|
||||
| Boundary | Decision |
|
||||
| --- | --- |
|
||||
| Many easy tasks | Batch bounded Luna/Terra slices; volume alone does not justify Astra. |
|
||||
| Ordinary cross-file bug | Terra; file count alone does not justify escalation. |
|
||||
| Difficult accounting discrepancy | Sol reasoning; parent owns the semantic decision. |
|
||||
| Coupled architecture across runtimes and tools | Consider Astra directly; keep shared decisions serial. |
|
||||
| Missing credential or inaccessible source | Resolve the evidence/environment gap; model escalation cannot supply access. |
|
||||
| Astra advertised only for task creation | Child support remains unproven; do not create another user task as a workaround. |
|
||||
| Named role fixes model/effort | Accept its supported binding or choose an overridable role; never attach prohibited fork overrides. |
|
||||
| Repository requires Sol parent | A Sol child cannot replace that owner; preserve the gate until an authorized supported change. |
|
||||
| Accepted output, unknown child model | Accept only after normal parent checks; do not attribute the result to the requested model. |
|
||||
|
||||
## Production Artifact Collection
|
||||
|
||||
- Luna: use known hosts/endpoints/paths and fixed selectors to read or download actual reports, JSON, logs or job outputs; verify size/hash, parseability, required fields and specified counts. Terra: locate the current output through bounded read-only investigation, diagnose retrieval failures, or correlate logs and artifacts by run/version. Do not launch both by default; use Terra only when discovery or diagnosis is needed.
|
||||
- Contract: name the production source, expected run/date/version, allowed read operations, bounded time/range/volume, local destination and acceptance checks. Reuse approved access and applicable server tools. Keep secrets out of prompts/receipts. Remote access stays read-only; explicitly allow local artifact writes even for a collector role.
|
||||
- Evidence: preserve the actual downloaded artifact and record source, retrieval time/timezone, producer run/version and generation time when exposed, local path, bytes/hash and check results. Unknown provenance stays unknown. Verify the expected run/freshness; mtime or successful download alone does not prove current output. Use immutable run identifiers where possible; detect/retry boundedly if files change during collection.
|
||||
- No fixture, cache or old snapshot may stand in for requested live output. If the current artifact is missing, stale, inconsistent or inaccessible, report that gap. Do not trigger jobs, regenerate artifacts, restart services, change permissions/configuration or deploy as part of collection; return such remediation to the parent under existing authorization.
|
||||
- P only for independent reads and disjoint local outputs; shared mutable SSH sessions or a changing multi-file output set require coordination/serialization. Parent reviews provenance, relevant artifact content and check results before semantic/final acceptance. Fetching is evidence collection, not production acceptance.
|
||||
|
||||
## Cost Calibration
|
||||
|
||||
Minimize total task usage: parent input/reasoning/output + all child input/reasoning/output + retries and reintegration. A smaller model can reduce price without reducing token count; additional agents can increase both. Do not assume pricing ratios or invent token measurements.
|
||||
|
||||
Use actual usage counters when exposed. Otherwise report observable proxies (characters loaded, number of children, full-history forks, repeated reads/checks) with their limitations. Record only available data needed for a routing decision; do not add telemetry work to every small task. Compare similar accepted tasks before changing defaults.
|
||||
|
||||
If context dominates, narrow inputs and avoid full-history forks. If output dominates, bound logs and returns. If reasoning/retries dominate, clarify acceptance and choose adequate capability/effort directly. Preserve required tests and material findings.
|
||||
|
||||
Official model positioning checked 2026-09-14: [Models](https://developers.openai.com/api/docs/models). The route table is an engineering starting strategy, not a measured cost ranking. Runtime tool schemas remain the source for actual child availability.
|
||||
@@ -0,0 +1,191 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Validate evidence structure, never runtime identity or semantic correctness."""
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import sys
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
REQUIRED = {
|
||||
"goal", "scope", "allowed_paths", "tool_receipts", "capability_source",
|
||||
"fallback_reason", "evidence", "changes", "verification", "risks", "escalate",
|
||||
}
|
||||
SOURCES = {"host-tool-schema", "child-rollout", "profile-probe", "fallback", "unknown"}
|
||||
STATES = {"created", "running", "idle", "completed", "interrupted", "failed", "output-returned"}
|
||||
|
||||
|
||||
def require(condition: bool, message: str) -> None:
|
||||
if not condition:
|
||||
raise ValueError(message)
|
||||
|
||||
|
||||
def nonempty(value: Any) -> bool:
|
||||
return isinstance(value, str) and bool(value.strip())
|
||||
|
||||
|
||||
def string_list(value: Any) -> bool:
|
||||
return isinstance(value, list) and all(nonempty(item) for item in value)
|
||||
|
||||
|
||||
def spawn_tool(value: str) -> bool:
|
||||
return value.split(".")[-1] == "spawn_agent"
|
||||
|
||||
|
||||
def validate(packet: Any) -> None:
|
||||
require(isinstance(packet, dict), "packet must be a JSON object")
|
||||
missing = REQUIRED - packet.keys()
|
||||
require(not missing, "missing fields: " + ", ".join(sorted(missing)))
|
||||
version = packet.get("schema_version", 1)
|
||||
require(type(version) is int and version in (1, 2), "schema_version must be 1 or 2")
|
||||
require(nonempty(packet["goal"]), "goal must be a non-empty string")
|
||||
require(isinstance(packet["scope"], dict), "scope must be an object")
|
||||
paths = packet["allowed_paths"]
|
||||
require(isinstance(paths, dict), "allowed_paths must be an object")
|
||||
for key in ("read", "write", "forbidden"):
|
||||
require(string_list(paths.get(key)), "allowed_paths." + key + " must be a string array")
|
||||
receipts = packet["tool_receipts"]
|
||||
require(isinstance(receipts, list), "tool_receipts must be an array")
|
||||
source = packet["capability_source"]
|
||||
require(isinstance(source, str) and source in SOURCES, "unsupported capability_source")
|
||||
fallback = packet["fallback_reason"]
|
||||
require(fallback is None or nonempty(fallback), "fallback_reason must be non-empty or null")
|
||||
require(source != "fallback" or nonempty(fallback), "fallback requires fallback_reason")
|
||||
require(isinstance(packet["evidence"], (dict, list)), "evidence must be an object or array")
|
||||
require(isinstance(packet["changes"], list), "changes must be an array")
|
||||
require(isinstance(packet["risks"], list), "risks must be an array")
|
||||
verification = packet["verification"]
|
||||
require(isinstance(verification, dict), "verification must be an object")
|
||||
require(isinstance(verification.get("commands"), list), "verification.commands must be an array")
|
||||
if version == 2:
|
||||
require(type(verification.get("deterministic")) is bool, "verification.deterministic must be boolean")
|
||||
for command in verification["commands"]:
|
||||
require(isinstance(command, dict), "each command must be an object")
|
||||
require(nonempty(command.get("command")), "command requires command text")
|
||||
require("exit_code" in command, "command requires exit_code (null if unknown)")
|
||||
code = command["exit_code"]
|
||||
require(code is None or type(code) is int, "exit_code must be integer or null")
|
||||
status = command.get("status")
|
||||
require(isinstance(status, str) and status in {"passed", "failed", "not_run", "unknown"},
|
||||
"command requires supported status")
|
||||
require(status != "passed" or code == 0, "passed command requires exit_code 0")
|
||||
require(status != "failed" or code is None or code != 0, "failed command contradicts exit_code 0")
|
||||
require(status != "not_run" or code is None, "not_run command cannot have exit_code")
|
||||
escalate = packet["escalate"]
|
||||
require(isinstance(escalate, dict) and type(escalate.get("required")) is bool,
|
||||
"escalate.required must be boolean")
|
||||
if escalate["required"]:
|
||||
require(nonempty(escalate.get("target")) and nonempty(escalate.get("reason")),
|
||||
"required escalation needs target and reason")
|
||||
for receipt in receipts:
|
||||
require(isinstance(receipt, dict), "each tool receipt must be an object")
|
||||
require(nonempty(receipt.get("tool")) and nonempty(receipt.get("status")),
|
||||
"tool receipt requires non-empty tool/status")
|
||||
if version == 1:
|
||||
_legacy(packet)
|
||||
else:
|
||||
_v2(packet)
|
||||
|
||||
|
||||
def _legacy(packet: dict) -> None:
|
||||
"""Retain unversioned v1 packets without interpreting them as v2 evidence."""
|
||||
receipts = packet["tool_receipts"]
|
||||
require(bool(receipts) or packet["capability_source"] == "fallback",
|
||||
"legacy empty tool_receipts require fallback")
|
||||
completed = []
|
||||
for receipt in receipts:
|
||||
if spawn_tool(receipt["tool"]) and receipt["status"] == "completed":
|
||||
for field in ("child_id", "model", "role", "effort"):
|
||||
require(nonempty(receipt.get(field)), "legacy completed receipt requires " + field)
|
||||
completed.append(receipt)
|
||||
require(packet["capability_source"] != "child-rollout" or bool(completed),
|
||||
"legacy child-rollout requires a completed child receipt")
|
||||
|
||||
|
||||
def _v2(packet: dict) -> None:
|
||||
execution = packet.get("execution")
|
||||
require(isinstance(execution, str) and execution in {"not_run", "local", "child"},
|
||||
"execution must be not_run, local or child")
|
||||
children = []
|
||||
for receipt in packet["tool_receipts"]:
|
||||
kind = receipt.get("kind")
|
||||
require(isinstance(kind, str) and kind in {"schema", "local", "child"}, "receipt requires kind")
|
||||
if kind != "child":
|
||||
require("child_id" not in receipt, "non-child receipt cannot claim child_id")
|
||||
require(not spawn_tool(receipt["tool"]) or kind == "schema",
|
||||
"spawn receipt must be child or schema inspection")
|
||||
continue
|
||||
children.append(receipt)
|
||||
require(nonempty(receipt.get("child_id")), "child receipt requires actual child_id")
|
||||
require(receipt["status"] in STATES, "unsupported child state")
|
||||
for section in ("requested", "observed"):
|
||||
identity = receipt.get(section)
|
||||
require(isinstance(identity, dict), "child receipt requires " + section)
|
||||
for key in ("model", "role", "effort"):
|
||||
require(nonempty(identity.get(key)), section + "." + key + " required; use unknown")
|
||||
origin = receipt.get("identity_source")
|
||||
require(isinstance(origin, str) and origin in {"host-tool-result", "unknown"},
|
||||
"identity_source must be host-tool-result or unknown")
|
||||
known = any(value != "unknown" for key, value in receipt["observed"].items()
|
||||
if key in {"model", "role", "effort"})
|
||||
if known:
|
||||
require(origin == "host-tool-result" and nonempty(receipt.get("identity_evidence")),
|
||||
"known observed identity requires host evidence")
|
||||
elif origin == "unknown":
|
||||
require(not receipt.get("identity_evidence"), "unknown identity cannot claim identity evidence")
|
||||
require((execution == "child") == bool(children), "execution must agree with child receipts")
|
||||
require(packet["capability_source"] != "child-rollout" or bool(children),
|
||||
"child-rollout requires a real child receipt")
|
||||
if execution == "not_run":
|
||||
require(all(r["kind"] == "schema" for r in packet["tool_receipts"]),
|
||||
"not_run may contain schema inspection only")
|
||||
require(not packet["changes"], "not_run cannot claim changes")
|
||||
require(all(c["status"] == "not_run" for c in packet["verification"]["commands"]),
|
||||
"not_run cannot claim executed commands")
|
||||
|
||||
|
||||
def self_test() -> None:
|
||||
packet = {
|
||||
"schema_version": 2, "execution": "not_run", "goal": "schema smoke check",
|
||||
"scope": {}, "allowed_paths": {"read": [], "write": [], "forbidden": []},
|
||||
"tool_receipts": [], "capability_source": "unknown", "fallback_reason": None,
|
||||
"evidence": {}, "changes": [], "verification": {"commands": [], "deterministic": True},
|
||||
"risks": [], "escalate": {"required": False},
|
||||
}
|
||||
validate(packet)
|
||||
packet["execution"] = "child"
|
||||
try:
|
||||
validate(packet)
|
||||
except ValueError:
|
||||
return
|
||||
raise AssertionError("child execution without receipts was accepted")
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("packet", nargs="?", help="JSON path, or - for stdin")
|
||||
parser.add_argument("--self-test", action="store_true")
|
||||
args = parser.parse_args()
|
||||
if not args.self_test and not args.packet:
|
||||
parser.error("provide a packet path or --self-test")
|
||||
try:
|
||||
if args.self_test:
|
||||
self_test()
|
||||
print("[evidence-packet] structural self-test ok")
|
||||
return 0
|
||||
raw = sys.stdin.read() if args.packet == "-" else Path(args.packet).read_text(encoding="utf-8")
|
||||
packet = json.loads(raw)
|
||||
validate(packet)
|
||||
version = packet.get("schema_version", 1)
|
||||
if version == 1:
|
||||
print("[evidence-packet] legacy v1; migrate new receipts to v2", file=sys.stderr)
|
||||
print("[evidence-packet] structure ok; identity and correctness not verified")
|
||||
return 0
|
||||
except (OSError, ValueError) as error:
|
||||
print("[evidence-packet] ERROR " + str(error), file=sys.stderr)
|
||||
return 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
Reference in New Issue
Block a user