5.7 KiB
Proactive fleet scheduling
Use for authorized complex parallel work. This is a decision procedure for the parent, not a background scheduler, permission grant or capacity override. Preserve the actual host schema, selected parent and final gates.
Discover and dispatch early
After minimum goal/baseline discovery, identify independent questions and useful implementation slices. Do not finish reading every subsystem before dispatching ready investigations. Once a slice's input contract is stable, it may enter implementation while unrelated investigation continues. Shared decisions stay with the parent; preliminary independent drafts are not accepted contracts.
Maintain a compact queue: task ID, dependency and contract baseline, file AND semantic owner, shared resources, acceptance, requested route, child ID, observed runtime state and integration state. Mark pending, ready, running, returned, accepted, blocked or superseded as task states; keep actual host state separate. Ready means the prerequisites, ownership and permitted actions are established and the output has independent value.
At intake, dependency resolution, a child return or a meaningful baseline change:
- Invalidate affected assumptions and identify ready slices. Prioritize work that unlocks dependencies, avoids costly wrong decisions or yields directly usable implementation.
- Check live capacity and integration load. Dispatch multiple independent ready tasks when useful parent work remains; avoid sequential startup followed by unnecessary waits. Every descendant counts against the actual host limit. Children may not recursively expand the fleet without a parent-assigned bounded delegation scope and capacity allocation.
- Continue the parent's independent critical work. Use incremental messages for running children and reuse suitable idle ones. Do not perform the same assigned investigation locally merely to stay busy.
- Inspect returned evidence and partial writes; integrate serially. Refill available capacity with ready work without waiting for all siblings. A returned result is not necessarily accepted or a released slot; follow actual lifecycle controls.
If the next action is immediately dependent on an unassigned task, keep it local. A previously independent task may become a dependency as work progresses; bounded waiting is then legitimate. No independent useful parent work means do not create a ceremonial side task to justify another spawn.
Choose a formation for the task
- Independent investigation: several focused investigators may cover distinct sources, consumers or falsifiable hypotheses. They should not all produce the same broad summary.
- Settled implementation: multiple Terra/Luna workers own complete independently checkable slices. Shared registries, schema and generated outputs require one owner even when paths differ.
- Difficult consequential decision: an Astra specialist analyzes the bounded uncertainty while the parent and other workers advance unaffected work. Gate dependent implementation on the parent's accepted contract. Read expert intervention for triggers and progress-based continuation.
- Independent review: assign a stable commit/artifact and the relevant requirements, without supplying the parent's preferred verdict. Review can overlap unrelated work; stale observations must be rechecked before acceptance.
There is no mandatory model ratio or permanent reviewer slot. Use several workers of the same adequate model when the tasks warrant it. If Astra is already the selected parent, do not add another Astra just to fulfill a role label. Reuse the parent capability unless a genuinely independent bounded perspective adds value.
Backpressure and recovery
Choose a working width up to actual available capacity based on ready independent work, parent integration capacity and shared-resource limits. Do not reserve slots mechanically or fill them with artificial work. Hardware, service rate limits and tool sessions can constrain useful width below the agent limit.
When returned patches are accumulating faster than the parent can inspect, or a returned contract blocks other slices, pause new dependent implementation and prioritize integration. When collisions, stale baselines, repeated rework or resource contention appear, stop affected dispatch and reduce width; interrupt obsolete/conflicting writers and inspect partial changes. Independent read-only work may continue if useful and resource-safe. Resume expansion after the cause is resolved and ready work remains.
Never run acceptance tests against a moving artifact and report them as final. Pin a commit/snapshot or wait until relevant writers stop; a worktree alone does not isolate external databases, generated directories or shared sessions. Before merge/delivery, resolve material counterexamples, verify the integrated baseline and audit the agent tree for remaining writers.
Evaluate effectiveness
Optimize accepted quality and end-to-end time subject to the user's cost constraints. Include parent work, all children, context duplication, retries, expert rounds and integration in cost measurements. Expensive early reasoning can be worthwhile when it prevents downstream rework; cheap high-volume workers can be wasteful when contracts are unsettled. Both are hypotheses until measured.
For a calibration exercise compare the same tasks and baselines at serial and increasing supported widths, varying granularity separately. Include decomposable and contract-heavy tasks, failures and repeated runs. Record quality, elapsed time, total exposed usage/cost, rework, and integration time; unknown counters or runtime model identity stay unknown. Do not claim speed/cost improvements from static scenarios or a successful single rollout.