返回 CodeWhale
SUBAGENTS.md
根目录 / docs / SUBAGENTS.md
1 # Fleet and sub-agents
2
3 > 阅读简体中文版:[zh_hans/SUBAGENTS.md](zh_hans/SUBAGENTS.md)
4
5 Fleet manages saved models and role assignments for these same sub-agents.
6 Use `agent` for an individual assignment and `workflow` for phases with
7 dependencies and completion checks. See [Workflow authoring](WORKFLOW_AUTHORING.md)
8 for plans that use the Fleet model shortlist.
9
10 Fleet roles are the user-facing vocabulary for delegated work: a parent
11 launches a focused `general`, `explore`, `planner`, `reviewer`, `implement`,
12 `test`, or `advisor` through `agent` and gets back an `agent_id`, declared
13 deliverables, and effective limits while the worker runs. The default receipt is
14 compact; request addressed detail when you need the transcript handle or ledger.
15 The internal runtime type is `FleetRole` (formerly
16 `SubAgentType`); the older role spellings (`worker`, `scout`, `plan`,
17 `review`, `builder`, `verifier`, `consultant`, `oracle`, …) remain accepted only as a persisted/deserialize
18 compatibility adapter during v0.9.x. New prompts and config should use fleet
19 names.
20
21 Child assignments run through the same `Engine::run_turn` planner, executor,
22 Session and approval inbox as ordinary turns. The private captured ChildGrant
23 narrows tool actions, shell, workspace, network and descendant depth at the final
24 execution boundary. The existing SubAgentManager still owns launch slots,
25 coordination, worker checkpoints and terminal delivery; it does not run another
26 model/tool loop. Work and the single bounded reporting turn share that configured
27 child Engine and its original deadline.
28
29 Transient provider failures use Core's dispatch and retry boundaries. An exhausted
30 budget or interrupted request preserves the recorded checkpoint and a continuation
31 handle. Each actually dispatched attempt carries its original session, route and
32 source identity: provider usage is billed once, then projected to the worker
33 ledger. A successful response without usage and an unknown request outcome retain
34 distinct coverage gaps; neither is priced as zero. For work that must survive
35 process restarts, sleep or remote execution, use Fleet or a Workflow-backed run.
36
37 Sub-agents inherit the parent's permitted tool registry, including `agent`
38 coordination. Spawning obeys one absolute depth ceiling: the root is depth 0,
39 its child is depth 1, and a child at `max_spawn_depth` cannot spawn again.
40 The operator default is 3, with a hard ceiling of 8. A role, saved profile, or
41 compatibility request can only narrow that ceiling. Recovery and transcript
42 forking retain the source's position and bounds; they do not buy another
43 generation. The removed `agent_open`/`agent_eval`/`agent_close` lifecycle tools
44 are absent from every registry.
45
46 Healthy children continue after an ordinary parent response. Their completion
47 returns through the existing Engine inbox and can wake the parent for another
48 normal turn. Explicit interruption or cancellation remains authoritative.
49 `detached: true` additionally opts a subtree out of parent-turn cancellation;
50 it does not remove child budgets or the headless host's deadline.
51
52 This doc covers roles and individual worker controls. Use `workflow` to coordinate
53 multiple assignments through the same worker runtime; see the per-role prompts in
54 `crates/tui/src/tools/subagent/mod.rs` (`*_AGENT_INTRO`) and the in-line
55 tool description.
56
57 ## Role taxonomy
58
59 The `type` field on `agent` selects a fleet posture for the child
60 (`agent_type` is accepted as a compatibility alias). Each role is a distinct
61 stance toward the work — not just a different label.
62
63 ## Maintainer posture
64
65 Sub-agents help Codewhale move faster, but the parent agent still owns the
66 maintainer decision. Use children to gather evidence, review patches, and run
67 verification while keeping the community posture in
68 [`AGENT_ETHOS.md`](AGENT_ETHOS.md): issues are open intake, PR gates are
69 review-load controls, and harvested work needs clear contributor credit.
70
71 When a child reviews community work, the parent should still inspect the PR
72 diff, linked issues, tests, and CI before merging, harvesting, closing, or
73 deferring it. A sub-agent's result is a working set, not a substitute for
74 stewardship.
75
76 | Role | Stance | Writes? | Network? | Shell posture | Typical use |
77 |---------------|----------------------------------------|---------|----------|---------------|----------------------------------------------|
78 | `general` | flexible; do whatever the parent says | yes | yes | yes | the default; multi-step tasks |
79 | `explore` | read-only; map the relevant code fast | no | yes | bounded inspection | "find every call site of `Foo`; check the PR with gh" |
80 | `planner` | analyse and produce a strategy | no | yes | read-only probes | "design the migration; don't execute" |
81 | `reviewer` | read-and-grade with severity scores | no | yes | bounded inspection | "audit this PR for bugs" |
82 | `implement` | land a specific change with min edit | yes | yes | yes | "rewrite `bar.rs::Foo::bar` to do X" |
83 | `test` | run tests / validation, report outcome | no | yes | bounded verification (no writes) | "verify the diff with the bounded test checks; report PASS/FAIL" |
84 | `advisor` | short-lived, high-reasoning counsel | no | yes | none | "what are we missing in this design?" |
85 | `custom` | explicit narrow tool allowlist | inherits | inherits | inherits | hand-picked tools on the parent's posture |
86
87 A role's default is what the role *intends*, and the parent's effective
88 posture is always the ceiling (a child never widens beyond its parent).
89 Read-only roles withhold **workspace writes** by intent; nothing else is
90 taken away by default — every role keeps network reads, and `custom`
91 inherits the parent's write/network/shell posture and is narrowed only by
92 its explicit tool list or the spawning call. The focused worker's header
93 states the effective posture (`scout · read-only · network · read-only
94 shell`) from the runtime's own permission snapshot.
95
96 **Delegation moves work, never authority.** A read-only parent may delegate
97 to `implement`, but the child's effective write, network, shell, and tool
98 permissions remain within the parent's live posture. Inspection roles can use
99 the classified read-only shell surface and, where native enforcement is
100 available, the explicit read-only analysis mode described below. A different
101 role name or `read_only` flag cannot grant a shell tool the caller lacks.
102 The clamp (`ChildAuthority::clamp` in `fleet/exact.rs`) intersects every field
103 with the narrower side. Deny lists are unioned, so
104 `inherit_disallowed_tools: false` cannot drop any operator or ancestor denial.
105 Resuming a saved worker intersects its saved posture with the current caller's
106 posture again. This containment is pinned by
107 `a_read_only_parents_delegation_never_widens_authority` in
108 `crates/tui/src/fleet/exact.rs` tests.
109
110 Inside the process, the resolved authority is one object —
111 `ChildGrant` in `crates/tui/src/worker_profile.rs`: `files`
112 (none/read/write), `shell` (none/inspect/verify/full), `network`, `desktop`
113 (never granted to a child), the named tool `surface`, the caller's explicit
114 `scope`, and remaining `spawn` depth. A role is a preset over that object
115 (`ChildGrant::for_role`); `ChildGrant::resolve` intersects it with the
116 parent-derived profile. The child's tool catalog, its dispatch refusals, and
117 its capability envelope all read the same fields — a tool that is visible is
118 callable, and a tool that is denied never appears.
119
120 The session's **permission posture** applies inside every child exactly as
121 it applies to the parent turn: under Auto-Review the same deterministic
122 floor and one-shot model guardian decide a worker's held calls (never a
123 prompt; an unavailable guardian denies, fail closed); under Ask a held call
124 the role cannot delegate is raised as an approval prompt in the parent's
125 UI and the worker waits visibly (`waiting for user`), or is denied with the
126 reason on hosts that cannot prompt; Full Access still fails closed on the
127 non-bypassable safety floor. Each decision nobody was prompted for is a
128 one-line note in that worker's transcript (visible when it is focused) and
129 an audit-log record. See `docs/MODES.md`.
130
131 Each role's full system prompt lives in
132 `crates/tui/src/tools/subagent/mod.rs` (search for
133 `*_AGENT_INTRO`). The prompt prefix loads automatically when the
134 child agent boots; the parent's assignment prompt becomes the first
135 turn's user message.
136
137 ## Context forking
138
139 `agent` starts fresh by default: the child gets its role prompt plus the
140 task you pass. Use `fork_context: true` when the child should continue from
141 the parent's current request prefix instead. (`fork_context` is not in the
142 advertised schema — it stays parse-accepted for compat callers, and
143 auto-forking for read-only roles continues unchanged.) In fork mode the runtime keeps the
144 parent prefill/prompt prefix byte-identical where available, appends a
145 structured state snapshot, then adds the sub-agent role instructions and task
146 at the tail. That preserves DeepSeek prefix-cache reuse while giving the child
147 the context needed for continuation, review, summarization, or compaction work.
148
149 Use fresh sessions for independent exploration. Use forked sessions when the
150 task depends on decisions, files, todos, or plan state already in the parent
151 transcript.
152
153 Forked state shows the parent's To-do snapshot — the sole Work surface, written
154 by `todo_write`. The child's `<codewhale:fork_state>` block carries the bounded
155 body rendered by `crates/tui/src/todo_snapshot.rs`, so a fork continues from the
156 parent's real progress position rather than a paraphrase. That To-do section is
157 resolved when the spawn happens, so a `todo_write` earlier in the same parent
158 turn is included.
159
160 **The list is shown once, at that spawn, and never re-sent.** No sub-agent
161 request re-states a To-do list, and neither does a parent request. Each agent
162 keeps its own private list (#4810); what it knows about that list comes from the
163 tool results its own `todo_write` calls returned, which are ordinary messages in
164 its own transcript. A worker therefore cannot read or write a parent's or a
165 sibling's list, and a forked child cannot mutate the snapshot it was handed or
166 keep reading later parent changes.
167
168 That same private list is what the child's in-transcript card shows. A
169 delegate card renders a bounded projection of **its own** agent's To-do — the
170 settled/total count, the in-progress item always included, up to three rows, and an
171 explicit `… +N more` when the bound elides the rest — built by
172 `card_todo_projection` from the same snapshot, priority order, and sanitizer the
173 model-facing body uses. A card only ever consumes an envelope whose `agent_id`
174 matches it, so a parent's list never appears under a child and no sibling's list
175 appears under another. An agent that has stated no work shows no To-do rows at
176 all rather than a placeholder task, and a terminal card keeps the last snapshot
177 its agent actually published. Fanout cards stay a dot grid and do not show child
178 To-do: with many workers behind one card there is no truthful place to hang a
179 single list. A child To-do appears only when the runtime already represents
180 that child as its own delegate card.
181
182 The durable Runtime ledger (projected through fleet task status) still owns
183 lifecycle state. `update_plan` is no
184 longer reachable by a model: `model_visible()` returns `false`
185 (`crates/tui/src/tools/plan.rs:408-413`), so it is filtered out of the API tool
186 list and never appears to a child. It survives only to replay older transcripts.
187 Strategy that used to go there now goes in the response body, and lifecycle
188 state goes in `todo_write`.
189
190 ## Worktree isolation
191
192 For parallel edit lanes, launch the child with `worktree: true`. Codewhale
193 creates a fresh git worktree and branch for that child, runs the child from the
194 isolated checkout, and reports the resulting workspace/branch in the returned
195 session projection and worker record. By default the branch is
196 `codex/agent-<name>-<id>` and the checkout lives beside the parent repo under
197 `.codewhale-worktrees/`, so the parent checkout stays clean.
198
199 Isolation is not write authority. A prompt-only start with no role/profile or
200 write declaration remains read-only, and read-only roles need no write scope.
201 Explicitly selected write-capable roles such as `general` and `implement`
202 inherit the parent's write ceiling and default to the workspace
203 (`write_roots: ["."]`) unless narrowed. Prefer explicit, disjoint `exact_files`
204 or `write_roots` for parallel work; `coordination_contracts` can reserve named
205 shared contracts. If only `deliverables` supplies a writer's scope, those files
206 become the exact-file scope.
207
208 `write_authority` is optional typed narrowing: `read_only` admits no write
209 scope, `workspace_write` uses the shared checkout, and `worktree_write`
210 requires actual worktree isolation. Incompatible role/scope declarations fail
211 before admission. Active overlapping shared claims fail before mutation; a
212 real isolated worktree may proceed in parallel. A `custom` role requires
213 explicit write-capable authority to claim writes; otherwise it starts
214 read-only.
215
216 Read-only is not file-only. A normally write-capable agent narrowed with
217 `write_authority: "read_only"` keeps the existing classifier-bounded inspection
218 shell when its parent permits it: Git history/status, search, and allowed `gh`
219 log reads, not arbitrary commands or test programs. Workflow read-only steps
220 likewise use the effective role's tools instead of imposing a second File-only
221 list; a `test` role retains its bounded verification interface. Explicit tool
222 allowlists, `deny_all_tools`, parent denials, network limits, and mutation checks
223 still apply. A parent without shell access cannot delegate it.
224
225 Optional fields:
226
227 - `worktree_branch`: exact branch to create.
228 - `worktree_base`: git ref to branch from; defaults to `HEAD`.
229 - `worktree_path`: exact checkout path. Relative and absolute paths must stay
230 under the default sibling `.codewhale-worktrees/<repo>/` root (symlinks are
231 resolved before the check). Any worktree request keeps the approval card,
232 even for a read-only role.
233
234 `cwd` may be combined with `worktree`: the requested directory becomes the
235 discovery anchor the repo root (and the new checkout) is resolved from
236 (`prepare_child_workspace`). Without `worktree`, `cwd` remains the manual
237 escape hatch for an already-created directory inside the parent workspace.
238
239 ### File deliverables and edit claims
240
241 Put required files in `deliverables`; keep the human outcome in
242 `expected_artifact`. For example, call `agent` with:
243
244 ```json
245 {
246 "action": "start",
247 "type": "implement",
248 "prompt": "Summarize the local routing evidence in reports/routing.md.",
249 "exact_files": ["reports/routing.md"],
250 "deliverables": ["reports/routing.md"],
251 "expected_artifact": "A concise report with source references and open gaps"
252 }
253 ```
254
255 At most 16 repo-relative file paths are accepted. Absolute paths, traversal,
256 repository metadata paths, and symlink traversal are refused. Completion checks
257 each file against the admitted scope and reports its path, status, and byte
258 count where available. The terminal statuses are `present`, `missing`, `empty`,
259 `not_file`, `out_of_scope`, `invalid_path`, and `unreadable`.
260 `present` means a nonempty regular file exists; it does not prove the report is
261 correct or that tests passed.
262
263 A missing or invalid required file sets `verification.status` to
264 `deliverable_missing` with the individual verdicts. Successful file checks can
265 produce `deliverables_present`; they do not turn a child self-report into an
266 independent quality gate. The completion notice includes the actual verdicts,
267 including when a worker fails or exhausts a budget.
268
269 Edit claims are checked separately against the spawn-time git HEAD and dirty
270 file contents. Explicit changed-file declarations can produce
271 `claim_mismatch` when a claimed file did not change, or when a successful
272 bounded write receipt changed a file the child did not declare. A peer's
273 change inside a worker's broad scope is not enough to attribute that write to
274 the worker. `path:LINE` and `path:LINE-LINE` evidence citations, including
275 sentence punctuation and Markdown links, never count as edit claims.
276
277 ### Read-only shell commands
278
279 Scout, reviewer and planner agents, agents narrowed with
280 `write_authority: "read_only"`, and durable Fleet workers with a read-only
281 shell grant all judge `bash` calls by the same read-only grammar:
282
283 - inspection programs: `ls`, `pwd`, `cat`, `head`, `tail`, `wc`, `which`,
284 `stat`, `file`, `du`, `df`, `grep`, `rg`, `fd`, `find` without `-exec` or
285 `-delete`, and `sed -n <range>p`;
286 - `git status`, `log`, `diff`, `show`, `ls-files`, `blame` and `grep`,
287 optionally after `-C <dir>` or `--no-pager`;
288 - the text filters `sort`, `uniq`, `cut`, `tr` and `comm`, and literal
289 `echo`/`printf`;
290 - with a network grant, `gh` issue/pr/release/repo/run/workflow view or list
291 reads.
292
293 `npm view` is refused even with a network grant: npm configuration can select
294 executable helpers, so package metadata reads are outside this bounded grammar.
295
296 Admitted commands can be joined with `|`, `&&`, `||` and `;`, for example
297 `git diff HEAD && echo '=== FILES ===' && ls -la`. A leading `cd <dir> &&`
298 sets the working directory, and that directory must be inside the workspace.
299 The only redirects are `2>/dev/null`, `>/dev/null` and `2>&1`. Quoted text is
300 data, so `rg 'a && b' src` is one search. Other redirects, `$` or backtick
301 expansion, subshells, backgrounding, inline environment assignments, and any
302 other program (such as `python`, `awk`, `jq` or `cargo`) are refused. Options
303 and path operands are still checked, and each `gh` or `npm` read needs the
304 network grant wherever it appears in the command.
305
306 A refused command comes back to the agent as an error result that names the
307 rule, for example
308 `[shell.readonly.command] program: `touch` is not a read-only inspection command`,
309 followed by what the agent can do instead. The agent keeps working, and
310 repeated refusals without progress end it as failed rather than completed.
311
312 ### Reading beside a writer
313
314 Read-only tools and classifier-approved shell reads can run while a peer owns
315 a shared write claim. For arbitrary analysis code, call `bash` with explicit
316 `read_only: true`:
317
318 ```json
319 {
320 "action": "run",
321 "read_only": true,
322 "command": "python3 -c \"import sqlite3; db = sqlite3.connect('file:cache/index.db?mode=ro', uri=True); print(db.execute('SELECT name FROM sqlite_schema').fetchall())\""
323 }
324 ```
325
326 This mode requires native filesystem read-only isolation and denies network
327 access. It accepts only foreground `run` with `command`, optional `cwd`, and
328 `timeout_ms`. Background or interactive modes, stdin, sandbox escalation, and
329 external execution backends are incompatible. If native enforcement is absent
330 or cannot be prepared, the call refuses before executing the command; the flag
331 never falls back to trusting a promise that the code only reads. Existing role,
332 tool, and ancestor policy restrictions still apply.
333
334 For a write refusal outside your own scope, `agent(action="claim", ...)` can
335 add permitted paths to your claim. It cannot take a live peer's claim. Wait for
336 that peer, choose disjoint bounded writes, or use a separate worktree for code
337 that needs writes. `action="release"` only clears claims whose owners are no
338 longer live; it is not a way to unlock another running worker's files.
339
340 ## Delegation briefs
341
342 The parent should pass a compact brief instead of a loose paragraph. Use the
343 structured `dependencies` and `acceptance` arrays for bounded prerequisite facts
344 and observable checks; keep the focused objective in `prompt`. Do not copy raw
345 parent reasoning or an unbounded transcript.
346
347 Default shape for the brief — a plain sentence beats the template when the
348 delegation is trivial:
349
350 ```
351 QUESTION:
352 SCOPE:
353 ALREADY_KNOWN:
354 EFFORT: quick | medium | thorough
355 STOP_CONDITION:
356 OUTPUT: VERDICT, EVIDENCE, GAPS, NEXT
357 ```
358
359 `scout` briefs default to quick, read-only investigation (no writes, but
360 network reach and the bounded verification surface are available for real
361 scouting). A quick scout usually needs only a handful of calls: orient,
362 search, read the decisive lines, and return — but stop at decisive evidence,
363 not at a number. There is no per-agent call cap to save; the runtime budgets
364 depth and concurrency, not curiosity. Do not repeat `ALREADY_KNOWN` work
365 unless evidence contradicts it. Builder and repair-style briefs should use
366 checkpoints before scope expansion or after repeated failures.
367
368 Good delegation prompt examples:
369
370 ```text
371 QUESTION: Does PR #N introduce release-risk behavior around provider routing?
372 SCOPE: PR #N diff, linked issue, provider routing tests, docs/PROVIDERS.md.
373 ALREADY_KNOWN: Branch is <release-branch>; workspace version is <live-version>.
374 EFFORT: medium
375 STOP_CONDITION: Return once you have either one BLOCKER/MAJOR issue or enough evidence for no MAJOR+ issues.
376 OUTPUT: VERDICT, EVIDENCE with file:line refs or PR refs, GAPS, NEXT.
377 ```
378
379 ```text
380 QUESTION: Where is the child-agent prompt assembled?
381 SCOPE: crates/tui/src/prompts*, crates/tui/src/tools/subagent/*.
382 ALREADY_KNOWN: The model-facing launcher is only `agent`; do not look for removed lifecycle tools.
383 EFFORT: quick
384 STOP_CONDITION: Stop after identifying the prompt source files and the function that wraps assignment text.
385 OUTPUT: VERDICT, EVIDENCE, GAPS, NEXT.
386 ```
387
388 ```text
389 QUESTION: Is the focused prompt/subagent test filter valid, and what fails if not?
390 SCOPE: cargo test -p codewhale-tui --lib --locked prompt; subagent filter if needed.
391 ALREADY_KNOWN: Do not fix failures; capture exact command, exit code, and first relevant assertion.
392 EFFORT: medium
393 STOP_CONDITION: Stop after one clean PASS or one reproducible failing assertion with command evidence.
394 OUTPUT: VERDICT, EVIDENCE, GAPS, NEXT.
395 ```
396
397 ### When to pick which role
398
399 - **`general`** — when the task is "do this whole thing", not "go
400 look", "design", or "verify". This is the right default; reach for
401 a more specific role only when the posture matters.
402 - **`explore`** — when the parent needs evidence before deciding what
403 to do next. Scouts are cheap and fast; open 2–3 in parallel
404 for independent regions.
405 They should orient first: confirm the project root, read relevant
406 `AGENTS.md`/`README.md` guidance in unfamiliar trees, search only the
407 likely scope, and return `path:line-range` evidence instead of a narrative
408 tour. The role name to use is `explore`.
409 - **`planner`** — when the parent has an objective but no executable
410 decomposition. Planners write artifacts (`todo_write` items,
411 strategy in the response body) but don't carry them out.
412 - **`reviewer`** — when there's already a change and the parent wants
413 it graded. Reviewers run under read-only posture, so the runtime
414 refuses patch attempts — describe the fix in the finding and the
415 parent dispatches a builder when the verdict is "fix it".
416 - **`implement`** — when the change is already specified and just
417 needs to land. Builders stay tightly scoped: minimum edit, no
418 drive-by refactoring, run a quick verification before handing back.
419 - **`test`** — when the parent needs an authoritative pass/fail
420 on the test suite or other validation. The verifier posture never
421 writes — the runtime refuses fix attempts — so capture the failing
422 assertion + stack and put fix candidates under RISKS for the parent
423 to dispatch. Shell is clamped to the bounded built-in verification
424 surface: Run tests/verifiers (pass `cwd` when the checks live in a
425 subdirectory), Git fetch for remote refs, Git merge_tree for merge
426 results. The write ceiling is read-only and unbounded shell forms
427 are refused (#5186). A refused probe is reported to the parent,
428 never worked around (#6298).
429 - **`advisor`** — when the operator wants a high-leverage second opinion
430 before cheaper execution continues. Consultants read enough to ground a
431 recommendation; their grant carries no writes and no shell, so the
432 runtime refuses both. `oracle` and
433 `consultant` remain accepted only when loading older requests or persisted
434 records; new prompts, receipts, and UI use `advisor`.
435 - **`custom`** — only when the parent needs to constrain the tool
436 set explicitly. Pass the allowlist via the `allowed_tools` field
437 on legacy/internal sub-agent records; the model-facing `agent` tool keeps the
438 public schema intentionally small.
439
440 ### Aliases
441
442 The model can spell each role multiple ways:
443
444 | Canonical | Aliases |
445 |---------------|------------------------------------------------------------------|
446 | `general` | `worker`, `default`, `general-purpose`, `general_purpose` |
447 | `explore` | `scout`, `explorer`, `exploration` |
448 | `planner` | `plan`, `planning`, `awaiter` |
449 | `reviewer` | `review`, `code-review`, `code_review` |
450 | `implement` | `builder`, `implementer`, `implementation` |
451 | `test` | `verifier`, `verify`, `verification`, `validator`, `tester` |
452 | `advisor` | `consultant`, `oracle` (compatibility input only) |
453 | `custom` | (none; explicit `allowed_tools` array required) |
454
455 All matching is case-insensitive. Unknown values produce a typed
456 error listing the accepted set, so the model can self-correct on
457 the next turn.
458
459 ## Concurrency cap
460
461 Up to **64** sub-agents run concurrently by default (`DEFAULT_MAX_SUBAGENTS`),
462 configurable via `[subagents].max_concurrent` in `~/.codewhale/config.toml` up to
463 the hard ceiling of **128** (`MAX_SUBAGENTS`). The session admits a bounded
464 queue of up to **1024** running plus queued sub-agents by default
465 (`MAX_SUBAGENT_ADMISSION`, `crates/tui/src/config/subagent_limits.rs:21`), so a turn can
466 request broad fan-out and let the manager drain it without creating an
467 unbounded population.
468
469 By default every admitted child may start immediately — there is no artificial
470 throttle beyond the rate-limit governor described below. Request the fan-out the work actually needs and let the runtime
471 queue and drain it; the caps above are enforcement, not a reason to
472 pre-refuse valid work. If you want gentler fan-out, lower `[subagents].launch_concurrency`
473 (how many direct children start at once); children beyond that limit **queue**
474 for a launch slot rather than bursting. `launch_concurrency` defaults to the
475 resolved `max_subagents` cap. (The pre-v0.8.61 `interactive_max_launch` key is
476 still accepted as a deprecated alias; the new key wins when both are set.)
477
478 High-fanout Workflows can tune that bounded population with `[subagents]
479 max_admitted` (aliases: `max_total`, `admission_limit`). That total ceiling
480 counts both **running** and **queued** agents, while `launch_concurrency` keeps
481 instantaneous execution bounded. Completed / failed / cancelled records persist
482 for inspection but don't occupy an admission slot. Agents that lost their
483 `task_handle` (e.g. across a process restart) also don't count against the cap.
484
485 ### Rate-limit governor
486
487 The one automatic throttle is the rate-limit governor. It watches provider
488 rate limits (HTTP 429) across a 60-second window. After repeated limits it
489 shrinks the number of launch slots; under a sustained burst it pauses new
490 launches entirely. Steady successes add slots back one at a time. It never
491 interrupts an agent that is already running, and quota exhaustion is not
492 treated as a throttle.
493
494 While the governor is holding launches back, it says so in two places:
495
496 - a queued agent's row gives the reason, for example
497 `launch slots throttled to 4/8 after 2 provider rate limit(s) in the last 60s`
498 or `launches paused after 4 provider rate limit(s) in the last 60s`, and the
499 time its wall budget ends;
500 - `GET /v1/agent-runs` returns a `governor` object next to `runs`, with
501 `launch_slots`, `max_launch_slots`, `paused`, `recent_rate_limits`, and a
502 `status` line while launches are held back. It describes launches made by the
503 runtime serving the request (Fleet runs).
504
505 Known limitation: the `/subagents` register header does not show the governor
506 line yet; the TUI receives agent lists from the Engine without governor state.
507
508 Provider profiles let one config stay aggressive for direct API routes while
509 keeping subscription or aggregator routes gentle. Every key under
510 `[subagents.providers.<provider>]` inherits from `[subagents]` when omitted.
511 Provider keys accept canonical names such as `deepseek`, `zai`, `openrouter`,
512 and aliases such as `glm` for Z.ai:
513
514 ```toml
515 [subagents]
516 # Global fallback for providers without a profile.
517 max_concurrent = 20
518 launch_concurrency = 20
519 max_admitted = 200
520 # Operator-selected Runtime delegation depth. The default is 3; this explicit
521 # value opts in above the default but remains below the hard ceiling of 8.
522 max_depth = 6
523 # Omitted or zero model-step budget is unbounded. Set a positive value only
524 # when an operator deliberately wants a per-child cap.
525 default_max_steps = 0
526 default_wall_time_secs = 1800
527
528 [subagents.providers.deepseek]
529 # Direct API key with room to fan out.
530 max_concurrent = 20
531 launch_concurrency = 20
532 max_admitted = 200
533
534 [subagents.providers.glm]
535 # Z.ai / GLM subscription-style route: keep pressure tight.
536 max_concurrent = 4
537 launch_concurrency = 3
538 max_admitted = 12
539 max_depth = 2
540 api_timeout_secs = 180
541 heartbeat_timeout_secs = 240
542
543 [subagents.providers.openrouter]
544 max_concurrent = 5
545 launch_concurrency = 3
546 max_admitted = 20
547
548 [subagents.providers.anthropic]
549 max_concurrent = 3
550 launch_concurrency = 2
551 max_admitted = 12
552 ```
553
554 Use `/config subagents status` to see both the global values and the active
555 provider's resolved fanout, depth, and timeout profile.
556
557 ## Advertised agent-tool fields
558
559 The model-facing `agent` schema exposes these controls:
560
561 | Purpose | Fields |
562 | --- | --- |
563 | Launch and route | `action`, `prompt`, `type`, `profile`, `name`, `model`, `model_strength`, `thinking` |
564 | Scope and outputs | `worktree`, `cwd`, `write_authority`, `write_roots`, `exact_files`, `coordination_contracts`, `deliverables`, `expected_artifact` |
565 | Narrow run limits | `max_steps`, `wall_time_secs` |
566 | Coordinate and recover | `agent_id`, `agent_ids`, `all_parked`, `message`, `until`, `detached`, `resume_from` |
567 | Inspect | `detail`, `offset`, `limit` |
568
569 `start` requires `prompt`. `message` requires a target and message;
570 `followup` requires a message and exactly one target form: `agent_id`/`name`,
571 `agent_ids`, or `all_parked: true`. `peek`, `interrupt`, and `cancel` require a
572 target. `claim` requires scope entries. These action requirements are validated
573 before execution.
574
575 `agent(action="roster")` reports each built-in role's resolved provider, model,
576 reasoning effort, known route limits and capability provenance. It uses the
577 same resolver as execution. An explicit saved profile wins first, followed by
578 a manual role pin in the current configuration, then a unique saved member
579 pinning that semantic role. Conflicting task `model` or `model_strength` choices
580 fail before admission. For an unpinned role, per-task `model` precedes
581 `model_strength`, then inherited role defaults and the session route.
582 When a Pod is selected, the `models` rows list its exact routes in saved order.
583 Use a listed `provider/model` selector for a task on an unpinned role; the session
584 model remains allowed. Off-list choices fail with the allowed routes, and a bare
585 model shared by multiple providers requires an exact selector. Without selected
586 models, current-provider overrides and `model_strength` retain their behavior;
587 foreign-provider requests fail. These choices do not change child authority.
588
589 The `profiles` rows expose saved members from the existing selected Fleet or
590 trusted config/personal/workspace/plugin layers, with bounded identities and the
591 same route/cost evidence. `profile="bug-hunter"` loads that member's instructions,
592 role, provider/model pin and depth limit. Conflicting type or model requests are
593 refused; explicit `thinking` overrides the saved tier. Missing providers, revoked
594 plugin authority and disabled project profiles fail before child admission.
595 Discovery never creates a profile or enrolls a model. These identity choices use
596 the existing child lifecycle; a saved profile alone does not create a continuing
597 Bot conversation or a computer lease.
598
599 Cost classes describe current uncached text input/output rates, not the total
600 price of a future task. Missing or routing-dependent prices remain unknown;
601 subscription/local routes are labelled not money metered. Discovery makes no
602 provider request and reports reachability as unverified.
603
604 **Parse-accepted but unadvertised (compat).** Other inputs remain accepted
605 for saved transcripts, ACP/MCP clients, fleet execution data, and
606 internal/operator compatibility. Runtime validates and intersects them with
607 live policy:
608
609 - delegation compatibility: `max_depth`, `maxDepth`, or `max_spawn_depth`;
610 values are restricted to 0 through the Runtime hard ceiling of 8 and only
611 narrow the inherited absolute ceiling. Model-facing calls inherit depth
612 from the operator and selected profile.
613 - workspace/isolation: `workspace_policy`, `fork_context`,
614 `worktree_path`, `worktree_branch`, `worktree_base`
615 - spawn contract: `deliberate`, `dependencies`, `acceptance`, `allowed_tools`
616 - lifecycle extras: `timeout_secs` (wait), `reason` (interrupt),
617 `include_archived` (status)
618
619 Compatibility input is not a way to widen inherited authority or remove a
620 finite budget.
621
622 ## Child budgets (steps, wall time, tokens)
623
624 `max_steps` and `wall_time_secs` are optional per-call limits.
625 Each can only narrow the applicable role, operator, parent, and saved-run
626 limits; the one exception is that `wall_time_secs` may raise the built-in
627 1800-second default (see below). Omission inherits those limits; explicit zero, null, negative, or
628 out-of-range values are rejected by the tool parser (schema minimum is 1).
629 Fleet file task-specs use a different convention — there, omitted-or-zero
630 means unbounded; see `docs/FLEET.md`.
631
632 `max_steps` counts model turns and accepts 1 through 2000. All roles default
633 to no model-turn cap unless an operator or ancestor supplies one; the internal
634 zero representation for that default never cancels a finite inherited cap.
635 `wall_time_secs` accepts 1 through 86400. Omitted, it defaults to 1800
636 seconds. An explicit value may go above that built-in default, for long
637 unattended work; an operator-configured `default_wall_time_secs` is both the
638 default and a ceiling, and role, parent, and saved-run deadlines still only
639 narrow. It covers model requests and tools. The effective absolute deadline
640 is persisted, and saved again when a queued agent launches, so a later
641 continuation is bounded by the deadline the agent actually worked to.
642
643 The work clock starts when the agent gets a launch slot. An agent that waits
644 in the launch queue waits at most `wall_time_secs`; if no slot opens in that
645 time it fails with a `never started` reason and zero steps, instead of being
646 reported as a run that used up its budget. An agent that does get a slot
647 after waiting receives its full `wall_time_secs` from that moment, still
648 bounded by any parent, saved-run or source deadline. So a parent can wait up
649 to about twice `wall_time_secs` for a queued agent. The queued row names the
650 reason for the wait and the time the agent stops waiting. If agents often
651 wait long, start fewer at once.
652
653 For example, a focused review can request:
654
655 ```json
656 {
657 "action": "start",
658 "type": "reviewer",
659 "prompt": "Review the parser diff and report concrete regressions.",
660 "max_steps": 12,
661 "wall_time_secs": 300
662 }
663 ```
664
665 The receipt's `effective_limits` is authoritative; a request for 300 seconds
666 cannot extend a parent's earlier deadline. A continuation keeps the source's
667 remaining steps, original deadline, and token history. A new ID, role, or
668 `resume_from` fork cannot reset those bounds.
669
670 ### Token accounting and partial results
671
672 Token budgets were retired in 0.9.14: token usage is tracked, never
673 enforced — runs are no longer stopped by token accounting. Legacy input that
674 still carries `token_budget` parses and is ignored; `max_steps` and
675 `wall_time_secs` remain the narrowable per-call limits.
676
677 The governor uses provider-reported input plus output tokens, not a local
678 estimate presented as a bill. Request output is capped to the remaining
679 allowance. Unknown prompt usage and requests already in flight can overshoot;
680 receipts retain the full reported usage. Missing usage remains unknown.
681 Worker records distinguish the worker's own token totals from shared
682 `budget_spent_tokens` and `budget_remaining_tokens`; do not sum a shared
683 pool once for every descendant.
684
685 The worker reserves room for one final report inside these limits: up to 10%
686 of a token allowance (at most 8192 tokens, only when at least 1024 can be
687 reserved), one turn when the step cap permits at least two, and up to 10% of
688 wall time (at most 10 seconds). Ordinary task execution stops before using
689 that reserve. Shared scopes hold back one token reserve for the scope;
690 reporting workers atomically claim remaining headroom so siblings cannot
691 independently reuse it. Continuation never refunds measured usage or resets
692 the original deadline.
693
694 The final reporting turn uses the worker's existing resolved provider and
695 model, with tools disabled and at most 1024 output tokens. It consolidates
696 bounded assistant notes and tool results into findings, evidence, produced
697 files, unfinished work and next steps. Estimated input cost counts against
698 its allowance. Provider transport retries remain inside the one logical
699 turn and its original wall-time deadline; no worker summary retry loop is
700 added. Token estimates are not billing receipts: unknown provider input and
701 requests already in flight can still overshoot, and actual usage is recorded.
702
703 The outcome stays `BudgetExhausted`, even when a useful report is obtained,
704 with the specific cause, checkpoint, measured usage and normal deliverable
705 verdicts. If the allowance is too small or already spent, earlier bounded
706 usage is unknown, the provider fails, or time expires, the worker returns
707 recorded partial text and says why a model report was unavailable.
708 Known missing response usage and attempts interrupted by timeout or cancellation
709 stay recorded across continuations and shared siblings; later known usage remains a
710 subtotal and cannot restore reporting headroom in that bounded scope.
711 Cancellation wins over reporting. Missing usage stays unknown. Exhausted
712 scopes reject further spawns or continuations; a partial report is not
713 successful completion.
714
715 ## Per-role models (#3018)
716
717 Children can run on a different model than the parent. Structured role pins,
718 the legacy model map, and convenience keys feed one override map. Structured
719 `[subagents.roles.<role>]` entries win over `[subagents.models]`, which wins over
720 the convenience keys. Keys are case-insensitive; within the structured table,
721 a canonical role key wins over its legacy alias:
722
723 ```toml
724 [subagents]
725 default_model = "deepseek-v4-flash" # fallback for every role
726 worker_model = "deepseek-v4-pro" # worker
727 scout_model = "deepseek-v4-flash" # scout
728 planner_model = "deepseek-v4-flash" # planner
729 reviewer_model = "deepseek-v4-pro" # reviewer
730 custom_model = "deepseek-v4-pro" # custom
731
732 [subagents.models]
733 # Free-form role → model map; any role alias accepted by agent works.
734 builder = "deepseek-v4-pro"
735
736 [subagents.roles.reviewer]
737 model = "deepseek/deepseek-v4-pro"
738 ```
739
740 These are manual pins for direct and Workflow `agent` starts. A task may restate
741 the same model or exact provider/model pair, but cannot change the pin with
742 `model` or `model_strength`. An explicit saved profile takes precedence over a
743 manual role pin. A type-only start also selects a unique saved role pin when
744 there is no manual override; ambiguous saved roles fail instead of choosing one.
745 Durable Fleet runs retain their selected member's frozen route.
746
747 A structured role pin may list approved replacement routes:
748
749 ```toml
750 [subagents.roles.reviewer]
751 model = "xai/grok-4.6"
752 replacements = ["deepseek/deepseek-v4-pro"]
753 ```
754
755 When the pinned route refuses the agent's **first** request (exhausted quota,
756 rejected credentials or authorization, or an unavailable model), the agent
757 retries that same request on the next listed route, keeping its role,
758 permissions, tools, scope and budgets. Listing a route authorizes sending the
759 agent's task to that provider, so each entry must name `provider/model`; at
760 most three are allowed and each is tried once. The route receipt records the
761 effective route, `route_source = "role.replacement"`, and a note with the
762 original route, the reason and the attempt. Replacement never happens after
763 the agent has run a tool, never for content-policy, context-length or
764 invalid-request errors, never for Codewhale's own permission denials, and never
765 for exact Fleet members or task-level `model` choices, which stay exact.
766
767 Structured role pins accept `provider/model`, preserving the configured provider's
768 exact identity and the complete model suffix. Unknown providers, empty pairs,
769 and cross-provider `auto` choices fail before admission. A bare structured model
770 inherits the session provider. For a namespaced model, qualify it explicitly,
771 for example `openrouter/deepseek/deepseek-v4-pro`. Legacy scalar and
772 `[subagents.models]` values keep their full provider-owned id, including slashes;
773 they do not change providers.
774
775 The v0.9.x convenience keys `explorer_model`, `awaiter_model`, and
776 `review_model` remain accepted as deprecated aliases so existing config files
777 do not break.
778
779 Model ids may be **any model the active provider accepts** — validation is
780 provider-aware and happens at spawn time, not load time. On the official
781 DeepSeek API only DeepSeek ids are accepted; every other provider passes the
782 id through to the provider API, which is the authority. A non-DeepSeek
783 example:
784
785 ```toml
786 provider = "moonshot"
787 model = "kimi-k2.7-code"
788
789 [subagents]
790 worker_model = "kimi-k2.6"
791 ```
792
793 Model ids are validated the same way when applied to a child route; an invalid
794 id on the official DeepSeek API fails the spawn with the accepted-id list
795 instead of an opaque provider 400.
796
797 With `/model auto`, sub-agent routing is provider-aware too: providers with a
798 known big/cheap pair (DeepSeek, and the hosted DeepSeek routes on NVIDIA NIM,
799 OpenRouter, Novita, SiliconFlow, SGLang, vLLM) route between that pair;
800 providers without a known cheap tier (e.g. Ollama, Moonshot) skip the
801 network router and keep children on the session model.
802
803 ## Per-profile provider routes (#3965)
804
805 `[subagents.models]` changes the child model within the active provider. A slash
806 in that legacy input does not grant another provider. To pin a different provider,
807 use a structured `[subagents.roles.<role>]` declaration as above, or use a
808 fleet/AgentProfile and select it with `profile` or its unique saved role.
809 The profile's explicit `provider` +
810 `model` fields win over the parent session route; omitting `provider` preserves
811 the existing inherit behavior.
812
813 Example: keep the parent session on DeepSeek, but run a formatter child on a
814 local LM Studio OpenAI-compatible endpoint:
815
816 ```toml
817 # ~/.codewhale/config.toml or workspace config
818 provider = "deepseek"
819
820 [providers.deepseek]
821 api_key = "YOUR_DEEPSEEK_KEY"
822
823 [providers.lm-studio]
824 kind = "openai-compatible"
825 base_url = "http://127.0.0.1:1234/v1"
826 api_key = "lm-studio"
827 model = "qwen-2.5-7b"
828 ```
829
830 ```toml
831 # .codewhale/agents/local-formatter.toml
832 id = "local-formatter"
833 role_hint = "formatter"
834 provider = "lm-studio"
835 model = "qwen-2.5-7b"
836 reasoning_effort = "off"
837
838 [instructions]
839 text = "Use small, local edits. Keep formatting changes mechanical."
840 ```
841
842 Then call `agent(profile: "local-formatter", prompt: "...")`. In-process
843 children build a client for `lm-studio`; fleet workers forward
844 `--provider lm-studio` to `codewhale exec`, which resolves the same
845 `[providers.lm-studio]` table. Unknown or unconfigured provider ids fail the
846 spawn rather than silently falling back to the parent provider.
847
848 ## Per-step API timeout (#1806, #1808)
849
850 Each sub-agent step wraps its DeepSeek `create_message` call in a
851 per-step timeout so a single stuck request can't pin the parent's
852 completion wakeup channel indefinitely. The default is `600` seconds.
853 A timed-out attempt is retried with exponential backoff (up to 5
854 retries) before the step interrupts with a preserved checkpoint.
855 Long-thinking children that legitimately exceed that, for example
856 heavy plan or review work behind `agent`, can extend the timeout in
857 `~/.codewhale/config.toml`:
858
859 ```toml
860 [subagents]
861 api_timeout_secs = 900 # 15 minutes; clamped to 1..=3600
862 ```
863
864 Values are clamped to `1..=3600`. `0` and `unset` keep the `600`
865 second default.
866
867 ## Stale-agent heartbeat (#2614)
868
869 Running agents also track manager-visible progress. If a child stops emitting
870 progress for the heartbeat window, the manager auto-cancels it, releases its
871 sub-agent slot, and keeps the cancelled record inspectable through the returned
872 transcript handle and persisted worker record. The default is 5 minutes
873 (resolved to at least 30 seconds above `api_timeout_secs`, so 630 seconds
874 with the 600-second default API timeout):
875
876 ```toml
877 [subagents]
878 heartbeat_timeout_secs = 300 # clamped to 30..=3600
879 ```
880
881 The effective heartbeat is kept at least 30 seconds above
882 `api_timeout_secs`, so a configured long model request is not cancelled before
883 its own request timeout can fire.
884
885 ## Lifecycle
886
887 Each opened session produces a record that progresses through:
888
889 ```
890 Pending → Running → (Completed | Failed(reason) | Cancelled | Interrupted(reason) | BudgetExhausted)
891 ```
892
893 An explicit interrupt, exhausted provider retries, or recovery of an orphaned
894 running record can leave an `Interrupted` worker with a checkpoint. Inspect
895 `needs_continuation` and the recorded reason; use `followup` for continuable
896 work. `BudgetExhausted` includes the specific token, step, or wall-time cause;
897 continuation cannot replenish an exhausted allowance.
898
899 `wait` observes workers. A timeout returns current outcomes and never parks,
900 cancels, or resumes them. `until: "completion"` returns when one child settles;
901 `until: "all"` joins the workers running when that call starts;
902 `until: "activity"` can return on progress. A later spawn is not silently added to an
903 earlier join.
904
905 An ordinary parent response leaves healthy children running. The same Engine
906 turn loop consumes their completion notices and can continue the parent.
907 Headless `codewhale exec` defers a successful final receipt until its existing
908 Engine reports no live children and no queued child completions. Its original
909 wall-clock deadline still bounds that settlement, including autonomous parent
910 turns. Cancellation, deadline exhaustion, a fatal event, or a lost Engine
911 channel stops settlement and returns the appropriate interrupted or failed
912 receipt with recorded partial usage. It does not report successful child
913 completion merely because the parent's first response ended.
914
915 ### Session boundaries (#405)
916
917 Each `SubAgentManager` instance assigns itself a fresh `session_boot_id` on
918 construction. Every new session stamps the agent with that id; the workspace
919 state file records it for restart recovery.
920
921 Work-bar/status projections focus on current-session agents by default.
922 Prior-session agents that are not still running are treated as archived records
923 so the model does not mistake stale work for live work. This is a
924 *prior-session* rule only: agents that finished in the CURRENT session keep
925 their work-bar rows for the rest of the session (quiet completion), and their
926 details still open from those rows.
927
928 Records that loaded from a pre-#405 persisted state file (no
929 `session_boot_id` field) classify as prior-session because the
930 manager can't match them to the current boot.
931
932 ## Run receipts, follow-up, and takeover
933
934 Each compatibility sub-agent has a persisted worker record in
935 `.codewhale/state/subagents.v1.json`. The record is the current run-ledger
936 slice for sub-agent lanes until those lanes are backed directly by the fleet
937 ledger: it stores `run_id`, objective, role/model,
938 workspace/branch, lifecycle events, artifact refs, follow-up target, takeover
939 target, usage provenance, and verification provenance.
940
941 The normal parent flow is to keep working and consume the completion event.
942 Default start and status receipts are compact; full snapshots and worker
943 records are diagnostic detail, not repeated in every response.
944
945 ### Continue an existing worker
946
947 `message` queues a note without waking the child. `followup` wakes a running
948 child or resumes a continuable checkpoint:
949
950 ```json
951 {"action":"followup","agent_id":"child-previous-id","message":"Continue the assignment using the recorded evidence."}
952 ```
953
954 Use the returned `agent_id` for subsequent waits and messages. The receipt's
955 `from` and `to` identify the original target and its current continuation.
956 The original receipt is retained. Retrying through an old ID follows the
957 persisted continuation chain and does not create a duplicate worker. If the
958 current successor is running, the follow-up is delivered there; if it has
959 already settled and cannot continue, the response says no message was
960 delivered. Duplicate workers are prevented, but repeated messages to a running
961 worker are still repeated messages.
962
963 For a batch, choose exactly one target form:
964
965 ```json
966 {"action":"followup","agent_ids":["child-a","child-b"],"message":"Continue the remaining checks."}
967 ```
968
969 ```json
970 {"action":"followup","all_parked":true,"message":"Continue the parked assignments."}
971 ```
972
973 Explicit batches accept up to 32 distinct IDs. `all_parked` selects parked
974 children you control and refuses more than 32 so you can choose explicit
975 batches. Bulk responses return separate `results` and `errors`; a failing
976 target does not roll back a successful continuation. Parent/descendant control
977 checks apply to both the addressed record and its current successor.
978
979 Use `start` with `resume_from` only to create a separate worker from a settled
980 child's transcript, for example to assign a new review. Each such start is a
981 new worker. Missing, running, or cross-workspace sources are refused; the
982 source's authority and budget bounds still apply. This is distinct from
983 continuing parked work with `followup`.
984
985 ### Compact status and full transcript retrieval
986
987 Unscoped `agent(action="status")` returns a session-scoped page bounded to
988 8 KiB. `offset` and `limit` page the roster; the default and maximum limit is
989 20. Follow `next_offset`, since the byte bound can return fewer rows than
990 requested. The model-facing roster wire uses one stable `columns` header and
991 an array of values per entry in `agents`; pair each row with the header instead
992 of reading it as an object. `null` means absent or unreported, and a measured
993 zero stays numeric `0`. The header is present even on an empty page.
994 Rows include worker and parent IDs, current depth, state,
995 elapsed time, own token total, recent activity, pending input, and continuation
996 lineage (`resumed_from` / `resumed_as`). Verification includes the verdict,
997 nonempty deliverable counts and a short warning when needed. Names, steps,
998 routes, effective limits (including maximum depth), and token breakdowns remain
999 available in the unchanged object projection when addressing one `agent_id`.
1000 Aggregate usage counts each worker's own reported tokens once and reports its
1001 coverage. Completion receipts additionally report measured descendant usage,
1002 deduplicate continuation lineage, and distinguish unknown usage from zero. A worker's
1003 `has_unreported_usage` and the descendant/subtree `unreported_usage_workers`
1004 counts identify missing responses even when later responses provide a measured subtotal.
1005
1006 Request one worker's detail when investigating a failure:
1007
1008 ```json
1009 {"action":"status","agent_id":"child-a","detail":true,"offset":0,"limit":20}
1010 ```
1011
1012 Addressed `peek` also accepts `detail: true`. Detail remains bounded to 32 KiB;
1013 message/event archives and deliverable verdicts are paged, and omission fields
1014 identify truncated detail. Use the returned typed `transcript_handle` with
1015 `handle_read` for the complete retained transcript. The handle's lookup
1016 coordinates are preserved even when diagnostic prose is omitted. Unscoped
1017 `detail: true` does not expand the entire roster into transcripts.
1018
1019 Artifacts are symbolic refs. Treat `result_summary` as a child self-report and
1020 inspect the specific `verification.status` and its evidence before relying on
1021 it. `usage.status` remains `unknown` until provider usage is reported, then
1022 becomes `reported` or `budget_exhausted` for a spent token scope. Neither a
1023 file's `present` verdict nor a completed lifecycle state proves a test gate.
1024
1025 ## Output contract
1026
1027 Non-scout sub-agents end with five Markdown headings, in this order:
1028
1029 ```
1030 ### SUMMARY one paragraph; what you did and what happened
1031 ### EVIDENCE path:line-range citations and key findings; one bullet each
1032 ### CHANGES files modified, with one-line descriptions; "None." if read-only
1033 ### RISKS what could go wrong / what the parent should double-check
1034 ### BLOCKERS what stopped you; "None." if you finished cleanly
1035 ```
1036
1037 Use `### HEADING` lines, with `EVIDENCE` before `CHANGES`. List edited
1038 repo-relative file paths under `### CHANGES`; blank lines before the bullets
1039 are allowed. Begin each bullet with its file path, followed by a description;
1040 quote paths containing spaces or literal trailing punctuation. The verifier
1041 also accepts older explicit `CHANGES:`,
1042 `Changed files:`, and `Files changed:` declarations. Evidence citations and
1043 paths under `RISKS` are not declarations of edits. The five-heading prompt
1044 contract is `SUBAGENT_OUTPUT_FORMAT` in
1045 `crates/tui/src/prompts/text.rs`. `prompt_documents_structured_subagent_briefs`
1046 in `crates/tui/src/prompts.rs` asserts every heading against it.
1047
1048 Scouts are the carve-out (#5189 F5): they end with `### SUMMARY` and
1049 `### EVIDENCE` only (`SUBAGENT_SCOUT_OUTPUT_FORMAT` in
1050 `crates/tui/src/prompts/text.rs`). `FleetRole::system_prompt` in
1051 `crates/tui/src/tools/subagent/mod.rs` injects the scout contract for
1052 `FleetRole::Scout` and the five-heading contract for every other role. A
1053 subagent test pins that scouts contain `## Output contract (scout)` and do
1054 not contain `### BLOCKERS`.
1055
1056 The parent reads `EVIDENCE` as a working set for the next turn, so
1057 scouts and reviewers should be precise here.
1058
1059 ## Memory and the `remember` tool (#489)
1060
1061 Sub-agents share the parent's native memory store when memory is enabled
1062 (`[memory] enabled = true` or `DEEPSEEK_MEMORY=on`). They can
1063 append durable notes via the `remember` tool — handy for a
1064 scout that discovers a project convention worth carrying across
1065 sessions, or a verifier that learns "this test is flaky".
1066
1067 `remember` takes a `scope` of `global` or `workspace`
1068 (`crates/tui/src/tools/remember.rs`, schema enum plus scope handling) and
1069 writes through `NativeMemoryStore` to `~/.codewhale/memory/global/MEMORY.md`
1070 or `~/.codewhale/memory/workspace/<id>/MEMORY.md`. Writes do not go through
1071 the standard write-approval flow. The legacy single-file `memory.md` path was
1072 removed in v0.9.4; see `docs/MEMORY.md` for the full layout.
1073
1074 ## Implementation notes
1075
1076 - Source: `crates/tui/src/tools/subagent/mod.rs`.
1077 - Persisted state: `<workspace>/.codewhale/state/subagents.v1.json`. Schema
1078 version `1` (forward-compatible — new optional fields use
1079 `#[serde(default)]`).
1080 - Settled records normally expire after `COMPLETED_AGENT_RETENTION`
1081 (default 1h), with a normal retained-record target of 256. Running /
1082 starting / waiting workers and the continuation identities and budget
1083 lineage needed by live work are preserved. Cleanup cannot discard an old
1084 ID while its continuation is still active or erase usage history needed
1085 to enforce an active scope.
1086 - `SubAgentRuntime::background_runtime()` starts from `child_runtime()` but
1087 replaces the turn-scoped child token with a fresh cancellation token, so
1088 parent turn cancellation does not stop detached background sessions.
1089 - The `is_running` check ignores agents whose `task_handle` is
1090 `None`; this avoids counting persisted-but-detached records
1091 toward the concurrency cap (#509).
1092 - `SharedSubAgentManager` is `Arc<RwLock<...>>` — read paths use
1093 read locks so `/agents` and the workbar projection don't block
1094 the main loop during multi-agent fan-out (#510).
1095
1096 Personal profiles use the same format at
1097 `$CODEWHALE_HOME/agents/<id>.toml` (normally `~/.codewhale/agents/`). For example:
1098
1099 ```toml
1100 # ~/.codewhale/agents/reasoner.toml
1101 base_role = "explore"
1102 provider = "openrouter"
1103 model = "qwen/qwen3.7-plus"
1104 reasoning_effort = "high"
1105
1106 [permissions]
1107 allow_shell = false
1108 trust = false
1109 ```
1110
1111 Select it with `agent(action: "start", profile: "reasoner", prompt: "...")`.
1112 The provider must also be configured in `config.toml`. The receipt names the
1113 resolved profile, its personal/project origin, provider/model and effective
1114 reasoning effort. Effort is normalized to the selected model's supported tiers;
1115 an explicit `thinking` request overrides the saved preference.
1116
1117 `allow_shell` and `trust` belong under `[permissions]`, not at the top level.
1118 A profile cannot grant `allow_shell = true`, `trust = true`, or disable approval.
1119 Use the appropriate `base_role` for the task; the parent session's live policy
1120 remains the authority ceiling. These profile fields are not a way to grant
1121 additional access.
1122
1123 A malformed, unreadable or duplicate profile now causes an explicit selection
1124 error, including when its name matches a built-in role. It never silently
1125 substitutes a lower roster layer. Repair the file and retry; profiles are reloaded
1126 for each launch. `agent(action: "roster")` reports affected profile identities and
1127 paths in `profile_load_issues` without exposing parser excerpts. Other valid
1128 profiles remain available, and a valid project override still wins over a broken
1129 personal definition. Fleet run creation performs the same check before storing
1130 a run or launching workers.
1131
1131 lines MARKDOWN