返回 CodeWhale
FLEET.md
根目录 / docs / FLEET.md
1 # Agent fleet
2
3 > Baca terjemahan bahasa Indonesia: [id/FLEET.md](id/FLEET.md)
4
5 Agent fleet is the local-first roster and member-selection layer for durable
6 multi-worker runs. It does not execute or authorize work. After fleet resolves
7 who should participate, the delegated coordinator launches a headless
8 `codewhale exec` run and the Runtime tracks it durably. See
9 [AGENT_RUNTIME.md](AGENT_RUNTIME.md) for how sub-agents, `exec`, and
10 fleet-backed workers converge on one runtime. In product language, a user may
11 still "open a sub-agent"; in architecture language, durable nested work uses a
12 fleet member identity with delegated runtime execution.
13
14 ## Naming and compatibility boundary
15
16 **Fleet** is the public product noun. The durable ledger, saved rosters, config
17 tables, and `--fleet` flag share that name:
18
19 | Surface | Canonical |
20 | --- | --- |
21 | CLI | `codewhale fleet …` |
22 | Slash command | `/fleet …` |
23
24 These shared names are load-bearing wherever changing them would break
25 existing workspaces, receipts, or scripts:
26
27 - the durable ledger `.codewhale/fleet.jsonl` and the log directories
28 `.codewhale/fleet/` and `.codewhale/fleet-host/`;
29 - saved rosters `fleets/<name>.toml` and their `schema = "fleet"` header, under
30 `$CODEWHALE_HOME/` or the workspace's `.codewhale/` (checked-in rosters at the
31 workspace root's `fleets/` are still read);
32 - the `[fleet]` config table (inline `[fleets.*]` tables were removed in 0.9.14; named fleets live in `fleets/<name>.toml` files);
33 - the `codewhale workflow run --fleet <name>` flag;
34 - wire, receipt, and control-plane operation ids such as `fleet.status`.
35
36 The rest of this document uses fleet as the public product noun and retains
37 these literal paths, keys, flags, and ids.
38
39 Use a fleet roster rather than anonymous short-lived `agent` fanout whenever a
40 delegated run needs stable member identities across retries, sleep/restart,
41 remote execution, receipts, or a ledgered audit trail. The initial CLI surface
42 is:
43
44 For a guided start-to-monitor walkthrough that combines fleet task specs with
45 Workflow authoring, see [fleet + Workflow Tutorial](FLEET_WORKFLOW_TUTORIAL.md).
46
47 ```sh
48 codewhale fleet init
49 codewhale fleet run tasks.json --check # validate only; nothing is created or launched
50 codewhale fleet run tasks.json --max-workers 4
51 codewhale fleet status
52 codewhale fleet inspect <worker-id>
53 codewhale fleet logs <worker-id>
54 codewhale fleet artifacts <worker-id>
55 codewhale fleet interrupt <worker-id>
56 codewhale fleet restart <worker-id>
57 codewhale fleet resume <run-id>
58 codewhale fleet stop --all
59 ```
60
61 `codewhale fleet resume <run-id>` is the restart-recovery verb: it replays the
62 ledger, reconciles any in-flight lease whose worker stopped heartbeating
63 (retrying within the task's budget, else failing and escalating per the alert
64 policy), and prints the post-resume status. It launches no new work and is
65 idempotent, so it is safe to run after a manager exit, laptop sleep, or runtime
66 restart.
67
68 Coordinator state for fleet-backed runs is stored under the workspace in
69 `.codewhale/fleet.jsonl`. Worker logs and adapter logs are stored under
70 `.codewhale/fleet/` and `.codewhale/fleet-host/`.
71
72 ## Public contract: identity, membership, and selection
73
74 **Fleet** = the user's model inventory: who is in the roster and which member is selected.
75
76 A public fleet identity consists only of:
77
78 - a stable member id and an optional user-facing name;
79 - a semantic role, such as `explore`, `implement`, or `reviewer` (the legacy
80 spellings `worker`, `scout`, `builder`, `verifier`, `consultant`, and `oracle`
81 are still accepted on input and map to `general`, `explore`, `implement`,
82 `test`, and `advisor`);
83 - an exact provider/model identity, or an explicit inherited route;
84 - visible roster state or origin.
85
86 Project or workspace trust, filesystem and network reach, secret access,
87 approval mode, sandboxing, tool authorization, and every other form of runtime
88 authority are separate delegated-coordination and Runtime policy inputs. They
89 are never fleet identity fields, and they never select or reroute a fleet
90 member. The Runtime applies and clamps those policies only after member
91 selection; if the selected member cannot run inside the effective envelope,
92 launch fails closed instead of choosing somebody else.
93
94 Natural-language member selection is deterministic. A caller may name:
95
96 - an exact member id, optionally as `member:<id>` or `id:<id>`;
97 - a unique user-facing member name, optionally as `name:<name>`;
98 - a unique semantic role, for example `explore` or `role:explore`;
99 - an exact pinned model id, for example `deepseek-v4-flash`, or its offline
100 display name, for example `DeepSeek V4 Flash`; or
101 - an exact `route:<provider>/<model>`.
102
103 An unqualified exact member id wins. Every other match succeeds only when it
104 identifies one distinct roster member. Multiple matches produce an ambiguity
105 error that names the candidates and asks for `member:<id>`; Codewhale never
106 picks whichever match happened to be listed first. Users do not need to know
107 an internal role label such as `explore`: a unique member name, display model, or
108 exact model id is equally valid. Saved v2 fleets store that optional human name
109 as `display_name` (the input alias `name` is also accepted); it must be one
110 trimmed printable line of at most 80 characters.
111
112 ### Your fleet as models
113
114 The same fleet file answers a third question: **which models has this person
115 put in their fleet?** Every exact `provider` + `model` pin in the selected
116 fleet — the operator route, each pinned member, and each explicitly marked
117 shortlist row — is a fleet model. Executable member rows supply the roles it
118 fills; a shortlist row has no role. All remain in the same fleet file. Legacy
119 members that omit `role` still use their id as the role identity.
120 Shortlist rows carry only model choices; reasoning, instructions, and capability
121 requirements belong on executable role members and are refused on shortlist rows.
122
123 - `/fleet models` prints the fleet: `provider/model · roles · price · context ·
124 tools`, facts read from the model catalog. With no selected fleet the line
125 reads "Your fleet is the session model only".
126 - `/fleet add <provider> <model> [role…]` adds a model (one member row per
127 role, or one `shortlist = true` row for a role-less add). The provider must be one you configured
128 and, when the catalog knows the provider, must serve that exact id.
129 A role member asked to run the fleet's own operator route inherits it
130 instead of pinning — the role follows when the operator moves; a pin on
131 any other route is the deliberate opt-out. Files that already pin the
132 operator route are read as inheritance.
133 With no fleet selected, a user-global fleet named `My fleet` is created and
134 selected first. `/fleet remove <provider> <model>` drops every row that pins
135 the route; the operator route is changed with `/fleet save`, not removed.
136 - In `/model`, `⇧F` adds or removes an explicit shortlist row for that exact
137 route. It preserves saved role pins and the operator route;
138 fleet models are listed first, labelled `fleet · <roles>`, ahead of your
139 own `⇧P` pins and the provider lists. `/models` prints the fleet before the
140 provider's list.
141
142 The operator model reads this list when it assigns sub-agents (design
143 `MODEL-ROUTING-CATALOG-20260901.md` §10, slice F2).
144 The model-facing roster resolves roles through the same route admission code
145 as a start. Explicit profiles and manual role pins remain authoritative; a
146 unique saved role pin also applies to a start naming only its role. Task model
147 choices are available for unpinned roles, constrained to the selected models
148 plus the session route. Shortlist model rows disclose their own exact route
149 independently of any role pin.
150
151 ### Interactive and persistent status
152
153 `/fleet status` and `codewhale fleet status` are the **same** command on two
154 surfaces. Both read the durable `.codewhale/fleet.jsonl` ledger for the
155 workspace, through one shared control-plane contract, and both report the same
156 verb id (`fleet.status`), read-vs-write authority, persistence scope, and
157 receipt. When the workspace has no ledger they say so with a typed reason
158 (`no_fleet_ledger`) instead of rendering an empty-looking "all clear" — and
159 neither creates the ledger as a side effect of reading it.
160
161 The current interactive session's sub-agents are a **different set**, and now
162 have their own name:
163
164 - `/fleet workers` (or `/subagents`, or Tab / `w` from the `/fleet` roster)
165 shows sub-agents attached to the current TUI session. It does not read the persistent ledger.
166 - `/fleet list|status|interrupt|resume` and `codewhale fleet
167 list|status|interrupt|resume` act on the durable ledger.
168 - `codewhale fleet restart <worker-id>` is CLI-only: it re-leases the task and
169 then drives the manager loop to completion. `/fleet restart` does not
170 silently do a smaller thing — it reports `surface_not_supported` and names
171 the CLI command.
172
173 The contract behind this — descriptors, availability reasons, exact-identity
174 targets, receipts, typed unknowns, and bounds — is documented in
175 [`docs/COMMAND_CONTROL_PLANE.md`](COMMAND_CONTROL_PLANE.md).
176
177 ## Authoring agent profiles (`/fleet setup`)
178
179 Agents: the durable artifact is the profile TOML described below; the
180 key-by-key walkthrough is the human interactive path.
181
182 `/fleet setup` (also `/fleet setup edit` / `new`) opens an in-TUI wizard for
183 authoring a reusable agent-team profile. Bare `/fleet` and the
184 `roster`/`roles`/`profiles`/`party` aliases open the selected fleet's member
185 roster. `/fleet saved` opens the named saved-fleet picker. `/fleet workers` opens the
186 current-session worker view; `/subagents` is a
187 compatibility shortcut for that view. For durable run history, use
188 `/fleet status` or the shell command `codewhale fleet status` described above —
189 they are the same command.
190
191 The wizard is progressive: you make one focused choice at a time — a **role**,
192 then a **model** (`inherit`, or a concrete model from *any configured
193 provider*, not only the one the parent session is currently using), then
194 **where the profile lives**, and finally a **review** of the member identity
195 and route. When the review also previews thinking, tools, approvals, or another
196 execution control, those rows summarize separate Runtime policy; they do not
197 become fleet identity or member selectors. The header shows "Saves to: …" on
198 every step — the choice you still have to make, or the exact resolved file
199 once you have made it. Nothing is written until you activate the save control
200 on the review step.
201
202 The **Destination** step is a focused two-option list:
203
204 - **This project** writes `<workspace>/.codewhale/agents/<role>.toml`. It
205 applies to this project only and takes precedence over a Personal profile
206 with the same id. When project profiles are disabled for the session
207 (`--no-project-config`) or the workspace folder is unavailable, the option is
208 shown disabled with that reason; the wizard never falls back to Personal on
209 its own.
210 - **Personal** writes `$CODEWHALE_HOME/agents/<role>.toml` and is available in
211 every project on this machine, except where a project has its own profile
212 with the same id.
213
214 For the highlighted option the step shows the exact file, whether saving would
215 create a new file or **replace an existing one**, and the precedence
216 consequence for the roster. The review step repeats those facts under
217 "Saves to" and names the final action by its effect — **Save to this
218 project**, **Save as Personal profile**, or **Replace …**. Replacing an
219 existing file asks for a second confirmation on the save control. Reopening a
220 saved member from `/fleet` starts from what is on disk: its member identity,
221 route, and save scope. Thinking (`inherit`, `off`, `low`, `medium`, `high`,
222 `max`, or `auto`) is adjusted on the review step, but remains a route
223 execution setting rather than part of the member's fleet identity.
224
225 Profile scope controls where a role definition is reusable; it does not widen
226 the authority of a running operation and is not a project-trust setting. To
227 coordinate several nearby repositories, start Codewhale from their shared
228 parent directory so that parent is the workspace. Project/workspace trust,
229 external paths, filesystem and network reach, secrets, approvals, sandboxing,
230 and tool authorization come from delegated-coordination and Runtime policy.
231 For nested delegation, Runtime intersects the requested child posture with the
232 live parent. For standalone `codewhale fleet` execution, Runtime instead uses
233 the bounded tool-authority envelope minted from the task's explicit write
234 scope together with live config, sandbox, and platform enforcement. Neither
235 path reads authority from the profile's storage scope or identity selector.
236 A worker whose envelope grants read-only shell access runs the same read-only
237 command grammar as an in-session read-only agent, including pipelines, chains
238 and a leading `cd` (see "Read-only shell commands" in `docs/SUBAGENTS.md`);
239 Admitted `gh` reads also need the envelope's network grant. `npm view` remains
240 outside the read-only grammar because npm configuration can select executable
241 helpers; a network grant does not authorize those helpers.
242
243 Picking a concrete model pins its provider explicitly: the saved profile records both
244 `model` and `provider` fields, so the route it names doesn't depend on
245 whichever provider happens to be active when the profile is later loaded.
246 Pressing **Enter** ("start") on the review step previews the exact starter
247 profile TOML inline on that same screen; nothing is written until you save it.
248 The `provider` field may be a built-in provider id such as `openrouter` or a
249 user-named OpenAI-compatible provider configured under `[providers.<name>]`
250 such as `lm-studio`; the launch path preserves that id and fails closed if the
251 provider is not configured.
252
253 Profiles are also how the model-facing `agent` tool selects a route: a child
254 either runs as a `profile` (whose saved route and thinking tier it uses
255 exactly) or inherits the operator's model. Per-task `model`, `model_strength`,
256 and `thinking` remain advertised for unpinned roles; saved profile and manual
257 role pins refuse overrides. See docs/SUBAGENTS.md for the advertised field
258 list and the parse-accepted compat list.
259
260 When a provider is configured, the review step also offers model-assisted
261 drafting behind an explicit preview-before-save gate:
262
263 - Press **`m`** to have your first configured model draft the profile. The
264 draft arrives sanitized and bounded. Separately, the Runtime keeps its
265 conservative execution floor (no shell or trust escalation and approval
266 required) regardless of what the model proposes.
267 - **Drafting is not saving.** The exact rendered TOML preview renders
268 inline on the review step (not in a separate scrollable viewer), so nothing
269 is saved until you press **`g`** or **Enter** to save (or press `m` again
270 to redraft). Saving writes the profile to the project or personal scope
271 shown in the preview.
272
273 ## Naming: Modes, Workflow, and fleet
274
275 These names describe different layers, not competing systems. Plan and Act are
276 the everyday work modes. Operate accepts ordinary messages and keeps the
277 parent's normal tool surface under the same approval, sandbox, shell, ask-rule,
278 and repository protections as Act. It prefers background fleet workers for
279 independent, parallel, isolated, or long-running work, but does not require a
280 worker for every executable step. Workflow is an optional orchestration overlay
281 for work that needs ordering, gates, shared budgets, replay, or deterministic
282 fan-in.
283
284 The short public vocabulary is:
285
286 - **Fleet** is the durable roster and deterministic member-selection surface.
287 It records member ids and names, semantic roles, provider/model identities,
288 and roster state. Fleet is also the name used by storage and wire formats.
289 - **Workflow** = what order the work follows: phases, gates, budgets, replay,
290 and fan-in.
291 - **Lane** = one running Workflow instance and its live progress.
292 - **Runtime** = where, how, and with what authority selected work executes.
293 Runtime owns the local or remote process, provider route, project/workspace
294 trust, filesystem, network, secrets, approvals, sandbox, tools, and API
295 boundary.
296
297 - **Workflow** is the repeatable plan and user-facing orchestration
298 overlay: a script/IR that decides which phases and agents run next, keeps
299 intermediate results out of the main conversation, and can be inspected or
300 rerun. A Workflow run should have a visible progress view and a clear active
301 header state instead of feeling like a hidden background task.
302 - **Fleet** is the durable roster and deterministic member-selection surface:
303 member ids and names, semantic roles, pinned or inherited provider/model
304 identities, and roster state. The delegated
305 coordinator and Runtime own launch concurrency, leases, heartbeats, logs,
306 receipts, tools, sandboxing, approvals, and authority.
307 - **High fan-out** is a behavior of a Workflow run, not a separate system:
308 when a phase needs many workers at once, Workflow dispatches them as a
309 fleet-backed run (durable workers, receipts, goal re-dispatch) rather than
310 reviving prompt-only sub-agent fanout.
311 - **Fan-in is explicit:** when the user needs one combined result, an owner
312 aggregates, verifies, and synthesizes the worker receipts. Independent tasks
313 may finish separately; dispatch is never presented as completion.
314
315 UI guidance: keep the main transcript calm. A Workflow run should appear as a
316 compact progress card plus workbar rows (the strip under the composer, or
317 a side workbar) with phase names, worker counts, receipts, and nested
318 indentation for child workers. Use the whale mark sparingly as an active
319 header/status signal; avoid repeating emoji-heavy rows for every worker.
320
321 ## Saved fleets and the Reasoning Router
322
323 A selected v2 fleet freezes each selected member's id, semantic role, provider,
324 and model identity into the durable run before a Workflow starts. Save the
325 fleet as `fleets/<name>.toml` under the workspace's `.codewhale/` (where the
326 fleet editor saves folder fleets) or under `$CODEWHALE_HOME`; a checked-in
327 `fleets/<name>.toml` at the workspace root is also read.
328 Models cannot replace those identity or route assignments at runtime:
329
330 ```toml
331 schema = "fleet"
332 schema_revision = 2
333 name = "release"
334
335 [operator]
336 provider = "deepseek"
337 model = "deepseek-v4-pro"
338
339 [[members]]
340 id = "implementer"
341 display_name = "Release Builder"
342 role = "implement"
343 provider = "zai"
344 model = "glm-5.2"
345
346 [[members]]
347 id = "advice"
348 role = "advisor"
349 provider = "openai"
350 model = "gpt-5.6"
351 ```
352
353 The workflow crate's older `schema = "exact"`, revision 1 files are migration
354 input only. Do not author revision-1 files; the selected roster and setup UI
355 read and write only `schema = "fleet"`, revision 2.
356
357 `workflow(fleet: "release")` runs a saved Fleet without selecting it. At
358 Workflow start, a member with no pin takes the Fleet's `[operator]` route or,
359 without one, the session route and reasoning tier; that frozen route is what
360 runs and what receipts name, and editing the file mid-run changes only the next
361 Workflow. If a saved Fleet and an older exact/legacy file share a name, the
362 Workflow refuses to guess; qualify the saved one as `user/<name>` or
363 `folder/<name>`. Members with `instructions` or `requires` cannot run in a
364 Workflow yet.
365
366 Reasoning is a separate route-execution decision, not fleet identity. The
367 optional Reasoning Router is a reusable Runtime service, not a fleet member.
368 Save one profile at `routers/<name>.toml` in any search root and reference it
369 from any number of fleets:
370
371 ```toml
372 name = "luna-low"
373 schema = "reasoning_router"
374 schema_revision = 1
375 provider = "openai"
376 model = "gpt-5.6-luna"
377 call_reasoning = "low"
378 ```
379
380 At runtime it may choose only the reasoning tier for an already-frozen worker
381 route. It cannot change the member, provider, model, or semantic role. The
382 Router call itself is capped at `off` or `low`; more expensive values are
383 rejected. A manually selected worker reasoning tier makes no Router call. Route
384 and reasoning receipts name the worker model and, when used, the Router's exact
385 provider/model so the operator can see which model did which job. If the same
386 bare Router or fleet name exists in more than one root, qualify it as
387 `codewhale_home/<name>`, `workspace/<name>` (the workspace's `.codewhale/`), or
388 `workspace_root/<name>` (the workspace root) instead of relying on shadowing.
389
390 Compatibility schemas may serialize `reasoning`, `permissions`, tool hints, or
391 other execution settings beside a member. Those values are not fleet identity,
392 member selectors, or active authority. A valid legacy `schema = "exact"` roster snapshot
393 retains its old `permissions` bytes only while verifying and replaying that
394 snapshot's recorded content hash; a fresh capture emits the authority-free
395 member shape. New-run validation rejects legacy roster
396 `security_policy` and worker `trust_level` fields; configure execution authority
397 through Runtime policy. The delegated coordinator resolves and durably freezes
398 the member first. Runtime then applies either the delegating parent's effective
399 ceiling or, for standalone fleet CLI work, Runtime execution configuration plus
400 live sandbox/platform enforcement. That boundary may
401 reduce or refuse the selected worker's execution surface, but it must never
402 choose a different member or route. See
403 [`docs/MODES.md`](MODES.md), [`docs/SUBAGENTS.md`](SUBAGENTS.md), and
404 [`docs/AGENT_RUNTIME.md`](AGENT_RUNTIME.md) for the enforcement contract.
405
406 Reasoning receipts record the requested tier *and* the tier the provider was
407 actually asked for. Those differ whenever a route cannot express the requested
408 one — Codewhale's route normalizer sends `high` for a requested `low` on most
409 routes, and Z.AI's GLM routes express only thinking on/off — so the receipt
410 reports the real request rather than the label that was selected. The value a
411 call actually carries is spelled by that route's own normalizer, not by the tier
412 label: an OpenAI Codex route is asked for `xhigh`, not `max`, and cannot be
413 asked for `off` at all.
414
415 A durable fleet CLI receipt keeps the selected profile id in
416 `effective_permissions.profile_id`, the resolved semantic role in
417 `resolved_route.role`, and the effective Runtime surface in the permission,
418 shell, and tool-scope fields. An exact Workflow launch receipt records
419 `member_role` separately from an optional Runtime `posture_role`, plus the
420 fingerprint of the effective authority envelope checked at the spawn boundary.
421 A member named `auditor` can therefore retain that identity while Runtime
422 reports a `custom` posture and independently proves the narrower surface it
423 enforced.
424
425 A Workflow start fails closed on anything decidable locally: an unresolvable
426 provider or model, a missing credential, a client that cannot be built for a
427 member's route, or an `auto` member with no usable Reasoning Router. Per-task
428 validation that the spawn boundary would refuse anyway — notably a write-capable
429 member with no declared `write_roots`/`exact_files`/`coordination_contracts` —
430 is checked before the Router is called, so an invalid task never spends a
431 routing request. If a spawn fails *after* a Router decision, the receipt is
432 still recorded: the tokens were spent, and any cross-provider disclosure already
433 happened.
434
435 ## Manager-owned Workflow fan-in
436
437 When parallel work must return one combined answer, prefer a manager-owned
438 Workflow over a flat `agent` fan-out. Default shape:
439
440 1. **Cast one manager** (operator or workflow orchestrator).
441 2. **Fan out** child tasks through `workflow` (`task()`, `parallel()`,
442 `pipeline()`, `phase()`) or a single manager session that owns the children.
443 3. **Wait** for child receipts or completion events.
444 4. **Aggregate and verify** load-bearing claims before treating them as facts.
445 5. **Synthesize** one result the operator can depend on.
446
447 Raw `agent` fan-out fits independent work with no combined result. When
448 results must be merged, compared, or verified, route through `workflow` so
449 the manager owns fan-in — that is what the shape above is for, not a ban
450 on simpler patterns when nothing needs combining.
451
452 ## Workflow on fleet
453
454 The intended high-capability path is agent-authored. When the main agent
455 decides a task needs more durable coordination than turn-by-turn sub-agent
456 calls, it drafts a Workflow script/IR, presents the run plan according to the
457 active permission mode, and the runtime compiles it into typed fleet work.
458
459 fleet remains the sub-agent roster and member-selection surface. It owns member
460 identity, membership, semantic roles, saved provider/model pins or inheritance,
461 and roster state. Workflow owns the orchestration plan:
462 branch, sequence, loop, expand, review, and reduce decisions. The delegated
463 coordinator and Runtime own slot admission, launch concurrency, the execution
464 ledger, and every authority decision. A workflow script receives no direct
465 shell, filesystem, network, provider-secret, cancellation, or TUI authority;
466 workers perform real work as `codewhale exec` processes under the effective
467 Runtime policy.
468
469 Default Workflow-to-fleet validation is intentionally bounded:
470
471 - 1,000 total worker agents per Workflow run;
472 - 16 live worker agents at once; larger populations queue (block) on the host's
473 per-run concurrency gate until a live slot frees, then route through fleet;
474 - Workflow IR structural nesting no deeper than 5;
475 - Runtime child delegation defaults to 3 levels and has an opt-in hard ceiling
476 of 8, independently of the Workflow document's structural depth;
477 - bounded loops only (`max_iterations` required);
478 - bounded dynamic expansion only (`max_children` plus a template required).
479
480 These are delegated-coordination population limits, not fleet identity and not
481 a demand to launch everything at once. A 1,000-agent Workflow should still
482 drain through the configured Runtime worker pool. They are also not model-step
483 budgets: omitted or zero `max_steps` remains unbounded. An explicit positive
484 `max_steps` may cap that task, while wall-clock timeouts, cancellation,
485 provider safeguards, heartbeats, and admission controls remain independent.
486
487 Recommended model layouts, such as a DeepSeek Pro orchestrator with Flash
488 workers in the first ring and cheaper workers farther out, are presets only.
489 Every slot can inherit the active model or carry an explicit model override.
490 Inheritance is literal: the model you select in `/model` is the **operator**
491 (the pinned first row in `/fleet roster`), and any worker whose task spec and
492 roster profile pin no model runs on that session model. Once a selected member
493 has an exact provider/model pin, the Runtime does not silently reroute that
494 identity because a policy input differs; it either runs that route inside the
495 effective envelope or fails closed. Route receipts record the requested and
496 resolved identity.
497
498 The setup UI should render this as an expanding grid: an orchestrator plus a
499 small number of visible sub-agent slots, with Right/Enter drilling into a slot's
500 next recursive ring rather than trying to show the whole tree at once.
501
502 ## Task Spec
503
504 `codewhale fleet run` accepts JSON or TOML. `codewhale fleet run <spec> --check`
505 runs every validation a real run performs (spec shape, roster members, agent
506 profiles, model routes) and prints the same warnings, then stops: no ledger is
507 created, no run is written, no worker starts, and nothing is spent. A minimal
508 JSON spec:
509
510 ```json
511 {
512 "name": "local smoke",
513 "tasks": [
514 {
515 "id": "lint",
516 "name": "Lint",
517 "instructions": "Run the lint check and report failures.",
518 "expected_artifacts": ["log"]
519 }
520 ]
521 }
522 ```
523
524 Workers are optional. If omitted, Codewhale creates local worker slots up to
525 `--max-workers`.
526
527 A spec file takes one of three shapes, chosen by its structure before any
528 field is read:
529
530 - a **document** — an object with `tasks` (and optionally `name`, `labels`,
531 `workers`, `usage_ceiling`);
532 - a **task array** — a bare JSON array of task objects;
533 - a **single task** — one task object with `id` / `instructions` at the top
534 level (JSON or TOML; a TOML file is never a task array).
535
536 Array and single-task files take their run name from the file name. Because
537 the shape is picked first, a malformed spec reports the real problem, for
538 example ``JSON spec document at tasks[1] (id "review"): missing field
539 `instructions` at line 7 column 5``. The checked-in
540 [`docs/examples/fleet-dogfood.toml`](examples/fleet-dogfood.toml) and the
541 tutorial's `tasks.json` are parsed by the test suite, so they stay valid.
542
543 Task specs are typed in Rust and keep verification data separate from worker
544 transcripts. Only the `worker` member/role reference participates in fleet
545 identity selection. The remaining execution fields are delegated-coordination
546 or Runtime inputs applied after the member is resolved. A task can declare:
547
548 - `id`, `name`, `description`, `objective`, and `instructions`
549 - `worker` role, tool profile, tools, and required capabilities
550 - `workspace` root, required files, writable paths, and environment allowlist
551 - `input_files`, extra `context`, `budget`, `timeout_seconds`, and `retry_policy`
552 - `expected_artifacts`, `scorer`, `tags`, and free-form `metadata`
553
554 None of those execution-policy fields becomes part of a fleet identity or an
555 alternate member selector. Omitted or zero `max_steps` means no model-step
556 ceiling; Codewhale must not synthesize a default step budget. (This is the
557 fleet file task-spec convention; model-facing `agent` calls differ — the tool
558 parser rejects an explicit zero. See `docs/SUBAGENTS.md`.) Explicit positive
559 step limits, timeouts, cancellation, provider safeguards, heartbeats, and
560 admission control are enforced independently by the delegated coordinator and
561 Runtime.
562
563 Workers write bounded artifact files under `.codewhale/fleet/` and ledger only
564 the artifact refs: kind, path, checksum, MIME type, and size. Receipts record
565 `pass`, `fail`, `partial`, `skip`, or `timeout`; failed receipts may also mark
566 the source as `transport`, `task`, or `verifier`. `codewhale fleet status`
567 surfaces those failure-source counts separately.
568
569 Deterministic built-in scorers are `exit_code`, `file_exists`, `regex_match`,
570 and `json_path`. Specs may also declare `command`,
571 `code_whale_verifier_prompt`, or `manual`; those record a partial receipt until
572 an explicit verifier pass completes.
573
574 ### Using Role Presets
575
576 Tasks can reference a semantic role name to select one unique roster member.
577 Built-in role names (`smoke-runner`, `reviewer`, `builder`, `read-only`) remain
578 available for compatibility, and custom roles may be defined in
579 `[fleet.roles]`.
580
581 ```json
582 {
583 "name": "smoke check",
584 "tasks": [
585 {
586 "id": "lint",
587 "name": "Lint check",
588 "instructions": "Run lint and report failures.",
589 "worker": { "role": "smoke-runner" },
590 "expected_artifacts": ["log"]
591 }
592 ]
593 }
594 ```
595
596 After identity resolution, compatibility role presets may provide tool,
597 timeout, or retry defaults to the delegated coordinator. Those defaults do not
598 grant authority, do not change which member was selected, and remain subject to
599 Runtime clamping. A task spec may request its execution settings explicitly:
600
601 ```json
602 {
603 "id": "deep-review",
604 "name": "Deep review",
605 "instructions": "Review the entire crate for soundness issues.",
606 "worker": {
607 "role": "reviewer",
608 "tools": ["cargo", "rg", "git"],
609 "capabilities": ["rust"]
610 },
611 "input_files": ["crates/**/*.rs"],
612 "budget": { "max_tokens": 32000 },
613 "expected_artifacts": ["log", "report"],
614 "scorer": { "kind": "regex_match", "path": ".codewhale/fleet/report.md", "pattern": "finding|all clear" }
615 }
616 ```
617
618 ### Multi-Task Run Example
619
620 A single fleet run can dispatch several independent tasks in parallel:
621
622 ```json
623 {
624 "name": "CI gate",
625 "tasks": [
626 {
627 "id": "check",
628 "name": "Compile check",
629 "instructions": "Run cargo check --workspace and report errors.",
630 "worker": { "role": "builder" },
631 "expected_artifacts": ["log"],
632 "scorer": { "kind": "exit_code" }
633 },
634 {
635 "id": "clippy",
636 "name": "Clippy lint",
637 "instructions": "Run cargo clippy --workspace and report warnings.",
638 "worker": { "role": "reviewer", "tools": ["cargo", "cargo-clippy"] },
639 "expected_artifacts": ["log"],
640 "scorer": { "kind": "exit_code" }
641 },
642 {
643 "id": "security",
644 "name": "Secret audit",
645 "instructions": "Search for plaintext secrets and report any matches.",
646 "worker": { "role": "read-only", "tools": ["rg"] },
647 "input_files": ["crates/**/*.rs"],
648 "expected_artifacts": ["log", "report"],
649 "retry_policy": { "max_attempts": 1 }
650 }
651 ]
652 }
653 ```
654
655 ## Alerts
656
657 fleet alerting is disabled by default. A caller must supply an enabled alert
658 config before anything is sent. Routes match typed fleet event classes, not log
659 strings:
660
661 - `stale`
662 - `restart_exhausted`
663 - `needs_human`
664 - `budget_exceeded`
665 - `verifier_failed`
666 - `run_completed`
667
668 Adapter config stores environment variable names, not secret values. Send-time
669 code resolves those names from the environment or a future secrets provider.
670 Ledger records store only audit labels such as `slack`, `webhook`, or
671 `pagerduty`; task specs persisted in the ledger redact webhook URLs and routing
672 keys.
673
674 Example alert config shape:
675
676 ```json
677 {
678 "enabled": true,
679 "dry_run": true,
680 "routes": [
681 {
682 "events": ["stale", "restart_exhausted", "verifier_failed"],
683 "adapter": "ops-slack"
684 },
685 {
686 "events": ["restart_exhausted"],
687 "adapter": "pager"
688 }
689 ],
690 "adapters": {
691 "ops-slack": {
692 "kind": "slack",
693 "webhook_env": "CODEWHALE_FLEET_SLACK_WEBHOOK",
694 "channel": "#codewhale-fleet"
695 },
696 "pager": {
697 "kind": "pager_duty",
698 "routing_key_env": "CODEWHALE_FLEET_PAGERDUTY_ROUTING_KEY",
699 "severity": "critical"
700 }
701 }
702 }
703 ```
704
705 Use dry-run to inspect a redacted adapter payload without sending:
706
707 ```sh
708 codewhale fleet alert-dry-run \
709 --event stale \
710 --run-id fleet-demo \
711 --worker-id fleet-demo-local-1 \
712 --task-id release-triage \
713 --reason "worker heartbeat stale since 2026-06-13T02:00:00Z" \
714 --adapter slack
715 ```
716
717 The payload includes the run id, worker id, task id, status, short reason, and
718 safe inspection commands such as `codewhale fleet status` and
719 `codewhale fleet inspect <worker-id>`. Endpoints, webhook secrets, and
720 PagerDuty routing keys are shown as `<redacted:env:...>`.
721
722 ## Status Surfaces
723
724 `codewhale fleet status` shows compact counts for queued, running, completed,
725 partial, failed, restarted, escalated, cancelled, stale, and verifier/transport
726 failure sources. `inspect` shows the worker state plus the current task
727 objective, role, host, heartbeat, latest event, artifact refs, latest error, and
728 alert state. `logs` prints bounded log artifact contents, and `artifacts` lists
729 artifact refs without embedding large payloads.
730
731 The Runtime API exposes the same ledger-backed projection behind the existing
732 runtime auth middleware:
733
734 ```text
735 GET /v1/fleet/runs
736 GET /v1/fleet/runs/{run_id}
737 GET /v1/fleet/runs/{run_id}/workers
738 GET /v1/fleet/workers/{worker_id}
739 POST /v1/fleet/workers/{worker_id}/interrupt
740 POST /v1/fleet/workers/{worker_id}/restart
741 POST /v1/fleet/runs/{run_id}/stop
742 ```
743
744 Action endpoints call the same manager controls as the CLI and record their
745 decisions in the fleet ledger.
746
747 ## Manager-Agent Runbook
748
749 Manager agents should treat fleet operations as typed, ledgered control-plane
750 work. Start with `codewhale fleet status`, then inspect one run or worker with
751 `codewhale fleet inspect <worker-id>`, `logs`, and `artifacts`. Use direct
752 reads of `.codewhale/fleet.jsonl`, host logs, or remote files only when the
753 typed CLI/API surface cannot provide the required evidence.
754
755 Classify the worker before taking action:
756
757 - `transient failure`: stale heartbeat, host timeout, interrupted transport,
758 retryable provider/network error, or an adapter status that can plausibly
759 recover without changing the task.
760 - `task failure`: the worker completed but produced an incorrect result,
761 domain failure, missing required artifact, or explicit task-level error.
762 - `verifier failure`: the worker result exists, but the scorer/verifier failed,
763 timed out, or disagrees with the receipt.
764 - `needs-human`: missing authority, secret request, destructive operation,
765 repeated restart exhaustion, ambiguous product decision, or conflicting
766 evidence that the manager cannot resolve from typed artifacts.
767
768 Choose one typed action:
769
770 - Restart a worker only when the failure is transient, retry budget remains,
771 the task is idempotent or retry-safe, and no permission or secret boundary is
772 involved: `codewhale fleet restart <worker-id>`.
773 - Interrupt or stop only when the current task is unsafe to continue or the
774 operator explicitly asks for cancellation: `codewhale fleet interrupt
775 <worker-id>` or `codewhale fleet stop --all`.
776 - Do not restart pure task failures by default; preserve artifacts and hand the
777 receipt to the task owner unless the task spec says retrying can produce new
778 evidence.
779 - For verifier failures, inspect scorer inputs and artifact refs first. If the
780 verifier cannot be corrected through typed fleet actions, escalate for human
781 review.
782 - For `needs-human`, draft an escalation instead of sending it unless alert
783 config explicitly authorizes sending.
784
785 Safe Slack or PagerDuty draft:
786
787 ```text
788 Codewhale fleet needs attention
789 Run: <run-id>
790 Worker: <worker-id>
791 Task: <task-id or unknown>
792 Classification: <transient failure | task failure | verifier failure | needs-human>
793 Reason: <one sentence, no secrets>
794 Latest typed evidence: codewhale fleet inspect <worker-id>; codewhale fleet artifacts <worker-id>
795 Safe log excerpt: <3 lines max or "see artifact <ref>">
796 Requested decision: <restart approval | verifier review | task owner review | permission decision>
797 ```
798
799 Post-run summaries should include the run id, workers checked, classification,
800 typed action taken or drafted, expected ledger effect, artifact refs reviewed,
801 and next owner. Keep summaries bounded; link artifact refs instead of copying
802 full logs or transcripts.
803
804 The bundled `fleet-manager` skill mirrors this runbook for manager agents. It
805 is a first-party system skill and should be discoverable through the normal
806 skill registry after system skills are installed or refreshed.
807
808 ## Host Adapters
809
810 The Runtime host-adapter boundary supports local child processes and explicit
811 SSH workers. Host choice is Runtime placement on a worker spec, not fleet member
812 identity or a member selector. It does not authenticate the host or grant
813 access. Adapters expose the same operations: start, read status, read bounded
814 logs, interrupt, restart, stop, and cleanup.
815
816 Local workers run as child processes with stdin closed and stdout/stderr written
817 to bounded host-adapter logs. They inherit only a small safe base environment
818 such as `PATH` and explicitly allowlisted variables.
819
820 SSH workers run through the system `ssh` client with `BatchMode=yes` and a
821 bounded connect timeout. Remote environment variables are sent with OpenSSH
822 `SendEnv`; values are not embedded in the local ssh argv or fleet logs.
823
824 Host keys must already be trusted: connections use `StrictHostKeyChecking=yes`
825 and never accept a new key automatically. An explicit `known_hosts` file limits
826 trust to that file; when it is omitted, OpenSSH uses its normal known-host stores.
827 Verify the host key before adding it to either store. The legacy
828 `host_key_fingerprint` field is unsupported and is rejected; migrate it to a
829 verified `known_hosts` entry. The `identity` file selects the client login key;
830 it does not verify the remote host.
831
832 Example SSH worker spec:
833
834 ```json
835 {
836 "id": "builder-1",
837 "name": "Builder 1",
838 "host": {
839 "kind": "ssh",
840 "host": "builder.example.com",
841 "user": "codewhale",
842 "port": 22,
843 "identity": "~/.ssh/codewhale_fleet",
844 "known_hosts": "~/.ssh/codewhale_fleet_known_hosts",
845 "working_directory": "/srv/codewhale/work",
846 "env_allowlist": ["CODEWHALE_PROFILE"],
847 "codewhale_binary": "/usr/local/bin/codewhale"
848 },
849 "capabilities": ["local", "linux", "tests"],
850 "max_concurrent_tasks": 1
851 }
852 ```
853
854 Defaults are intentionally conservative:
855
856 - no hosted control plane or cloud provisioning is enabled;
857 - SSH requires an explicit host, working directory, and Codewhale binary path;
858 - secret-like environment names such as `TOKEN`, `SECRET`, `PASSWORD`,
859 `API_KEY`, and `PRIVATE_KEY` are rejected from adapter allowlists;
860 - secrets should remain in Codewhale config providers or remote host config,
861 not in task instructions, argv, or fleet logs.
862
863 ## Runtime policy and authority are not fleet identity
864
865 fleet does not define a project/workspace trust level, filesystem or network
866 reach, secret access, approval mode, sandbox, tool set, or execution authority.
867 Those belong to delegated-coordination and Runtime policy. This separation is
868 load-bearing:
869
870 - member resolution considers only member id/name, semantic role,
871 provider/model identity, and roster state;
872 - the selected identity is frozen before any authority policy is evaluated;
873 - the Runtime applies the live parent ceiling when one exists; standalone
874 fleet CLI launches instead carry an explicit bounded authority envelope,
875 and both paths remain subject to live sandbox and platform enforcement;
876 - no trust, permission, capability, secret, sandbox, approval, or tool-policy
877 value may select another member or silently change its provider/model route;
878 and
879 - receipts report requested and effective Runtime posture separately from the
880 fleet member identity.
881
882 Older persisted configuration and protocol shapes may still contain fields such as
883 `security_policy`, `trust_level`, `permissions`, `capability_grants`, secret
884 references, host authentication, environment allowlists, or tool profiles.
885 They remain deserializable for ledger replay, but new fleet run creation rejects
886 `security_policy` and worker `trust_level` rather than pretending they grant
887 authority. Their presence in old data does not make them fleet variables or
888 grants. The active Runtime remains the final authority and fails closed when a
889 requested operation cannot be enforced.
890
891 For current enforcement behavior, use [Modes](MODES.md),
892 [Sub-agents](SUBAGENTS.md), [Agent Runtime](AGENT_RUNTIME.md), and the
893 [Command Control Plane](COMMAND_CONTROL_PLANE.md). Keep secret values out of
894 task instructions, arguments, logs, and receipts; adapter and Runtime layers
895 must continue to redact or reject them independently of fleet selection.
896
897 ## Child grants: 0.10.1 scope and the 0.11 rework (#6298)
898
899 A child's authority is one grant object, `ChildGrant`
900 (`crates/tui/src/worker_profile.rs`). It shipped in v0.10.0 (966ef974e1,
901 #5633). It has `files` (none / read / write), `shell` (none / inspect /
902 verify / full), `network`, `desktop`, a tool `surface`, the caller's explicit
903 `scope`, and `spawn`. Roles are presets over it (`ChildGrant::for_role`), and
904 `ChildGrant::resolve` intersects the preset with the parent-derived profile
905 and the caller's scope. The child's tool catalog, its dispatch refusals, and
906 its capability envelope all read the same grant, so a tool the child can see
907 is a tool it can call. `desktop` is in no preset.
908
909 `ShellGrant::Verify` also shipped in v0.10.0: the Verifier preset gets the
910 bounded built-in verification surface (default workspace checks, pure test
911 selection, bounded Git fetch and merge-tree) instead of a shell grammar.
912
913 **Other fixes shipped by v0.10.0** (on top of the grant):
914
915 - Children never inherit desktop or computer-control tools (b5e48cd31, #6296).
916 - A bounded verify surface for Git: `fetch` against a configured remote name
917 and a read-only `merge_tree` (b89349286).
918 - Refusals name the sanctioned alternative and tell a child to report a
919 blocked probe to its parent instead of working around it (23747acea).
920 - One reasoning vocabulary (c2bc1244d). Token budgets are tracked but never
921 enforced (a7a8bdb33).
922
923 **0.10.1 scope.** This release adds no new grant-model code. #6298 is
924 re-scoped to the remainder below, which lands in 0.11 as its own slices.
925
926 **0.11 remainder** (one slice at a time):
927
928 1. **A `verify` shell mode for builds.** `cargo test`/`check` run under an
929 explicit, bounded write scope (`target/`, refs), so a verifier can run the
930 builds it is handed without a full shell.
931 2. **Classified tool families that fail closed.** MCP and desktop tools form a
932 labeled family. A child gets that family only when the spawn grants it with
933 a reason, and an unclassified tool is not granted.
934 3. **Legible grants.** The role picker, roster, and receipts show the effective
935 grant, model, and thinking tier in plain words.
936
937 Related work is tracked in #6015, #5633, #6194, and #6232.
938
938 lines MARKDOWN