返回 CodeWhale
AGENT_RUNTIME.md
根目录 / docs / AGENT_RUNTIME.md
1 # The Codewhale Agent Runtime — one durable substrate, familiar launchers
2
3 > 阅读简体中文版:[zh_hans/AGENT_RUNTIME.md](zh_hans/AGENT_RUNTIME.md)
4
5 This document explains how sub-agents, the headless `exec` path, Agent Fleet,
6 and Runtime relate. These concepts had drifted into *two* parallel "worker"
7 systems. The fix is to make the **Runtime worker run** the durable execution
8 primitive: fleet owns Agent identity, membership, and selection; Runtime owns
9 execution, authority, and lifecycle. "Sub-agent" remains useful product
10 vocabulary for a nested role, but it must not imply a separate execution
11 substrate with weaker lifecycle semantics. It also answers the open direction
12 question in #2972 ("how much Claude Code convergence is right?").
13
14 ## The core idea
15
16 There is exactly **one** thing that runs detached Agent work: a **headless
17 Runtime worker** with a durable execution lifecycle. It is a model loop with
18 the full, authority-gated tool surface that can, in turn, delegate child work
19 through the same lifecycle. Everything else is a way to select, launch, or
20 observe that one Runtime.
21
22 ```
23 ┌──────────────────────────────────────┐
24 │ headless Runtime │
25 │ execution · authority · lifecycle │
26 │ can spawn child workers │
27 └───────────────┬──────────────────────┘
28 │
29 ┌───────────────┴──────────────────────┐
30 │ one durable execution substrate │
31 └───────┬───────────────┬──────────────┘
32 │ launches │ launches
33 ┌──────────────┴──────┐ ┌─────┴────────────────┐ ┌──────────────────────┐
34 │ TUI turn │ │ `codewhale exec` │ │ Agent fleet │
35 │ interactive, in-proc │ │ headless CLI │ │ identity · membership │
36 │ │ │ full tools · stream │ │ · selection │
37 └─────────────────────┘ └──────────────────────┘ └──────────┬───────────┘
38 │
39 └─ selects a Runtime worker
40 ```
41
42 - A **sub-agent** is the user-facing name for a *nested assignment* with a role
43 (`explore`, `reviewer`, `implement`, `test`, ... — see the canonical seven in
44 `docs/SUBAGENTS.md`). It should be backed by
45 the same Runtime worker lifecycle used for a fleet-selected Agent. `agent` is
46 the model-facing launcher, not a second runtime.
47 - **`codewhale exec`** is the headless front door: usable by anyone at any time
48 (CI, scripts, another agent), full tools, emits a `stream-json` event stream,
49 and can spawn sub-agents. It is *the* runtime with a CLI on it.
50 - A **fleet-selected Agent** executes as a Runtime `codewhale exec` run. fleet
51 supplies identity, membership, and selection; it does not re-implement
52 execution. Runtime owns the durable ledger, scheduling/leasing/retry,
53 authority, local or SSH transport, and terminal lifecycle.
54
55 So "fleet vs sub-agent" is not a choice between execution substrates. fleet
56 answers **who** is eligible and selected, Runtime answers **how and where** the
57 authorized work executes, and sub-agent remains the role/UX vocabulary for a
58 nested assignment.
59
60 ## The cutover rule
61
62 If a detached `agent` child can fail on a one-off provider timeout with no
63 retry while an equivalent Runtime worker would retry and preserve ledger
64 evidence, then the cutover is incomplete. Treat that as a Codewhale Runtime
65 gap, not as normal "sub-agent behavior".
66
67 The compatibility `agent` runtime now retries transient provider header,
68 stream, and timeout failures with backoff before marking a worker interrupted;
69 when retries are exhausted it preserves a checkpoint and returns a continuation
70 handle. The remaining convergence work is to keep that lifecycle durable across
71 process restarts, remote execution, and full Runtime-ledger scheduling.
72
73 The target rule is:
74
75 - durable or long-running work goes through the Runtime worker lifecycle;
76 - `agent` should enqueue
77 or observe a Runtime worker run instead of owning an independent
78 lifecycle;
79 - in-process children are allowed only as a small compatibility/latency
80 optimization, and they must expose the same terminal states, retry semantics,
81 receipts, and inspection handles as the durable Runtime path.
82
83 In product language it is fine to say "open a sub-agent". In architecture
84 language that means "start a nested Runtime worker with this role", optionally
85 using a member selected from fleet.
86
87 ## Why this shape (and why it fixes the lag)
88
89 The motivating problem: spawning many in-process sub-agents made the TUI lag,
90 because each child cloned a heavy runtime and rebuilt the whole tool registry,
91 *and* the TUI rendered a full card/transcript per child.
92
93 Surveying Claude Code, Codex, and Kimi, the thing that keeps an orchestrator
94 light at high fanout is **not** a process boundary — all three run sub-agents
95 in-process. It is **isolation + a compact event stream**:
96
97 - a child's transcript **never** flows back into the parent — the parent gets a
98 result summary and a small lifecycle event stream;
99 - the UI renders **counts** (`2 running / 3 done`), not a child session per
100 worker;
101 - each worker's tool surface is built directly from a **role/capability
102 profile**, not "build everything then filter".
103
104 "Headless" therefore means *the execution is not shaped like the UI* — it does
105 **not** mean fewer abilities. A headless worker keeps the full toolset and can
106 spawn sub-agents.
107
108 When the work also needs to be **durable** (survive the TUI closing, a laptop
109 sleeping) or **remote** (SSH), Runtime runs the worker out-of-process as
110 `codewhale exec`. fleet may supply the selected Agent identity, but Runtime
111 retains execution authority and lifecycle ownership. The heavy construction
112 then lives in another process entirely, so the orchestrator stays smooth
113 regardless of fanout, and the run survives restarts — the day-scale autonomy
114 goal of #3154.
115
116 ## One recursion axis
117
118 A worker runs at `spawn_depth = 0` and may spawn children while
119 `spawn_depth + 1 ≤ max_spawn_depth`, so a budget of `N` affords `N` nested
120 delegation levels. Sub-agents and fleet-selected Runtime workers share **one**
121 axis, sourced from `codewhale_config`:
122
123 - `DEFAULT_SPAWN_DEPTH = 3` — the default budget for both standalone sub-agents
124 and fleet-selected Runtime workers (so they cannot drift into "two moving
125 targets");
126 - `MAX_SPAWN_DEPTH_CEILING = 8` — the opt-in cap that every configured Runtime
127 value, including fleet execution `max_spawn_depth`, clamps to.
128
129 The model-facing `agent` schema intentionally omits `max_depth`. The parser
130 still accepts `max_depth`, `maxDepth`, and `max_spawn_depth` for saved
131 transcripts, ACP/MCP clients, and internal compatibility callers, and rejects
132 values above 8. Current model-authored calls inherit the Runtime configuration
133 instead of negotiating recursion depth in the tool schema.
134
135 Workflow IR has a separate default structural validation limit of five nested nodes.
136 That limit constrains the orchestration document's shape; it does not grant or
137 consume Runtime child-delegation depth.
138
139 The root worker always runs even at budget 0; the budget gates *child*
140 delegation. The default affords at least three nested levels.
141
142 ## Event vocabulary
143
144 The Runtime execution ledger persists the worker's own event stream rather than
145 a separate, simulated taxonomy. Compatibility APIs and types still expose this
146 through the `Fleet...` prefix. `codewhale exec --output-format stream-json`
147 emits
148 `{"type": "content" | "tool_use" | "tool_result" | "sandbox_denied" |
149 "workflow_event" | "session_capture" | "turn_usage" | "metadata" | "done" |
150 "error"}` lines, which map onto the Runtime ledger's compatibility type
151 `FleetWorkerEventPayload` (`RunningTool`, `WorkflowEvent`, `Running`,
152 `Completed`, `Failed`, …). `workflow_event` carries the typed
153 run/phase/task/gate receipt while a Workflow is in flight and is retained as a
154 typed `WorkflowEvent` in the Runtime execution ledger; the enclosing Runtime
155 worker still owns the terminal `done` or `error`. One vocabulary, two surfaces.
156
157 `session_capture` is emitted once, when the exec run persisted its transcript
158 as a saved session, and carries the recoverable id in exactly one place:
159
160 ```json
161 {"type": "session_capture", "schema": "codewhale.exec-stream", "schema_version": 1,
162 "content": "<redacted:…>", "saved_session_id": "01J…"}
163 ```
164
165 - `saved_session_id` is the raw saved-session id, emitted only after a
166 successful save. For local Fleet workers, the parent assigns a fresh ID and
167 shares the Runtime's existing session directory. The executor advertises
168 `FleetReceipt.saved_session_id` only when that exact ID was reported and its
169 saved transcript can be loaded. A client with Runtime API access can then
170 read the reply through `GET /v1/sessions/{id}`. SSH workers retain their
171 excerpt and remote log, but do not advertise an unavailable local session
172 link. An ID is a lookup key, not a substitute for Runtime authentication.
173 - `content` is the same redacted fingerprint the terminal `metadata.session_id`
174 carries, so a captured `metadata` receipt stays safe to log on its own and
175 the two events can still be correlated. `metadata.resume_command` therefore
176 names this field (`codewhale exec --resume <session_capture.saved_session_id>`)
177 rather than carrying the id itself.
178
179 The terminal `metadata` receipt also carries the worker's visible final answer:
180 `visible_final_answer_chars` is the real character count of the final
181 assistant reply, and `visible_final_answer_excerpt` is a bounded (4,000
182 characters, `...` when cut), secret-redacted excerpt of it, omitted when the
183 current turn produced no visible answer. Resumed turns never reuse an older
184 reply, and failed/interrupted receipts may carry partial current-turn text;
185 the receipt status remains authoritative. The Runtime executor reads the excerpt from
186 this receipt — never from the streamed `content` deltas, which are the run
187 thinking out loud — and attaches it to `Completed.summary` and, for a task
188 with no scorer and no file artifact, to the receipt notes as the task's
189 deliverable. Lifecycle event labels and worker inspection summaries show a
190 short excerpt; the event `payload` and the receipt keep the full excerpt.
191
192 `turn_usage` is the per-model-call usage receipt, emitted once per model
193 request (turn-step) when the provider reported usage for that call:
194
195 ```json
196 {"type": "turn_usage", "schema": "codewhale.exec-stream", "schema_version": 1,
197 "turn": 1, "input_tokens": 1200, "output_tokens": 180,
198 "reasoning_tokens": 90, "prompt_cache_hit_tokens": 900,
199 "prompt_cache_miss_tokens": 300, "prompt_cache_write_tokens": 0,
200 "reasoning_replay_tokens": 40, "duration_ms": 1834}
201 ```
202
203 - `turn` is the 1-based index of the model call within the exec run;
204 `input_tokens`, `output_tokens`, and `duration_ms` are always present.
205 - Optional token fields are **omitted** when the provider does not report
206 them — never emitted as null and never backfilled with zeros. Field names
207 mirror the terminal `metadata` receipt: `prompt_cache_hit_tokens` is the
208 provider's cache-read count (Anthropic `cache_read_input_tokens`),
209 `prompt_cache_write_tokens` the cache-creation count
210 (`cache_creation_input_tokens`). `reasoning_tokens` appears only for
211 provider paths that report it (OpenAI-compatible
212 `completion_tokens_details` / Responses `output_tokens_details`; Anthropic
213 does not report a thinking-token count). `reasoning_replay_tokens` is a
214 client-side estimate for DeepSeek V4 interleaved-thinking replays.
215 - When a provider reports no usage at all for a call, the whole event is
216 skipped for that call. Latency/convergence analysis should sum
217 `turn_usage` events instead of inferring per-step tokens from wall time;
218 the terminal `metadata` receipt still carries the cumulative totals.
219
220 ## Convergence with Claude Code (#2972)
221
222 Codewhale should converge with Claude Code on **shape**, not on branding:
223
224 - **Adopt**: a headless runtime with a real CLI/SDK front door; sub-agents as
225 isolated runs that return summaries (not transcripts); a compact, event-driven
226 fanout projection; capability/role tool profiles; the skills ecosystem
227 (#2743); structured run receipts.
228 - **Keep distinct**: Codewhale branding and first-class DeepSeek/GLM/MiniMax/
229 multi-provider support; the local-first **Agent fleet** as the identity,
230 membership, and selection layer; durable local/SSH execution and authority in
231 Runtime; Workflow as the ordering overlay.
232 - **Do not** fork execution semantics per surface. The TUI, `agent`,
233 `exec`, and the Runtime API must all drive the *same* Runtime and observe the
234 *same* event stream. fleet selections are passed to that Runtime rather than
235 creating a second execution path — divergence there is what produced the
236 "two moving targets" this document exists to prevent.
237
238 The litmus test for any new agent surface: *does it launch and observe the one
239 runtime, or does it invent a second one?* Only the former is allowed.
240
241 ## Historical note: what remained after v0.9.0
242
243 Archived roadmap snapshot — live state is the issue tracker, not this list.
244 Refreshed 2026-08-17 from a full audit of the older 0.9-era documents. Those
245 plans are evidence, not a second source of truth. v0.9.0 consolidated the
246 underwater shell, message-first Operate, permission postures, the wired
247 Workflow engine and durable run journal, Lane CLI/runtime, setup with
248 `operate_ready`, constitution rebalance, and ProviderLake/Models.dev. The
249 remaining work belongs to later releases:
250
251 1. **Rebrand completion** — the `deepseek`/`deepseek-tui` binary shims and
252 shim release assets were removed in v0.9.0; the remaining obligation is the
253 Homebrew `codewhale` formula rollout (`docs/REBRAND.md`).
254 2. **Operate as a value stream** — a control-board surface over the underwater
255 shell (WIP, queue age, bottleneck); phase history (#4039); Workrooms Phase 2
256 (#3209/#3210) as the inbox substrate;
257 receipt reconciliation.
258 3. **Flow control** — real WIP limits and visible queues (#4015, #4016),
259 reconciled with the shipped 16-concurrent/1k-run access model (#4292).
260 4. **fleet identity and Runtime/Workflow convergence residuals** — live
261 tmux/verifier-gate dogfood closing #4175/#4177/#4178/#4179; fleet consuming
262 canonical AgentProfiles and selecting members while Runtime owns execution;
263 Conductor/topology (#4010, #4012) as stretch.
264 5. **TTC design implementation** (design doc in `codewhale-ops`) — approved and now unblocked after v0.9.0.
265 6. **HarnessProfile completion** — the status/UX display lane
266 (`docs/rfcs/HARNESS_PROFILE_CUTLINE.md`).
267 7. **File decomposition, landed** — the v0.9.0-era offenders were split out:
268 `main.rs` is now a thin stub and `ui.rs` has been decomposed into focused
269 modules under `crates/tui/src/tui/` (~3.9k lines today; the
270 `docs/rfcs/FILE_DECOMPOSITION_0_9_0.md` figures are the 0.9.0-era snapshot).
271 The remaining work is the "thin TUI over core" north star tracked in
272 `POST_0_9_1_SEAMS.md`.
273
274 Explicitly deferred by their own documents: external workflow memory (boundary
275 only), automatic harness evolution, hosted workrooms, `constitution_modules`
276 (needs sign-off), permission profiles (#3211, needs design), and plan-ceiling
277 probing (needs a product decision).
278
279 ## Public launch contract for an external harness (#4641)
280
281 An external evaluation harness (for example a future Verifiers v1 built-in
282 harness) embeds Codewhale by launching the public `codewhale exec` front door
283 against an interception endpoint it owns. Codewhale owns only its **launch
284 contract**; the harness owns interception, traces, model-call timing, token
285 accounting, retries, rollout limits, and runtime orchestration. Do not add a
286 harness runtime, trace parser, or receipt schema to Codewhale.
287
288 A reproducible headless launch uses only existing generic surfaces:
289
290 - an explicit temporary config that names the route and the credential
291 **environment variable**, never the secret itself:
292
293 ```toml
294 provider = "openai"
295
296 [providers.openai]
297 base_url = "" # the harness fills in its interception endpoint
298 model = "" # the harness fills in the target model
299 api_key_env = "VF_CODEWHALE_API_KEY"
300 ```
301
302 - `CODEWHALE_HOME` set to a fresh per-run directory;
303 - `CODEWHALE_SECRET_BACKEND=file`;
304 - `CODEWHALE_MCP_CONFIG` pointing to a generated per-run MCP JSON file that
305 contains only the task servers the harness supplies
306 (`{"mcpServers":{"task-tools":{"url":""}}}`; the `mcpServers` alias and
307 URL-based Streamable HTTP / SSE transports already exist);
308 - `CODEWHALE_MEMORY=false` and `CODEWHALE_TELEMETRY=false`. The 0.9.12 source
309 defaults usage counting on with an opt-out. Every sealed harness explicitly
310 sets the run-scoped kill switch so a test cannot collect or send from a
311 fresh or reused home. Ordinary enabled sessions send aggregate counts to an endpoint
312 (`https://telemetry.codewhale.net/v1/telemetry`, the shipped default) rather
313 than to a local file. It is a hard floor — an explicit "off" in the
314 environment beats `--telemetry true` and `telemetry = true` in config. Set
315 `CODEWHALE_TELEMETRY_ENDPOINT=` (empty) instead if a harness wants an enabled
316 home to keep buffering locally without contacting anything. See
317 [`docs/TELEMETRY.md`](TELEMETRY.md);
318 - `CODEWHALE_ALLOW_INSECURE_HTTP=1` **only** when the harness supplies a
319 trusted `http://` interception endpoint (container/tunnel endpoints are not
320 always loopback);
321 - `--append-system-prompt` and `--disallowed-tools` when the caller supplies
322 them.
323
324 The interception secret stays in the child environment (resolved through the
325 route's `api_key_env`); it is never written into argv, the route config, logs,
326 the `stream-json` stream, or any generated file.
327
328 The exact argument order is:
329
330 ```sh
331 codewhale \
332 --config .vf-codewhale/config.toml \
333 --workspace . \
334 --no-project-config \
335 --skip-onboarding \
336 exec \
337 --auto \
338 --sandbox danger-full-access \
339 --output-format stream-json \
340 -- "<task prompt>"
341 ```
342
343 `--no-project-config` must appear **before** the subcommand (like
344 `--skip-onboarding`). The public dispatcher parses it and forwards it ahead of
345 the TUI subcommand; `Exec` then skips the workspace-specific
346 `[workspace]`/`[projects]` user-config overlay so the config surface depends
347 only on the explicit `--config`. `crates/tui/tests/integration/verifiers_harness_contract.rs`
348 is the provider-free acceptance lock for this contract.
349
350 ### Future upstream checklist (out of scope here — do not run)
351
352 Actually adding Codewhale as a built-in harness lives in the external Verifiers
353 repository; the public, immutable Codewhale GitHub Releases with checksum
354 manifests it needs have existed since v0.9.1 (latest published release is
355 v0.9.13, published 2026-09-14; the workspace source version is 0.9.13).
356 That upstream change is expected to be limited to a new
357 `verifiers/v1/harnesses/codewhale/` package plus its test-matrix and docs
358 registration, with `CodewhaleHarnessConfig` pinning the target release,
359 `setup()` downloading and verifying the released archive, and `launch()`
360 writing the temporary route/MCP files above and calling `runtime.run_program(...)`.
361
362 Holdouts, explicitly **not** performed by this contract work: tagging,
363 publishing, or creating a Codewhale release; opening or submitting the upstream
364 Verifiers PR; running its credentialed E2E matrix; or claiming
365 runtime/architecture support before the exact released archive has run in that
366 upstream runtime.
367
367 lines MARKDOWN