返回 CodeWhale
HOOKS.md
根目录 / docs / HOOKS.md
1 # Hooks
2 > 阅读简体中文版:[zh_hans/HOOKS.md](zh_hans/HOOKS.md)
3
4 Hooks run a shell command when the Codewhale **TUI** reaches a lifecycle
5 point. They are plain processes: they receive context through environment
6 variables, some receive a JSON payload on stdin, and three of them can steer
7 what Codewhale does next.
8
9 This page is the authoritative reference for what is implemented today.
10 Configuration syntax that overlaps with the rest of `config.toml` lives in
11 [CONFIGURATION.md](CONFIGURATION.md); this file is the event-by-event
12 contract.
13
14 ## Scope
15
16 Hooks fire in the interactive TUI and in the engine turn loop, which the
17 Runtime API threads behind the desktop app and web drive as well.
18
19 | Surface | Fires hooks |
20 | --- | --- |
21 | `codewhale` / `codew` interactive TUI | yes |
22 | `codewhale exec` (headless one-shot) | opt-in: `--hooks` fires `tool_call_before` and `shell_env` |
23 | Runtime API threads (desktop app, web) | yes: `tool_call_before`, `shell_env`, `tool_call_after`, `on_error`; `GET /v1/hooks` lists the set |
24 | the `codewhale` CLI dispatcher and its subcommands | no |
25 | app-server / ACP | no |
26 | the `workflow` tool and sub-agent *internals* | no — but the TUI fires `subagent_spawn` / `subagent_complete` around them |
27 | public API | there is none |
28
29 The `crates/hooks` event-sink crate in this repository is an unrelated
30 internal mechanism. It shares no configuration, no event names, and no
31 contract with the hooks described here.
32
33 ### `codewhale exec --hooks`
34
35 Headless runs fire no hooks by default — a CI job should not start paging an
36 on-call rotation merely because a config exists. `codewhale exec --hooks`
37 opts the run in. The engine-side events are `tool_call_before` (exit code 2
38 still denies the call; `ask` resolves fail-closed because nothing can prompt
39 headlessly) and `shell_env`. UI-driven events such as `session_start`,
40 `message_submit`, and `turn_end` do not fire — they live in the interactive
41 shell, not the turn loop. Fleet worker subprocesses never fire operator
42 hooks. Independently of this flag, `permissions.toml` typed rules already
43 apply to `exec` — the run drives the same turn loop, and a `deny` blocks in
44 every mode.
45
46 ## Quick start
47
48 ```toml
49 # ~/.codewhale/config.toml
50 [hooks]
51 enabled = true
52
53 [[hooks.hooks]]
54 name = "announce"
55 event = "session_start"
56 command = "echo 'Codewhale session started'"
57 ```
58
59 Run `/hooks` in the TUI to list what is configured, whether the global switch
60 is on, and any entry that was rejected at load. Run `/hooks events` for the
61 event names.
62
63 ## Configuration
64
65 ```toml
66 [hooks]
67 enabled = true # global switch; false suppresses every hook
68 default_timeout_secs = 30 # see the timeout note below
69 working_dir = "/path/to/dir" # default: the session workspace
70
71 [[hooks.hooks]]
72 event = "tool_call_before" # required; one of the 15 names below
73 command = "~/.codewhale/hooks/gate.sh" # required; `sh -c` on Unix, `cmd /C` on Windows
74 name = "gate" # optional label for /hooks and log lines
75 timeout_secs = 30 # optional, default 30
76 background = false # optional; foreground inside the hook worker
77 continue_on_error = true # optional, default true
78 condition = { type = "tool_name", name = "bash" } # optional
79 ```
80
81 `timeout_secs` note, stated as implemented: when `[hooks].default_timeout_secs`
82 is set it **overrides** every hook's own `timeout_secs`, it does not merely
83 supply a default for hooks that omit one. Leave it unset if you want per-hook
84 timeouts to apply. `/hooks list` shows the timeout the runtime will actually
85 apply, and names the override when one is in force.
86
87 `default_timeout_secs = 0` is **rejected at load**. Because the value replaces
88 every hook's own `timeout_secs`, a zero there would expire every hook in the
89 config immediately — including a `tool_call_before` gate, which then denies
90 every matching tool call. The override is ignored, per-hook `timeout_secs`
91 applies, the hooks themselves still load, and the rejection is reported by
92 `/hooks list` under *configuration problems*. Per-hook `timeout_secs = 0` is
93 rejected too, but that only drops the one hook that wrote it.
94
95 Hooks run with the workspace (or `working_dir`) as the current directory.
96
97 ### Timeouts
98
99 The timeout applies to **foreground and background hooks alike**. When it
100 expires:
101
102 - the hook's whole process group is killed — Unix process groups, Windows Job
103 Objects — so a hook that spawns children does not outlive its budget;
104 - the child is then reaped, so nothing is normally left detached or zombied;
105 - a foreground hook's result is `success = false`, `exit_code = None`, empty
106 `stdout`/`stderr`, and `error = "Hook timed out after Ns"`;
107 - a background hook's timeout is logged at `warn` under the `hooks` target.
108 Nothing is reported to the caller, because the caller stopped waiting the
109 moment it submitted the hook.
110
111 **Termination is best-effort, and the bound that is guaranteed is Codewhale's,
112 not the OS's.** The kill can fail to land — a process wedged in an
113 uninterruptible state on Unix, a `TerminateJobObject` a protected process
114 survives on Windows — and no user-space program can promise otherwise. What
115 Codewhale does guarantee is that it stops waiting: the containment handle is
116 released (which re-signals the Unix process group and closes the kill-on-close
117 Windows Job Object) and the reap gets one short bounded window. If the child
118 still cannot be confirmed dead, that is logged at `warn` and the foreground
119 result says so — `error = "hook could not be reaped after its timeout"` rather
120 than the stronger timeout wording. So a timed-out hook never blocks the turn,
121 but treat "killed" as best-effort rather than absolute.
122
123 ### Background hooks
124
125 `background = true` describes real scheduling, not just a config flag. A
126 background hook is **submitted, never awaited**:
127
128 - it is enqueued without blocking into a fixed 32-entry supervisor queue,
129 drained by two persistent workers that apply the timeout above; saturation
130 or supervisor loss is a failed submission, and no invocation creates its
131 own detached supervisor thread;
132 - it receives the same environment variables and the same stdin JSON payload
133 as the foreground form of that event — the payload contract does not change,
134 only the steering does;
135 - its stdout and stderr are discarded (`Stdio::null()`), so it can never
136 return a verdict;
137 - the `HookResult` the runtime hands its caller is flagged as a background
138 submission and carries no exit code. Steering code reads
139 `observed_exit_code()`, which is `None` for a background hook, so a
140 background hook can never allow, deny, ask, or rewrite anything.
141
142 `shell_env` ignores `background` entirely — its stdout *is* the contract, so
143 it always runs in the foreground. `/hooks list` reports that as a
144 configuration warning, and does not label the hook `[bg]`.
145
146 Observer-only UI events are submitted with non-blocking `try_send` to one
147 32-entry queue drained by two persistent workers. A configured foreground
148 observer is still awaited in config order inside a worker, but the terminal
149 event loop never waits on its process and never creates a thread per event.
150 Queue saturation or dispatcher loss drops that observer event and produces an
151 event-specific error toast that survives an agent's ordinary progress-status
152 update. Steering events retain their gate or transform semantics:
153 fresh/queued `message_submit` dispatch reports through a bounded result
154 channel, same-turn steering runs that transform on the blocking worker before
155 calling the engine steer path, and `tool_call_before` / `shell_env` execute on
156 the engine or tool worker rather than the terminal event loop.
157
158 ### The hook process environment
159
160 A hook command inherits the environment of the Codewhale process, plus the
161 `DEEPSEEK_*` variables for its event. Codewhale does not filter that
162 inheritance, so treat a hook exactly as you would treat any command you type
163 in the same shell that launched Codewhale: whatever is exported there is
164 visible to it.
165
166 This is *not* true of the command a `shell_env` hook feeds — see
167 [`shell_env`](#shell_env) for the bounded allowlist that governs a **local**
168 `exec_shell`, and for what changes when an external sandbox backend is
169 configured instead (the backend owns its base environment, and your
170 `shell_env` values are transmitted to it).
171
172 ### Conditions
173
174 | Condition | Matches | Supported on |
175 | --- | --- | --- |
176 | `{ type = "always" }` | every invocation (also the default when omitted) | every event |
177 | `{ type = "tool_name", name = "bash" }` | exact tool name; `*` globs are supported, e.g. `mcp__*`. The shell tool's spellings `bash`, `Bash`, and `exec_shell` are aliases: a condition naming any one matches all three | `tool_call_before`, `tool_call_after`, `shell_env`, `on_error` |
178 | `{ type = "tool_category", category = "shell" }` | tool category | `tool_call_before`, `tool_call_after`, `shell_env`, `on_error` |
179 | `{ type = "mode", mode = "plan" }` | the context's mode string, case-insensitive | every event **except** `shell_env` |
180 | `{ type = "exit_code", code = 1 }` | the exit code the tool actually reported | `tool_call_after`, `on_error` |
181 | `{ type = "all", conditions = [...] }` | every nested condition | every event |
182 | `{ type = "any", conditions = [...] }` | at least one nested condition | every event |
183
184 Three rules keep conditions from lying:
185
186 - **`exit_code` needs a real exit code.** It matches only when the event
187 actually observed a process exit code — `tool_call_after`, or `on_error` for
188 a tool failure, in both cases for a process-backed tool such as `bash`.
189 A command that exits nonzero reports its code too, even though `bash`
190 returns it as a failed call. A timed-out or killed command usually has no
191 exit code; `DEEPSEEK_TOOL_STATUS` says which it was.
192 A tool that reports no exit code never matches an `exit_code` condition; the
193 condition is not satisfied by a default, a zero, or a success flag. The value
194 is a 64-bit integer, so a Windows crash code such as `3221225477`
195 (`0xC0000005`) is matchable.
196 - **Tool-scoped `on_error` hooks are supported.** `on_error` fires for
197 transport and capacity errors *and* for tool failures; the tool-failure
198 firing carries the tool name, call id, result, and reported exit code. A
199 `tool_name` / `tool_category` / `exit_code` condition on `on_error` is
200 therefore a valid configuration. An `on_error` firing with no tool behind it
201 simply does not match such a condition — it is skipped at dispatch, not
202 rejected at load.
203 - **Unsupported conditions are rejected at load.** A condition that references
204 context its event never carries can never match, and a hook wearing one is
205 silently inert — the dangerous form of that is a `deny` gate the operator
206 believes is armed. Codewhale drops those hooks at load, logs the reason
207 under the `hooks` tracing target, and shows them in `/hooks list` as
208 `rejected:`. Nested predicates inside `all` / `any` are checked too. A hook
209 with `timeout_secs = 0` or an empty `command` is rejected the same way.
210 Rejection is **per entry**: a broken hook never takes another one with it,
211 even when the two share a `name` or are both unnamed.
212
213 ### Project-local hooks
214
215 A repository may ship `<workspace>/.codewhale/hooks.toml` using the same shape,
216 but only its `[[hooks]]` entries are merged — a project file cannot change
217 `enabled`, `default_timeout_secs`, or `working_dir`, which always come from your
218 own config. Because hooks are executable configuration, project hooks load
219 **only** after both workspace trust and separate approval of the exact hooks file
220 in user-owned config. Use `/hooks review` to inspect the commands and digest,
221 then `/hooks approve <digest>` to enable those bytes on the next session. Review
222 any scripts the commands call too. A file change requires another approval.
223 `/hooks revoke` blocks future and queued launches; it does not stop commands
224 already running. Session `/trust on` alone does not enable project hooks.
225 Approved project hooks are appended
226 after global hooks, so they run last and win `updatedInput` ties. A malformed
227 trusted project file logs a warning and Codewhale falls back to global hooks
228 only. Validation runs over the merged set, so a rejected project hook is
229 reported the same way a rejected global one is.
230
231 ## The 15 events
232
233 | Event | Fires | Steering |
234 | --- | --- | --- |
235 | `session_start` | once, after the engine is up and before the first draw | observer |
236 | `session_end` | once, on graceful shutdown | observer |
237 | `turn_end` | after a turn completes and post-turn state is updated | observer |
238 | `message_submit` | before a submitted message reaches history or the model | **can replace or block the text** |
239 | `tool_call_before` | before each tool call executes | **can allow / deny / ask, rewrite input, add context** |
240 | `tool_call_after` | after each tool result settles, including completions the transcript does not redraw | observer |
241 | `mode_change` | on every applied Plan/Work/Operate transition (`Act` is a compatibility alias for Work) | observer |
242 | `on_error` | on transport, capacity, and auth errors, and on tool failures | observer |
243 | `subagent_spawn` | when a sub-agent starts | observer |
244 | `subagent_complete` | when a sub-agent completes, fails, or is cancelled | observer |
245 | `shell_env` | immediately before each `exec_shell` invocation | **contributes environment variables** |
246 | `session_idle` | when the session settles back to idle after a turn or a wait — no prompt, approval, or continuation outstanding | observer |
247 | `session_error` | when a turn ends in a terminal failure; transient tool failures the agent absorbs never fire it | observer |
248 | `waiting_for_user` | when the agent starts waiting on you: an approval prompt opens, a `request_user_input` question is presented, or a goal continuation is parked between passes | observer |
249 | `session_busy` | when an idle or waiting session begins or resumes work; startup and repeated observations of the same state stay silent | observer |
250
251 `waiting_for_user`'s payload carries `reason`: `approval`, `user_input`, or
252 `goal_continuation`. All three state events carry `from`/`to` transition fields;
253 `session_idle` also carries `last_turn_status` when known, and `session_error` carries
254 the bounded terminal `error` text. Busy, idle, and waiting map onto the session
255 states the control socket's `status` verb already publishes
256 (`idle` / `in_progress` / `waiting`), so a hook and a supervisor never
257 disagree about what the session is doing. Hook authors that want opencode's
258 grace-period semantics for error alerts should debounce inside the hook —
259 `session_error` already excludes absorbed, transient failures, and a turn
260 that fails and is retried by the operator fires again only if the retry also
261 ends failed.
262
263 ### What "observer" means, exactly
264
265 Observer means Codewhale ignores the hook's **result**: stdout is discarded, a
266 non-zero exit is logged as a warning, and nothing about the turn, the tool
267 result, the sub-agent, or the error changes because of it.
268
269 Observer does **not** mean side-effect-free. An observer hook is an arbitrary
270 shell command running with your credentials. It can write files, push commits,
271 page an on-call rotation, or delete the workspace. The only thing it cannot do
272 is change what Codewhale itself does next.
273
274 The steering allowlist is exactly three events — `message_submit`,
275 `tool_call_before`, `shell_env` — and it is asserted by a test over every
276 variant, so a new event defaults to observer.
277
278 ### Session identity
279
280 Every event in one TUI session carries the same `DEEPSEEK_SESSION_ID`. The id
281 is minted once at launch, in the form `sess_xxxxxxxx`, and it survives a
282 workspace switch and a trust decision that adds project hooks — both reload
283 the hook set without starting a new session. Engine-fired `tool_call_before`
284 reports the same id as the UI-fired events, so tool records correlate with the
285 session records around them.
286
287 `session_end` fires after the queued startup-default writes have been drained
288 and while the app is still live, so it observes the settled end state rather
289 than a half-torn-down one.
290
291 ## Environment variables
292
293 Every hook receives the subset of these that applies to its event. The
294 `DEEPSEEK_` prefix is retained for compatibility with hooks written before the
295 rebrand.
296
297 | Variable | Set for | Notes |
298 | --- | --- | --- |
299 | `DEEPSEEK_SESSION_ID` | every event except `shell_env` | `sess_xxxxxxxx`, stable for the whole session |
300 | `DEEPSEEK_WORKSPACE` | every event except `shell_env` | absolute workspace path |
301 | `DEEPSEEK_MODEL` | every event except `shell_env` | active model id |
302 | `DEEPSEEK_MODE` | every event except `shell_env` | see the mode-spelling note below |
303 | `DEEPSEEK_TOTAL_TOKENS` | UI-fired events | session token total at fire time |
304 | `DEEPSEEK_MESSAGE` | `message_submit`, `subagent_*` | truncated at 5 000 bytes with a `...[truncated]` marker |
305 | `DEEPSEEK_ERROR` | `on_error` | error message, truncated at 5 000 bytes |
306 | `DEEPSEEK_PREVIOUS_MODE` | `mode_change` | mode label before the change |
307 | `DEEPSEEK_TOOL_NAME` | `tool_call_before`, `tool_call_after`, `shell_env`, `on_error` (tool failures) | |
308 | `DEEPSEEK_TOOL_CALL_ID` | `tool_call_before`, `tool_call_after`, `on_error` (tool failures) | engine call id; correlates before/after/error for one call |
309 | `DEEPSEEK_TOOL_ARGS` | `tool_call_before`, `shell_env` | tool input JSON preview, capped at 10 000 bytes |
310 | `DEEPSEEK_TOOL_RESULT` | `tool_call_after`, `on_error` (tool failures) | truncated at 10 000 bytes |
311 | `DEEPSEEK_TOOL_SUCCESS` | `tool_call_after`, `on_error` (tool failures) | `true` / `false` |
312 | `DEEPSEEK_TOOL_EXIT_CODE` | `tool_call_after` and `on_error` **when the tool reported one** | absent otherwise — never synthesized; set for a failing command as well as a passing one; 64-bit, so Windows crash codes such as `3221225477` survive |
313 | `DEEPSEEK_TOOL_STATUS` | `tool_call_after` and `on_error` **when a shell tool reported one** | `completed`, `failed`, `timed_out`, `killed`, or `running` (moved to the background); absent for other tools |
314 | `DEEPSEEK_TOOL_EXECUTION_RECEIPT` | `tool_call_after` and `on_error` **for a settled, local, foreground shell run** | complete JSON, at most 32 KiB, or absent; see [Execution receipt](#execution-receipt) |
315 | `DEEPSEEK_SESSION_COST` | when cost is supplied | USD, six decimal places |
316
317 ### Execution receipt
318
319 `DEEPSEEK_TOOL_EXECUTION_RECEIPT` says what a shell tool (`bash`, `Bash`,
320 `exec_shell`) actually ran. The before-hook input is not the same thing: a
321 `tool_call_before` hook can rewrite it. The receipt is built from what the
322 process manager recorded when it spawned the process, after admission and
323 any rewrite.
324
325 ```json
326 {"schema_version":1,"command":"printf hello","cwd":"/absolute/workspace","state":"completed","scope":"local","exit_code":0,"stdout":"hello","stderr":"","stdout_truncated":false,"stderr_truncated":false,"output_kind":"separate"}
327 ```
328
329 | Field | Meaning |
330 | --- | --- |
331 | `command` | the admitted shell source handed to the shell, not the shell executable or its argv wrapper |
332 | `cwd` | the canonical absolute path of the directory the process started in: symlinks are resolved, so a directory has one spelling whether or not the call passed `cwd`; it is resolved before spawn and that same path is handed to the OS |
333 | `state` | `completed` for an observed exit, including a nonzero one; `interrupted` for a signal, kill, cancel, or timeout |
334 | `scope` | always `local` in schema 1 |
335 | `exit_code` | the observed integer, or `null`; never synthesized from `state` |
336 | `stdout`, `stderr` | previews of the tool's retained output, which may already omit early process output; long previews keep their first and last bytes around a `[receipt preview truncated]` marker |
337 | `stdout_truncated`, `stderr_truncated` | `true` when the tool's own output capture or the preview dropped bytes |
338 | `output_kind` | `separate` for `Bash` / `exec_shell`; `combined` for lowercase `bash`, whose stdout and stderr share one pipe — `stdout` then holds the combined preview and `stderr` is empty |
339
340 The rules are conservative:
341
342 - **Exact or absent.** `command` and `cwd` are never truncated. If either is
343 over 8 KiB, contains NUL, or the directory is relative, not UTF-8, or cannot
344 resolve before spawn, the receipt is left out. So is a run whose end the shell
345 tool could not observe (the OS wait itself failed): its state is unknown,
346 and the receipt does not guess it. Previews shrink until the serialized JSON fits 32 KiB;
347 if it still cannot fit, the receipt is left out rather than cut.
348 - **Absence means nothing.** It implies neither success nor failure. An inherited
349 `DEEPSEEK_TOOL_EXECUTION_RECEIPT` is cleared before applying the current call's context.
350 - **Scope.** A receipt is built only while a `tool_call_after` or `on_error`
351 hook is configured, and only for a settled, pipe-backed, unsandboxed, local
352 foreground run. Background launches, a foreground run moved to `/jobs`,
353 PTY (`tty` / `combined_output`) and interactive sessions, OS-sandboxed and
354 external-backend execution, the read-only shell's hardened argv, Windows,
355 a PowerShell shell on any platform (it wraps the source or runs it from a
356 temporary script), and calls refused before execution have none.
357 - **Hooks only.** The receipt is not kept in the durable Runtime API item
358 record; that record already carries the tool output.
359 - It is set for a failed run as well as a passing one, so `on_error` for a
360 failed shell call carries it too. Every other variable is unchanged.
361
362 For `tool_call_after`, the same execution evidence is also delivered as a
363 versioned JSON document on stdin, in foreground and background form:
364
365 ```json
366 {"schema_version":1,"event":"tool_call_after","tool_name":"bash","session_id":"session-id","tool_call_id":"call-id","session_id_truncated":false,"tool_call_id_truncated":false,"tool_name_truncated":false,"execution_receipt":{"schema_version":1,"command":"printf hello","cwd":"/absolute/workspace","command_truncated":false,"cwd_truncated":false,"execution":"started","completion":"completed","exit_code":0,"stdout":"hello","stderr":"","stdout_truncated":false,"stderr_truncated":false,"output_mode":"combined"}}
367 ```
368
369 The stdin projection leaves the environment receipt above unchanged.
370 `completion` is the observed `completed`, `failed`, `killed`, or `timed_out`
371 status; a nonzero exit is `failed`. `exit_code` remains a signed 64-bit integer
372 or `null`. `output_mode` carries the existing receipt's `output_kind`:
373 `separate` or `combined`. With combined output, empty `stderr` does not mean
374 the command wrote nothing to stderr.
375
376 The complete document is capped at 64 KiB. Correlation identifiers are capped
377 at 1,024 UTF-8 bytes plus a truncation marker and carry individual flags;
378 missing identifiers are `null`. The shell name is exact. Execution `command`
379 and `cwd` retain the exact-or-absent rule above, so their truncation flags are
380 always false. Output truncation flags include both capture and preview loss.
381
382 Only a native shell call with a valid, settled receipt gets this stdin
383 document. Unsupported or unobserved paths get no document, and absence still
384 means unknown. Hooks remain observers: their stdout cannot allow, deny, or
385 rewrite the completed call, and background hooks are not awaited. `on_error`
386 continues to receive the environment receipt only.
387
388 **Mode-spelling note.** UI-fired events (`session_start`, `session_end`,
389 `message_submit`, `tool_call_after`, `mode_change`, `on_error`, `turn_end`,
390 `subagent_*`, `session_busy`, `session_idle`, `session_error`, `waiting_for_user`)
391 set `DEEPSEEK_MODE` to the UI label — `ACT`, `PLAN`, `OPERATE`.
392 `tool_call_before` fires inside the engine and uses the engine's own mode
393 spelling (`Agent`, `Plan`, `Operate`). `mode` conditions compare
394 case-insensitively, so `{ type = "mode", mode = "plan" }` matches both, but a
395 hook that string-matches `$DEEPSEEK_MODE` exactly should accept both spellings.
396
397 **`shell_env` is the narrow one.** It receives only `DEEPSEEK_TOOL_NAME` and
398 `DEEPSEEK_TOOL_ARGS` — no session id, workspace, model, or mode. A
399 `{ type = "mode", … }` condition on a `shell_env` hook is therefore rejected at
400 load; scope those with `tool_name` or `tool_category` instead.
401
402 ## Steering events
403
404 ### `message_submit`
405
406 Receives JSON on stdin and may rewrite or block the submitted text.
407
408 ```json
409 {
410 "event": "message_submit",
411 "text": "original user text",
412 "text_bytes": 18,
413 "text_original_bytes": 18,
414 "text_truncated": false,
415 "session_id": "sess_12345678",
416 "workspace": "/path/to/workspace",
417 "mode": "ACT",
418 "model": "deepseek-chat",
419 "total_tokens": 1234
420 }
421 ```
422
423 The complete serialized stdin document is capped at 32 KiB. `text` is the
424 largest deterministic UTF-8 prefix that fits after JSON escaping and bounded
425 metadata are included. `text_original_bytes` records the producer's full byte
426 length, `text_bytes` records the retained prefix, and `text_truncated` states
427 whether they differ. This same boundary applies to immediate input, restored
428 queue entries, merged steers, and text produced by an earlier hook.
429
430 - exit `0` printing `{"text": "..."}` with a non-empty string replaces the text
431 - exit `0` with empty stdout, or JSON without `text`, leaves the text unchanged
432 - `{"text": ""}` or a replacement over 32 000 characters is invalid stdout,
433 logged and ignored
434 - exit `2` blocks the submission before history or dispatch; a structured
435 `reason` field supplies a bounded, redacted message shown in the TUI.
436 Unstructured stdout/stderr/error output is never copied into the denial
437 - other non-zero exits follow `continue_on_error`: `true` warns and continues,
438 `false` blocks the submission
439 - `background = true` makes the hook observer-only — it still receives this
440 bounded payload on stdin, but it cannot transform or block
441
442 Multiple `message_submit` hooks run in config order and each sees the previous
443 hook's output.
444
445 ### `tool_call_before`
446
447 Receives the tool context in environment variables and may print a JSON
448 decision on stdout with exit `0`:
449
450 ```json
451 {
452 "decision": "allow",
453 "reason": "human-readable explanation, used for deny",
454 "updatedInput": { "command": "ls -la" },
455 "additionalContext": "text appended to the tool result for the model"
456 }
457 ```
458
459 - `deny` blocks the tool; the model gets a permission-denied result carrying
460 `reason`
461 - `ask` forces the interactive approval prompt in Ask and Auto-Review. Full
462 Access does not open tool-approval prompts, so `ask` does not downgrade it
463 - `updatedInput` must be an object no larger than 32 KiB serialized and
464 replaces the tool input; last hook wins
465 - `additionalContext` is appended to the tool result as `[hook context] ...`;
466 multiple hooks concatenate
467 - `reason` and `additionalContext` are bounded and sanitized before use: each
468 field is capped at 2 000 characters, the concatenated context for one tool
469 call is capped at 8 000, control characters are stripped (so hook stdout
470 cannot repaint the TUI or forge structure in the transcript), and a clipped
471 value carries a `…[truncated]` marker. What a hook adds to the turn's context
472 budget is therefore bounded no matter what it prints
473 - exit `2` is a legacy hard deny and wins regardless of stdout
474 - empty stdout, non-JSON stdout, and JSON without `decision` all mean allow
475 - precedence across matching hooks: no-verdict-with-`continue_on_error = false`
476 > deny > ask > allow
477 - `background = true` hooks are submitted and never awaited, so they have no
478 verdict and cannot steer; Codewhale logs a warning when one is configured
479 for this event
480
481 **A gate that could not answer is not permission.** If a foreground
482 `tool_call_before` hook produces no verdict — it timed out, the process could
483 not be started, or a strict process exited non-zero without an explicit JSON
484 decision — and *that hook* is configured with
485 `continue_on_error = false`, the tool call is denied. Strictness is read off
486 the hook that actually ran, not off the event: a strict `write_file` gate whose
487 condition did not match an `exec_shell` call has no say in whether that call
488 proceeds, and a lenient hook's timeout never denies just because some other
489 strict hook exists in config. Every no-verdict outcome is logged either way.
490
491 The denial message names the hook and the reason and nothing else: the hook
492 name is truncated, the detail is truncated, control characters are stripped,
493 and a spawn failure is reported by error kind (`NotFound`,
494 `PermissionDenied`, …) rather than by echoing the command line or the resolved
495 interpreter path.
496
497 ### `shell_env`
498
499 Runs synchronously before each `exec_shell` and its stdout is parsed as
500 `KEY=VALUE` lines. A leading `export ` is stripped, `#` comment lines and blank
501 lines are skipped, and a matching pair of surrounding single or double quotes is
502 removed from the value. Later hooks override earlier ones. Use it for ephemeral
503 credentials, per-skill `PATH` adjustments, or short-lived tokens.
504
505 `background` is ignored for this event: the hook always runs in the foreground
506 because its stdout is the contract.
507
508 An entry a shell cannot carry is dropped rather than allowed to break the tool
509 call: an empty name, a name containing whitespace, `=`, a control character, or
510 a NUL; a value containing a NUL; a value over 32 KiB; and anything past 256 KiB
511 of accumulated output from one hook. Each drop is logged by key name only. A
512 `shell_env` hook is an ordinary process whose stdout can contain anything —
513 "the hook printed something odd" must never become "the `exec_shell` call
514 aborted".
515
516 **Exactly what the shell command ends up with — local execution.** When
517 `exec_shell` runs the command locally (the default), it does not inherit
518 Codewhale's ambient environment. Its environment is built as:
519
520 1. a sanitized fixed allowlist of parent variables — `PATH`, `HOME`, `USER`,
521 `LANG` and the other `LC_*`/locale entries, `TERM`, `SHELL`, `TMPDIR`,
522 proxy variables, color/terminal entries such as `NO_COLOR`, `CARGO_HOME`/`RUSTUP_HOME`/`RUSTUP_TOOLCHAIN`, the Windows system and MSVC toolchain entries, and other platform keys (the full list lives in `crates/tui/src/child_env.rs`) — and nothing else. Variables
523 outside that allowlist, including anything that looks like a secret, are
524 dropped;
525 2. then the `KEY=VALUE` pairs your `shell_env` hooks produced, applied on top.
526 These are explicit values you configured, so they win over the allowlist.
527
528 So a `shell_env` hook is the supported way to get a credential into one local
529 `exec_shell` invocation. Ambient secrets exported in the terminal that launched
530 Codewhale are **not** forwarded to a local `exec_shell` on their own.
531
532 **With an external sandbox backend configured, the allowlist above is not the
533 contract.** If `exec_shell` is routed to a configured sandbox/execution
534 backend, Codewhale does not construct the process environment at all: it hands
535 the command and your `shell_env` values to the backend as extra environment
536 variables, and the **backend owns its own base environment**. What is present
537 besides your values — an image's baked-in variables, the backend's own
538 injections, whatever a remote runner exports — is determined by that backend,
539 not by the list above. Do not assume the local allowlist applies there.
540
541 Disclosure, because it is the part that matters for a hook that emits
542 credentials: **`shell_env` values are transmitted to the configured backend.**
543 For a remote or containerized backend that means the values leave this machine
544 and are subject to that backend's logging, retention, and access controls.
545 Codewhale's own audit log still records key names only, but that says nothing
546 about what the backend does with the values. If a `shell_env` hook emits a
547 secret, scope it to a backend you trust with that secret — for example by
548 conditioning the hook, or by not configuring an external backend for sessions
549 where those hooks are active.
550
551 Resolved **key names — never values** — are written to `~/.codewhale/audit.log`
552 so a session can be reconciled afterwards. A hook that fails or times out
553 contributes no variables and does not abort the shell call.
554
555 ```toml
556 [[hooks.hooks]]
557 name = "aws-creds"
558 event = "shell_env"
559 command = "aws-vault export my-profile --format=env"
560 condition = { type = "tool_category", category = "shell" }
561 ```
562
563 ## Structured observer payloads
564
565 `turn_end`, `subagent_spawn`, `subagent_complete`, `session_busy`, `session_idle`,
566 `session_error`, and `waiting_for_user` receive JSON on stdin in addition to the
567 environment variables. Their stdout is ignored. Background forms of these
568 events receive the same payload on stdin.
569
570 `tool_call_after` also receives JSON on stdin for a settled native shell call
571 with a tracked [execution receipt](#execution-receipt), in foreground and
572 background form. Other tool calls have no stdin document.
573
574 The remaining observer events — `session_start`, `session_end`,
575 `mode_change`, `on_error` — receive environment variables only, with no stdin
576 payload, in both foreground and background form.
577
578 ### Session state transitions
579
580 The first observed state is recorded silently, whether idle, busy, or waiting.
581 Repeating the same state emits nothing. For a turn that pauses for user input
582 and then completes, the transition hooks receive these payloads in submission
583 order:
584
585 | Event | JSON stdin |
586 | --- | --- |
587 | `session_busy` | `{"from":"idle","to":"in_progress"}` |
588 | `waiting_for_user` | `{"from":"in_progress","to":"waiting","reason":"user_input"}` |
589 | `session_busy` | `{"from":"waiting","to":"in_progress"}` |
590 | `session_idle` | `{"from":"in_progress","to":"idle","last_turn_status":"completed"}` |
591
592 The dispatcher has two workers, so command completion order is not guaranteed.
593 `session_error` is a separate terminal-failure event, with `status` and `error`
594 fields rather than `from` and `to`.
595
596 ### `turn_end`
597
598 Fires after post-turn state, usage totals, cost accounting, notifications,
599 receipts, and queue recovery have been updated, and before queued follow-up
600 dispatch — so the payload can report the queued count without a hook being able
601 to change what is sent next.
602
603 ```json
604 {
605 "event": "turn_end",
606 "session_id": "sess_12345678",
607 "workspace": "/path/to/workspace",
608 "mode": "ACT",
609 "created_at": "2026-07-12T10:30:00+00:00",
610 "model_backed": true,
611 "provider": "deepseek",
612 "billing_surface": null,
613 "model": "deepseek-chat",
614 "turn_id": "turn_12345678",
615 "status": "completed",
616 "error": null,
617 "duration_ms": 1834,
618 "usage": {
619 "input_tokens": 1200,
620 "output_tokens": 180,
621 "prompt_cache_hit_tokens": 900,
622 "prompt_cache_miss_tokens": 300,
623 "prompt_cache_write_tokens": 0,
624 "reasoning_tokens": null,
625 "reasoning_replay_tokens": null
626 },
627 "totals": {
628 "session_tokens": 1380,
629 "conversation_tokens": 1380,
630 "input_tokens": 1200,
631 "output_tokens": 180
632 },
633 "tool_count": 2,
634 "queued_message_count": 1,
635 "stop_hook_active": false
636 }
637 ```
638
639 `created_at` anchors time-window pricing. `provider` and `model` identify the
640 effective route for model-backed turns. `billing_surface` is an optional,
641 non-secret classification of the endpoint that served the turn (recognized
642 StepFun routes emit `stepfun-payg` or `stepfun-plan`); the raw base URL is never
643 written to hook records. Shell-only, manual-compaction, and purge completions
644 have no matching `TurnStarted`, so they report `model_backed: false`, a `null`
645 provider, and a synthetic `lifecycle_<uuid>` turn id. `stop_hook_active` is
646 always `false` today; it reserves room for re-entry protection.
647
648 ### `subagent_spawn` / `subagent_complete`
649
650 ```json
651 {
652 "event": "subagent_complete",
653 "agent_id": "agent_1",
654 "session_id": "sess_12345678",
655 "workspace": "/path/to/workspace",
656 "mode": "ACT",
657 "model": "deepseek-chat",
658 "total_tokens": 1234,
659 "result_preview": "bounded preview of the result",
660 "result_truncated": false,
661 "status": "completed"
662 }
663 ```
664
665 `subagent_spawn` carries `prompt_preview` / `prompt_truncated` instead, and no
666 `status`. Both payloads are bounded on purpose: previews are truncated rather
667 than shipping full prompts or results. These hooks are observer-only — failures
668 do not affect sub-agent scheduling, prompts, or results, and `continue_on_error`
669 has no effect because later matching hooks always run.
670
671 ## Failure behavior
672
673 - A non-zero exit is logged at `warn` under the `hooks` tracing target with the
674 hook name, event, exit code, duration, and a generic failure category. Raw
675 stdout/stderr/error text is not persisted in the log receipt.
676 - For `execute`-path events, `continue_on_error = false` stops later hooks for
677 that event; except on `tool_call_before` (above) it does not roll back the
678 action that fired them.
679 - Structured observer events (`turn_end`, `subagent_*`, `session_busy`,
680 `session_idle`, `session_error`, `waiting_for_user`) always continue to the
681 next matching hook.
682 - Observer events use a bounded persistent dispatcher. Queue-full and
683 dispatcher-unavailable submissions are not retried silently; the TUI keeps
684 an event-specific error toast separate from the ordinary status line.
685 - A hook that exceeds its timeout has its whole process group killed and is
686 then reaped, foreground or background — best-effort, with a bounded reap
687 wait; see [Timeouts](#timeouts).
688
689 ## Security notes
690
691 - Hooks are arbitrary shell commands from your own config; treat
692 `~/.codewhale/config.toml` as executable.
693 - Project-supplied hooks require exact-file approval in addition to workspace trust in
694 user-owned config.
695 - Hook commands inherit Codewhale's own environment. A local `exec_shell` does
696 not — see [`shell_env`](#shell_env).
697 - `shell_env` audit records contain key names only. That covers Codewhale's own
698 logging; with an external sandbox backend configured, the values themselves
699 are transmitted to that backend and are then subject to its handling.
700 - With an external sandbox backend, the local parent-variable allowlist does
701 not apply — the backend owns its base environment.
702 - Payload previews, tool arguments/results, error messages, captured stdout and
703 stderr, replacement messages, and steering objects are bounded so hook input
704 or output cannot become an unbounded copy of the transcript.
705 - Nothing Codewhale persists in a denial echoes the stdin payload, hook
706 environment, raw stdout/stderr/error, command line, or a resolved filesystem
707 path. `/hooks list` shows a sanitized, single-line command preview capped at
708 60 characters; it is not a verbatim copy. Structured denial reasons are
709 bounded and redact path-, argument-, command-, and secret-like tokens,
710 including quoted or `key=value` forms and `Authorization: Bearer …`.
711
711 lines MARKDOWN