| 1 | # Hooks |
| 2 | > 阅读简体中文版:[zh_hans/HOOKS.md](zh_hans/HOOKS.md) |
| 3 | |
| 4 | Hooks run a shell command when the Codewhale **TUI** reaches a lifecycle |
| 5 | point. They are plain processes: they receive context through environment |
| 6 | variables, some receive a JSON payload on stdin, and three of them can steer |
| 7 | what Codewhale does next. |
| 8 | |
| 9 | This page is the authoritative reference for what is implemented today. |
| 10 | Configuration syntax that overlaps with the rest of `config.toml` lives in |
| 11 | [CONFIGURATION.md](CONFIGURATION.md); this file is the event-by-event |
| 12 | contract. |
| 13 | |
| 14 | ## Scope |
| 15 | |
| 16 | Hooks fire in the interactive TUI and in the engine turn loop, which the |
| 17 | Runtime API threads behind the desktop app and web drive as well. |
| 18 | |
| 19 | | Surface | Fires hooks | |
| 20 | | --- | --- | |
| 21 | | `codewhale` / `codew` interactive TUI | yes | |
| 22 | | `codewhale exec` (headless one-shot) | opt-in: `--hooks` fires `tool_call_before` and `shell_env` | |
| 23 | | Runtime API threads (desktop app, web) | yes: `tool_call_before`, `shell_env`, `tool_call_after`, `on_error`; `GET /v1/hooks` lists the set | |
| 24 | | the `codewhale` CLI dispatcher and its subcommands | no | |
| 25 | | app-server / ACP | no | |
| 26 | | the `workflow` tool and sub-agent *internals* | no — but the TUI fires `subagent_spawn` / `subagent_complete` around them | |
| 27 | | public API | there is none | |
| 28 | |
| 29 | The `crates/hooks` event-sink crate in this repository is an unrelated |
| 30 | internal mechanism. It shares no configuration, no event names, and no |
| 31 | contract with the hooks described here. |
| 32 | |
| 33 | ### `codewhale exec --hooks` |
| 34 | |
| 35 | Headless runs fire no hooks by default — a CI job should not start paging an |
| 36 | on-call rotation merely because a config exists. `codewhale exec --hooks` |
| 37 | opts the run in. The engine-side events are `tool_call_before` (exit code 2 |
| 38 | still denies the call; `ask` resolves fail-closed because nothing can prompt |
| 39 | headlessly) and `shell_env`. UI-driven events such as `session_start`, |
| 40 | `message_submit`, and `turn_end` do not fire — they live in the interactive |
| 41 | shell, not the turn loop. Fleet worker subprocesses never fire operator |
| 42 | hooks. Independently of this flag, `permissions.toml` typed rules already |
| 43 | apply to `exec` — the run drives the same turn loop, and a `deny` blocks in |
| 44 | every mode. |
| 45 | |
| 46 | ## Quick start |
| 47 | |
| 48 | ```toml |
| 49 | # ~/.codewhale/config.toml |
| 50 | [hooks] |
| 51 | enabled = true |
| 52 | |
| 53 | [[hooks.hooks]] |
| 54 | name = "announce" |
| 55 | event = "session_start" |
| 56 | command = "echo 'Codewhale session started'" |
| 57 | ``` |
| 58 | |
| 59 | Run `/hooks` in the TUI to list what is configured, whether the global switch |
| 60 | is on, and any entry that was rejected at load. Run `/hooks events` for the |
| 61 | event names. |
| 62 | |
| 63 | ## Configuration |
| 64 | |
| 65 | ```toml |
| 66 | [hooks] |
| 67 | enabled = true # global switch; false suppresses every hook |
| 68 | default_timeout_secs = 30 # see the timeout note below |
| 69 | working_dir = "/path/to/dir" # default: the session workspace |
| 70 | |
| 71 | [[hooks.hooks]] |
| 72 | event = "tool_call_before" # required; one of the 15 names below |
| 73 | command = "~/.codewhale/hooks/gate.sh" # required; `sh -c` on Unix, `cmd /C` on Windows |
| 74 | name = "gate" # optional label for /hooks and log lines |
| 75 | timeout_secs = 30 # optional, default 30 |
| 76 | background = false # optional; foreground inside the hook worker |
| 77 | continue_on_error = true # optional, default true |
| 78 | condition = { type = "tool_name", name = "bash" } # optional |
| 79 | ``` |
| 80 | |
| 81 | `timeout_secs` note, stated as implemented: when `[hooks].default_timeout_secs` |
| 82 | is set it **overrides** every hook's own `timeout_secs`, it does not merely |
| 83 | supply a default for hooks that omit one. Leave it unset if you want per-hook |
| 84 | timeouts to apply. `/hooks list` shows the timeout the runtime will actually |
| 85 | apply, and names the override when one is in force. |
| 86 | |
| 87 | `default_timeout_secs = 0` is **rejected at load**. Because the value replaces |
| 88 | every hook's own `timeout_secs`, a zero there would expire every hook in the |
| 89 | config immediately — including a `tool_call_before` gate, which then denies |
| 90 | every matching tool call. The override is ignored, per-hook `timeout_secs` |
| 91 | applies, the hooks themselves still load, and the rejection is reported by |
| 92 | `/hooks list` under *configuration problems*. Per-hook `timeout_secs = 0` is |
| 93 | rejected too, but that only drops the one hook that wrote it. |
| 94 | |
| 95 | Hooks run with the workspace (or `working_dir`) as the current directory. |
| 96 | |
| 97 | ### Timeouts |
| 98 | |
| 99 | The timeout applies to **foreground and background hooks alike**. When it |
| 100 | expires: |
| 101 | |
| 102 | - the hook's whole process group is killed — Unix process groups, Windows Job |
| 103 | Objects — so a hook that spawns children does not outlive its budget; |
| 104 | - the child is then reaped, so nothing is normally left detached or zombied; |
| 105 | - a foreground hook's result is `success = false`, `exit_code = None`, empty |
| 106 | `stdout`/`stderr`, and `error = "Hook timed out after Ns"`; |
| 107 | - a background hook's timeout is logged at `warn` under the `hooks` target. |
| 108 | Nothing is reported to the caller, because the caller stopped waiting the |
| 109 | moment it submitted the hook. |
| 110 | |
| 111 | **Termination is best-effort, and the bound that is guaranteed is Codewhale's, |
| 112 | not the OS's.** The kill can fail to land — a process wedged in an |
| 113 | uninterruptible state on Unix, a `TerminateJobObject` a protected process |
| 114 | survives on Windows — and no user-space program can promise otherwise. What |
| 115 | Codewhale does guarantee is that it stops waiting: the containment handle is |
| 116 | released (which re-signals the Unix process group and closes the kill-on-close |
| 117 | Windows Job Object) and the reap gets one short bounded window. If the child |
| 118 | still cannot be confirmed dead, that is logged at `warn` and the foreground |
| 119 | result says so — `error = "hook could not be reaped after its timeout"` rather |
| 120 | than the stronger timeout wording. So a timed-out hook never blocks the turn, |
| 121 | but treat "killed" as best-effort rather than absolute. |
| 122 | |
| 123 | ### Background hooks |
| 124 | |
| 125 | `background = true` describes real scheduling, not just a config flag. A |
| 126 | background hook is **submitted, never awaited**: |
| 127 | |
| 128 | - it is enqueued without blocking into a fixed 32-entry supervisor queue, |
| 129 | drained by two persistent workers that apply the timeout above; saturation |
| 130 | or supervisor loss is a failed submission, and no invocation creates its |
| 131 | own detached supervisor thread; |
| 132 | - it receives the same environment variables and the same stdin JSON payload |
| 133 | as the foreground form of that event — the payload contract does not change, |
| 134 | only the steering does; |
| 135 | - its stdout and stderr are discarded (`Stdio::null()`), so it can never |
| 136 | return a verdict; |
| 137 | - the `HookResult` the runtime hands its caller is flagged as a background |
| 138 | submission and carries no exit code. Steering code reads |
| 139 | `observed_exit_code()`, which is `None` for a background hook, so a |
| 140 | background hook can never allow, deny, ask, or rewrite anything. |
| 141 | |
| 142 | `shell_env` ignores `background` entirely — its stdout *is* the contract, so |
| 143 | it always runs in the foreground. `/hooks list` reports that as a |
| 144 | configuration warning, and does not label the hook `[bg]`. |
| 145 | |
| 146 | Observer-only UI events are submitted with non-blocking `try_send` to one |
| 147 | 32-entry queue drained by two persistent workers. A configured foreground |
| 148 | observer is still awaited in config order inside a worker, but the terminal |
| 149 | event loop never waits on its process and never creates a thread per event. |
| 150 | Queue saturation or dispatcher loss drops that observer event and produces an |
| 151 | event-specific error toast that survives an agent's ordinary progress-status |
| 152 | update. Steering events retain their gate or transform semantics: |
| 153 | fresh/queued `message_submit` dispatch reports through a bounded result |
| 154 | channel, same-turn steering runs that transform on the blocking worker before |
| 155 | calling the engine steer path, and `tool_call_before` / `shell_env` execute on |
| 156 | the engine or tool worker rather than the terminal event loop. |
| 157 | |
| 158 | ### The hook process environment |
| 159 | |
| 160 | A hook command inherits the environment of the Codewhale process, plus the |
| 161 | `DEEPSEEK_*` variables for its event. Codewhale does not filter that |
| 162 | inheritance, so treat a hook exactly as you would treat any command you type |
| 163 | in the same shell that launched Codewhale: whatever is exported there is |
| 164 | visible to it. |
| 165 | |
| 166 | This is *not* true of the command a `shell_env` hook feeds — see |
| 167 | [`shell_env`](#shell_env) for the bounded allowlist that governs a **local** |
| 168 | `exec_shell`, and for what changes when an external sandbox backend is |
| 169 | configured instead (the backend owns its base environment, and your |
| 170 | `shell_env` values are transmitted to it). |
| 171 | |
| 172 | ### Conditions |
| 173 | |
| 174 | | Condition | Matches | Supported on | |
| 175 | | --- | --- | --- | |
| 176 | | `{ type = "always" }` | every invocation (also the default when omitted) | every event | |
| 177 | | `{ type = "tool_name", name = "bash" }` | exact tool name; `*` globs are supported, e.g. `mcp__*`. The shell tool's spellings `bash`, `Bash`, and `exec_shell` are aliases: a condition naming any one matches all three | `tool_call_before`, `tool_call_after`, `shell_env`, `on_error` | |
| 178 | | `{ type = "tool_category", category = "shell" }` | tool category | `tool_call_before`, `tool_call_after`, `shell_env`, `on_error` | |
| 179 | | `{ type = "mode", mode = "plan" }` | the context's mode string, case-insensitive | every event **except** `shell_env` | |
| 180 | | `{ type = "exit_code", code = 1 }` | the exit code the tool actually reported | `tool_call_after`, `on_error` | |
| 181 | | `{ type = "all", conditions = [...] }` | every nested condition | every event | |
| 182 | | `{ type = "any", conditions = [...] }` | at least one nested condition | every event | |
| 183 | |
| 184 | Three rules keep conditions from lying: |
| 185 | |
| 186 | - **`exit_code` needs a real exit code.** It matches only when the event |
| 187 | actually observed a process exit code — `tool_call_after`, or `on_error` for |
| 188 | a tool failure, in both cases for a process-backed tool such as `bash`. |
| 189 | A command that exits nonzero reports its code too, even though `bash` |
| 190 | returns it as a failed call. A timed-out or killed command usually has no |
| 191 | exit code; `DEEPSEEK_TOOL_STATUS` says which it was. |
| 192 | A tool that reports no exit code never matches an `exit_code` condition; the |
| 193 | condition is not satisfied by a default, a zero, or a success flag. The value |
| 194 | is a 64-bit integer, so a Windows crash code such as `3221225477` |
| 195 | (`0xC0000005`) is matchable. |
| 196 | - **Tool-scoped `on_error` hooks are supported.** `on_error` fires for |
| 197 | transport and capacity errors *and* for tool failures; the tool-failure |
| 198 | firing carries the tool name, call id, result, and reported exit code. A |
| 199 | `tool_name` / `tool_category` / `exit_code` condition on `on_error` is |
| 200 | therefore a valid configuration. An `on_error` firing with no tool behind it |
| 201 | simply does not match such a condition — it is skipped at dispatch, not |
| 202 | rejected at load. |
| 203 | - **Unsupported conditions are rejected at load.** A condition that references |
| 204 | context its event never carries can never match, and a hook wearing one is |
| 205 | silently inert — the dangerous form of that is a `deny` gate the operator |
| 206 | believes is armed. Codewhale drops those hooks at load, logs the reason |
| 207 | under the `hooks` tracing target, and shows them in `/hooks list` as |
| 208 | `rejected:`. Nested predicates inside `all` / `any` are checked too. A hook |
| 209 | with `timeout_secs = 0` or an empty `command` is rejected the same way. |
| 210 | Rejection is **per entry**: a broken hook never takes another one with it, |
| 211 | even when the two share a `name` or are both unnamed. |
| 212 | |
| 213 | ### Project-local hooks |
| 214 | |
| 215 | A repository may ship `<workspace>/.codewhale/hooks.toml` using the same shape, |
| 216 | but only its `[[hooks]]` entries are merged — a project file cannot change |
| 217 | `enabled`, `default_timeout_secs`, or `working_dir`, which always come from your |
| 218 | own config. Because hooks are executable configuration, project hooks load |
| 219 | **only** after both workspace trust and separate approval of the exact hooks file |
| 220 | in user-owned config. Use `/hooks review` to inspect the commands and digest, |
| 221 | then `/hooks approve <digest>` to enable those bytes on the next session. Review |
| 222 | any scripts the commands call too. A file change requires another approval. |
| 223 | `/hooks revoke` blocks future and queued launches; it does not stop commands |
| 224 | already running. Session `/trust on` alone does not enable project hooks. |
| 225 | Approved project hooks are appended |
| 226 | after global hooks, so they run last and win `updatedInput` ties. A malformed |
| 227 | trusted project file logs a warning and Codewhale falls back to global hooks |
| 228 | only. Validation runs over the merged set, so a rejected project hook is |
| 229 | reported the same way a rejected global one is. |
| 230 | |
| 231 | ## The 15 events |
| 232 | |
| 233 | | Event | Fires | Steering | |
| 234 | | --- | --- | --- | |
| 235 | | `session_start` | once, after the engine is up and before the first draw | observer | |
| 236 | | `session_end` | once, on graceful shutdown | observer | |
| 237 | | `turn_end` | after a turn completes and post-turn state is updated | observer | |
| 238 | | `message_submit` | before a submitted message reaches history or the model | **can replace or block the text** | |
| 239 | | `tool_call_before` | before each tool call executes | **can allow / deny / ask, rewrite input, add context** | |
| 240 | | `tool_call_after` | after each tool result settles, including completions the transcript does not redraw | observer | |
| 241 | | `mode_change` | on every applied Plan/Work/Operate transition (`Act` is a compatibility alias for Work) | observer | |
| 242 | | `on_error` | on transport, capacity, and auth errors, and on tool failures | observer | |
| 243 | | `subagent_spawn` | when a sub-agent starts | observer | |
| 244 | | `subagent_complete` | when a sub-agent completes, fails, or is cancelled | observer | |
| 245 | | `shell_env` | immediately before each `exec_shell` invocation | **contributes environment variables** | |
| 246 | | `session_idle` | when the session settles back to idle after a turn or a wait — no prompt, approval, or continuation outstanding | observer | |
| 247 | | `session_error` | when a turn ends in a terminal failure; transient tool failures the agent absorbs never fire it | observer | |
| 248 | | `waiting_for_user` | when the agent starts waiting on you: an approval prompt opens, a `request_user_input` question is presented, or a goal continuation is parked between passes | observer | |
| 249 | | `session_busy` | when an idle or waiting session begins or resumes work; startup and repeated observations of the same state stay silent | observer | |
| 250 | |
| 251 | `waiting_for_user`'s payload carries `reason`: `approval`, `user_input`, or |
| 252 | `goal_continuation`. All three state events carry `from`/`to` transition fields; |
| 253 | `session_idle` also carries `last_turn_status` when known, and `session_error` carries |
| 254 | the bounded terminal `error` text. Busy, idle, and waiting map onto the session |
| 255 | states the control socket's `status` verb already publishes |
| 256 | (`idle` / `in_progress` / `waiting`), so a hook and a supervisor never |
| 257 | disagree about what the session is doing. Hook authors that want opencode's |
| 258 | grace-period semantics for error alerts should debounce inside the hook — |
| 259 | `session_error` already excludes absorbed, transient failures, and a turn |
| 260 | that fails and is retried by the operator fires again only if the retry also |
| 261 | ends failed. |
| 262 | |
| 263 | ### What "observer" means, exactly |
| 264 | |
| 265 | Observer means Codewhale ignores the hook's **result**: stdout is discarded, a |
| 266 | non-zero exit is logged as a warning, and nothing about the turn, the tool |
| 267 | result, the sub-agent, or the error changes because of it. |
| 268 | |
| 269 | Observer does **not** mean side-effect-free. An observer hook is an arbitrary |
| 270 | shell command running with your credentials. It can write files, push commits, |
| 271 | page an on-call rotation, or delete the workspace. The only thing it cannot do |
| 272 | is change what Codewhale itself does next. |
| 273 | |
| 274 | The steering allowlist is exactly three events — `message_submit`, |
| 275 | `tool_call_before`, `shell_env` — and it is asserted by a test over every |
| 276 | variant, so a new event defaults to observer. |
| 277 | |
| 278 | ### Session identity |
| 279 | |
| 280 | Every event in one TUI session carries the same `DEEPSEEK_SESSION_ID`. The id |
| 281 | is minted once at launch, in the form `sess_xxxxxxxx`, and it survives a |
| 282 | workspace switch and a trust decision that adds project hooks — both reload |
| 283 | the hook set without starting a new session. Engine-fired `tool_call_before` |
| 284 | reports the same id as the UI-fired events, so tool records correlate with the |
| 285 | session records around them. |
| 286 | |
| 287 | `session_end` fires after the queued startup-default writes have been drained |
| 288 | and while the app is still live, so it observes the settled end state rather |
| 289 | than a half-torn-down one. |
| 290 | |
| 291 | ## Environment variables |
| 292 | |
| 293 | Every hook receives the subset of these that applies to its event. The |
| 294 | `DEEPSEEK_` prefix is retained for compatibility with hooks written before the |
| 295 | rebrand. |
| 296 | |
| 297 | | Variable | Set for | Notes | |
| 298 | | --- | --- | --- | |
| 299 | | `DEEPSEEK_SESSION_ID` | every event except `shell_env` | `sess_xxxxxxxx`, stable for the whole session | |
| 300 | | `DEEPSEEK_WORKSPACE` | every event except `shell_env` | absolute workspace path | |
| 301 | | `DEEPSEEK_MODEL` | every event except `shell_env` | active model id | |
| 302 | | `DEEPSEEK_MODE` | every event except `shell_env` | see the mode-spelling note below | |
| 303 | | `DEEPSEEK_TOTAL_TOKENS` | UI-fired events | session token total at fire time | |
| 304 | | `DEEPSEEK_MESSAGE` | `message_submit`, `subagent_*` | truncated at 5 000 bytes with a `...[truncated]` marker | |
| 305 | | `DEEPSEEK_ERROR` | `on_error` | error message, truncated at 5 000 bytes | |
| 306 | | `DEEPSEEK_PREVIOUS_MODE` | `mode_change` | mode label before the change | |
| 307 | | `DEEPSEEK_TOOL_NAME` | `tool_call_before`, `tool_call_after`, `shell_env`, `on_error` (tool failures) | | |
| 308 | | `DEEPSEEK_TOOL_CALL_ID` | `tool_call_before`, `tool_call_after`, `on_error` (tool failures) | engine call id; correlates before/after/error for one call | |
| 309 | | `DEEPSEEK_TOOL_ARGS` | `tool_call_before`, `shell_env` | tool input JSON preview, capped at 10 000 bytes | |
| 310 | | `DEEPSEEK_TOOL_RESULT` | `tool_call_after`, `on_error` (tool failures) | truncated at 10 000 bytes | |
| 311 | | `DEEPSEEK_TOOL_SUCCESS` | `tool_call_after`, `on_error` (tool failures) | `true` / `false` | |
| 312 | | `DEEPSEEK_TOOL_EXIT_CODE` | `tool_call_after` and `on_error` **when the tool reported one** | absent otherwise — never synthesized; set for a failing command as well as a passing one; 64-bit, so Windows crash codes such as `3221225477` survive | |
| 313 | | `DEEPSEEK_TOOL_STATUS` | `tool_call_after` and `on_error` **when a shell tool reported one** | `completed`, `failed`, `timed_out`, `killed`, or `running` (moved to the background); absent for other tools | |
| 314 | | `DEEPSEEK_TOOL_EXECUTION_RECEIPT` | `tool_call_after` and `on_error` **for a settled, local, foreground shell run** | complete JSON, at most 32 KiB, or absent; see [Execution receipt](#execution-receipt) | |
| 315 | | `DEEPSEEK_SESSION_COST` | when cost is supplied | USD, six decimal places | |
| 316 | |
| 317 | ### Execution receipt |
| 318 | |
| 319 | `DEEPSEEK_TOOL_EXECUTION_RECEIPT` says what a shell tool (`bash`, `Bash`, |
| 320 | `exec_shell`) actually ran. The before-hook input is not the same thing: a |
| 321 | `tool_call_before` hook can rewrite it. The receipt is built from what the |
| 322 | process manager recorded when it spawned the process, after admission and |
| 323 | any rewrite. |
| 324 | |
| 325 | ```json |
| 326 | {"schema_version":1,"command":"printf hello","cwd":"/absolute/workspace","state":"completed","scope":"local","exit_code":0,"stdout":"hello","stderr":"","stdout_truncated":false,"stderr_truncated":false,"output_kind":"separate"} |
| 327 | ``` |
| 328 | |
| 329 | | Field | Meaning | |
| 330 | | --- | --- | |
| 331 | | `command` | the admitted shell source handed to the shell, not the shell executable or its argv wrapper | |
| 332 | | `cwd` | the canonical absolute path of the directory the process started in: symlinks are resolved, so a directory has one spelling whether or not the call passed `cwd`; it is resolved before spawn and that same path is handed to the OS | |
| 333 | | `state` | `completed` for an observed exit, including a nonzero one; `interrupted` for a signal, kill, cancel, or timeout | |
| 334 | | `scope` | always `local` in schema 1 | |
| 335 | | `exit_code` | the observed integer, or `null`; never synthesized from `state` | |
| 336 | | `stdout`, `stderr` | previews of the tool's retained output, which may already omit early process output; long previews keep their first and last bytes around a `[receipt preview truncated]` marker | |
| 337 | | `stdout_truncated`, `stderr_truncated` | `true` when the tool's own output capture or the preview dropped bytes | |
| 338 | | `output_kind` | `separate` for `Bash` / `exec_shell`; `combined` for lowercase `bash`, whose stdout and stderr share one pipe — `stdout` then holds the combined preview and `stderr` is empty | |
| 339 | |
| 340 | The rules are conservative: |
| 341 | |
| 342 | - **Exact or absent.** `command` and `cwd` are never truncated. If either is |
| 343 | over 8 KiB, contains NUL, or the directory is relative, not UTF-8, or cannot |
| 344 | resolve before spawn, the receipt is left out. So is a run whose end the shell |
| 345 | tool could not observe (the OS wait itself failed): its state is unknown, |
| 346 | and the receipt does not guess it. Previews shrink until the serialized JSON fits 32 KiB; |
| 347 | if it still cannot fit, the receipt is left out rather than cut. |
| 348 | - **Absence means nothing.** It implies neither success nor failure. An inherited |
| 349 | `DEEPSEEK_TOOL_EXECUTION_RECEIPT` is cleared before applying the current call's context. |
| 350 | - **Scope.** A receipt is built only while a `tool_call_after` or `on_error` |
| 351 | hook is configured, and only for a settled, pipe-backed, unsandboxed, local |
| 352 | foreground run. Background launches, a foreground run moved to `/jobs`, |
| 353 | PTY (`tty` / `combined_output`) and interactive sessions, OS-sandboxed and |
| 354 | external-backend execution, the read-only shell's hardened argv, Windows, |
| 355 | a PowerShell shell on any platform (it wraps the source or runs it from a |
| 356 | temporary script), and calls refused before execution have none. |
| 357 | - **Hooks only.** The receipt is not kept in the durable Runtime API item |
| 358 | record; that record already carries the tool output. |
| 359 | - It is set for a failed run as well as a passing one, so `on_error` for a |
| 360 | failed shell call carries it too. Every other variable is unchanged. |
| 361 | |
| 362 | For `tool_call_after`, the same execution evidence is also delivered as a |
| 363 | versioned JSON document on stdin, in foreground and background form: |
| 364 | |
| 365 | ```json |
| 366 | {"schema_version":1,"event":"tool_call_after","tool_name":"bash","session_id":"session-id","tool_call_id":"call-id","session_id_truncated":false,"tool_call_id_truncated":false,"tool_name_truncated":false,"execution_receipt":{"schema_version":1,"command":"printf hello","cwd":"/absolute/workspace","command_truncated":false,"cwd_truncated":false,"execution":"started","completion":"completed","exit_code":0,"stdout":"hello","stderr":"","stdout_truncated":false,"stderr_truncated":false,"output_mode":"combined"}} |
| 367 | ``` |
| 368 | |
| 369 | The stdin projection leaves the environment receipt above unchanged. |
| 370 | `completion` is the observed `completed`, `failed`, `killed`, or `timed_out` |
| 371 | status; a nonzero exit is `failed`. `exit_code` remains a signed 64-bit integer |
| 372 | or `null`. `output_mode` carries the existing receipt's `output_kind`: |
| 373 | `separate` or `combined`. With combined output, empty `stderr` does not mean |
| 374 | the command wrote nothing to stderr. |
| 375 | |
| 376 | The complete document is capped at 64 KiB. Correlation identifiers are capped |
| 377 | at 1,024 UTF-8 bytes plus a truncation marker and carry individual flags; |
| 378 | missing identifiers are `null`. The shell name is exact. Execution `command` |
| 379 | and `cwd` retain the exact-or-absent rule above, so their truncation flags are |
| 380 | always false. Output truncation flags include both capture and preview loss. |
| 381 | |
| 382 | Only a native shell call with a valid, settled receipt gets this stdin |
| 383 | document. Unsupported or unobserved paths get no document, and absence still |
| 384 | means unknown. Hooks remain observers: their stdout cannot allow, deny, or |
| 385 | rewrite the completed call, and background hooks are not awaited. `on_error` |
| 386 | continues to receive the environment receipt only. |
| 387 | |
| 388 | **Mode-spelling note.** UI-fired events (`session_start`, `session_end`, |
| 389 | `message_submit`, `tool_call_after`, `mode_change`, `on_error`, `turn_end`, |
| 390 | `subagent_*`, `session_busy`, `session_idle`, `session_error`, `waiting_for_user`) |
| 391 | set `DEEPSEEK_MODE` to the UI label — `ACT`, `PLAN`, `OPERATE`. |
| 392 | `tool_call_before` fires inside the engine and uses the engine's own mode |
| 393 | spelling (`Agent`, `Plan`, `Operate`). `mode` conditions compare |
| 394 | case-insensitively, so `{ type = "mode", mode = "plan" }` matches both, but a |
| 395 | hook that string-matches `$DEEPSEEK_MODE` exactly should accept both spellings. |
| 396 | |
| 397 | **`shell_env` is the narrow one.** It receives only `DEEPSEEK_TOOL_NAME` and |
| 398 | `DEEPSEEK_TOOL_ARGS` — no session id, workspace, model, or mode. A |
| 399 | `{ type = "mode", … }` condition on a `shell_env` hook is therefore rejected at |
| 400 | load; scope those with `tool_name` or `tool_category` instead. |
| 401 | |
| 402 | ## Steering events |
| 403 | |
| 404 | ### `message_submit` |
| 405 | |
| 406 | Receives JSON on stdin and may rewrite or block the submitted text. |
| 407 | |
| 408 | ```json |
| 409 | { |
| 410 | "event": "message_submit", |
| 411 | "text": "original user text", |
| 412 | "text_bytes": 18, |
| 413 | "text_original_bytes": 18, |
| 414 | "text_truncated": false, |
| 415 | "session_id": "sess_12345678", |
| 416 | "workspace": "/path/to/workspace", |
| 417 | "mode": "ACT", |
| 418 | "model": "deepseek-chat", |
| 419 | "total_tokens": 1234 |
| 420 | } |
| 421 | ``` |
| 422 | |
| 423 | The complete serialized stdin document is capped at 32 KiB. `text` is the |
| 424 | largest deterministic UTF-8 prefix that fits after JSON escaping and bounded |
| 425 | metadata are included. `text_original_bytes` records the producer's full byte |
| 426 | length, `text_bytes` records the retained prefix, and `text_truncated` states |
| 427 | whether they differ. This same boundary applies to immediate input, restored |
| 428 | queue entries, merged steers, and text produced by an earlier hook. |
| 429 | |
| 430 | - exit `0` printing `{"text": "..."}` with a non-empty string replaces the text |
| 431 | - exit `0` with empty stdout, or JSON without `text`, leaves the text unchanged |
| 432 | - `{"text": ""}` or a replacement over 32 000 characters is invalid stdout, |
| 433 | logged and ignored |
| 434 | - exit `2` blocks the submission before history or dispatch; a structured |
| 435 | `reason` field supplies a bounded, redacted message shown in the TUI. |
| 436 | Unstructured stdout/stderr/error output is never copied into the denial |
| 437 | - other non-zero exits follow `continue_on_error`: `true` warns and continues, |
| 438 | `false` blocks the submission |
| 439 | - `background = true` makes the hook observer-only — it still receives this |
| 440 | bounded payload on stdin, but it cannot transform or block |
| 441 | |
| 442 | Multiple `message_submit` hooks run in config order and each sees the previous |
| 443 | hook's output. |
| 444 | |
| 445 | ### `tool_call_before` |
| 446 | |
| 447 | Receives the tool context in environment variables and may print a JSON |
| 448 | decision on stdout with exit `0`: |
| 449 | |
| 450 | ```json |
| 451 | { |
| 452 | "decision": "allow", |
| 453 | "reason": "human-readable explanation, used for deny", |
| 454 | "updatedInput": { "command": "ls -la" }, |
| 455 | "additionalContext": "text appended to the tool result for the model" |
| 456 | } |
| 457 | ``` |
| 458 | |
| 459 | - `deny` blocks the tool; the model gets a permission-denied result carrying |
| 460 | `reason` |
| 461 | - `ask` forces the interactive approval prompt in Ask and Auto-Review. Full |
| 462 | Access does not open tool-approval prompts, so `ask` does not downgrade it |
| 463 | - `updatedInput` must be an object no larger than 32 KiB serialized and |
| 464 | replaces the tool input; last hook wins |
| 465 | - `additionalContext` is appended to the tool result as `[hook context] ...`; |
| 466 | multiple hooks concatenate |
| 467 | - `reason` and `additionalContext` are bounded and sanitized before use: each |
| 468 | field is capped at 2 000 characters, the concatenated context for one tool |
| 469 | call is capped at 8 000, control characters are stripped (so hook stdout |
| 470 | cannot repaint the TUI or forge structure in the transcript), and a clipped |
| 471 | value carries a `…[truncated]` marker. What a hook adds to the turn's context |
| 472 | budget is therefore bounded no matter what it prints |
| 473 | - exit `2` is a legacy hard deny and wins regardless of stdout |
| 474 | - empty stdout, non-JSON stdout, and JSON without `decision` all mean allow |
| 475 | - precedence across matching hooks: no-verdict-with-`continue_on_error = false` |
| 476 | > deny > ask > allow |
| 477 | - `background = true` hooks are submitted and never awaited, so they have no |
| 478 | verdict and cannot steer; Codewhale logs a warning when one is configured |
| 479 | for this event |
| 480 | |
| 481 | **A gate that could not answer is not permission.** If a foreground |
| 482 | `tool_call_before` hook produces no verdict — it timed out, the process could |
| 483 | not be started, or a strict process exited non-zero without an explicit JSON |
| 484 | decision — and *that hook* is configured with |
| 485 | `continue_on_error = false`, the tool call is denied. Strictness is read off |
| 486 | the hook that actually ran, not off the event: a strict `write_file` gate whose |
| 487 | condition did not match an `exec_shell` call has no say in whether that call |
| 488 | proceeds, and a lenient hook's timeout never denies just because some other |
| 489 | strict hook exists in config. Every no-verdict outcome is logged either way. |
| 490 | |
| 491 | The denial message names the hook and the reason and nothing else: the hook |
| 492 | name is truncated, the detail is truncated, control characters are stripped, |
| 493 | and a spawn failure is reported by error kind (`NotFound`, |
| 494 | `PermissionDenied`, …) rather than by echoing the command line or the resolved |
| 495 | interpreter path. |
| 496 | |
| 497 | ### `shell_env` |
| 498 | |
| 499 | Runs synchronously before each `exec_shell` and its stdout is parsed as |
| 500 | `KEY=VALUE` lines. A leading `export ` is stripped, `#` comment lines and blank |
| 501 | lines are skipped, and a matching pair of surrounding single or double quotes is |
| 502 | removed from the value. Later hooks override earlier ones. Use it for ephemeral |
| 503 | credentials, per-skill `PATH` adjustments, or short-lived tokens. |
| 504 | |
| 505 | `background` is ignored for this event: the hook always runs in the foreground |
| 506 | because its stdout is the contract. |
| 507 | |
| 508 | An entry a shell cannot carry is dropped rather than allowed to break the tool |
| 509 | call: an empty name, a name containing whitespace, `=`, a control character, or |
| 510 | a NUL; a value containing a NUL; a value over 32 KiB; and anything past 256 KiB |
| 511 | of accumulated output from one hook. Each drop is logged by key name only. A |
| 512 | `shell_env` hook is an ordinary process whose stdout can contain anything — |
| 513 | "the hook printed something odd" must never become "the `exec_shell` call |
| 514 | aborted". |
| 515 | |
| 516 | **Exactly what the shell command ends up with — local execution.** When |
| 517 | `exec_shell` runs the command locally (the default), it does not inherit |
| 518 | Codewhale's ambient environment. Its environment is built as: |
| 519 | |
| 520 | 1. a sanitized fixed allowlist of parent variables — `PATH`, `HOME`, `USER`, |
| 521 | `LANG` and the other `LC_*`/locale entries, `TERM`, `SHELL`, `TMPDIR`, |
| 522 | proxy variables, color/terminal entries such as `NO_COLOR`, `CARGO_HOME`/`RUSTUP_HOME`/`RUSTUP_TOOLCHAIN`, the Windows system and MSVC toolchain entries, and other platform keys (the full list lives in `crates/tui/src/child_env.rs`) — and nothing else. Variables |
| 523 | outside that allowlist, including anything that looks like a secret, are |
| 524 | dropped; |
| 525 | 2. then the `KEY=VALUE` pairs your `shell_env` hooks produced, applied on top. |
| 526 | These are explicit values you configured, so they win over the allowlist. |
| 527 | |
| 528 | So a `shell_env` hook is the supported way to get a credential into one local |
| 529 | `exec_shell` invocation. Ambient secrets exported in the terminal that launched |
| 530 | Codewhale are **not** forwarded to a local `exec_shell` on their own. |
| 531 | |
| 532 | **With an external sandbox backend configured, the allowlist above is not the |
| 533 | contract.** If `exec_shell` is routed to a configured sandbox/execution |
| 534 | backend, Codewhale does not construct the process environment at all: it hands |
| 535 | the command and your `shell_env` values to the backend as extra environment |
| 536 | variables, and the **backend owns its own base environment**. What is present |
| 537 | besides your values — an image's baked-in variables, the backend's own |
| 538 | injections, whatever a remote runner exports — is determined by that backend, |
| 539 | not by the list above. Do not assume the local allowlist applies there. |
| 540 | |
| 541 | Disclosure, because it is the part that matters for a hook that emits |
| 542 | credentials: **`shell_env` values are transmitted to the configured backend.** |
| 543 | For a remote or containerized backend that means the values leave this machine |
| 544 | and are subject to that backend's logging, retention, and access controls. |
| 545 | Codewhale's own audit log still records key names only, but that says nothing |
| 546 | about what the backend does with the values. If a `shell_env` hook emits a |
| 547 | secret, scope it to a backend you trust with that secret — for example by |
| 548 | conditioning the hook, or by not configuring an external backend for sessions |
| 549 | where those hooks are active. |
| 550 | |
| 551 | Resolved **key names — never values** — are written to `~/.codewhale/audit.log` |
| 552 | so a session can be reconciled afterwards. A hook that fails or times out |
| 553 | contributes no variables and does not abort the shell call. |
| 554 | |
| 555 | ```toml |
| 556 | [[hooks.hooks]] |
| 557 | name = "aws-creds" |
| 558 | event = "shell_env" |
| 559 | command = "aws-vault export my-profile --format=env" |
| 560 | condition = { type = "tool_category", category = "shell" } |
| 561 | ``` |
| 562 | |
| 563 | ## Structured observer payloads |
| 564 | |
| 565 | `turn_end`, `subagent_spawn`, `subagent_complete`, `session_busy`, `session_idle`, |
| 566 | `session_error`, and `waiting_for_user` receive JSON on stdin in addition to the |
| 567 | environment variables. Their stdout is ignored. Background forms of these |
| 568 | events receive the same payload on stdin. |
| 569 | |
| 570 | `tool_call_after` also receives JSON on stdin for a settled native shell call |
| 571 | with a tracked [execution receipt](#execution-receipt), in foreground and |
| 572 | background form. Other tool calls have no stdin document. |
| 573 | |
| 574 | The remaining observer events — `session_start`, `session_end`, |
| 575 | `mode_change`, `on_error` — receive environment variables only, with no stdin |
| 576 | payload, in both foreground and background form. |
| 577 | |
| 578 | ### Session state transitions |
| 579 | |
| 580 | The first observed state is recorded silently, whether idle, busy, or waiting. |
| 581 | Repeating the same state emits nothing. For a turn that pauses for user input |
| 582 | and then completes, the transition hooks receive these payloads in submission |
| 583 | order: |
| 584 | |
| 585 | | Event | JSON stdin | |
| 586 | | --- | --- | |
| 587 | | `session_busy` | `{"from":"idle","to":"in_progress"}` | |
| 588 | | `waiting_for_user` | `{"from":"in_progress","to":"waiting","reason":"user_input"}` | |
| 589 | | `session_busy` | `{"from":"waiting","to":"in_progress"}` | |
| 590 | | `session_idle` | `{"from":"in_progress","to":"idle","last_turn_status":"completed"}` | |
| 591 | |
| 592 | The dispatcher has two workers, so command completion order is not guaranteed. |
| 593 | `session_error` is a separate terminal-failure event, with `status` and `error` |
| 594 | fields rather than `from` and `to`. |
| 595 | |
| 596 | ### `turn_end` |
| 597 | |
| 598 | Fires after post-turn state, usage totals, cost accounting, notifications, |
| 599 | receipts, and queue recovery have been updated, and before queued follow-up |
| 600 | dispatch — so the payload can report the queued count without a hook being able |
| 601 | to change what is sent next. |
| 602 | |
| 603 | ```json |
| 604 | { |
| 605 | "event": "turn_end", |
| 606 | "session_id": "sess_12345678", |
| 607 | "workspace": "/path/to/workspace", |
| 608 | "mode": "ACT", |
| 609 | "created_at": "2026-07-12T10:30:00+00:00", |
| 610 | "model_backed": true, |
| 611 | "provider": "deepseek", |
| 612 | "billing_surface": null, |
| 613 | "model": "deepseek-chat", |
| 614 | "turn_id": "turn_12345678", |
| 615 | "status": "completed", |
| 616 | "error": null, |
| 617 | "duration_ms": 1834, |
| 618 | "usage": { |
| 619 | "input_tokens": 1200, |
| 620 | "output_tokens": 180, |
| 621 | "prompt_cache_hit_tokens": 900, |
| 622 | "prompt_cache_miss_tokens": 300, |
| 623 | "prompt_cache_write_tokens": 0, |
| 624 | "reasoning_tokens": null, |
| 625 | "reasoning_replay_tokens": null |
| 626 | }, |
| 627 | "totals": { |
| 628 | "session_tokens": 1380, |
| 629 | "conversation_tokens": 1380, |
| 630 | "input_tokens": 1200, |
| 631 | "output_tokens": 180 |
| 632 | }, |
| 633 | "tool_count": 2, |
| 634 | "queued_message_count": 1, |
| 635 | "stop_hook_active": false |
| 636 | } |
| 637 | ``` |
| 638 | |
| 639 | `created_at` anchors time-window pricing. `provider` and `model` identify the |
| 640 | effective route for model-backed turns. `billing_surface` is an optional, |
| 641 | non-secret classification of the endpoint that served the turn (recognized |
| 642 | StepFun routes emit `stepfun-payg` or `stepfun-plan`); the raw base URL is never |
| 643 | written to hook records. Shell-only, manual-compaction, and purge completions |
| 644 | have no matching `TurnStarted`, so they report `model_backed: false`, a `null` |
| 645 | provider, and a synthetic `lifecycle_<uuid>` turn id. `stop_hook_active` is |
| 646 | always `false` today; it reserves room for re-entry protection. |
| 647 | |
| 648 | ### `subagent_spawn` / `subagent_complete` |
| 649 | |
| 650 | ```json |
| 651 | { |
| 652 | "event": "subagent_complete", |
| 653 | "agent_id": "agent_1", |
| 654 | "session_id": "sess_12345678", |
| 655 | "workspace": "/path/to/workspace", |
| 656 | "mode": "ACT", |
| 657 | "model": "deepseek-chat", |
| 658 | "total_tokens": 1234, |
| 659 | "result_preview": "bounded preview of the result", |
| 660 | "result_truncated": false, |
| 661 | "status": "completed" |
| 662 | } |
| 663 | ``` |
| 664 | |
| 665 | `subagent_spawn` carries `prompt_preview` / `prompt_truncated` instead, and no |
| 666 | `status`. Both payloads are bounded on purpose: previews are truncated rather |
| 667 | than shipping full prompts or results. These hooks are observer-only — failures |
| 668 | do not affect sub-agent scheduling, prompts, or results, and `continue_on_error` |
| 669 | has no effect because later matching hooks always run. |
| 670 | |
| 671 | ## Failure behavior |
| 672 | |
| 673 | - A non-zero exit is logged at `warn` under the `hooks` tracing target with the |
| 674 | hook name, event, exit code, duration, and a generic failure category. Raw |
| 675 | stdout/stderr/error text is not persisted in the log receipt. |
| 676 | - For `execute`-path events, `continue_on_error = false` stops later hooks for |
| 677 | that event; except on `tool_call_before` (above) it does not roll back the |
| 678 | action that fired them. |
| 679 | - Structured observer events (`turn_end`, `subagent_*`, `session_busy`, |
| 680 | `session_idle`, `session_error`, `waiting_for_user`) always continue to the |
| 681 | next matching hook. |
| 682 | - Observer events use a bounded persistent dispatcher. Queue-full and |
| 683 | dispatcher-unavailable submissions are not retried silently; the TUI keeps |
| 684 | an event-specific error toast separate from the ordinary status line. |
| 685 | - A hook that exceeds its timeout has its whole process group killed and is |
| 686 | then reaped, foreground or background — best-effort, with a bounded reap |
| 687 | wait; see [Timeouts](#timeouts). |
| 688 | |
| 689 | ## Security notes |
| 690 | |
| 691 | - Hooks are arbitrary shell commands from your own config; treat |
| 692 | `~/.codewhale/config.toml` as executable. |
| 693 | - Project-supplied hooks require exact-file approval in addition to workspace trust in |
| 694 | user-owned config. |
| 695 | - Hook commands inherit Codewhale's own environment. A local `exec_shell` does |
| 696 | not — see [`shell_env`](#shell_env). |
| 697 | - `shell_env` audit records contain key names only. That covers Codewhale's own |
| 698 | logging; with an external sandbox backend configured, the values themselves |
| 699 | are transmitted to that backend and are then subject to its handling. |
| 700 | - With an external sandbox backend, the local parent-variable allowlist does |
| 701 | not apply — the backend owns its base environment. |
| 702 | - Payload previews, tool arguments/results, error messages, captured stdout and |
| 703 | stderr, replacement messages, and steering objects are bounded so hook input |
| 704 | or output cannot become an unbounded copy of the transcript. |
| 705 | - Nothing Codewhale persists in a denial echoes the stdin payload, hook |
| 706 | environment, raw stdout/stderr/error, command line, or a resolved filesystem |
| 707 | path. `/hooks list` shows a sanitized, single-line command preview capped at |
| 708 | 60 characters; it is not a verbatim copy. Structured denial reasons are |
| 709 | bounded and redact path-, argument-, command-, and secret-like tokens, |
| 710 | including quoted or `key=value` forms and `Authorization: Bearer …`. |
| 711 |