返回 CodeWhale
CONFIGURATION.md
根目录 / docs / CONFIGURATION.md
1 # Configuration
2
3 codewhale reads configuration from a TOML file plus environment variables.
4 At process startup it may also load literal built-in-provider credentials from
5 a workspace-local `.env` file. Use the tracked `.env.example` as the template;
6 copy it to `.env`, then add only credential values.
7
8 A workspace is not configuration authority. Codewhale therefore ignores
9 config/profile/home paths, provider/model/base-URL routing, MCP/plugin state,
10 approval/sandbox/shell posture, executable paths, runtime settings, and every
11 other non-credential `.env` entry. Variable expansion is rejected so a
12 repository cannot substitute an ambient secret into a credential value. Use
13 `config.toml`, CLI flags, or values exported by the launching shell for those
14 explicit control-plane settings. `.env` is read from a stable regular-file
15 handle, is capped at 1 MiB, and symbolic links, reparse points, and multiply
16 linked files are rejected.
17
18 ## Reading and checking configuration from the CLI
19
20 `codewhale config get <key>` reads scalar keys, whole tables such as `tools`,
21 and nested paths such as `tools.user_input_timeout_seconds`. Displayed tables
22 and nested values apply the same recursive credential redaction as `config dump`.
23
24 `config set` supports its named scalar keys and the provider, route, and
25 notification commands. It refuses a key that nothing reads and suggests the
26 nearest real key (`config set calm_mod on` names `calm_mode`). A settings.toml
27 key such as `calm_mode` or `tool_collapse` is validated and written to the
28 user-global settings.toml, not config.toml; `config get` reads it back from
29 there even when an old config.toml copy (which nothing reads) is still present,
30 and names that copy so you can `config unset` it. Settings keys have no project
31 scope, so `--project` refuses them. A config.toml key is checked against the
32 type its reader expects and stored as a TOML boolean or number where the
33 reader needs one (`yolo = true`, `max_subagents = 4`); `reasoning_effort`
34 accepts the same aliases as `/effort` (#6563). Other dotted writes fail before
35 modifying the file and name the TOML table to edit. For example, set a tools
36 timeout in the file as:
37
38 ```toml
39 [tools]
40 user_input_timeout_seconds = 0
41 ```
42
43 `codewhale config doctor` checks credential presence and endpoint shape, and
44 warns about each config.toml root key that nothing reads: a settings.toml key
45 left in config.toml is named as misplaced, and any other unread key gets a
46 did-you-mean. Keys read by any runtime reader are not reported. The warnings do
47 not fail the check, and a clean result does not validate every runtime setting's
48 value (#6083, #6563).
49
50 ## Constitution, project instructions, and repo authority
51
52 Codewhale has several instruction surfaces. They are deliberately separate so a
53 personal constitution, repo policy, project instructions, and runtime security
54 controls do not blur together.
55
56 - **Bundled global Constitution** — the compiled base law in the binary. It is
57 the default floor for every session.
58 - **User-global constitution** — the normal guided setup output. Manage it with
59 `/constitution` or `/setup`; Codewhale stores structured data at
60 `$CODEWHALE_HOME/constitution.json` (default `~/.codewhale/constitution.json`)
61 and renders it into a separate `<codewhale_user_constitution>` prose block.
62 This can express preferences and stop conditions, but it does not change
63 runtime approval policy, sandbox, shell, network, trust, or MCP permissions.
64 - **Repo-local constitution** — optional project policy in
65 `.codewhale/constitution.json`, described below.
66 - **`AGENTS.md`** — cross-agent **project instructions** (prose). This is the
67 canonical file for "how should an agent work in this repo." Run `/init` to
68 scaffold one. `CLAUDE.md` and `.claude/instructions.md` are read as
69 compatibility fallbacks.
70 - **Memory and handoffs** — recalled state. Useful, but lower authority than
71 constitutions and project instructions.
72
73 ### Managing the user-global constitution (`/setup` and `/constitution`)
74
75 The bundled **working agreement** is the safe default and no longer adds a
76 required first-run screen. Customize it later through `/constitution` or the
77 progressive `/setup` guide. Provider/model readiness, workspace trust, and
78 runtime posture stay separate from this guidance.
79
80 On the **Constitution** step:
81
82 - **`1`–`6`** tune the guided draft. **`G`** previews it, and **`G`** again
83 ratifies and saves a fresh structured `constitution.json`.
84 - **`A`** (shown only when a provider is configured) asks your first configured
85 model to draft the constitution. Drafting is **not** saving: the draft is
86 rendered through the same preview and you still press **`G`** to ratify
87 before anything persists.
88 - **`K`** keeps your existing loaded constitution unchanged (shown only when a
89 valid file is already present).
90 - **`U`** (or `/constitution bundled`) records the bundled/default law.
91
92 `/constitution` (alias `/law`) is the primary management surface once you are
93 set up. Subcommands: `status` (the default), `preview`, `review`, `repo` (the
94 repo-local law block), `explain`, `edit`/`guided`, `repair`, `posture`, and
95 `bundled`. Managing the constitution never changes runtime approval, sandbox,
96 shell, network, trust, default mode, or MCP authority — those stay in runtime
97 posture/config.
98
99 Each repo can carry two distinct, complementary files:
100
101 - **`AGENTS.md`** — ordinary project working instructions.
102 - **`.codewhale/constitution.json`** — Codewhale-specific **repo authority /
103 prioritization policy**: when local sources conflict, which should Codewhale
104 trust first, and what to verify before claiming a task is done. `.codewhale/`
105 lives inside the repo (like `.github/`). Example:
106
107 ```json
108 {
109 "schema_version": 1,
110 "authority": [
111 "current user request",
112 "live code and tests",
113 "GitHub issue/PR details",
114 "AGENTS.md",
115 "memory",
116 "old handoffs"
117 ],
118 "protected_invariants": [
119 "do not break old-session transcript replay"
120 ],
121 "branch_policy": "PRs target the integration branch, not main",
122 "verification_policy": {
123 "before_claiming_done": ["run focused tests", "read changed files back"]
124 },
125 "escalate_when": [
126 "a destructive action was not explicitly authorized"
127 ]
128 }
129 ```
130
131 All fields are optional. When present, the file is rendered into the system
132 prompt as concise prose in a higher-authority block. Legacy `WHALE.md` files
133 are ignored and reported as migration-only diagnostics.
134
135 Each `protected_invariants` entry may be either a plain string (advisory
136 prose, the historical shape) or an object carrying path globs, which is
137 additionally **mechanically enforced** in the tool gate. See
138 [Enforced repo-law invariants](#enforced-repo-law-invariants) below.
139
140 This is the **repo-local law** layer in Codewhale's hierarchy: *bundled global
141 Constitution* → *user-global constitution* (`$CODEWHALE_HOME/constitution.json`,
142 rendered as prose) → *repo constitution* (`.codewhale/constitution.json`, this
143 file) → *AGENTS/project instructions* → *memory and handoffs* → *current
144 request and live evidence for the active turn*. Runtime policy
145 (permissions/sandbox/cost limits enforced in code) is separate from all of
146 these prompt layers. The repo constitution gives project decision rules; it
147 does not replace the bundled Constitution, the user-global constitution, or
148 the current user request.
149
150 > **`WHALE.md` is deprecated.** It overlapped confusingly with `AGENTS.md`.
151 > Codewhale no longer reads `WHALE.md` as project or global context. If one is
152 > present, setup/context diagnostics report it as ignored so you can migrate it.
153 > Move ordinary instructions to `AGENTS.md` and Codewhale-specific authority
154 > policy to `.codewhale/constitution.json`. Personal standing guidance belongs
155 > in `/constitution` / `$CODEWHALE_HOME/constitution.json`. (The global
156 > Codewhale Constitution shipped in the model prompt is a separate thing and is
157 > unaffected.)
158
159 ### Enforced repo-law invariants
160
161 By default a `protected_invariants` entry is advisory prose: it is rendered into
162 the prompt as guidance the agent should honor, but nothing stops a write. An
163 entry written as an **object with `paths`** is different — it compiles into a
164 mechanical write hold that the engine's tool gate evaluates before the write
165 runs. The law becomes mechanism, not just a request.
166
167 An enforced entry has this shape:
168
169 ```json
170 {
171 "schema_version": 1,
172 "protected_invariants": [
173 "Keep DeepSeek support first-class.",
174 {
175 "text": "The wire format is frozen; protocol changes need a human.",
176 "paths": ["crates/protocol/**"],
177 "action": "block"
178 },
179 {
180 "text": "Release notes need human review.",
181 "paths": ["CHANGELOG.md"],
182 "action": "ask"
183 }
184 ]
185 }
186 ```
187
188 - `text` — required. The reason surfaced on the hold. An empty `text` is skipped.
189 - `paths` — workspace-relative globs (globset syntax, e.g. `crates/protocol/**`,
190 `**/secrets.toml`, `CHANGELOG.md`). An object with no usable `paths` stays
191 advisory-only despite the object shape.
192 - `action` — optional, defaults to `ask`. `ask` force-prompts in Ask and
193 Auto-Review; in Full Access it denies the protected write without opening a
194 modal. `block` **denies the write outright** in every posture.
195
196 Semantics:
197
198 - **Tighten-only.** The schema has no allow/widen shape, so law can only *add*
199 holds — a crafted constitution can never grant authority or weaken a gate
200 above it.
201 - **Not bypassable by mode.** Like the built-in safety floor, an `ask` hold
202 force-prompts in Ask and Auto-Review. Full Access never opens approval
203 modals, so the same hold fails closed as a hard block; `block` always denies.
204 Mode cannot turn a hold off.
205 - **Repo-local only.** Only the repo's `.codewhale/constitution.json`
206 participates. The user-global constitution stays advisory prose and never
207 reaches this mechanism.
208 - **Fails safe.** A missing file, parse error, or invalid glob degrades to
209 fewer or zero rules — never a hold on unprotected paths and never a poisoned
210 gate. Across matches the strongest action wins, so `block` outranks `ask`.
211 - **Leaves a receipt.** Every hold emits a `tool.repo_law_decision` tool-audit
212 event naming the invariant, the matched path, and the source file; the
213 approval/denial reason names the invariant too.
214
215 **Coverage is deliberately limited.** Holds are evaluated only for the write
216 tools `write_file`, `edit_file`, `apply_patch`, and `fim_edit`, and only
217 against the filesystem targets named in their inputs (`path`/`target`/
218 `destination`/`file_path`, `changes[].path`, and unified-diff /
219 `apply_patch`-envelope headers). A shell command that writes a protected path is **not** held by
220 repo law — those writes are still governed by the ordinary approval, sandbox,
221 and shell-write gates, not by this mechanism.
222
223 ### Expert full base-prompt override (#3638)
224
225 The global Constitution (the base system prompt, normally compiled in from
226 `crates/tui/src/prompts/text.rs` as `BASE_PROMPT`) can be replaced per-user
227 without rebuilding. This is
228 an expert escape hatch, not the normal `/constitution` guided setup output.
229 Because this is a prompt trust boundary, it takes **two deliberate steps** — a
230 file alone is not enough:
231
232 1. Drop the replacement at `~/.codewhale/prompts/constitution.md` (under
233 `$CODEWHALE_HOME` when set).
234 2. Set the explicit opt-in flag `CODEWHALE_ALLOW_BASE_PROMPT_OVERRIDE=1`
235 (`true`/`on`/`yes` also accepted).
236
237 If the file exists but the flag is unset, the override is **ignored** (with a
238 log line pointing to the flag) and the bundled Constitution stays in place.
239 This is intended for repurposing the TUI beyond software engineering — e.g.
240 long-form writing or document review — where the engineering-oriented base
241 prompt is a poor fit. It is loaded once at startup; a **missing or empty file
242 is a no-op**, so existing installs keep the bundled prompt.
243
244 Scope is deliberately narrow: only the byte-stable **base prompt segment** is
245 overridable. Mode deltas, the approval policy, the tool taxonomy, Context
246 Management, and the Compaction Relay are still owned by Codewhale's runtime
247 assembly, so an override **cannot remove safety-relevant guidance** (sandbox,
248 approvals) — it only swaps the task/voice framing. To customize ordinary
249 personal behavior, prefer `/constitution`; to customize per-repo behavior,
250 prefer `AGENTS.md` + `.codewhale/constitution.json` above.
251
252 ## Where It Looks
253
254 Default config path:
255
256 - `~/.codewhale/config.toml`
257 - Legacy fallback: `~/.deepseek/config.toml`
258
259 Overrides:
260
261 - CLI: `codewhale --config /path/to/config.toml`
262 - Env: `CODEWHALE_CONFIG_PATH=/path/to/config.toml`
263 - Legacy env alias: `DEEPSEEK_CONFIG_PATH=/path/to/config.toml`
264
265 If both are set, `--config` wins. Environment variable overrides are applied after the file is loaded.
266
267 ### TUI editability audit
268
269 Inside the TUI, run `/config audit` to see which documented keys can be changed
270 from the current session, which ones can also be persisted, and which ones stay
271 file-only or restart-only. The audit includes current values for the high-impact
272 runtime controls such as `approval_policy`, `allow_shell`,
273 `stream_chunk_timeout_secs`, `base_url`, `mcp_config_path`, and the
274 `[subagents]` concurrency/depth/timeout keys.
275
276 Use the command's "Command / reason" column as the source of truth before
277 editing by hand. For example, `/config approval_mode on-request --save` writes
278 top-level `approval_policy = "on-request"`, while provider base URLs are saved
279 but still require restarting the model client.
280
281 ### User workspace entries
282
283 Interactive Agent sessions expose shell tools by default with approval gating
284 unless you explicitly disable them. For a shell opt-in that should live in the
285 user's global config for noninteractive or durable-task profiles rather than in
286 the repository, add a workspace-scoped entry:
287
288 ```toml
289 [workspace.'/absolute/path/to/project']
290 allow_shell = true
291 ```
292
293 The entry applies only when the launched workspace path matches the table key.
294 The legacy `[projects."/absolute/path/to/project"]` table is also accepted for
295 this user-owned override.
296
297 In interactive mode, the per-project overlay
298 `<workspace>/.codewhale/config.toml` is applied after this user entry. A
299 project-level `allow_shell = false` can still tighten the session; project-level
300 `allow_shell = true` is ignored.
301
302 ### Per-project overlay (#485)
303
304 When the TUI starts in a workspace that contains a regular-file
305 `<workspace>/.codewhale/config.toml`, the safe values declared in that file are
306 merged on top of the global config. Legacy
307 `<workspace>/.deepseek/config.toml` files are still read when the Codewhale path
308 is absent. Symlinked project config files are rejected. This lets a repo suggest
309 a model or tighten local safety posture without touching the user's
310 `~/.codewhale/config.toml`. Pass `--no-project-config` to skip the overlay for
311 one launch.
312
313 Supported keys in the project overlay (top-level fields only):
314
315 | Key | Effect |
316 |---|---|
317 | `model` | override `default_text_model` |
318 | `reasoning_effort` | force `"high"` / `"max"` for a complex repo |
319 | `approval_policy` | only values that tighten the user's current permission posture |
320 | `sandbox_mode` | only values that tighten the user's current sandbox posture |
321 | `max_subagents` | clamp sub-agent concurrency for a constrained repo (clamped to 1..=128) |
322 | `allow_shell` | `false` can disable shell access; `true` is ignored |
323
324 The overlay is intentionally narrow — it covers the fields a repo
325 maintainer is most likely to want to standardize across contributors.
326 Credential, endpoint, provider-selection, MCP config, hooks, skills,
327 retry, hotbar bindings, and `instructions = [...]` settings stay user-global.
328 If a repo-local config declares `api_key`, `base_url`, `providers`, `provider`,
329 `mcp_config_path`, `notes_path`, `hotbar`, `allow_shell = true`, or `instructions`,
330 Codewhale ignores that key and keeps the user's global setting.
331
332 The consolidated `codewhale` runtime uses one config file for DeepSeek auth
333 and model defaults. `codewhale auth set --provider deepseek` saves
334 the key to `~/.codewhale/config.toml` (migrating legacy `~/.deepseek/config.toml`
335 on first launch when needed), and `codewhale --model deepseek-v4-flash` is
336 forwarded to the TUI as `CODEWHALE_MODEL`. The dispatcher no longer writes the
337 `DEEPSEEK_*` twins of these variables; a `DEEPSEEK_*` value you set yourself is
338 still read as a legacy alias when the `CODEWHALE_*` one is unset.
339
340 `codewhale login` signs in to the Codewhale account — it is the same browser
341 device flow as `codewhale account login`, not a provider-key command. Provider
342 credentials are configured exclusively through `codewhale auth set
343 --provider <provider>`.
344
345 That provider credential is distinct from the optional managed-product
346 account. `codewhale account login` starts the Codewhale browser device flow;
347 `codewhale account status` and `codewhale account logout` inspect or remove the
348 session for the selected `--profile`. Account sessions prefer the OS
349 credential manager and fall back automatically to the private `0600`
350 Codewhale secrets file when no credential manager is available (headless
351 hosts, SSH, containers).
352 `codewhale account keys list|set|remove` manages the
353 signed-in account's BYOK vault without displaying secret values. The older
354 `codewhale cloud ...` spelling remains a command alias.
355
356 ### Portable config bundles
357
358 `codewhale config export --portable [--project] [--out FILE]` writes a
359 portable, secret-free bundle of your configuration: sorted TOML with
360 credential and machine-specific keys (API keys, base URLs, socket paths)
361 dropped, never a redacted placeholder in their place. Without `--out` the
362 bundle goes to stdout. Typed tables, arrays, numbers, booleans, and datetimes
363 remain typed. Machine-bound authority is deliberately non-portable: project
364 trust overlays, credential readers, auto-running hooks, executable LSP
365 definitions, and local path bindings are omitted rather than copied to a new
366 host. So is trust posture (`yolo`, `allow_shell`, `approval_policy`,
367 `sandbox_mode`, `sandbox_network_access`, and the sandbox read-denylist
368 paths) and local endpoints or executables (`[lifecycle_outbox]`,
369 `[extension_host]`, `[control_socket]`). URLs that carry a credential —
370 userinfo, a token-named query parameter, or a Slack/Discord/Teams webhook
371 path — are treated as credentials.
372
373 `codewhale config import <FILE|HTTPS_URL|-> [--dry-run] [--yes] [--project]`
374 applies a bundle. The envelope is strict (`schema_version = 1`, kind
375 `codewhale.portable-config`; unknown fields fail). Import prints a
376 deterministic plan — added / changed / skipped / conflicting / rejected —
377 then asks for consent unless `--yes` is given; headless use requires it.
378 Credential-shaped entries are rejected by key name and by value shape;
379 rejections name the field, never a value. Remote bundles come from HTTPS
380 only (loopback http excepted) with a 5 MiB cap. Application backs up the
381 target document to `<config>.bundle-backup-<timestamp>-<random>`, rolls back on any
382 failure, and re-importing an applied bundle changes nothing.
383
384 Import also rejects the non-portable authority classes omitted by export,
385 including nested/camel/dotted credential keys and cookie headers. This keeps a
386 hand-authored or remote bundle from reintroducing machine trust, local
387 credential access, or automatically executable commands that a local export
388 would refuse to carry. Structured tables are deep-merged: a portable model or
389 preference update does not erase target-local provider credentials, endpoints,
390 hooks, or executable definitions that were deliberately omitted from the
391 bundle. Arrays and scalar values still replace the corresponding portable
392 value.
393
394 Sections map to scope: `[project]` entries land only in the workspace
395 document (`--project`, which must target an actual workspace config),
396 `[global]` only in the user-global one; `preferences`, `profiles`, and
397 `plugins` apply at either scope. A global bundle operation refuses a workspace
398 document just as a project operation refuses the user-global document.
399
400 ### Credential read precedence (#5197)
401
402 Credential reads are **folder-independent by default**: every layer below is
403 user-global or process-scoped, so a key saved in one repo resolves identically
404 in every other repo. Repo-local config never carries credential material — the
405 project overlay above reads only its allowlisted keys and ignores `api_key`,
406 and a credential write aimed at a workspace-scoped config is rescoped to the
407 user-global `~/.codewhale/config.toml` (#5045, #5193).
408
409 For the active provider, the runtime resolves the API key in this exact order
410 (first match wins):
411
412 1. **Route-specific auth contract.** Routes whose `auth_mode` disables API
413 keys stop here with no credential. OAuth routes use their explicitly
414 granted token: the official `openai-codex` plan route accepts only
415 Codewhale-owned credentials from `codewhale auth chatgpt`, bound to its
416 issued registration and verified account. Imported Codex CLI and process
417 tokens cannot authorize this route;
418 `[providers.xai] auth_mode = "oauth"` reads Codewhale's own xAI
419 device-login store (or a consent-granted Grok CLI file).
420 2. **Explicit CLI key.** `--api-key` forwarded with its source marker wins
421 over every saved slot.
422 3. **Config file `api_key`.** The `[providers.<name>] api_key` table slot for
423 the active provider. `deepseek-CN` also reads `[providers.deepseek]`. An
424 older top-level `api_key` is read as `[providers.deepseek] api_key` (see
425 [Legacy top-level `base_url` and `api_key`](#legacy-top-level-base_url-and-api_key)).
426 File-owned keys stay bound to their file-owned endpoint: when the
427 environment replaces the route's base URL with a custom host, the saved
428 key is not sent there.
429 4. **`api_key_env` binding.** `[providers.<name>] api_key_env = "VAR"` reads
430 the named environment variable. For custom providers an unset or empty
431 binding is a loud error, not a silent fallback (#5104).
432 5. **Secret store.** The durable per-provider slot written by
433 `codewhale auth set` (file-backed under `~/.codewhale/secrets/` by
434 default; the OS keyring only when explicitly selected). Skipped for named
435 custom routes, self-hosted providers, custom endpoints other than an
436 explicitly authenticated loopback, and routes whose `auth_mode` needs no
437 key.
438 6. **Ambient environment.** The provider's own variable
439 (`DEEPSEEK_API_KEY`, `OPENROUTER_API_KEY`, `MOONSHOT_API_KEY`, …).
440 Ambient keys are only ever sent to the provider's official endpoint and
441 are skipped under the same conditions as the secret store.
442 7. **Keyless fallback.** Self-hosted providers and loopback endpoints may run
443 with no credential; every other route fails with provider-specific setup
444 guidance.
445
446 Legacy compatibility: `~/.deepseek/config.toml` is migrated into
447 `~/.codewhale/config.toml` on first launch, `DEEPSEEK_*` environment
448 variables remain accepted aliases for the `CODEWHALE_*` forms, and
449 `DEEPSEEK_SECRET_BACKEND` is the legacy alias for `CODEWHALE_SECRET_BACKEND`.
450
451 Run `codewhale auth status` to inspect the active provider's config
452 file, OS keyring backend, environment variable, winning source, and last-four
453 label without printing the key itself. The command only probes the active
454 provider's keyring entry.
455
456 For hosted, generic OpenAI-compatible, self-hosted, OpenAI Responses, or native
457 Anthropic providers, set `provider = "<id>"` or pass
458 `codewhale --provider <id>`. The canonical provider IDs are `deepseek`,
459 `nvidia-nim`, `openai`, `atlascloud`, `wanjie-ark`, `volcengine`,
460 `openrouter`, `orcarouter`, `xiaomi-mimo`, `novita`, `fireworks`,
461 `siliconflow`, `arcee`, `siliconflow-CN`, `moonshot`, `sglang`, `vllm`,
462 `ollama`, `ollama-cloud`, `huggingface`, `modelscope`, `together`, `qianfan`, `openai-codex`,
463 `anthropic`, `openmodel`, `zai`, `stepfun`, `minimax`, `deepinfra`,
464 `sakana`, `longcat`, `opencode-go`, `opencode-zen`, `meta`, `xai`,
465 `mistral`, `telecomjs`, `modelstudio-token-plan`, `google`,
466 `edenai`, `concentrate`, `codewhale`, and `custom` (a user-defined OpenAI-compatible endpoint via
467 `[providers.<name>]`).
468 For the provider-by-provider registry, including wire protocol, auth variables,
469 default base URLs, model IDs, and capability metadata, see
470 [PROVIDERS.md](PROVIDERS.md).
471 The facade saves provider credentials to the shared user config and forwards
472 the resolved key, base URL, provider, and model to the TUI process. Use
473 `codewhale auth set --provider nvidia-nim --api-key "YOUR_NVIDIA_API_KEY"` or
474 `codewhale auth set --provider openai --api-key "YOUR_OPENAI_COMPATIBLE_API_KEY"` or
475 `codewhale auth set --provider atlascloud --api-key "YOUR_ATLASCLOUD_API_KEY"` or
476 `codewhale auth set --provider wanjie-ark --api-key "YOUR_WANJIE_API_KEY"` or
477 `codewhale auth set --provider xiaomi-mimo --api-key "YOUR_XIAOMI_KEY"` or
478 `codewhale auth set --provider fireworks --api-key "YOUR_FIREWORKS_API_KEY"` or
479 `codewhale auth set --provider siliconflow --api-key "YOUR_SILICONFLOW_API_KEY"` or
480 `codewhale auth set --provider arcee --api-key "YOUR_ARCEE_API_KEY"` or the
481 matching provider ID from [PROVIDERS.md](PROVIDERS.md) to save provider keys
482 through the facade. The generic `openai` provider defaults
483 to `https://api.openai.com/v1`, accepts `OPENAI_BASE_URL`, and defaults to
484 `gpt-5.6`. A custom OpenAI-compatible gateway can still select its own model
485 explicitly. `atlascloud` defaults to
486 `https://api.atlascloud.ai/v1`, accepts `ATLASCLOUD_BASE_URL`, and uses
487 `deepseek-ai/deepseek-v4-flash` as its default model. `wanjie-ark` targets
488 Wanjie Ark's OpenAI-compatible endpoint at
489 `https://maas-openapi.wanjiedata.com/api/v1`, defaults to `deepseek-reasoner`,
490 and passes model IDs through unchanged because Wanjie model access is
491 account-scoped. SGLang, vLLM, and Ollama are
492 self-hosted and can run without an API key by default. Ollama defaults to
493 `http://localhost:11434/v1` and sends model tags such as `codewhale-coder:1.3b`
494 or `qwen2.5-coder:7b` unchanged. Self-hosted providers and loopback custom
495 URLs (`localhost`, `127.0.0.1`, `[::1]`, `0.0.0.0`) do not read the secret store
496 unless API-key auth is explicitly requested; use an env var or config-file key
497 when a local server does require bearer auth.
498 Ollama Cloud is the separate hosted `ollama-cloud` provider. It defaults to
499 `https://ollama.com/v1` and `gpt-oss:120b`; save its key with
500 `codewhale auth set --provider ollama-cloud`. Ambient auth reads
501 `OLLAMA_CLOUD_API_KEY` first, then Ollama's official `OLLAMA_API_KEY`.
502 SiliconFlow defaults to `https://api.siliconflow.com/v1`, accepts
503 `SILICONFLOW_BASE_URL`, and uses `deepseek-ai/DeepSeek-V4-Pro` by default.
504 `provider = "siliconflow-CN"` selects the China regional default
505 `https://api.siliconflow.cn/v1` with the `[providers.siliconflow_cn]` table and
506 `SILICONFLOW_API_KEY` credential slot.
507 Arcee AI defaults to `https://api.arcee.ai/api/v1`, accepts `ARCEE_BASE_URL`,
508 and uses `trinity-large-thinking` by default for Codewhale agent work.
509 `trinity-large-preview` is also listed as a direct Arcee API model; OpenRouter's
510 `arcee-ai/trinity-large-thinking` remains the OpenRouter namespaced form, while
511 the direct Arcee provider uses the bare `trinity-large-thinking` ID. Direct
512 Arcee large-model API calls are tracked as 256K-context BF16 serving; Thinking
513 is reasoning-capable, while Preview is not marked as a thinking model.
514
515 ### OpenRouter vendor pinning
516
517 OpenRouter serves each model through several upstream vendors, and Codewhale
518 can pin requests to a vendor with `[providers.openrouter] vendor` (#6007):
519
520 ```toml
521 provider = "openrouter"
522 [providers.openrouter]
523 model = "deepseek/deepseek-v4-pro"
524 vendor = "deepinfra" # copy the vendor slug from the model's OpenRouter page
525 ```
526
527 This sends `"provider": {"order": ["deepinfra"], "allow_fallbacks": false}`
528 on OpenRouter requests. A base slug can match multiple endpoint variants;
529 copy a full slug such as `deepinfra/turbo` to select one variant. An unavailable
530 pin fails at OpenRouter. Codewhale's separate `fallback_providers` setting can
531 still switch the whole route after a recoverable error.
532
533 The pin applies across OpenRouter models, including auxiliary requests on that
534 route. Set `vendor = ""` to clear it. Reload config or restart to apply edits;
535 requests already in flight keep their captured route. Other providers do not
536 inherit the pin. `/preview-request` shows the primary request's routing fields.
537
538 Model strings still pass through verbatim: `:floor` sorts by price, `:nitro`
539 sorts by throughput, and `@preset/my-team-preset` references an account preset.
540 See OpenRouter's [provider routing](https://openrouter.ai/docs/guides/routing/provider-selection)
541 and [presets](https://openrouter.ai/docs/guides/features/presets) documentation.
542 Codewhale does not fetch per-vendor endpoint prices or availability; pinned
543 usage reports a routing-dependent unknown cost instead of a catalog estimate.
544
545 ### Custom OpenAI-Compatible Gateways
546
547 For a single third-party service that implements the OpenAI Chat Completions
548 API, the simplest setup is the built-in `openai` provider name pointed at the
549 gateway:
550
551 ```toml
552 provider = "openai"
553 default_text_model = "your-model-id"
554
555 [providers.openai]
556 api_key = "YOUR_OPENAI_COMPATIBLE_API_KEY"
557 base_url = "https://your-gateway.example/v1"
558 ```
559
560 Put the endpoint under `[providers.openai]`, not the legacy top-level
561 `base_url`, so the OpenAI-compatible provider receives it. `default_text_model`
562 is the model ID sent to the gateway; `[providers.openai].model` can be used as
563 the OpenAI-provider-specific override.
564
565 If you keep several OpenAI-compatible gateways, or need a stable name for an
566 AgentProfile provider pin, define a user-named custom provider table:
567
568 ```toml
569 provider = "lm-studio"
570
571 [providers.lm-studio]
572 kind = "openai-compatible"
573 base_url = "http://127.0.0.1:1234/v1"
574 api_key = "lm-studio"
575 model = "qwen-2.5-7b"
576 ```
577
578 Custom provider names may be selected with `provider = "<name>"`,
579 `--provider <name>`, or an AgentProfile `provider = "<name>"` when the matching
580 `[providers.<name>]` table exists.
581
582 StepFun has a first-class provider entry, so keep Coding Plan credentials and
583 base URL scoped to `[providers.stepfun]`:
584
585 ```toml
586 provider = "stepfun"
587
588 [providers.stepfun]
589 api_key = "YOUR_STEPFUN_API_KEY"
590 base_url = "https://api.stepfun.ai/step_plan/v1"
591 model = "step-3.7-flash"
592 ```
593
594 `/provider` setup asks which StepFun billing route the key belongs to —
595 pay-as-you-go (`https://api.stepfun.ai/v1`) or a Step Plan subscription
596 (`https://api.stepfun.ai/step_plan/v1`) — and validates the key against the
597 endpoint you pick before saving it. The answer is written to
598 `[providers.stepfun].base_url` and nowhere else. If that key already holds a
599 base URL Codewhale does not recognize as one of those two routes, the question
600 is skipped and your value is left untouched.
601
602 Alibaba Bailian / Model Studio DashScope Qwen routes use the same OpenAI
603 provider shape:
604
605 ```toml
606 provider = "openai"
607
608 [providers.openai]
609 api_key = "YOUR_DASHSCOPE_API_KEY"
610 base_url = "https://dashscope-intl.aliyuncs.com/compatible-mode/v1"
611 model = "qwen-plus"
612 context_window = 1000000
613 ```
614
615 Use the regional DashScope `compatible-mode/v1` base URL that matches the
616 region of your API key. Codewhale keeps `qwen-plus` scoped to the `openai`
617 provider route and does not infer a different provider from the model prefix.
618 The same rule applies to all provider-prefixed model strings: a prefix such as
619 `deepseek-ai/...` or `deepseek/...` is a provider-owned wire ID under the
620 selected provider, not an automatic switch to the DeepSeek provider.
621 Set `context_window` to the gateway/model's real total context window when it
622 differs from Codewhale's static model metadata. See
623 [Context length (context window)](#context-length-context-window) for the full
624 resolution order and for how to check which value is in effect.
625
626 If the gateway accepts `POST /chat/completions` but rejects
627 `/v1/chat/completions`, set a provider-local `path_suffix`:
628
629 ```toml
630 [providers.openai]
631 base_url = "https://your-gateway.example/v1"
632 path_suffix = "/chat/completions"
633 ```
634
635 The suffix applies only to chat-completion requests. Model listing and
636 DeepSeek beta paths keep their built-in routing so a generic gateway override
637 does not accidentally rewrite `/models` or `/beta/completions`.
638
639 For private gateways with broken or intercepted certificates, use
640 `SSL_CERT_FILE` with a trusted CA bundle. The legacy provider-table key
641 `insecure_skip_tls_verify = true` is still parsed so `codewhale doctor` can
642 report stale configs, but provider clients reject it instead of disabling TLS
643 certificate verification.
644
645 Local HTTP endpoints such as Ollama, SGLang, and vLLM are allowed by default
646 when they use localhost or loopback addresses. For a non-local `http://`
647 gateway, launch with `DEEPSEEK_ALLOW_INSECURE_HTTP=1` only on a trusted network:
648
649 ```bash
650 DEEPSEEK_ALLOW_INSECURE_HTTP=1 codewhale
651 ```
652
653 Third-party OpenAI-compatible gateways that need extra request headers can set
654 `http_headers = { "X-Model-Provider-Id" = "your-model-provider" }` at the top
655 level or under a provider table such as `[providers.deepseek]`. When configured,
656 codewhale sends those custom headers on model API requests. The equivalent
657 environment override is `DEEPSEEK_HTTP_HEADERS`, using comma-separated
658 `name=value` pairs such as
659 `X-Model-Provider-Id=your-model-provider,X-Gateway-Route=dev`. `Authorization`
660 and `Content-Type` are managed by the client and are not overridden by this
661 setting.
662
663 ### Vision Model
664
665 Codewhale's chat provider and `image_analyze` tool are configured separately.
666 The main chat path remains the selected text/tool provider; image analysis runs
667 through `[vision_model]` when the `vision_model` feature is enabled.
668
669 Xiaomi's current image-understanding docs include `mimo-v2.5` for image input.
670 To use MiMo for `image_analyze`, configure the vision model explicitly:
671
672 ```toml
673 [features]
674 vision_model = true
675
676 [vision_model]
677 model = "mimo-v2.5"
678 api_key = "YOUR_XIAOMI_KEY"
679 base_url = "https://api.xiaomimimo.com/v1"
680 ```
681
682 The example above uses Xiaomi MiMo's pay-as-you-go OpenAI-compatible endpoint.
683 If you are using a Token Plan key (`tp-...`) for `[vision_model]`, you must set
684 `base_url` explicitly because this generic OpenAI-compatible block does not
685 auto-select MiMo endpoints. Use
686 `https://token-plan-sgp.xiaomimimo.com/v1` for Singapore accounts,
687 `https://token-plan-cn.xiaomimimo.com/v1` for China-region accounts, or
688 `https://token-plan-ams.xiaomimimo.com/v1` for Europe/Amsterdam accounts.
689
690 ### Auto Model Routing (`[auto]`, `[auto.router]`)
691
692 With `model = "auto"`, each turn runs on your **declared default model** unless
693 you have opted into something else. Auto never guesses a cheaper or stronger
694 model from how a request is worded; the old keyword-and-length heuristic was
695 removed (`auto_route_declared_fallback` in `crates/tui/src/model_routing.rs`).
696 Two optional layers change that:
697
698 - an `[auto.router]` classifier, which you write down, that picks a model per
699 turn; and
700 - `[auto] cost_saving`, which prefers the active provider's fast sibling.
701
702 With neither set, Auto is local and free: the turn uses the default model and
703 no classifier call is made.
704
705 **There is no default classifier.** With `[auto.router]` unset, no classifier
706 call happens, whatever keys you hold. Holding a DeepSeek key used to elect
707 `deepseek-v4-flash` automatically. That was removed because it spent tokens on
708 a route the user never chose and privileged one provider over the rest
709 (`AutoRouterConfig` in `crates/tui/src/config.rs`). Electing a network
710 classifier is now something you write down.
711
712 Point the classifier at any configured provider with `[auto.router]`:
713
714 ```toml
715 [auto.router]
716 provider = "zai"
717 model = "glm-5-turbo"
718 thinking = "off" # optional; defaults to off
719 timeout_secs = 4 # optional; default 4, 0 = default, capped at 300
720 ```
721
722 A classifier call happens only when `[auto.router]` names both `provider` and
723 `model` *and* that provider has a key:
724 `router_available = router_configured && has_api_key_for(...)` in
725 `ModelInventory::from_config` (`crates/tui/src/model_inventory.rs`). If either
726 condition fails, or the classifier call errors or times out, the local
727 fallback decides: the default model, or the fast sibling under `cost_saving`.
728 The turn's route receipt (Turn Inspector, Ctrl+Alt+O or `/turn inspect`, "Model
729 route + tokens/cost") records which path was taken, and a
730 router you configured that cannot run or fails (missing key, HTTP error,
731 timeout, invalid answer) is shown as `Auto router: failing — …` rather than
732 silently ignored.
733
734 #### Set up model routing
735
736 `/router` (also `/model router`) opens one Router setup view with these
737 presets. Each writes only `[auto.router]`; none is ever chosen for you.
738
739 | Preset | What it writes | Cost and privacy |
740 | --- | --- | --- |
741 | `/router jev` | Jev, TypeSafe's decision model, over OpenRouter (`typesafe/jev-1.13`) or TypeSafe direct, whichever key you have | About $0.00002 per turn ($0.042 per million input tokens, output free). Your latest request and up to six recent context lines go to OpenRouter → TypeSafe (or TypeSafe). |
742 | `/router fast` | The active provider's runnable fast tier with thinking off | Your existing key; the classifier sees the same request text. |
743 | `/router off` | Removes `[auto.router]` | No router call; Auto turns use the default model, or the fast tier while `[auto] cost_saving = true` (Off leaves that setting alone). |
744 | `/router custom` | Nothing; prints the TOML to edit | — |
745
746 Choosing Jev or Fast makes **one test call** with a fixed sample request and
747 shows the tier it picked, the probabilities and confidence, the latency and the
748 provider-reported cost. `Enter` (or `/router save <preset>`) then saves through
749 the normal config writer; `Esc` discards it. TypeSafe paused new signups on
750 2026-09-22, so OpenRouter is the default route for new users. A TypeSafe key is
751 read from `TYPESAFE_API_KEY`, the `typesafe` secret-store entry, or
752 `[providers.typesafe] api_key` / `api_key_env`.
753
754 #### Decision routers (`kind = "decision"`)
755
756 A decision router asks a non-generative decision model typed Choice questions
757 for the active provider's `fast` and `strong` tiers and thinking level, then
758 reads calibrated probabilities. No prose is parsed.
759
760 ```toml
761 [auto.router]
762 kind = "decision" # default "chat"
763 provider = "openrouter" # or "typesafe"
764 model = "typesafe/jev-1.13" # "~typesafe/jev-latest" also works; TypeSafe direct: "jev-latest"
765 timeout_secs = 2
766 min_confidence = 0.5 # default 0.5, clamped to 0..1
767 ```
768
769 - The router is called only when the active provider has a runnable strong/fast
770 pair; otherwise there is no call and no spend.
771 - An answer below `min_confidence` takes the local fallback. Under
772 `[auto] cost_saving`, a `strong` answer also needs a probability of at least
773 0.75, or the turn stays on the fast tier.
774 - An unknown `kind`, or a decision `provider` other than `openrouter` /
775 `typesafe`, leaves the router unconfigured and shown as failing.
776 - `[auto.router] thinking` is ignored for decision routers. Both routes settle tokens through
777 the originating session's routed-usage ledger. TypeSafe is a named Custom route
778 with unknown billing; its price is never borrowed from the active chat provider.
779 Provider-reported cost remains verbatim on the decision receipt.
780 - OpenRouter uses `POST /api/alpha/decisions`; TypeSafe uses `POST /v1/systemone`.
781 The shared client validates Choice, Noul and Score against the offered questions
782 and bounds responses to 256 KiB. Malformed answers fail open with usage retained.
783
784 #### Shadow Decision Gate (experimental, off by default)
785
786 The Superfast Decision Gate asks a decision model three typed questions about
787 the latest user message — does it need a tool, can it be answered from the
788 conversation, and what is its intent — and logs a conservative recommendation.
789 It is shadow-only: it never changes routing, never skips or delays the model
790 call, and fails open on any error, timeout or malformed answer. It uses the same
791 System One client as the decision router above; there is no separate HTTP
792 client. It is configured from the environment and reads it when a turn starts:
793
794 ```sh
795 SUPERFAST_ENABLED=1 # off unless set
796 SUPERFAST_PROVIDER=typesafe # or openrouter; required when enabled
797 SUPERFAST_BASE_URL=http://localhost:8000/v1 # optional TypeSafe-route base, e.g. self-hosted
798 SUPERFAST_MODEL=jev-latest # default jev-latest / ~typesafe/jev-latest
799 SUPERFAST_TIMEOUT_MS=150 # 1..=10000, default 150
800 ```
801
802 - Enabling the gate never picks an endpoint by itself: without
803 `SUPERFAST_PROVIDER` nothing is sent and a warning is logged.
804 - The key comes from the same place the decision router reads it. The TypeSafe
805 route always authenticates, so a self-hosted server that ignores auth still
806 needs a placeholder `TYPESAFE_API_KEY`.
807 - Only the latest user message is sent, truncated to 4,000 characters and
808 redacted of configured secrets. The log (target `superfast`) carries the
809 route, failure class and latency, never prompt text.
810 - The gate retains the originating turn's accounting owner and cancellation.
811 Its tokens settle through the shared ledger; missing usage or a cancelled/timed
812 out request after dispatch creates an explicit coverage gap. Unknown TypeSafe
813 pricing is recorded as unpriced rather than free.
814 - Provider-reported cost survives rejected answers and late responses in bounded
815 receipts on the originating turn or session. These are diagnostic evidence;
816 unknown pricing never becomes an authoritative dollar total. Incomplete token
817 counters record a coverage gap while preserving the reported raw evidence.
818
819 The wire contracts are documented in [TypeSafe's OpenAPI schema](https://api.typesafe.ai/openapi.json)
820 and [OpenRouter's Decisions examples](https://openrouter.ai/blog/insights/what-is-jev/).
821
822 The Decision Gate concept and reference implementation are by Andrea Bruno,
823 released under CC BY 4.0:
824 [harness-superfast](https://github.com/Andrea-Bruno/harness-superfast).
825
826 Two `[auto]` keys shape routing (`AutoConfig` in `crates/tui/src/config.rs`):
827
828 ```toml
829 [auto]
830 cost_saving = false # default false
831 cross_provider = false # default false
832 ```
833
834 - **`cost_saving`** (default `false`). Without a classifier, Auto pins the
835 active provider's validated fast sibling instead of the default model. A
836 provider with no runnable fast sibling stays on the default. With a
837 classifier, the classifier is told to prefer the fast tier for routine or
838 ambiguous work and to pick the strong tier only for clearly agentic,
839 multi-step, architecture, security or debugging work. Cost-saving never
840 switches provider just to save money.
841 - **`cross_provider`** (default `false`). Auto stays on the provider the session
842 is configured to use. The classifier is only shown that provider's models,
843 and the fallback never leaves it. Setting `cross_provider = true` lets the
844 classifier choose among every runnable provider. There is no interactive
845 toggle for `cross_provider`; it has to be set in config. A decision router
846 always chooses within the active provider.
847
848 To bootstrap MCP and skills directories at their resolved paths, run `codewhale setup`.
849 To only scaffold MCP, run `codewhale mcp init`.
850
851 Note: `setup`, `doctor`, `mcp`, `features`, `sessions`, `resume`/`fork`, `exec`,
852 `review`, and `eval` are all available from the installed `codewhale` command.
853 The consolidated dispatcher also provides `auth`, `config`, `model`, `thread`, `sandbox`,
854 `app-server`, `mcp-server`, `completions`, `login`/`logout`, `account`,
855 `metrics`, `update`, `lane`, `workflow`, and `web`. Plain prompts enter the
856 in-process TUI runtime. Release installers expose the same bytes as `codew`.
857
858 ### Startup Update Checks
859
860 By default, the TUI starts a background check for the latest stable Codewhale
861 release and shows a short toast only when a newer release is available and the
862 official release assets are complete. The check never blocks startup, never
863 blocks a turn, and fails silently when offline.
864
865 Disable the startup check entirely for air-gapped, corporate-proxy, or managed
866 desktop environments:
867
868 ```toml
869 [update]
870 check_for_updates = false
871 ```
872
873 #### Throttling
874
875 The answer is cached in `~/.codewhale/update-check.json` and reused for
876 `check_interval_hours` (default `1`). Only the *network request* is throttled —
877 the notice still appears on every launch while an update is outstanding. Set `0`
878 to check on every launch.
879
880 ```toml
881 [update]
882 check_interval_hours = 1
883 ```
884
885 A failed check is not cached, so an outage does not suppress the notice until
886 the interval elapses.
887
888 #### Automatic suppression
889
890 Checks are skipped, without contacting the network, when any of these is set to
891 a non-falsey value:
892
893 | Variable | Why |
894 | --- | --- |
895 | `CODEWHALE_NO_UPDATE_CHECK` | Explicit opt-out. |
896 | `NO_UPDATE_NOTIFIER` | The cross-CLI convention, honored for compatibility. |
897 | `CI`, `CONTINUOUS_INTEGRATION`, `GITHUB_ACTIONS`, `GITLAB_CI`, `BUILDKITE`, `CIRCLECI`, `JENKINS_URL`, `TEAMCITY_VERSION`, `TF_BUILD` | Automated build; nobody is at the terminal. |
898
899 Values of `""`, `0`, `false`, `no`, and `off` do not count as set, so a
900 `CI=false` export does not disable checks for ordinary users.
901
902 #### Which update command is offered
903
904 Codewhale never installs anything on its own — it only tells you an update
905 exists. The command it names depends on how the running binary was installed,
906 detected from its path:
907
908 | Install | Command offered |
909 | --- | --- |
910 | GitHub release binary (including Termux) | `codewhale update` |
911 | npm (`node_modules` on the path) | `npm install -g codewhale@latest` |
912 | Homebrew (`Cellar` / `linuxbrew` prefix) | `brew upgrade codewhale` |
913 | `cargo install` (`~/.cargo/bin`) | `cargo install codewhale-cli --locked --force` |
914
915 For package-managed installs the notice also warns against `codewhale update`:
916 replacing a binary Homebrew or npm owns leaves the manager describing a version
917 that is no longer on disk, and the next upgrade silently reverts you.
918
919 Override the detection with `CODEWHALE_INSTALL_METHOD=npm|homebrew|cargo|binary`
920 if you relocated the binary somewhere the path heuristics cannot read.
921
922 To redirect the startup check, set `update_uri` to an internal endpoint that
923 returns GitHub-compatible latest-release JSON. Minimal mirror metadata with a
924 `tag_name` field is accepted; if `assets` are present, Codewhale requires the
925 same uploaded asset set as the official release before showing the toast.
926
927 ```toml
928 [update]
929 check_for_updates = true
930 update_uri = "https://internal.mirror.example/codewhale/releases/latest"
931 ```
932
933 When `update_uri` is not set, startup checks honor release mirror environment
934 variables such as `CODEWHALE_RELEASE_BASE_URL` before falling back to the
935 official GitHub API endpoint. If a configured `update_uri` cannot be fetched or
936 parsed and a release mirror env var is set, the TUI falls back to that mirror
937 instead of failing startup.
938
939 ## Workshop output budgets
940
941 By default an oversized tool result uses bounded spillover: a result
942 under the byte threshold stays inline, a larger one gets a head/tail
943 preview plus a session artifact the model can read back. There is no
944 synthesis sub-agent and no per-call `raw = true` escape.
945
946 `[workshop] large_output_threshold_tokens` and
947 `[workshop.per_tool_thresholds]` (exact model-visible tool names such as
948 `bash`) only take effect when the process opts in to adaptive evidence
949 routing with `CODEWHALE_ADAPTIVE_OUTPUT_ROUTING=1`; they set where a result
950 stops being inline and becomes handle-only evidence.
951
952 Two optional byte ceilings (#5367) apply either way. They raise the
953 model-visible floor and never lower it:
954
955 - `read_result_max_bytes` — cap for a single `read` / `read_file`
956 result. Absent keeps the compile-time defaults (100000 bytes for
957 `read`, which has no line cap; 16KiB / 500 lines for `read_file`).
958 For `read` this is the middle of a three-layer budget: the model's
959 own per-call `max_bytes` (hard maximum 500000) raises the budget for
960 one call, this setting raises the floor for the whole process, and
961 either way 2MiB is the absolute ceiling. Highest wins; neither layer
962 can lower a budget the other granted.
963 - `tool_result_max_bytes` — cap for a generic tool result after
964 spillover. Absent keeps the 12K-character compact floor (48K on
965 windows ≥500K tokens). Hard cap is 2MiB.
966
967 ## Context length (context window)
968
969 Also called context size, context limit, max context, or window. This is the
970 total token window Codewhale budgets against, and it drives the header/footer
971 context percent, the auto-compaction trigger, context-pressure checks, and the
972 request output cap. If Codewhale compacts at 128K on a model you know serves a
973 1M window, this is the setting to change (#5134).
974
975 **See what is in effect, and where the value came from.** Every one of these
976 prints the resolved window *and* its source:
977
978 - `/status` — a `Context window:` row with the percent and token counts, and a
979 `Window source:` row naming the provenance and the exact key that overrides
980 it.
981 - `/config` → Provider — `Context window` (your override, or `(not set)`) and
982 `Effective context window` (`1048576 tokens · configured`). Typing
983 `context length` in the `/config` filter jumps straight to them.
984 - `/context report` — `Window: 1048576 tokens (12.4% used, ...; source: configured)`.
985 - `/context json` — machine-readable `context_window_tokens` and
986 `context_window_source`.
987
988 **Change it** with the provider-table key `context_window`:
989
990 ```toml
991 [providers.moonshot]
992 context_window = 1048576
993 ```
994
995 or from the CLI:
996
997 ```bash
998 codewhale config set providers.moonshot.context_window 1048576
999 codewhale config unset providers.moonshot.context_window # back to automatic
1000 ```
1001
1002 Use the table for the provider you are actually on (`providers.openai`,
1003 `providers.deepseek`, `providers.moonshot`, …); `/status` names it for you. The
1004 value is a positive token count for the route's *total* window.
1005
1006 When one gateway fronts models with heterogeneous windows, scope the override
1007 to an exact wire model id with `[providers.<name>.model_context_windows]`:
1008
1009 ```toml
1010 [providers.command_code]
1011 context_window = 204800
1012
1013 [providers.command_code.model_context_windows]
1014 "MiniMaxAI/MiniMax-M2.5" = 204800
1015 "google/gemini-3.1-flash-lite" = 1000000
1016 ```
1017
1018 Keys are the exact wire model ids the route sends (dotted and `org/model`
1019 spellings both work as TOML keys when quoted); each value must be a positive
1020 token count. A matching entry beats the provider-level `context_window` for
1021 that model only — every other model on the provider still resolves against
1022 `context_window` and the rungs below. From the CLI:
1023
1024 ```bash
1025 codewhale config set 'providers.command_code.model_context_windows."MiniMaxAI/MiniMax-M2.5"' 204800
1026 codewhale config unset 'providers.command_code.model_context_windows."MiniMaxAI/MiniMax-M2.5"'
1027 ```
1028
1029 ### How the effective window is resolved
1030
1031 First match wins, and the source label each surface prints is exactly this
1032 rung:
1033
1034 1. `configured (per-model)` — a `[providers.<name>.model_context_windows]`
1035 entry keyed by the route's exact wire model id. A hard override for that
1036 model only; it never rewrites another model's window.
1037 2. `configured` — `[providers.<name>] context_window` in `config.toml`. A hard
1038 override for every model on the provider: nothing below it can raise or
1039 lower the result. Read-time aliases:
1040 `contextWindow`, `context_window_tokens`, `contextWindowTokens`,
1041 `context_length`, `contextLength`.
1042 3. `provider-reported` — route-scoped 1M metadata a provider actually reported
1043 for the Kimi Code `k3` route, when it was observed within the last 24 hours.
1044 4. `static Kimi Code safe floor` — 262,144 tokens for Kimi Code memberships,
1045 because 1M access is plan-gated (Allegretto and above).
1046 5. `catalog` — the bundled route catalog (hand-curated offerings first, then
1047 the bundled Models.dev rows). The official `openai-codex` roster supplies
1048 account-specific model choices, not context-window metadata. Refresh it
1049 with `codewhale models --update --provider openai-codex`; Codewhale does
1050 not read `$CODEX_HOME` for this route.
1051 6. `model-name hint` — an `_Nk` suffix parsed from the model name itself
1052 (`qwen3-32b-256k` → 256,000), vendor-agnostic. A naming convention the
1053 serving engine may not honor is not a fact about the route, so this rung
1054 sits *below* the catalog: any catalog row for the same id beats it (#5441).
1055 7. `fallback` — the static per-provider capability table: 200,000 for
1056 Anthropic-wire routes, 128,000 for `openai-codex`, 8,192 for Ollama,
1057 otherwise Codewhale's static per-model metadata, and finally 128,000 when
1058 the model is unknown.
1059
1060 ### What "(unverified)" means
1061
1062 The `model-name hint` and `fallback` rungs still drive real budgets — the
1063 compaction trigger, the context meter, and the output reservation all use the
1064 number — but they are guesses, not capabilities anyone checked. Every surface
1065 that renders one of these windows appends `(unverified)` to its source label
1066 (the status line, the context-pressure message, `/status`, `/config`, and the
1067 model picker chip), so a window you did not configure and no provider reported
1068 can never read as a verified limit (#5239, #5441). The `context_window` and
1069 `model_context_windows` provider-table keys above are the fix: a configured
1070 window is a hard override and renders as `configured` (or
1071 `configured (per-model)`) with no marker.
1072
1073 Output ceilings follow the same rule (#5440): an Anthropic-family model the
1074 catalog does not describe keeps the 64K Messages floor as its clamp, labeled
1075 `unverified` (or an "assumed floor") instead of `documented`. The official
1076 ChatGPT plan preview does not accept `max_output_tokens`; Codewhale sends no
1077 output-token cap on that route. A local request budget is not an upstream
1078 limit or a guarantee about the completed response length.
1079
1080 There is no environment variable for the context window; the provider-table
1081 `context_window` and per-model `model_context_windows` keys are the user
1082 knobs. They are the right ones to set when a gateway or self-hosted runtime
1083 serves a window Codewhale's catalog does not model — per-model when only some
1084 of a provider's routes differ, provider-wide when they all do. Codewhale will
1085 not invent a window it cannot justify — it falls back to a conservative value,
1086 labels it `fallback`, and marks it `(unverified)` at every surface that shows
1087 it.
1088
1089 ### Adjacent knobs
1090
1091 - `auto_compact_threshold_percent` (settings.toml; also accepted as
1092 `auto_compact_threshold`; `10`–`100`, default `80`): the share of the full
1093 route context window at which auto-compaction fires, clamped so it can never
1094 cross the spendable input ceiling after output reservation and headroom.
1095 Editable from `/config`. Raising the window without touching this raises the
1096 absolute compaction point along with it.
1097 - `auto_compact` (settings.toml, on/off): turns automatic compaction off
1098 entirely; `/compact` and Ctrl+L stay available.
1099 - `[compaction] summary_instructions` and
1100 `[compaction] retained_user_message_tokens` (config.toml): standing
1101 summarizer instructions and the verbatim user-message retention budget. See
1102 the `compaction.*` entry in the config-key reference below.
1103 - `CODEWHALE_MAX_OUTPUT_TOKENS` (environment variable; legacy alias
1104 `DEEPSEEK_MAX_OUTPUT_TOKENS`): overrides the requested output cap. Without an
1105 override, Codewhale starts at the safe `65536` request cap and intersects it
1106 with any smaller documented model or route ceiling; a catalog `max_output`
1107 such as DeepSeek V4's 384K remains a capability ceiling, not the amount every
1108 response requests. Explicit overrides are preserved within the resolved
1109 route context window and any route output ceiling, and preflight/emergency
1110 budgeting reserves the same effective value that can reach the wire. A
1111 separately documented route input ceiling also clamps preflight and
1112 compaction even when the total context window is larger. A blank canonical
1113 variable falls through to a nonblank legacy value; a nonblank invalid or
1114 zero canonical value is authoritative and falls back to the safe automatic
1115 default instead of activating a stale legacy setting. There is no
1116 `max_output_tokens` key in `config.toml`.
1117
1118 Before compaction replaces conversation history, Codewhale durably saves the
1119 original messages to the session's `artifacts/context-transfer-<id>.json` and,
1120 when a model summary is produced, its handoff to the matching `.md` file.
1121 These use the existing session artifact store and persistence redaction.
1122 A failed write aborts compaction without replacing context. Pruning-only passes
1123 save the original messages without making an extra model call. Pressure metadata
1124 shows estimated input tokens and the configured trigger; it is an estimate, not
1125 an exact promise about a provider's remaining context.
1126
1127 Compaction history is available through `codewhale metrics` (or `--json`) and
1128 `audit.log` in the Codewhale home. Completed passes record their trigger,
1129 summary/pruning path, message and estimated-token counts, effective threshold,
1130 and summarizer token usage. Automatic refusals are recorded once per turn with
1131 their reason. These are local diagnostics, not provider invoice totals; earlier
1132 artifacts are not retroactively counted. A text-mode `exec` that attempts
1133 compaction saves its owning session at turn completion so its recovery artifacts
1134 remain discoverable. A killed process may leave artifacts without that final
1135 session snapshot; the audit writer reports I/O failures instead of inventing data.
1136
1137 See [Settings File](#settings-file-persistent-ui-preferences) for the
1138 compaction settings and [Token Quantities and
1139 Drivers](#token-quantities-and-drivers) for what each displayed token number
1140 actually measures.
1141
1142 ## Profiles
1143
1144 You can define multiple profiles in the same file:
1145
1146 ```toml
1147 default_text_model = "deepseek-flash"
1148
1149 [providers.deepseek]
1150 api_key = "PERSONAL_KEY"
1151
1152 [profiles.work.providers.deepseek]
1153 api_key = "WORK_KEY"
1154 base_url = "https://api.deepseek.com/beta"
1155
1156 [profiles.nvidia-nim]
1157 provider = "nvidia-nim"
1158 default_text_model = "deepseek-ai/deepseek-v4-pro"
1159
1160 [profiles.nvidia-nim.providers.nvidia_nim]
1161 api_key = "NVIDIA_KEY"
1162 base_url = "https://integrate.api.nvidia.com/v1"
1163
1164 [profiles.fireworks]
1165 provider = "fireworks"
1166 default_text_model = "accounts/fireworks/models/deepseek-v4-pro"
1167
1168 [profiles.siliconflow]
1169 provider = "siliconflow"
1170 default_text_model = "deepseek-ai/DeepSeek-V4-Pro"
1171
1172 [profiles.siliconflow.providers.siliconflow]
1173 base_url = "https://api.siliconflow.com/v1"
1174
1175 [profiles.openai-compatible]
1176 provider = "openai"
1177
1178 [profiles.openai-compatible.providers.openai]
1179 base_url = "https://openai-compatible.example/v4"
1180 model = "glm-5"
1181
1182 [profiles.atlascloud]
1183 provider = "atlascloud"
1184
1185 [profiles.atlascloud.providers.atlascloud]
1186 base_url = "https://api.atlascloud.ai/v1"
1187 model = "deepseek-ai/deepseek-v4-flash"
1188
1189 [profiles.sglang]
1190 provider = "sglang"
1191
1192 [profiles.sglang.providers.sglang]
1193 base_url = "http://localhost:30000/v1"
1194 model = "deepseek-ai/DeepSeek-V4-Pro"
1195
1196 [profiles.vllm]
1197 provider = "vllm"
1198
1199 [profiles.vllm.providers.vllm]
1200 base_url = "http://localhost:8000/v1"
1201 model = "deepseek-ai/DeepSeek-V4-Pro"
1202
1203 [profiles.ollama]
1204 provider = "ollama"
1205
1206 [profiles.ollama.providers.ollama]
1207 base_url = "http://localhost:11434/v1"
1208 model = "qwen2.5-coder:7b" # any tag `ollama list` shows
1209
1210 [profiles.ollama-cloud]
1211 provider = "ollama-cloud"
1212
1213 [profiles.ollama-cloud.providers.ollama_cloud]
1214 base_url = "https://ollama.com/v1"
1215 model = "gpt-oss:120b"
1216 ```
1217
1218 Select a profile with:
1219
1220 - CLI: `codewhale --profile work`
1221 - Env: `DEEPSEEK_PROFILE=work`
1222
1223 If a profile is selected but missing, codewhale exits with an error listing available profiles.
1224
1225 ## Legacy top-level `base_url` and `api_key`
1226
1227 Older releases kept DeepSeek's endpoint and key at the top of `config.toml`,
1228 and each reader decided for itself which other routes inherited them. They
1229 now live in provider tables only. An older file keeps working unchanged:
1230 every load reads the top-level keys as if they were already in their table,
1231 by one rule, per file and per `[profiles.<name>]`:
1232
1233 1. `provider = "custom"` with no `[providers.custom]` table: the endpoint, key
1234 and a copy of the model become `[providers.custom]`.
1235 2. An endpoint on another vendor's official host (for example
1236 `integrate.api.nvidia.com`, `xiaomimimo.com`, `openrouter.ai`, or the
1237 ChatGPT Codex endpoint) belongs to that vendor's table, so a DeepSeek route
1238 never sends requests there.
1239 3. Anything else belongs to `[providers.deepseek]`, which DeepSeek-CN also
1240 reads for its endpoint and key.
1241
1242 The key goes with DeepSeek (or the literal custom route). It follows the
1243 endpoint to another vendor only when the same table sets `provider` to that
1244 vendor, the host is that vendor's own, and the vendor's table has no key. When
1245 no `provider` is set, a NIM host still selects `nvidia-nim` and
1246 `api.deepseeki.com` still selects `deepseek-cn`. A `[vision_model]` without a
1247 key of its own keeps using the top-level key. A profile's top-level
1248 `base_url` now overrides the base file's `[providers.deepseek] base_url`
1249 (before, the base table silently won).
1250
1251 When a top-level value and its table disagree, the table's `base_url` and the
1252 top-level `api_key` are used, which is what the runtime sent before.
1253
1254 Loading never rewrites the file. The next save Codewhale makes (for example
1255 `codewhale config set`, `/config ... --save`, or `auth set`) moves the keys,
1256 keeps comments, writes a one-time credential-free copy of the old file to
1257 `config.toml.pre-migrate.bak`, and prints one line saying what moved. A
1258 disagreeing pair is never resolved in the file on its own; both values stay
1259 until you choose, and `codewhale config doctor` / `codewhale doctor` report
1260 it (sources only, never values). `codewhale config get|set|unset base_url`
1261 (and `api_key`) address the active provider's table.
1262
1263 To tidy the file yourself:
1264
1265 ```bash
1266 codewhale config migrate --dry-run # show what would move
1267 codewhale config migrate # move it (backup first)
1268 codewhale config migrate --prefer top-level # resolve a conflict: keep the top-level value
1269 codewhale config migrate --prefer table # resolve a conflict: keep the table value
1270 ```
1271
1272 ## Environment Variables
1273
1274 Most runtime environment variables override config values. API-key variables are
1275 fallbacks after saved config and keyring credentials.
1276
1277 The three user-facing slots — provider, model, base URL — expose `CODEWHALE_*`
1278 aliases. When both forms are set the `CODEWHALE_*` value wins; the
1279 `DEEPSEEK_*` form is kept for older shells:
1280
1281 - `CODEWHALE_PROVIDER` (preferred) / `DEEPSEEK_PROVIDER` (legacy alias) —
1282 `deepseek|deepseek-anthropic|nvidia-nim|openai|atlascloud|wanjie-ark|volcengine|openrouter|xiaomi-mimo|novita|fireworks|siliconflow|arcee|siliconflow-CN|moonshot|sglang|vllm|ollama|ollama-cloud|huggingface|modelscope|together|qianfan|openai-codex|anthropic|openmodel|zai|stepfun|minimax|deepinfra|mistral`
1283 - `CODEWHALE_MODEL` (preferred) / `DEEPSEEK_MODEL` (legacy alias) — default model for the active provider
1284 - `CODEWHALE_BASE_URL` (preferred) / `DEEPSEEK_BASE_URL` (legacy alias) — base URL for the active provider
1285
1286 `CODEWHALE_BASE_URL` applies to the **active** route only. A request pinned to
1287 another provider — a subagent or fleet child, a routed tool, the per-turn
1288 auto-router, a picker preview — resolves its endpoint from that provider's own
1289 `[providers.<table>]`, then its provider-scoped variable (`MOONSHOT_BASE_URL`,
1290 `OPENAI_BASE_URL`, …), then that provider's default. It never inherits the
1291 active session's host, and a custom route with no configured `base_url` fails
1292 closed on a loopback placeholder rather than borrowing another provider's
1293 endpoint. The environment writes the value into the active identity's own
1294 table: DeepSeek-CN falls back to `[providers.deepseek]` for its endpoint and
1295 key, but never to a DeepSeek value the environment addressed to DeepSeek alone.
1296 A managed-config overlay that supplies or reselects the effective
1297 route's endpoint takes the generic override away from every route.
1298
1299 Remaining variables:
1300
1301 - `DEEPSEEK_API_KEY`
1302 - `DEEPSEEK_ANTHROPIC_BASE_URL`
1303 - `DEEPSEEK_HTTP_HEADERS` (custom model request headers, comma-separated `name=value` pairs)
1304 - `DEEPSEEK_DEFAULT_TEXT_MODEL` (extra legacy alias of `DEEPSEEK_MODEL`)
1305 - `DEEPSEEK_STREAM_IDLE_TIMEOUT_SECS` (stream idle timeout in seconds; default `900`, clamped to `1..=3600`)
1306 - `DEEPSEEK_STREAM_OPEN_TIMEOUT_SECS` (connection setup + response-header wait in seconds; default `45`, clamped to `5..=300`; distinct from the per-chunk idle timeout; `stream.open_timeout_secs` (legacy `tui.stream_open_timeout_secs` fallback) wins when positive)
1307 - `CODEWHALE_CACHE_MAXIMAL` (`1`/`true`/`on`/`yes`) — cache-maximal context mode (#528). When on, the Repo Working Set block materializes the **full current contents** of the top active files into the system prompt each turn (deterministic order, byte-bounded), instead of only listing their paths. The block stays byte-stable while those files are unchanged so DeepSeek's KV prefix cache keeps hitting; editing a file cache-misses from its block onward. Off by default (path list only). Byte caps default to 24 KB per file / 96 KB total.
1308 - `NVIDIA_API_KEY` or `NVIDIA_NIM_API_KEY` (when provider is `nvidia-nim`)
1309 - `NVIDIA_NIM_BASE_URL`, `NIM_BASE_URL`, or `NVIDIA_BASE_URL`
1310 - `NVIDIA_NIM_MODEL`
1311 - `OPENAI_API_KEY`
1312 - `OPENAI_BASE_URL`
1313 - `OPENAI_MODEL`
1314 - `ATLASCLOUD_API_KEY`
1315 - `ATLASCLOUD_BASE_URL`
1316 - `ATLASCLOUD_MODEL`
1317 - `WANJIE_ARK_API_KEY`, `WANJIE_API_KEY`, or `WANJIE_MAAS_API_KEY`
1318 - `WANJIE_ARK_BASE_URL`, `WANJIE_BASE_URL`, or `WANJIE_MAAS_BASE_URL`
1319 - `WANJIE_ARK_MODEL`, `WANJIE_MODEL`, or `WANJIE_MAAS_MODEL`
1320 - `VOLCENGINE_API_KEY`, `VOLCENGINE_ARK_API_KEY`, or `ARK_API_KEY`
1321 - `VOLCENGINE_BASE_URL`, `VOLCENGINE_ARK_BASE_URL`, or `ARK_BASE_URL`
1322 - `VOLCENGINE_MODEL` or `VOLCENGINE_ARK_MODEL`
1323 - `OPENROUTER_API_KEY`
1324 - `OPENROUTER_BASE_URL`
1325 - `OPENROUTER_MODEL`
1326 - `XIAOMI_MIMO_TOKEN_PLAN_API_KEY`, `MIMO_TOKEN_PLAN_API_KEY`, `XIAOMI_MIMO_API_KEY`, `XIAOMI_API_KEY`, or `MIMO_API_KEY`
1327 - `XIAOMI_MIMO_BASE_URL` or `MIMO_BASE_URL`
1328 - `XIAOMI_MIMO_MODEL` or `MIMO_MODEL`
1329 - `XIAOMI_MIMO_MODE` or `MIMO_MODE` (`token-plan-sgp`, `token-plan-cn`,
1330 `token-plan-ams`, or `pay-as-you-go`)
1331 - `NOVITA_API_KEY`
1332 - `NOVITA_BASE_URL`
1333 - `NOVITA_MODEL`
1334 - `FIREWORKS_API_KEY`
1335 - `FIREWORKS_BASE_URL`
1336 - `FIREWORKS_MODEL`
1337 - `HUGGINGFACE_API_KEY` or `HF_TOKEN` (`HF_TOKEN` is a fallback alias accepted when provider is `huggingface`)
1338 - `MODELSCOPE_API_KEY`
1339 - `HUGGINGFACE_BASE_URL` or `HF_BASE_URL`
1340 - `HUGGINGFACE_MODEL` or `HF_MODEL`
1341 - `SILICONFLOW_API_KEY`
1342 - `SILICONFLOW_BASE_URL`
1343 - `SILICONFLOW_MODEL`
1344 - `ARCEE_API_KEY`
1345 - `ARCEE_BASE_URL`
1346 - `ARCEE_MODEL`
1347 - `TOGETHER_API_KEY`
1348 - `TOGETHER_BASE_URL`
1349 - `TOGETHER_MODEL`
1350 - `QIANFAN_API_KEY` or `BAIDU_QIANFAN_API_KEY`
1351 - `QIANFAN_BASE_URL` or `BAIDU_QIANFAN_BASE_URL`
1352 - `QIANFAN_MODEL` or `BAIDU_QIANFAN_MODEL`
1353 - `OPENAI_CODEX_ACCESS_TOKEN` or `CODEX_ACCESS_TOKEN` (legacy process tokens; ignored by the official ChatGPT plan route)
1354 - `OPENAI_CODEX_BASE_URL` or `CODEX_BASE_URL`
1355 - `OPENAI_CODEX_MODEL` or `CODEX_MODEL`
1356 - `OPENAI_CODEX_ACCOUNT_ID` or `CODEX_ACCOUNT_ID` (legacy identity hints; ignored by the official ChatGPT plan route)
1357 - `CODEWHALE_CHATGPT_NEW_ACCOUNT=1` explicitly replaces the single selected ChatGPT registration after a validated browser sign-in
1358 - `ANTHROPIC_API_KEY`
1359 - `ANTHROPIC_BASE_URL`
1360 - `ANTHROPIC_MODEL`
1361 - `ZAI_API_KEY` or `Z_AI_API_KEY`
1362 - `ZAI_BASE_URL` or `Z_AI_BASE_URL`
1363 - `ZAI_MODEL` or `Z_AI_MODEL`
1364 - `STEPFUN_API_KEY` or `STEP_API_KEY`
1365 - `STEPFUN_BASE_URL` or `STEP_BASE_URL`
1366 - `STEPFUN_MODEL` or `STEP_MODEL`
1367 - `MINIMAX_API_KEY`
1368 - `MINIMAX_BASE_URL`
1369 - `MINIMAX_MODEL`
1370 - `DEEPINFRA_API_KEY` or `DEEPINFRA_TOKEN`
1371 - `DEEPINFRA_BASE_URL`
1372 - `DEEPINFRA_MODEL`
1373 - `MISTRAL_API_KEY`
1374 - `MISTRAL_BASE_URL`
1375 - `MISTRAL_MODEL`
1376 - `MOONSHOT_API_KEY` or `KIMI_API_KEY`
1377 - `MOONSHOT_BASE_URL` or `KIMI_BASE_URL`
1378 - `MOONSHOT_MODEL`, `KIMI_MODEL_NAME`, or `KIMI_MODEL`
1379 - `SGLANG_BASE_URL`
1380 - `SGLANG_MODEL`
1381 - `SGLANG_API_KEY` (optional; many localhost SGLang servers do not require auth)
1382 - `VLLM_BASE_URL`
1383 - `VLLM_MODEL`
1384 - `VLLM_API_KEY` (optional; many localhost vLLM servers do not require auth)
1385 - `OLLAMA_BASE_URL`
1386 - `OLLAMA_MODEL`
1387 - `OLLAMA_API_KEY` (optional; many localhost Ollama servers do not require auth)
1388 - `OLLAMA_CLOUD_BASE_URL`
1389 - `OLLAMA_CLOUD_MODEL`
1390 - `OLLAMA_CLOUD_API_KEY` (preferred Cloud key; `OLLAMA_API_KEY` is the official fallback)
1391 For every product-level `CODEWHALE_*` variable below, the matching legacy
1392 `DEEPSEEK_*` name is still read as a compatibility fallback; when both are set,
1393 the `CODEWHALE_*` value wins.
1394
1395 - `CODEWHALE_LOG_LEVEL` or `RUST_LOG` (`info`/`debug`/`trace` enables lightweight verbose logs)
1396 - `CODEWHALE_SKILLS_DIR`
1397 - `CODEWHALE_MCP_CONFIG`
1398 - `CODEWHALE_NOTES_PATH`
1399 - `CODEWHALE_MEMORY` (`1|on|true|yes|y|enabled` turns user memory on)
1400 - `CODEWHALE_MEMORY_PATH`
1401 - `CODEWHALE_TELEMETRY` / `DEEPSEEK_TELEMETRY` (legacy alias) — anonymous usage
1402 counting is on by default in the current 0.9.12 source, with a disclosure
1403 naming Codewhale and PostHog and an easy durable opt-out. Prior explicit
1404 declines remain off. Accepts `0|1|true|false|yes|no|on|off|enabled|
1405 disabled`. An explicit "off" is a **floor**: it beats `--telemetry true` and
1406 `telemetry = true` in config, and a value this list cannot read also resolves
1407 to off, because a typo in a kill switch must never resolve to "on". See
1408 [`TELEMETRY.md`](TELEMETRY.md).
1409 - `CODEWHALE_TELEMETRY_ENDPOINT` / `DEEPSEEK_TELEMETRY_ENDPOINT` (legacy alias)
1410 — `https://`, or plain `http://` only for loopback. Overrides the config file.
1411 Unset selects the shipped default,
1412 `https://telemetry.codewhale.net/v1/telemetry`; setting it to the **empty
1413 string** routes batches to a local dry-run file and contacts nobody. Either
1414 way it only decides where a session sends — it cannot override an opt-out.
1415 - `CODEWHALE_ALLOW_SHELL` (`1`/`true` enables)
1416 - `CODEWHALE_APPROVAL_POLICY` (`on-request|untrusted|never`)
1417 - `CODEWHALE_SANDBOX_MODE` (`read-only|workspace-write|danger-full-access|external-sandbox`)
1418 - `CODEWHALE_NO_NEW_PRIVS` (`0`/`false`/`no`/`off`/`disabled` opts out) — Linux only. The
1419 TUI process sets the kernel's irreversible no-new-privileges flag at startup
1420 as defense-in-depth, which blocks `sudo`/`su`/setuid helpers for Codewhale's
1421 whole process tree. The flag is already skipped when the startup sandbox
1422 mode resolves to `danger-full-access` (#5723), so this variable is the
1423 explicit override in the remaining cases: set it to a falsey value before
1424 launching if you administer through Codewhale as a wheel-group user under a
1425 narrower posture and need escalation to work (#5413), or set a truthy value
1426 to force the flag on even under `danger-full-access`. Unset, the posture
1427 decides; the other startup hardening (no ptrace, no core dumps) always
1428 stays on.
1429 - `CODEWHALE_MANAGED_CONFIG_PATH`
1430 - `CODEWHALE_REQUIREMENTS_PATH`
1431 - `CODEWHALE_MAX_SUBAGENTS` (clamped to `1..=128`)
1432 - `CODEWHALE_TASKS_DIR` (runtime task queue/artifact storage, default
1433 `~/.codewhale/tasks`, with legacy `~/.deepseek/tasks` fallback when only the
1434 legacy directory exists)
1435 - `CODEWHALE_RUNTIME_DIR` (override the runtime thread store root). Interactive
1436 sessions default to `$CODEWHALE_HOME/sessions/<session-id>/runtime` so each
1437 Codewhale process owns its own store (#5630). The store is single-owner: a
1438 second process on the **same** root fails at startup. Set this variable to
1439 share one store across processes, or when the runtime API server should use a
1440 stable non-session path. Unset, the API/server path remains
1441 `$CODEWHALE_HOME/tasks/runtime`. Legacy alias: `DEEPSEEK_RUNTIME_DIR`.
1442 - `CODEWHALE_ALLOW_INSECURE_HTTP` (`1`/`true` allows non-local `http://` base URLs; default is reject)
1443 - `CODEWHALE_FORCE_HTTP1` (`1|true|yes|on` pins the HTTP client to HTTP/1.1, disabling HTTP/2; useful on Windows or behind proxies that mishandle long-lived H2 streams)
1444 - `CODEWHALE_HOME` (override the base data directory; defaults to `~/.codewhale`).
1445 If you previously exported `DEEPSEEK_HOME`, rename it to `CODEWHALE_HOME`;
1446 the old env var is not used for new Codewhale state paths.
1447 - `CODEWHALE_RELEASE_BASE_URL` (release asset mirror used by `codewhale update`
1448 and by TUI startup update checks when `[update].update_uri` is not set, or as
1449 a fallback when that configured URI cannot be fetched)
1450 - `CODEWHALE_AUTOMATIONS_DIR` (override the automations storage directory; uses
1451 `~/.codewhale/automations` by default, with legacy `~/.deepseek/automations`
1452 fallback when only the legacy directory exists)
1453 - `NO_ANIMATIONS` (`1|true|yes|on` forces `low_motion = true` and
1454 `fancy_animations = false` at startup, regardless of the saved
1455 settings; see [`docs/ACCESSIBILITY.md`](./ACCESSIBILITY.md)).
1456 - `SSL_CERT_FILE` — corporate-proxy / TLS-inspecting MITM users
1457 point this at a PEM bundle (or single DER cert) and the cert(s)
1458 get added alongside the platform's system trust store. Failures
1459 log a warning and continue — the existing system roots still
1460 apply.
1461
1462 ### Instruction sources (`instructions = [...]`, #454)
1463
1464 Add a list of additional system-prompt sources that get
1465 concatenated, in declared order, alongside the auto-loaded
1466 `AGENTS.md`:
1467
1468 ```toml
1469 instructions = [
1470 "./AGENTS.md",
1471 "~/.codewhale/global.md",
1472 "~/team/agents-shared.md",
1473 ]
1474 ```
1475
1476 Rules:
1477
1478 - Paths run through `expand_path` so `~` and env vars work.
1479 - Each file is capped at 100 KiB; oversized files are
1480 truncated with a `[…elided]` marker rather than skipped.
1481 - Missing files are skipped with a tracing warning so a stale
1482 entry doesn't fail the launch.
1483 - Only user-owned config, profiles, and managed config may set this array.
1484 Project config (`<workspace>/.codewhale/config.toml`, or legacy
1485 `<workspace>/.deepseek/config.toml`) ignores `instructions` so a cloned repo
1486 cannot choose arbitrary local files to place into the prompt.
1487
1488 ### Hooks
1489
1490 Hooks are a **TUI runtime feature**. They fire from the interactive TUI and
1491 the engine turn loop it drives; `codewhale exec`, the CLI subcommands, the
1492 app-server / ACP surfaces, and the `workflow` tool do not fire them.
1493
1494 [`docs/HOOKS.md`](HOOKS.md) is the authoritative reference for all eleven hook
1495 events — their firing points, environment variables, stdin payloads, timeout
1496 and background semantics, and which three of them can steer Codewhale. The
1497 sections below cover the configuration surface and the steering contracts in
1498 more depth.
1499
1500 Two contract points worth reading there before writing a hook:
1501
1502 - `background = true` means **submitted and never awaited**. The hook still
1503 gets the documented stdin payload and the same timeout, but it has no exit
1504 code and cannot steer.
1505 - A condition that references context its event never carries (an `exit_code`
1506 condition outside `tool_call_after` / `on_error`, a `mode` condition on
1507 `shell_env`, a tool condition on a non-tool event) is **rejected at load**,
1508 logged, and shown in `/hooks list`. It does not silently never match.
1509 Rejection is per entry, so a broken hook never drops another one that merely
1510 shares its `name` or is likewise unnamed.
1511
1512 ### `/hooks` listing
1513
1514 Run `/hooks` (or `/hooks list`) inside the TUI to see every
1515 configured lifecycle hook grouped by event, including each
1516 hook's name, command preview, effective timeout, and condition. When
1517 `[hooks].default_timeout_secs` is set it replaces every per-hook
1518 `timeout_secs`, and the listing shows that effective value and names the
1519 override rather than echoing the per-hook number. A
1520 `default_timeout_secs = 0` is rejected at load — it would expire every hook
1521 in the config immediately — so the override is ignored, per-hook
1522 `timeout_secs` applies, the listing shows that per-hook value with no
1523 override provenance, and the rejection appears under `configuration
1524 problems`. The
1525 `[hooks].enabled` flag's state is shown at the top so it's
1526 obvious when hooks are globally suppressed, and any entry rejected
1527 at load is listed under `configuration problems` with the reason.
1528 Hooks are configured under `[[hooks.hooks]]` entries — see
1529 [`docs/HOOKS.md`](HOOKS.md) for the full schema.
1530
1531 ### Mutable `message_submit` hooks
1532
1533 `message_submit` hooks run before a submitted message is added to
1534 history or sent to the model. Unlike observer-only lifecycle hooks,
1535 non-background `message_submit` hooks can replace or block the
1536 submitted text.
1537
1538 ```toml
1539 [[hooks.hooks]]
1540 event = "message_submit"
1541 command = "~/.codewhale/hooks/inject-context.sh"
1542 timeout_secs = 2
1543 continue_on_error = true
1544 ```
1545
1546 The hook receives JSON on stdin:
1547
1548 ```json
1549 {
1550 "event": "message_submit",
1551 "text": "original user text",
1552 "text_bytes": 18,
1553 "text_original_bytes": 18,
1554 "text_truncated": false,
1555 "session_id": "sess_12345678",
1556 "workspace": "/path/to/workspace",
1557 "mode": "agent",
1558 "model": "deepseek-chat",
1559 "total_tokens": 1234
1560 }
1561 ```
1562
1563 The entire serialized document is capped at 32 KiB. Codewhale retains the
1564 largest UTF-8-safe `text` prefix that fits after JSON escaping and bounded
1565 metadata, and the three `text_*` fields make truncation explicit. Immediate
1566 messages, restored queue entries, merged steers, and prior-hook replacements
1567 all cross this same serialization boundary.
1568
1569 If the hook exits `0` and prints JSON with a non-empty string `text` field,
1570 that value replaces the submitted text:
1571
1572 ```json
1573 { "text": "replacement user text" }
1574 ```
1575
1576 Exit `0` with empty stdout, or stdout JSON without `text`, leaves
1577 the current text unchanged. A JSON `text` field must not be empty;
1578 `{"text":""}` is treated as invalid stdout and ignored. Exit `2`
1579 blocks the submission before the turn starts; a structured `reason` field can
1580 provide the bounded, redacted status message shown in the TUI. Raw stdout,
1581 stderr, and process-error text are not copied into denial receipts.
1582 Other non-zero exits follow the hook's `continue_on_error` setting.
1583 Timeouts and spawn failures are also surfaced as transient TUI status
1584 messages when `continue_on_error = true` lets submission continue.
1585
1586 Multiple `message_submit` hooks run in config order, and each hook
1587 receives the text produced by the previous hook. Hooks marked
1588 `background = true` are observer-only and cannot transform or block
1589 the message — they still receive the same stdin payload and the same
1590 environment, they are simply never awaited. Existing environment
1591 variables remain available.
1592 `shell_env` hooks keep their existing `KEY=VALUE` stdout contract;
1593 JSON stdout contracts exist for `message_submit` (above) and
1594 `tool_call_before` (below).
1595
1596 ### `tool_call_before` decision hooks
1597
1598 `tool_call_before` hooks run before each tool call executes. In
1599 addition to the legacy hard deny (exit code `2`, which always wins
1600 regardless of stdout), a foreground hook may print a JSON decision on
1601 stdout with exit code `0`:
1602
1603 ```json
1604 {
1605 "decision": "allow" | "deny" | "ask",
1606 "reason": "human-readable explanation (used for deny)",
1607 "updatedInput": { "command": "ls -la" },
1608 "additionalContext": "text appended to the tool result for the model"
1609 }
1610 ```
1611
1612 All fields are optional. Empty stdout, non-JSON stdout, and JSON
1613 without a `decision` field behave exactly as before (allow). An
1614 unrecognized `decision` string logs a fixed warning without echoing the
1615 untrusted value and is treated as allow.
1616
1617 - `deny` blocks the tool; the model receives a permission-denied tool
1618 result containing `reason`.
1619 - `ask` forces the interactive approval prompt in Ask and Auto-Review even for
1620 tools that would otherwise auto-run. Full Access does not open tool-approval
1621 prompts, so hook `ask` does not downgrade that posture.
1622 - `updatedInput` must be a JSON object; it replaces the tool input
1623 before execution. When several hooks supply it, the last hook wins.
1624 - `additionalContext` is appended to the tool result sent back to the
1625 model as `[hook context] ...`. Multiple hooks' contexts are
1626 concatenated.
1627
1628 When multiple hooks match, precedence is deny > ask > allow. Hooks
1629 marked `background = true` cannot steer tool calls — they are
1630 submitted and never awaited, so they have no verdict to contribute.
1631
1632 A foreground hook that produced no verdict at all — it hit its timeout, the
1633 process could not be started, or a strict process exited non-zero without an
1634 explicit JSON decision — is not treated as permission. If
1635 *that* hook is configured with `continue_on_error = false`, the outcome
1636 denies the tool call and the denial names the hook and a bounded reason.
1637 Strictness is read off the hooks that actually matched this call, so a
1638 strict gate scoped to another tool cannot deny it, and a lenient hook's
1639 timeout does not deny merely because a strict hook exists elsewhere in
1640 config. Under the default `continue_on_error = true` the outcome is
1641 logged and the call proceeds.
1642
1643 `reason` and `additionalContext` are capped (2 000 characters per field,
1644 8 000 for the concatenated context of one call) and stripped of control
1645 characters before they reach the TUI or the model.
1646
1647 Example deny hook:
1648
1649 ```toml
1650 [[hooks.hooks]]
1651 event = "tool_call_before"
1652 command = '''echo '{"decision":"deny","reason":"blocked by project policy"}' '''
1653 condition = { type = "tool_name", name = "exec_shell" }
1654 ```
1655
1656 Example ask hook (force approval for every MCP tool):
1657
1658 ```toml
1659 [[hooks.hooks]]
1660 event = "tool_call_before"
1661 command = '''echo '{"decision":"ask"}' '''
1662 condition = { type = "tool_name", name = "mcp__*" }
1663 ```
1664
1665 Example input rewrite:
1666
1667 ```toml
1668 [[hooks.hooks]]
1669 event = "tool_call_before"
1670 command = "~/.codewhale/hooks/clamp-shell-timeout.sh"
1671 condition = { type = "tool_name", name = "exec_shell" }
1672 ```
1673
1674 where the script reads the hook context, then prints
1675 `{"updatedInput": {...}}` with the adjusted arguments.
1676
1677 `tool_name` conditions support `*` globs: `mcp__*` matches every MCP
1678 tool (e.g. `mcp__github__create_issue`) but not built-ins like
1679 `read_file`; exact names keep matching exactly. Other regex
1680 metacharacters in the pattern are matched literally.
1681
1682 ### Project-local hooks
1683
1684 Repositories can ship policy in `<workspace>/.codewhale/hooks.toml`,
1685 using the same shape as the `[hooks]` table (top-level fields plus
1686 `[[hooks]]` entries). Project hooks are executable shell
1687 configuration, so Codewhale only loads them after the workspace has
1688 been trusted in user-owned config through the trust prompt or a
1689 `[projects."<workspace>"] trust_level = "trusted"` entry. Session
1690 `/trust on` mode does not enable repo-supplied hooks by itself, and
1691 repo-local legacy markers such as `.deepseek/trusted` do not enable
1692 project hooks. Once trusted, project hooks are appended after global
1693 hooks from `config.toml`, so they run last and, for `updatedInput`,
1694 win ties. A malformed trusted project file logs a warning and startup
1695 falls back to global hooks only.
1696
1697 ```toml
1698 # .codewhale/hooks.toml
1699 [[hooks]]
1700 event = "tool_call_before"
1701 command = '''echo '{"decision":"deny","reason":"no shell in this repo"}' '''
1702 condition = { type = "tool_name", name = "exec_shell" }
1703 ```
1704
1705 ### Turn-end observer hooks
1706
1707 `turn_end` hooks observe the end of each model turn after post-turn
1708 state, usage totals, cost accounting, notifications, receipts, and
1709 queue recovery have been updated. They receive JSON on stdin and are
1710 observer-only: stdout is ignored, failures are logged as warnings, and
1711 the hook cannot block user input, mutate the transcript, or change the
1712 next queued follow-up.
1713
1714 Observer-only UI events share one 32-entry queue and two persistent workers;
1715 the terminal loop uses non-blocking submission and does not create a thread per
1716 event. A full queue or unavailable dispatcher drops that observer event and is
1717 kept as an event-specific error toast, independent of ordinary agent/turn
1718 status text.
1719
1720 ```toml
1721 [[hooks.hooks]]
1722 event = "turn_end"
1723 command = "~/.codewhale/hooks/turn-audit.sh"
1724 timeout_secs = 2
1725 continue_on_error = true
1726 ```
1727
1728 The payload includes common hook metadata plus post-turn accounting:
1729
1730 ```json
1731 {
1732 "event": "turn_end",
1733 "session_id": "sess_12345678",
1734 "workspace": "/path/to/workspace",
1735 "mode": "agent",
1736 "created_at": "2026-07-12T10:30:00+00:00",
1737 "model_backed": true,
1738 "provider": "deepseek",
1739 "model": "deepseek-chat",
1740 "billing_surface": null,
1741 "turn_id": "turn_12345678",
1742 "status": "completed",
1743 "error": null,
1744 "duration_ms": 1834,
1745 "usage": {
1746 "input_tokens": 1200,
1747 "output_tokens": 180,
1748 "prompt_cache_hit_tokens": 900,
1749 "prompt_cache_miss_tokens": 300,
1750 "prompt_cache_write_tokens": 0,
1751 "reasoning_tokens": null,
1752 "reasoning_replay_tokens": null
1753 },
1754 "totals": {
1755 "session_tokens": 1380,
1756 "conversation_tokens": 1380,
1757 "input_tokens": 1200,
1758 "output_tokens": 180
1759 },
1760 "tool_count": 2,
1761 "queued_message_count": 1,
1762 "stop_hook_active": false
1763 }
1764 ```
1765
1766 `created_at` anchors time-window pricing; `provider` and `model` identify the
1767 effective route used for model-backed turns. `billing_surface` is an optional,
1768 non-secret classification derived from the endpoint that actually served the
1769 turn. Recognized StepFun routes emit `stepfun-payg` or `stepfun-plan`; the raw
1770 base URL is never written to hook or runtime records. Runtime `TurnRecord`
1771 exports call the same field `effective_billing_surface`, which `scorecard`
1772 accepts as an alias. This keeps subscription quota separate from token-priced
1773 usage. Unrecognized and custom endpoints remain `null` and unpriced.
1774
1775 Shell-only lifecycle completions set `model_backed` to `false` and may report a
1776 `null` provider; offline scorecards exclude those records from model token and
1777 cost totals. Completion-only shell, manual-compaction, and purge events that do
1778 not have a matching `TurnStarted` retain the observer notification with a
1779 synthetic `lifecycle_<uuid>` turn id and the time the completion was observed.
1780
1781 For `interrupted` or `failed` turns, `status` reflects that terminal
1782 state and `error` carries the engine error string when one is available.
1783 `stop_hook_active` is reserved for future re-entry protection and is
1784 currently always `false`.
1785
1786 ### Sub-agent lifecycle hooks
1787
1788 `subagent_spawn` and `subagent_complete` hooks observe sub-agent lifecycle
1789 events. They receive bounded JSON metadata on stdin and are observer-only:
1790 hook failures are logged as warnings and do not block sub-agent scheduling,
1791 change prompts, or change results. For these observer events,
1792 `continue_on_error` has no effect: later matching hooks still run even when an
1793 earlier hook exits non-zero.
1794
1795 ```toml
1796 [[hooks.hooks]]
1797 event = "subagent_complete"
1798 command = "~/.codewhale/hooks/subagent-audit.sh"
1799 timeout_secs = 2
1800 continue_on_error = true
1801 ```
1802
1803 `subagent_spawn` receives:
1804
1805 ```json
1806 {
1807 "event": "subagent_spawn",
1808 "agent_id": "agent_12345678",
1809 "session_id": "sess_12345678",
1810 "workspace": "/path/to/workspace",
1811 "mode": "agent",
1812 "model": "deepseek-chat",
1813 "total_tokens": 1234,
1814 "prompt_preview": "bounded prompt preview",
1815 "prompt_truncated": false
1816 }
1817 ```
1818
1819 `subagent_complete` receives the same common fields plus terminal metadata:
1820
1821 ```json
1822 {
1823 "event": "subagent_complete",
1824 "agent_id": "agent_12345678",
1825 "session_id": "sess_12345678",
1826 "workspace": "/path/to/workspace",
1827 "mode": "agent",
1828 "model": "deepseek-chat",
1829 "total_tokens": 1234,
1830 "status": "completed",
1831 "result_preview": "bounded result preview",
1832 "result_truncated": false
1833 }
1834 ```
1835
1836 Previews are capped before delivery so lifecycle hooks do not receive full
1837 sub-agent prompts, transcripts, or unbounded results. Use the transcript handle
1838 returned by `agent` when full sub-agent details are needed.
1839
1840 ### Running-turn input
1841
1842 Composer shortcuts keep the same role throughout a session:
1843
1844 - **Enter** sends when idle and queues a next-turn follow-up while busy. The
1845 behavior does not change before versus after the provider's first token.
1846 - With an empty composer and queued follow-ups visible, **Enter** sends the
1847 oldest queued follow-up into the active turn now.
1848 - **Ctrl+Enter** (or **Cmd+Enter** when the terminal forwards it) explicitly
1849 steers the active turn. It sends normally when idle.
1850 - By default, **Shift+Enter**, **Alt+Enter**, and **Ctrl+J** insert a newline.
1851 - Set `composer_multiline_mode = true` to make **Enter** insert a newline and
1852 **Shift+Enter** send instead. **Alt+Enter**, **Ctrl+J**, and supported
1853 **Ctrl+Enter** / **Cmd+Enter** behavior is unchanged.
1854 - **Ctrl+G** and **Ctrl+S** only stash drafts; they never send or steer.
1855
1856 ### Composer stash (`/stash`, Ctrl+G / Ctrl+S)
1857
1858 Press **Ctrl+G** in the composer to park the current draft to
1859 `~/.codewhale/composer_stash.jsonl`. `/stash list` shows parked
1860 drafts with one-line previews and timestamps; `/stash pop`
1861 restores the most recently parked draft (LIFO); `/stash clear`
1862 wipes the file. Capped at 200 entries; multiline drafts round-trip intact.
1863 **Ctrl+S** remains an alias in terminals that forward it; Cursor and VS Code
1864 reserve Ctrl+S for Save, so Ctrl+G is the portable default.
1865
1866 ## Settings File (Persistent UI Preferences)
1867
1868 codewhale also stores user preferences in:
1869
1870 - `~/.codewhale/settings.toml` on new installs
1871 - `~/.deepseek/settings.toml` or the legacy platform config-dir
1872 `deepseek/settings.toml` when an existing settings file is present
1873
1874 Notable settings include `auto_compact`, which uses a model-aware default-on
1875 policy for known context windows up to the 1M-token V4 class. Automatic
1876 compaction runs before the active model limit and carries the compacted summary
1877 forward into the next request. The trigger defaults to
1878 `auto_compact_threshold_percent = 80`. Users who prefer manual continuity can
1879 persist `auto_compact = false`; manual `/compact` / Ctrl+L remains available.
1880 You can inspect or update these from the TUI with `/settings` and `/config`
1881 (interactive editor).
1882
1883 Common settings keys:
1884
1885 - `theme` (`system`, `terminal`, `underwater`, `underwater-retro`,
1886 `shoreline`, `shoreline-light`, `dark`, `light`, `grayscale`,
1887 `catppuccin-mocha`, `tokyo-night`, `dracula`, `gruvbox-dark`, `claude`,
1888 `matrix`, `solarized-light`, `uwu`; default `underwater`): `underwater` is
1889 the dark navy fresh-install default, `shoreline` is its warm charcoal
1890 alternative and `shoreline-light` the light variant, `system` follows terminal
1891 background detection, `dark`/`light` use the Codewhale Whale pair,
1892 `terminal` inherits the host terminal, `grayscale` is the low-opinion
1893 black/white theme, and the named community presets apply across the TUI.
1894 Aliases such as `whale`, `mono`, `black-white`, `tokyonight`, and `gruvbox`
1895 are accepted. In Whale, cobalt blue owns action/focus, seafoam owns live
1896 work, Signal Gold owns human decisions and the whale, coral owns warnings,
1897 rose owns danger, violet owns Operate, and green remains completed/verified.
1898 Text labels, markers, and motion policy carry the same states when color is
1899 unavailable; color is never the only cue.
1900 User-authored overlays live only at `~/.codewhale/themes/<name>.json` (or
1901 `$CODEWHALE_HOME/themes/<name>.json`) and are selected with
1902 `/theme custom:<name>`. The filename is a bounded slug, symlinks and files
1903 over 64 KiB are refused, colors must be `#RRGGBB`, and unknown fields fail
1904 validation. `/theme schema` prints the embedded JSON Schema and `/theme path`
1905 shows the exact directory. An overlay names one compiled `base` theme and
1906 changes only listed semantic colors; it cannot include or read another file.
1907 Open `/theme` to browse valid overlays, preview them live, and keep the
1908 active `custom:<name>` selector when the picker is opened without moving.
1909 - `auto_compact` (on/off, model-aware default on for known context windows
1910 unless explicitly configured)
1911 - `auto_compact_threshold_percent` (10-100, default `80`): pre-send
1912 auto-compaction threshold used only when `auto_compact` is enabled.
1913 - `paste_burst_detection` (on/off, default on): fallback rapid-key paste
1914 detection for terminals that do not emit bracketed-paste events. This is
1915 independent of terminal bracketed-paste mode.
1916 - `work_surface_placement` (`bottom`, `top`, `left`, `right`, or `off`;
1917 default `bottom`): places the workbar — Tasks / To-do / Workers — under the
1918 composer (the default bottom workbar), above the transcript, in a side
1919 workbar, or hides it entirely (`off`). Side choices fall back to the top
1920 layout on narrow terminals without changing the saved preference. Set it
1921 live with `/config work_surface_placement right --save` (or `left` / `top` /
1922 `bottom` / `off`).
1923 - `rail_panel` (`tasks`, `agents`, `background`, `files`, `notepad`,
1924 `context`, `git`, `price`; default `tasks`, alias key `rail`): which panel
1925 the workbar shows. Panel selection is orthogonal to placement. `tasks` is
1926 the full live work list (to-dos, then sub-agents); `agents` narrows to the
1927 sub-agent rows; `background` lists background shells and automations;
1928 `files` lists touched files; `notepad` shows the workspace notes; `context`
1929 is a read-only session-facts list; `git` shows branch status; `price`
1930 shows cost. In every panel except `context`, rows are selectable and
1931 clickable and open their detail surface. `Alt+!`/`Alt+@`/`Alt+#`/`Alt+$`
1932 switch panels live.
1933 - `work_surface_top_height` (2–16) and `work_surface_side_width` (26–80):
1934 ceilings for the top strip's height and the side workbar's width. Both are
1935 normally persisted by dragging the divider rather than edited by hand; the
1936 strip still auto-fits its content below the ceiling.
1937 - `focus_texture` (`off`, `scrim`, or `grain`; default `off`): focus-context
1938 texture for modal views. `scrim` dims the already-rendered background
1939 outside the focused modal toward the theme surface; `grain` sprinkles
1940 sparse dots over blank cells there. The texture is static (no time
1941 component, so it is unaffected by `low_motion`), never writes over a cell
1942 that carries text, and preserves the 4.5:1 body-text contrast floor
1943 wherever both colors are resolvable. It is skipped entirely on frames
1944 below the ambient-life minimum size and when the focused modal already
1945 covers 90% or more of the frame. Set it live with
1946 `/config focus_texture scrim --save`.
1947 - `mention_menu_limit` (integer, default `128`): maximum number of
1948 `@`-mention popup candidates retained before the composer renders the
1949 visible window. The visible rows still depend on terminal height.
1950 - `mention_walk_depth` (integer, default `10`): maximum workspace depth for
1951 `@`-mention completion walks. Set to `0` for unlimited depth in deeply
1952 nested workspaces; keep the default in very large repos unless needed.
1953 - `mention_menu_behavior` (`fuzzy`, `browser`; default `fuzzy`): controls how
1954 `@`-mention completions are populated. `fuzzy` searches the workspace and
1955 applies mention frecency. `browser` lists only the immediate children of the
1956 currently typed directory segment in deterministic alphabetical order.
1957 - `show_thinking` (on/off)
1958 - `thinking_default_expanded` (on/off, default off): renders thinking blocks
1959 expanded initially when `show_thinking` is enabled. Space still toggles the
1960 selected block, and it decides only the blocks you have not touched: a block
1961 you expanded or collapsed yourself keeps that state if you change this
1962 setting later. This is useful in SSH/tmux environments where the Space
1963 binding may be intercepted.
1964 - `thinking_preview_lines` (integer, default `2`): how many body rows a
1965 **collapsed** completed thought still shows. `0` is header-only; `10` is
1966 the older dump. Live streaming preview is unchanged. Expand a block with
1967 Space, or set `thinking_default_expanded` to open every block.
1968 - `help_expand_groups` (on/off, default off): start Help/shortcuts with every
1969 group expanded. Default folds the long tail (Grok-style); type-to-filter
1970 still unfolds matches.
1971 - `pin_last_prompt` (on/off, default on): pin the last user prompt at the top
1972 of the transcript viewport after it scrolls off.
1973 - `show_tool_details` (on/off)
1974 - `inline_diffs` (`full`, `summary`, or `off`; default `full`): controls the
1975 inline presentation of successful structured File mutations. `full` shows a
1976 bounded red/green diff and semantic statistics, `summary` keeps only the
1977 statistics, and `off` keeps the calm changed-file outcome. All three retain
1978 the exact applied change in the selected File receipt's Alt/Option+V detail.
1979 Failure and cancellation never render a successful diff. Save the choice
1980 with `/config inline_diffs <mode> --save`.
1981 - `locale` (`auto`, `en`, `ja`, `zh-Hans`, `zh-Hant`, `pt-BR`, `es-419`, `vi`,
1982 `ko`; default `auto`): UI chrome locale. `auto` checks `LC_ALL`,
1983 `LC_MESSAGES`, then `LANG`; unsupported locale selections resolve to English.
1984 Every shipped pack holds full `en.json` parity, so no string falls back
1985 to English. The runtime also exposes the resolved locale in the system
1986 prompt as the fallback natural language for V4 reasoning and replies when the
1987 latest user message is ambiguous. Clear user language still takes priority;
1988 Chinese turns should produce Chinese `reasoning_content` and Chinese final
1989 replies even when the resolved locale is English.
1990 - `background_color` (`#RRGGBB`, `RRGGBB`, or `default`): optional main TUI
1991 background color applied to the root, header, transcript, and footer
1992 surfaces while preserving panel contrast.
1993 - `cost_currency` (`usd`, `cny`; default `usd`): currency used by the footer,
1994 context panel, `/cost`, `/tokens`, and long-turn notification summaries. The
1995 aliases `rmb` and `yuan` normalize to `cny`.
1996 - `default_mode` (`agent`, `plan`, or `operate`; legacy values are accepted for migration but are not live mode vocabulary)
1997 - `launch_screen` (legacy, migration-only): this historical `on`/`off` value
1998 is still accepted when reading an existing settings file, but it no longer
1999 changes behavior and is omitted from new saves. A fresh interactive launch
2000 always opens Tideline Startup; only an explicit resume or an explicit
2001 initial prompt enters the live session directly.
2002 - `sidebar_focus` (legacy, migration-only): the classic right sidebar this key
2003 configured was removed in the 0.9.4 rail unification. The key is still read
2004 once so old settings carry forward, then folds into the live keys:
2005 `pinned`/`work`/`plan`/`todos` become `rail_panel = "pinned"`,
2006 `agents`/`subagents` become `rail_panel = "agents"`, `context`/`session`
2007 become `rail_panel = "context"`, `tasks`/`auto` (the old default) become the
2008 `tasks` panel, `sessions` enables `sessions_rail`, and `hidden` turns the
2009 workbar off via `work_surface_placement = "off"`. An explicit `rail_panel`
2010 in the file always wins over the migrated value. Configure the workbar with
2011 `rail_panel` and `work_surface_placement`, not this key.
2012 - `sessions_rail` (`on`/`off`; default `off`): show the persistent Sessions
2013 list in the workbar. Rows list this workspace's recent
2014 non-archived sessions, newest first, with the active one marked; activating a
2015 row opens the session picker preselected on it (`/sessions open <id>`), so
2016 resume keeps its single implementation. Rows are projected from cached
2017 session metadata — the list never reads a transcript per frame, and never
2018 contacts a provider.
2019 - `session_auto_resume` (`on`/`off`; default `off`): reattach to this
2020 workspace's most recent session when Codewhale starts. Off by default so
2021 plain `codewhale` keeps starting fresh. `--resume`, `--continue`, and
2022 `--fresh` always take precedence. When it is on, startup still refuses to
2023 resume a session that is archived, fails to load, or is recorded against a
2024 different workspace; each of those falls back to a fresh transcript and says
2025 which session was skipped and why. It applies to the interactive launch only
2026 — `codewhale "<prompt>"` and `codewhale exec` are never silently prefixed
2027 with a prior conversation.
2028 - `max_input_history` (number of submitted input history entries; cleared
2029 drafts are also kept locally for composer history search). Note the spelling:
2030 the serde field on disk is `max_input_history`
2031 (`crates/tui/src/settings.rs:426`, default 100). `max_history` is the key
2032 name accepted by `/config set` and `settings.set()` (`settings.rs:1388`), not
2033 a settings.toml key — writing `max_history` into the file is silently
2034 ignored.
2035 - `default_model` (model name override)
2036
2037 `/task digest` (alias `/tasks digest`) renders the canonical Work Graph
2038 operations and four-state To-do list as plain text, running work first. It
2039 reads the same snapshots as the styled Work surface and owns no parallel
2040 progress state.
2041
2042 Plan and Work are the everyday visible modes in the UI; Operate is an explicit
2043 preview entry while its Workflow control surface is still being built. Switch
2044 between them with `/mode`. For compatibility, older settings files with
2045 `default_mode = "normal"` still load as `agent`.
2046
2047 Localization scope is tracked in [LOCALIZATION.md](LOCALIZATION.md). The v0.7.6
2048 core pack covers high-visibility TUI chrome only; provider/tool schemas,
2049 personality prompts, and full documentation remain English unless explicitly
2050 translated later.
2051
2052 Readability semantics:
2053
2054 - Selection uses a unified style across transcript, composer menus, and modals.
2055 - Footer hints use a dedicated semantic role (`FOOTER_HINT`) so hint text stays readable across themes.
2056
2057 ### Token Quantities and Drivers
2058
2059 DeepSeek V4 prefix caching makes token labels matter. These quantities are kept
2060 separate:
2061
2062 | Quantity | Meaning | Allowed to drive |
2063 |---|---|---|
2064 | Active request input estimate | Conservative estimate of the next request's live system prompt and transcript payload. | Header/footer context percent, auto-compaction trigger, opt-in Flash seam trigger, and emergency overflow preflight. |
2065 | Reserved response headroom | The effective request cap plus `1024` safety tokens on every route. Normal no-override requests start at `65536`; a smaller route/provider ceiling narrows that value, and an explicit output override raises it only within the resolved route window and output ceiling. The identical cap reaches the wire and drives preflight; reasoning effort does not add a second hidden reservation. A separately published route input ceiling independently clamps the spendable input budget. | Emergency overflow budget checks only. |
2066 | Cumulative API usage | Provider-reported input plus output tokens summed across completed API calls; multi-tool turns may count the same stable prefix more than once. | Session usage and approximate cost telemetry only. |
2067 | Prompt cache hit/miss | Provider cache telemetry for the most recent call when available. | Cache-hit display and cost estimation only; never compaction or seam triggers. |
2068 | Context percent | Active request input estimate divided by the model context window. | Display only; it mirrors the active-input basis used by context safeguards. |
2069 | Cost estimate | Approximate spend from provider usage and configured DeepSeek rates. | Display only. |
2070
2071 For known context-window models, including 1M-class V4 models, replacement
2072 compaction is enabled by default unless the user explicitly configures
2073 `auto_compact = false`. It fires at the active model's compaction threshold and
2074 replaces old history with recent user context followed by one ordinary
2075 checkpoint message. The standing system prompt remains unchanged. Unknown model
2076 ids remain opt-in.
2077
2078 ### Command Migration Notes
2079
2080 If you are upgrading from older releases:
2081
2082 - Old: `/codewhale`
2083 New: `/links` (aliases: `/dashboard`, `/api`)
2084 - Old: `/set model deepseek-reasoner`
2085 New: `/config` and edit the `model` row to `deepseek-v4-pro` or `deepseek-v4-flash`
2086 - Old: visible `Normal` mode or `default_mode = "normal"`
2087 New: use `Agent` / `default_mode = "agent"`; legacy `normal` still maps to `agent`
2088 - Old: discover `/set` in slash UX/help
2089 New: use `/config` for editing and `/settings` for read-only inspection
2090
2091 ## Key Reference
2092
2093 ### Kimi Code membership model IDs
2094
2095 The exact `https://api.kimi.com/coding/v1` endpoint accepts `k3`, `k3-256k`,
2096 `kimi-for-coding`, and `kimi-for-coding-highspeed`. Use `k3-256k` for a fixed
2097 262,144-token K3 window; use bare `k3` with `context_window = 1048576` only
2098 when the membership plan includes the 1M entitlement. Both K3 ids use the same
2099 reasoning contract, and all four membership ids omit generic sampling fields.
2100
2101 ### Core keys (used by the TUI/engine)
2102
2103 - `provider` (string, optional): `deepseek` (default), `deepseek-anthropic`, `nvidia-nim`, `openai`, `atlascloud`, `wanjie-ark`, `volcengine`, `openrouter`, `xiaomi-mimo`, `novita`, `fireworks`, `siliconflow`, `arcee`, `siliconflow-CN`, `moonshot`, `sglang`, `vllm`, `ollama`, `ollama-cloud`, `huggingface`, `modelscope`, `together`, `qianfan`, `openai-codex`, `anthropic`, `openmodel`, `zai`, `stepfun`, `minimax`, `deepinfra`, `sakana`, `longcat`, `opencode-go`, `meta`, `mistral`, `telecomjs`, `xai`, `orcarouter`, `modelstudio-token-plan`, `google`, `edenai`, or `custom`. Legacy `deepseek-cn` configs are still accepted as an alias for `deepseek`; DeepSeek uses the same official host [`https://api.deepseek.com`](https://api-docs.deepseek.com/) worldwide. `deepseek-anthropic` targets DeepSeek's Anthropic Messages-compatible endpoint at `https://api.deepseek.com/anthropic` using `DEEPSEEK_API_KEY`; `nvidia-nim` targets NVIDIA's NIM-hosted DeepSeek endpoints through `https://integrate.api.nvidia.com/v1`; `openai` targets a generic OpenAI-compatible endpoint, defaulting to `https://api.openai.com/v1`; `atlascloud` targets AtlasCloud's OpenAI-compatible endpoint at `https://api.atlascloud.ai/v1`; `wanjie-ark` targets Wanjie Ark's OpenAI-compatible endpoint at `https://maas-openapi.wanjiedata.com/api/v1`; `volcengine` targets Volcengine Ark's OpenAI-compatible coding endpoint at `https://ark.cn-beijing.volces.com/api/coding/v3`; `openrouter` targets `https://openrouter.ai/api/v1`; `xiaomi-mimo` targets Xiaomi MiMo's OpenAI-compatible endpoint, using `https://token-plan-sgp.xiaomimimo.com/v1` by default for Token Plan keys (`tp-...`) and `https://api.xiaomimimo.com/v1` for pay-as-you-go keys. For Token Plan accounts outside the Singapore default, set `base_url` explicitly or use `mode = "token-plan-cn"` for China and `mode = "token-plan-ams"` for Europe/Amsterdam; `novita` targets `https://api.novita.ai/openai/v1`; `fireworks` targets `https://api.fireworks.ai/inference/v1`; `siliconflow` targets SiliconFlow, defaulting to `https://api.siliconflow.com/v1`; `arcee` targets Arcee AI's OpenAI-compatible endpoint at `https://api.arcee.ai/api/v1`; `siliconflow-CN` targets the SiliconFlow China regional endpoint through `[providers.siliconflow_cn]`; `moonshot` targets Moonshot/Kimi, defaulting to `https://api.moonshot.ai/v1`; `sglang` targets a self-hosted OpenAI-compatible endpoint, defaulting to `http://localhost:30000/v1`; `vllm` targets a self-hosted vLLM OpenAI-compatible endpoint, defaulting to `http://localhost:8000/v1`; `ollama` targets Ollama's OpenAI-compatible endpoint, defaulting to `http://localhost:11434/v1`; `huggingface` targets Hugging Face Inference Providers at `https://router.huggingface.co/v1`; `modelscope` targets ModelScope's OpenAI-compatible inference API at `https://api-inference.modelscope.cn/v1`; `together` targets Together AI at `https://api.together.xyz/v1`; `qianfan` targets Baidu Qianfan at `https://api.baiduqianfan.ai/v1`; `openai-codex` targets official ChatGPT plan inference with Codewhale-owned OAuth at `https://api.openai.com/v1`; `anthropic` targets Claude's native Messages API; `openmodel` targets OpenModel's Anthropic-compatible Messages API at `https://api.openmodel.ai`; `zai` targets Z.ai at `https://api.z.ai/api/coding/paas/v4`; `stepfun` targets StepFun at `https://api.stepfun.ai/v1`; `minimax` targets MiniMax at `https://api.minimax.io/v1`; `deepinfra` targets DeepInfra at `https://api.deepinfra.com/v1/openai`; `sakana` targets Sakana AI Fugu at `https://api.sakana.ai/v1`; `longcat` targets Meituan LongCat at `https://api.longcat.chat/openai/v1`; `opencode-go` targets the subscription-backed OpenCode Go model-aware route (Chat Completions, Responses, or Messages according to the documented model) at `https://opencode.ai/zen/go/v1`; `meta` targets Meta Model API; `mistral` targets Mistral AI's OpenAI-compatible endpoint at `https://api.mistral.ai/v1`; `telecomjs` targets TelecomJS TokenHub at `https://aigw.telecomjs.com/v1`; and `xai` targets xAI's API-key or OAuth route.
2104 - `opencode-zen` (string provider value): selects the model-aware OpenCode Zen gateway through `[providers.opencode_zen]`. The default base URL is `https://opencode.ai/zen/v1`, the default model is `gpt-5.6`, and credentials come from `api_key`, `OPENCODE_ZEN_API_KEY`, or fallback `OPENCODE_API_KEY`—never ChatGPT/Codex OAuth. `OPENCODE_ZEN_BASE_URL` and `OPENCODE_ZEN_MODEL` are accepted. The selected model's wire comes from the curated Zen snapshot, then from the AI SDK package its Models.dev `opencode` row names (refresh with `codewhale models --update`): GPT, Grok, and Muse Spark use Responses; Claude and most Qwen rows use Anthropic Messages; DeepSeek, MiniMax, GLM, Kimi, `qwen3.8-max`, and the free rows use Chat Completions. Gemini, Models.dev rows marked `deprecated`, and models no loaded catalog lists fail closed because Codewhale has no proven supported wire contract for them. See the exact current model groups in [`PROVIDERS.md`](PROVIDERS.md#opencode-zen-protocol-catalog).
2105 - `minimax-anthropic` (string provider value): selects MiniMax's Anthropic-compatible Messages route through `[providers.minimax_anthropic]`. The default Base URL is `https://api.minimax.io/anthropic`; set `https://api.minimaxi.com/anthropic` for China. Keep the `/anthropic` suffix because Codewhale appends `/v1/messages`. The route uses `MINIMAX_API_KEY` and defaults to `MiniMax-M3`; `MiniMax-M2.7` is also registered. Official M3 input modalities are text, image, and video, with adaptive or disabled thinking. M2.7 is text-only and always keeps thinking enabled.
2106 - `api_key` (string, required for hosted providers): must be non-empty for DeepSeek/hosted providers (or set the provider API key env var). Self-hosted SGLang, vLLM, and local `ollama` can omit it. `ollama-cloud` requires a key saved for that provider or supplied by `OLLAMA_CLOUD_API_KEY`, then `OLLAMA_API_KEY`.
2107 - `auth_mode` (string, optional provider-table key): selects a provider-specific authentication contract. Kimi Code membership uses `auth_mode = "api_key"` (or omit the field), a key created in the [Kimi Code console](https://www.kimi.com/code/console), `base_url = "https://api.kimi.com/coding/v1"`, and bare `model = "k3"` for K3. Codewhale gives that route a safe 262,144-token baseline; set `context_window = 1048576` only when the Kimi Code plan includes 1M access (Allegretto and above). `k3[1m]` is a Claude Code-only convention, not an API model ID, and Codewhale rejects it instead of silently changing the wire model or assuming an entitlement. `model = "kimi-for-coding"` remains the valid K2.7 compatibility route available to all Kimi Code members. Legacy `auth_mode = "kimi_oauth"` fails closed with API-key guidance and never probes, reads, refreshes, or rewrites `kimi_cli`/`kimi_code_cli` credential files. First-class OAuth requires Codewhale's own vendor-registered client identity and remains tracked in #4417.
2108 - `base_url` (string, optional, `[providers.<name>]` key; see [Legacy top-level `base_url` and `api_key`](#legacy-top-level-base_url-and-api_key) for the older top-level spelling): defaults to `https://api.deepseek.com/beta` for DeepSeek's OpenAI-compatible Chat Completions API, including legacy `provider = "deepseek-cn"` configs. Other defaults are `https://api.deepseek.com/anthropic` for `deepseek-anthropic`, `https://integrate.api.nvidia.com/v1` for `nvidia-nim`, `https://api.openai.com/v1` for `openai`, `https://api.atlascloud.ai/v1` for `atlascloud`, `https://maas-openapi.wanjiedata.com/api/v1` for `wanjie-ark`, `https://ark.cn-beijing.volces.com/api/coding/v3` for `volcengine`, `https://openrouter.ai/api/v1` for `openrouter`, `https://token-plan-sgp.xiaomimimo.com/v1` for `xiaomi-mimo` when the API key starts with `tp-...` and `https://api.xiaomimimo.com/v1` otherwise, `https://api.novita.ai/openai/v1` for `novita`, `https://api.fireworks.ai/inference/v1` for `fireworks`, `https://api.siliconflow.com/v1` for `siliconflow`, `https://api.siliconflow.cn/v1` for `siliconflow-CN`, `https://api.arcee.ai/api/v1` for `arcee`, `https://api.moonshot.ai/v1` for `moonshot`, `https://api.minimax.io/v1` for `minimax`, `https://api.openmodel.ai` for `openmodel`, `https://api.z.ai/api/coding/paas/v4` for `zai`, `https://api.stepfun.ai/v1` for `stepfun`, `https://api.deepinfra.com/v1/openai` for `deepinfra`, `https://api.sakana.ai/v1` for `sakana`, `https://router.huggingface.co/v1` for `huggingface`, `https://api-inference.modelscope.cn/v1` for `modelscope`, `https://api.together.xyz/v1` for `together`, `https://api.baiduqianfan.ai/v1` for `qianfan`, `https://api.openai.com/v1` for `openai-codex`, `https://api.anthropic.com` for `anthropic`, `https://api.mistral.ai/v1` for `mistral`, `http://localhost:30000/v1` for `sglang`, `http://localhost:8000/v1` for `vllm`, `http://localhost:11434/v1` for `ollama`, and `https://ollama.com/v1` for `ollama-cloud`. Set `base_url = "https://token-plan-cn.xiaomimimo.com/v1"` for China-region Xiaomi MiMo Token Plan accounts or `base_url = "https://token-plan-ams.xiaomimimo.com/v1"` for Europe/Amsterdam accounts. Mistral-specific reasoning fields and polymorphic replay are enabled only on the documented first-party HTTPS `/v1` hosts; a custom Mistral base URL keeps generic Chat semantics. Set `https://api.deepseek.com` or `https://api.deepseek.com/v1` explicitly to opt out of DeepSeek beta features.
2109 - `ollama-cloud` route: select `provider = "ollama-cloud"`, configure `[providers.ollama_cloud]` when overriding the default `https://ollama.com/v1` / `gpt-oss:120b` tuple, and save a key from [Ollama account settings](https://ollama.com/settings/keys) with `codewhale auth set --provider ollama-cloud`. Ambient precedence is `OLLAMA_CLOUD_API_KEY`, then `OLLAMA_API_KEY`; arbitrary Ollama model IDs pass through unchanged.
2110 - Legacy Ollama Cloud migration: a released `provider = "ollama"` config whose normalized `[providers.ollama].base_url` is exactly `https://ollama.com/v1` is upgraded to the `ollama-cloud` runtime identity in memory. Only that exact tuple may read its old `ollama` provider table and secret slot. The config and secrets are never rewritten, and neighboring paths, HTTP downgrades, lookalike hosts, or an explicit `ollama-cloud` selection never consume the fallback.
2111 - `telecomjs` base URL and catalog: `[providers.telecomjs]` defaults to `https://aigw.telecomjs.com/v1`; `TELECOMJS_BASE_URL` overrides it. With `TELECOMJS_API_KEY`, `/models` refreshes a key-scoped catalog without mixing rows into another provider.
2112 - `edenai` gateway: select `provider = "edenai"`; `[providers.edenai]` defaults to `https://api.edenai.run/v3` and `deepseek/deepseek-v4-pro`. `EDENAI_API_KEY`, `EDENAI_BASE_URL`, and `EDENAI_MODEL` are accepted. Use `EDENAI_BASE_URL = "https://api.eu.edenai.run/v3"` for Eden AI's documented EU endpoint; the default `deepseek/deepseek-v4-pro` is only listed on the global catalog, so pair the EU endpoint with an EU-listed model such as `qwen/deepseek-v4-pro` via `EDENAI_MODEL` or `model`. The provider refreshes Eden AI's `/models` catalog, but leaves model-specific reasoning controls untouched because the gateway spans multiple model families.
2113 - `codewhale` (Codewhale API): select `provider = "codewhale"`; `[providers.codewhale]` defaults to `https://api.codewhale.net/v1` and `deepseek/deepseek-v4-pro`. The credential is a Codewhale account API key (`cwc_key_…`) with the `models:infer` scope, read from `CODEWHALE_API_KEY` or the `codewhale` secret-store slot; `codewhale account api-keys create --name <name> --use` mints one and saves it locally. `CODEWHALE_API_BASE` overrides the origin and must be HTTPS except on loopback. The model catalog is the account's own authenticated `GET /v1/models`: ids are `provider/model` and each row states its wire (`chat-completions` → `/v1/chat/completions`, `anthropic-messages` → `/v1/messages`, `responses` → `/v1/responses`). Connect the underlying provider keys with `codewhale account keys set <provider>`.
2114 - `concentrate` gateway: select `provider = "concentrate"`; `[providers.concentrate]` defaults to `https://api.concentrate.ai/v1` and `deepseek-v4-pro` over the OpenAI Responses wire. `CONCENTRATE_API_KEY`, `CONCENTRATE_BASE_URL`, and `CONCENTRATE_MODEL` are accepted. Model ids pass through verbatim (`gpt-5.6-sol`, `openai/gpt-5.6-sol`, or `concentrate/auto` for the gateway router). BYOK only; see [PROVIDERS.md](PROVIDERS.md#concentrate-notes).
2115 - `mistral` model and reasoning contract: `[providers.mistral]` defaults to `mistral-code-latest`; `MISTRAL_MODEL` overrides it and the generic `CODEWHALE_MODEL` override wins when both are set. The current picker also lists `mistral-medium-latest`, `mistral-small-latest`, and `mistral-large-latest`. On exact first-party HTTPS `/v1` routes, Medium and Small accept only `reasoning_effort = "none" | "high"` and replay polymorphic thinking blocks. Deprecated native Magistral IDs may still be configured explicitly, remain always-reasoning, and never receive the adjustable effort field.
2116 - `context_window` (integer, optional provider-table key): override the total context window for the active `[providers.<name>]` route when an OpenAI-compatible gateway, hosted model alias, or self-hosted runtime has a different limit than Codewhale's static model table. For example, `[providers.openai] context_window = 1000000` lets an OpenAI-compatible DashScope/Qwen route budget against a 1M-token window instead of the conservative fallback. For Kimi Code K3, keep `model = "k3"` and set `[providers.moonshot] context_window = 1048576` only when the membership plan includes 1M access; otherwise omit it to retain the 262,144-token safe baseline. The value must be greater than 0 and affects prompt context notes, compaction thresholds, context-pressure checks, and request output caps. Full resolution order, and how to see which rung produced the current window: [Context length (context window)](#context-length-context-window).
2117 - `path_suffix` (string, optional provider-table key): override the chat-completions path for OpenAI-compatible gateways that do not serve `/v1/chat/completions`. For example, `[providers.openai] path_suffix = "/chat/completions"` sends chat requests to the unversioned base URL plus `/chat/completions`; `models` and `beta/*` requests keep their normal routing.
2118 - `reasoning_stream_style` (string, optional provider-table key): override how streaming reasoning is separated from answer text for the active provider route. Use `separate_field` for `reasoning_content` / `reasoning` deltas, `inline_tags` for gateways that stream `<think>...</think>` inside `delta.content`, or `none` to render incoming content exactly as answer text. When unset, every Chat Completions route uses `separate_field` (Mistral's first-party route uses its typed thinking blocks); set `none` only for a gateway that streams its answer inside `reasoning_content`.
2119 - `[providers.<name>.auth]` (table, optional): provider-scoped auth source metadata. `source = "command"` stores a command argv plus optional `timeout_ms`; `source = "secret"` stores a `secret_id`. This slice lets provider readiness, `/provider`, and doctor JSON report the auth source class without exposing command argv output or secret values; executing commands and resolving external secret material is handled by the follow-up resolver work.
2120 - `insecure_skip_tls_verify` (bool, optional provider-table key): legacy compatibility key, disabled by default. When true on the active provider table, provider clients reject the configuration instead of skipping TLS certificate verification. Use `SSL_CERT_FILE` for corporate or private CA bundles; `codewhale doctor` reports stale uses of this setting.
2121 - `default_text_model` (string, optional): defaults to `deepseek-flash` for DeepSeek and `deepseek-anthropic`, `gpt-5.6` for OpenAI, `grok-4.6` for xAI, `deepseek-ai/deepseek-v4-pro` for NVIDIA NIM, `deepseek-ai/deepseek-v4-flash` for AtlasCloud, `deepseek-reasoner` for Wanjie Ark, `DeepSeek-V4-Pro` for Volcengine Ark, `deepseek/deepseek-v4-pro` for OpenRouter and Novita, `mimo-v2.5-pro` for Xiaomi MiMo, `accounts/fireworks/models/deepseek-v4-pro` for Fireworks, `deepseek-ai/DeepSeek-V4-Pro` for SiliconFlow and DeepInfra, `trinity-large-thinking` for Arcee AI, `kimi-k2.7-code` for Moonshot, `MiniMax-M3` for MiniMax, `GLM-5.3` for Z.ai, `step-3.7-flash` for StepFun, `ernie-4.0-turbo-8k` for Qianfan, `fugu` for Sakana AI, `deepseek-ai/DeepSeek-V4-Pro` for SGLang/vLLM, no fixed id for local Ollama (Codewhale adopts a chat-capable tag from the live local catalog — coder or tool-capable tags first, then the largest context — and never an embedding or reranker model), and `gpt-oss:120b` for Ollama Cloud. Hugging Face and Together AI both default to `deepseek-ai/DeepSeek-V4-Pro`; `openai-codex` defaults to `gpt-5.6`; `anthropic` defaults to `claude-sonnet-4-6`; `openmodel` defaults to `deepseek-v4-flash`. Current public DeepSeek IDs include `deepseek-v4-pro` and `deepseek-flash` (V4.1 Flash, shipped as the unversioned id), both with 1M context windows, 384K max output, and thinking mode enabled by default. DeepSeek's live pricing/model page now labels the Pro backend `DeepSeek-V4-Pro-0813`; the callable API ID remains `deepseek-v4-pro`, so Codewhale does not send the backend label or the Claude Code-specific `deepseek-v4-pro[1m]` selector. DeepSeek retires `deepseek-chat` and `deepseek-reasoner` on July 24, 2026; direct first-party routes migrate both to `deepseek-v4-flash`, with omitted reasoning settings preserving their former non-thinking (`off`) and thinking (`high`) intent. Explicit `reasoning_effort` wins, and provider-owned ids on Wanjie Ark, aggregators, self-hosted runtimes, and custom endpoints are not globally rewritten. SiliconFlow retains its own mapping: `deepseek-reasoner` and `deepseek-r1` select its Pro model while `deepseek-chat` and `deepseek-v3` select Flash. Provider-specific mappings translate `deepseek-v4-pro` / `deepseek-v4-flash` to each provider's model ID where supported. OpenRouter also recognizes recent large IDs such as `arcee-ai/trinity-large-thinking`, `minimax/minimax-m3`, `minimax/minimax-m2.7`, `xiaomi/mimo-v2.5-pro`, `qwen/qwen3.6-flash`, `qwen/qwen3.6-35b-a3b`, `qwen/qwen3.6-max-preview`, `qwen/qwen3.6-27b`, `qwen/qwen3.6-plus`, `qwen/qwen3.7-max`, `google/gemma-4-31b-it`, `moonshotai/kimi-k2.7-code`, `moonshotai/kimi-k2.6`, `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free`, and `nvidia/nemotron-3-ultra-550b-a55b`; direct Arcee uses bare IDs such as `trinity-large-thinking` and `trinity-large-preview`; direct Moonshot recognizes `kimi-k3`, `kimi-k2.7-code`, and `kimi-k2.6`. The exact Kimi Code endpoint recognizes bare `k3` for K3 and `kimi-for-coding` for K2.7; those membership IDs are distinct from the direct Moonshot IDs and are never rewritten across routes. Direct MiniMax recognizes `MiniMax-M3` and the documented M2.x chat model IDs; direct Z.ai recognizes `GLM-5.3` (the default), `GLM-5.2`, `GLM-5.1`, and `GLM-5-Turbo`, and OpenRouter recognizes the matching `z-ai/glm-5.1`, `z-ai/glm-5.2`, `z-ai/glm-5.3`, and `z-ai/glm-5-turbo` IDs — `GLM-5.3` has been live on the Z.ai Coding Plan since 2026-08-13; it inherits its catalog metadata from `GLM-5.2` until Z.ai publishes distinct 5.3 numbers and carries no price, and an explicit `GLM-5.2` selection keeps its own id; direct Sakana recognizes `fugu` and `fugu-ultra-20260615`; direct Xiaomi MiMo recognizes chat IDs `mimo-v2.5-pro`, `mimo-v2.5-pro-ultraspeed`, and `mimo-v2.5`, while TTS IDs are selected through `codewhale speech` / `tts`. Generic `openai`, `atlascloud`, `wanjie-ark`, `xiaomi-mimo`, `arcee`, `moonshot`, `minimax`, `openmodel`, `zai`, `stepfun`, `qianfan`, `sakana`, local Ollama, and Ollama Cloud model IDs are passed through unchanged after known aliases are normalized. OpenRouter and SiliconFlow provider configs with a custom `base_url` also preserve explicit model values, which lets OpenAI-compatible gateways accept bare model IDs. Use `/models` or `codewhale models` to discover live IDs from your configured endpoint. `CODEWHALE_MODEL` overrides this for a single process; `DEEPSEEK_MODEL` is the legacy alias.
2122 - TelecomJS uses `deepseek-v4-pro` only as a conservative pre-refresh fallback. Once its key-scoped `/models` catalog is available, the picker uses those live rows; Codewhale omits unsupported reasoning request fields on this route.
2123 - `reasoning_effort` (string, optional): `auto`, `off`, `minimal`, `low`, `medium`, `high`, `xhigh`, `ultra` (`ultracode` is an alias of `ultra`), or `max`; defaults to the configured UI tier. DeepSeek Platform receives top-level `thinking` / `reasoning_effort` fields. Ollama Cloud's OpenAI-compatible Chat Completions route preserves its documented `none` / `low` / `medium` / `high` / `max` ladder (`off` is sent as `none`; `xhigh` and `ultracode` normalize to `max`). Direct xAI `grok-4.7` and `grok-4.6` on exact `https://api.x.ai/v1` receive top-level `reasoning_effort = "low" | "medium" | "high" | "xhigh"` (the ladder comes from the bundled catalog row; `grok-4.5` maps `xhigh` to `high`, and rows without a documented effort such as `grok-4.3` get no field); Grok reasoning cannot be disabled, so `off` normalizes to `high`, `max`/`ultracode` to `xhigh`, and `auto` leaves the field omitted so xAI's documented default `high` applies. A custom xAI-compatible `base_url` does not inherit that dialect. Direct Moonshot `kimi-k3` on exact `https://api.moonshot.ai/v1` is always-thinking and receives only top-level `reasoning_effort = "low" | "high" | "max"`; `off` normalizes to `low`, and `medium` to `high`. Kimi Code membership `k3` on exact `https://api.kimi.com/coding/v1` instead receives nested `thinking.effort`, and its `off` setting also normalizes to enabled `low`. Normal dispatched `auto` uses Codewhale's auto-reasoning selector and sends a concrete route-normalized tier; only an omitted reasoning setting leaves the provider default in control. Neighboring gateways and model/endpoint combinations retain the generic Moonshot contract. OpenAI Codex normalizes stale `off` to `low` and sends `max` / `ultracode` as Responses `xhigh`. Z.ai receives documented `thinking` controls and treats enabled thinking as the GLM coding high/max lane. NVIDIA NIM receives equivalent settings through `chat_template_kwargs`.
2124 - `verbosity` (string, optional): `normal` or `concise`. `normal` keeps the
2125 default conversational prompt. `concise` appends a prompt discipline block
2126 for direct, low-chatter output; CLI noninteractive commands (`exec` and
2127 `eval`) default to `concise` unless config/env/CLI overrides it.
2128 Override per process with `CODEWHALE_VERBOSITY` or the legacy
2129 `DEEPSEEK_VERBOSITY` alias.
2130 - `telemetry` (bool, optional): anonymous usage counting, **`true` by default
2131 in the current 0.9.12 source**. Notice version `5` names Codewhale and PostHog
2132 and describes the opt-out policy; no acceptance is invented for a default
2133 user. Existing explicit declines remain off. An explicit `false` here is the
2134 durable *opt-out*: it deletes the random install id, truncates buffered and
2135 dry-run events, and leaves a tombstone reasserted while the setting is false.
2136 It is a floor: `--telemetry true` and `CODEWHALE_TELEMETRY=1` lose to it.
2137 Use `/settings` or `codewhale config set telemetry true` to turn counting
2138 back on explicitly for new sessions through the existing privacy transition. `CODEWHALE_TELEMETRY`
2139 (legacy alias `DEEPSEEK_TELEMETRY`) and `--telemetry false` provide a run-scoped
2140 kill switch that stops collection and delivery without erasing the owner's
2141 state. A repo-local `.codewhale/config.toml` cannot set this preference.
2142 `codewhale config telemetry` shows the disclosure; `codewhale config get
2143 telemetry` reports preference and privacy status. Full schema and opt-out
2144 behavior: [`TELEMETRY.md`](TELEMETRY.md).
2145 - `telemetry_endpoint` (string, optional): where batches are POSTed. Leaving it
2146 unset selects the shipped default,
2147 **`https://telemetry.codewhale.net/v1/telemetry`** — the first-party ingest
2148 service described in [`TELEMETRY.md`](TELEMETRY.md), whose source is in
2149 `telemetry-ingest/`. This key decides only *where* a permitted session sends;
2150 it cannot override an opt-out. Setting it
2151 to the **empty string** is how you stay enabled and contact nobody: each batch
2152 is then written to `$CODEWHALE_HOME/telemetry/dryrun.jsonl` and no HTTP client
2153 is constructed at all, so you can read exactly what would have been sent. Any
2154 other value replaces the default outright. `https://` is required; plain
2155 `http://` is accepted only for loopback hosts, and there is no environment
2156 variable that overrides that refusal. A rejected endpoint turns telemetry off
2157 for the run rather than falling back to plaintext or to the default. Override
2158 per process with `CODEWHALE_TELEMETRY_ENDPOINT` (legacy alias
2159 `DEEPSEEK_TELEMETRY_ENDPOINT`), where an empty value means the same "contact
2160 nobody". A repo-local `.codewhale/config.toml` cannot set it.
2161 - `allow_shell` (bool, optional): in interactive TUI Agent sessions, omitting
2162 this keeps shell tools available with approval prompts; setting it to `false`
2163 hides shell tools. Headless, durable-task, and other noninteractive profiles
2164 keep the conservative omitted-field default and require `allow_shell = true`
2165 to expose shell. Plan mode always hides shell; Full Access enables shell and
2166 auto-approval.
2167 - `approval_policy` (string, optional): `on-request`, `untrusted`, or `never`. Runtime `approval_mode` editing in `/config` also accepts `on-request` and `untrusted` aliases.
2168 - `[approval] default_selection` (string, optional): which option an approval
2169 card highlights when it first appears — `deny` (default) or `allow_once`.
2170 `deny` means a reflexive Enter on a card you have not read refuses the call.
2171 Set `allow_once` to restore the pre-v0.9.6 Enter-to-approve muscle memory
2172 (#5293). It moves the highlight only: which calls are prompted for is still
2173 `approval_policy` plus the rules in `permissions.toml`.
2174
2175 ```toml
2176 [approval]
2177 default_selection = "allow_once"
2178 ```
2179 - `[approval] timeout_seconds` (integer, optional): bound how long an
2180 approval may wait — the TUI card and the Runtime API approvals the desktop
2181 app and web use alike. When the window elapses the call is refused
2182 (fail-closed) and recorded as a timeout, not as the operator's denial.
2183 Omitted or `0` waits indefinitely, which is the default everywhere: no
2184 approval is ever denied on your behalf unless you set this. Values above
2185 24h clamp with a warning (#6101).
2186
2187 ```toml
2188 [approval]
2189 timeout_seconds = 300
2190 ```
2191 - `sandbox_mode` (string, optional): `read-only`, `workspace-write`, `danger-full-access`, `external-sandbox`.
2192 Platform support is not identical. macOS uses Seatbelt when its runtime
2193 probe succeeds. Linux uses bubblewrap only when `prefer_bwrap = true` and
2194 `/usr/bin/bwrap` is executable; without that opt-in it reports no OS command
2195 sandbox. Windows does not currently advertise an OS sandbox; its planned helper contract starts
2196 with process-tree containment only and must not be described as read-only
2197 filesystem isolation, workspace-write enforcement, network blocking,
2198 registry isolation, or AppContainer isolation until those are implemented.
2199 - The cross-layer relationship between mode admission, hooks, registered tool
2200 requirements, typed rules, auto-review, repo law, human approval, and the
2201 execution sandbox is defined in
2202 [Authorization order](AUTHORIZATION_ORDER.md).
2203 - **Read deny-list.** Every sandbox posture — `read-only` included — grants
2204 read access to the whole filesystem; the postures differ in what they may
2205 *write* and whether they may reach the network. The read deny-list narrows
2206 that:
2207 - `sandbox_read_denylist_defaults` (bool, default `true`): apply the built-in
2208 credential-store set — `~/.ssh`, `~/.gnupg`, cloud credential directories
2209 (`~/.aws`, `~/.config/gcloud`, `~/.azure`, `~/.kube`, …), `~/.netrc`,
2210 `~/.npmrc`, `~/.git-credentials`, macOS keychains, browser profiles,
2211 Codewhale's own secret stores, and `.env` files (but not `.env.example`
2212 and friends). Ordinary source, `Cargo.toml`, `~/.gitconfig`, `~/.cargo`,
2213 and `~/.npm` stay readable so builds and tests still work. Set `false` to
2214 restore the pre-0.9.12 full-disk-read behavior.
2215 - `sandbox_denied_read_paths` (list of paths): additional denied subpaths.
2216 `~` expands. These can never be exempted.
2217 - `sandbox_read_denylist_exempt` (list of paths): subtract a path from the
2218 *built-in defaults* when a project genuinely needs it. Deny wins over
2219 allow: this never reopens anything in `sandbox_denied_read_paths`.
2220
2221 Exemption granularity is **whole-rule**, not per-file. An exempt path
2222 removes a built-in rule only when the rule's own path is at or below it,
2223 so exempting `~/.ssh/config` does nothing: the `~/.ssh` rule still denies
2224 it, because `~/.ssh` does not lie within `~/.ssh/config`. To reopen that
2225 one file you must exempt `~/.ssh` itself — which also reopens the private
2226 keys next to it. That is the documented tradeoff: there is no shipped way
2227 to narrow a built-in rule to "everything except one file"; copy what you
2228 need out of the denied tree instead. The one name-shaped rule, `.env`
2229 files, is exempted by *name* rather than by path: any exempt entry whose
2230 file name is exactly `.env` — bare `.env`, `~/.env`,
2231 `some/project/.env` — disables the entire `.env` filename rule, i.e.
2232 every `.env` and `.env.<name>` on disk rather than one project's.
2233 (`.env.example` and friends are never denied, so they need no exemption.)
2234
2235 Enforced at two points: sandboxed shell commands (Seatbelt last-match-wins
2236 `deny file-read*` rules; bubblewrap masks each path) and Codewhale's own
2237 in-process tools, which the OS sandbox never wraps — `read_file` / `read` /
2238 `read_media` for contents, `list_dir` / `file_search` for *enumeration*
2239 (listing a denied directory, or searching one, is refused just as Seatbelt
2240 blocks its readdir; a name search rooted above a denied tree skips entries
2241 inside it). A refused read is always an explicit error, never an empty
2242 result, and the error names the path as the caller spelled it rather than a
2243 symlink target's real location.
2244
2245 **This is defense-in-depth, not a security boundary.** It does not stop a
2246 hardlink to a denied file, a secret already copied into the workspace, an
2247 indirect read (`ssh-agent`, `security find-generic-password`, `aws sts …`),
2248 reads under `danger-full-access` shell commands, reads by MCP servers or
2249 other unwrapped child processes, or exfiltration of anything that *was*
2250 read. Keep least-privilege credentials and short-lived tokens doing the real
2251 work.
2252 - `permissions.toml` (sibling file, optional): typed permission rule records
2253 loaded next to `config.toml`, for example `~/.codewhale/permissions.toml`.
2254 This active user file is the only permission-rule source today; project
2255 config overlays do not load a project-local `permissions.toml`. A rule's
2256 optional `workspace` field is its repository scope, not a second source.
2257 Manually authored `[[rules]]` entries accept `tool`, optional `command` or
2258 `path`, optional absolute `workspace`, optional `command_exact = true`, and
2259 optional `action = "deny" | "ask" | "allow"`; omitted `action` defaults to
2260 `"ask"`. `workspace` limits a rule to that repository, while
2261 `command_exact = true` changes a command rule from the historical
2262 arity-aware prefix match to a complete-command match. `deny` blocks matching
2263 invocations before mode-based
2264 approval handling, `allow` skips approval for matching invocations, and
2265 `ask` forces approval only in modes that can prompt. Outside the TUI
2266 auto-approve path, a matching `ask` rule under `approval_policy = "never"`
2267 is rejected because no prompt can be shown. In Full Access / auto-approval sessions,
2268 `ask` rules do not downgrade the session into prompting or blocking; explicit
2269 `deny` rules still block according to the current execution-policy logic.
2270
2271 In a supported approval card, press `S` to allow the request once and append
2272 exact `action = "ask"` rules to this file. For eligible safe requests, choose
2273 **Always allow this exact rule in this repo** (shortcut `P`) to append an
2274 `action = "allow"` rule with the current absolute `workspace` scope.
2275 Remembered shell grants set `command_exact = true`, so later commands with
2276 extra arguments do not inherit the grant. File and patch grants retain the
2277 exact workspace-relative paths produced by the existing validation path.
2278 Supported saves are intentionally narrow:
2279 `exec_shell` stores the exact approved command string; `write_file` and
2280 `edit_file` store the exact workspace-relative file path; `apply_patch`
2281 stores one exact workspace-relative `path` rule per validated touched file
2282 from apply-patch preflight. Existing exec command matching remains
2283 arity-aware for manually authored prefix rules; approval-card allow grants
2284 use complete-command matching. File paths are normalized to the same
2285 workspace-relative form used by runtime matching.
2286
2287 `read_file` rules can still be authored manually when you want future reads
2288 of a specific path to ask, allow, or deny, but the approval UI does not save
2289 `read_file` rules. Commands classified as requiring approval or dangerous,
2290 critical approval cards, and repo-law prompts cannot save allow grants and
2291 continue to require review.
2292
2293 `/permissions` (or `/permissions list`) is the narrow rule-management
2294 surface. It lists each numbered rule with the active user-file source, its
2295 exact effective matcher (tool-wide, command prefix, exact command, or exact
2296 normalized path), global or repository scope, and whether that scope
2297 applies in the current workspace. `/config ask-rules` remains a compatibility
2298 entry to the same list.
2299
2300 Deletion is review-gated: `/permissions remove <number>` only previews the
2301 selected rule and prints a confirmation command. That command carries an
2302 opaque token bound to the exact file bytes and rule index; if another writer
2303 changes `permissions.toml`, confirmation fails instead of deleting a rule
2304 that moved into the old position. Confirmed removal and approval-card appends
2305 share the adjacent `permissions.toml.lock`, preserve unrelated TOML comments
2306 and formatting, and atomically replace the file. The running TUI reloads the
2307 user ruleset without clearing session-only approvals.
2308
2309 This editor intentionally does not create or rewrite rules, persist deny
2310 choices from approval cards, expand globs, or create broad
2311 directory/recursive rules. Author those supported exact/prefix records
2312 manually when needed.
2313 - `[[hotbar]]` (array of tables, optional): user-owned 1-8 slot bindings for
2314 the TUI hotbar. Each entry has `slot`, `action`, and optional `label`.
2315 The Hotbar is hidden by default (#3807): omitting `hotbar` and setting
2316 `hotbar = []` both mean no hotbar. `/hotbar on` writes the eight default
2317 slots (`slash.workflow`, `slash.goal`, `slash.auto`, `mode.plan`,
2318 `mode.agent`, `mode.operate`, `palette.open`, `sidebar.toggle`) as explicit
2319 `[[hotbar]]` tables. When one or more `[[hotbar]]` tables are present, that
2320 list is the whole bar; missing slots stay empty. Invalid slots outside `1..=8` are skipped with a warning, duplicate
2321 slots use the later entry, and unknown action IDs are kept so the UI can show
2322 a disabled/unknown cell instead of silently deleting user config. Trusted
2323 user config, profiles, and managed config replace the whole list; project
2324 overlays cannot change hotbar bindings. Setup or wizard flows that persist
2325 hotbar bindings write this same schema to the resolved `~/.codewhale/config.toml`
2326 path, preserving legacy `~/.deepseek/config.toml` only when that fallback file
2327 is already the active config.
2328
2329 ```toml
2330 [[hotbar]]
2331 slot = 1
2332 action = "mode.plan"
2333 label = "Plan"
2334
2335 [[hotbar]]
2336 slot = 2
2337 action = "session.compact"
2338 ```
2339 - `[auto_review]` (table, optional): tool-call review policy — a deterministic floor plus a model guardian tier.
2340 This layer sits on top of the existing permission posture; it can hold or block a
2341 tool call, but it is not an auto-push, auto-merge, or hosted review service.
2342 Block rules are checked first, then the built-in safety floor, then allow
2343 rules. In Ask, a safety hold opens approval; in Auto-Review, Full Access,
2344 or a non-interactive `never` posture it fails closed as a hard block. The
2345 safety floor still covers publish-like actions and destructive
2346 background/headless actions even if an allow rule matches.
2347
2348 ```toml
2349 [[auto_review.allow]]
2350 id = "read-only-inspection"
2351 action_kind = "read"
2352 reason = "Read-only inspection is safe to run automatically."
2353
2354 [[auto_review.block]]
2355 id = "no-release-publish"
2356 action_kind = "publish"
2357 reason = "Release and publish actions require maintainer review."
2358 ```
2359
2360 Rule matchers are exact `tool` and/or `action_kind`. At least one matcher is
2361 required. `action_kind` accepts the six decision-relevant kinds `read`,
2362 `write`, `shell`, `external`, `publish`, and `destructive`. Invalid names
2363 fail config validation instead of silently broadening into another policy
2364 class. In block rules, the old names remain conservative compatibility
2365 aliases: `network`, `git`, `mcp_action`, `browser`, and `unknown` map to
2366 `external`; `secret` maps to `destructive`; and `mcp_read` maps to `read`.
2367 Retired narrow kinds in allow rules fail validation rather than widening to
2368 a broader class. The retired `text_contains` matcher likewise fails
2369 validation instead of silently broadening an old intent-dependent rule.
2370 Fallback holds in interactive Auto-Review escalate to one stateless guardian
2371 request. The request contains the exact held call and deterministic
2372 observations as separate JSON fields. Conversation history, skill
2373 instructions, attached file contents, and other expanded model context are
2374 excluded. The guardian does not infer user intent or compute an authorization
2375 score. It exposes no tools and returns a risk level, allow/deny, and a
2376 rationale. High or critical risk cannot auto-run even if the model says
2377 allow. An oversized exact call is denied rather than truncated. Exactly one
2378 reviewer request is made; incomplete or malformed output, timeout,
2379 cancellation, provider failure, or an empty rationale all fail closed. The
2380 deterministic floor is never model-reviewed, and headless adapters use the
2381 deterministic-only tier. The pinned Codex, Kimi, and DeepSeek source
2382 boundaries are linked from [Permission Posture](MODES.md#permission-posture).
2383 Reviewer outcomes emit `tool.auto_review` audit events with
2384 `gate = "guardian"`.
2385
2386 Auto-review decisions emit `tool.auto_review` audit events with
2387 `gate = "deterministic"` when tool
2388 audit logging is enabled. Future PreToolUse/PostToolUse hooks can add
2389 observer input around this layer, but the configured auto-review policy is
2390 evaluated before a tool call is allowed to proceed.
2391 - `managed_config_path` (string, optional): managed config file loaded after user/env config.
2392 - `requirements_path` (string, optional): requirements file used to enforce allowed approval/sandbox values.
2393 - `max_subagents` (int, optional): defaults to `64` and is clamped to `1..=128`.
2394 - `subagents.*` (optional compatibility table): manual per-role model pins
2395 for direct and Workflow `agent` starts. An explicit saved profile wins,
2396 then a manual role pin, then a unique saved role pin. Conflicting tool
2397 `model` or `model_strength` choices are refused before admission. Unpinned
2398 roles allow task model/strength choices before inherited defaults.
2399 `[subagents.roles.<role>] model = "provider/model"` folds into the existing override
2400 map and wins over `[subagents.models]`, then the convenience keys. Structured
2401 canonical role keys win over legacy aliases. Only this structured syntax
2402 separates the explicit provider from the model suffix; unknown providers
2403 fail before admission. Bare structured model ids inherit the active provider.
2404 Legacy scalar/map values preserve namespaced provider-owned ids unchanged.
2405 Supported convenience keys are
2406 `default_model`, `worker_model`, `scout_model`, `planner_model`,
2407 `reviewer_model`, `custom_model`, `max_concurrent`, `max_admitted`,
2408 `launch_concurrency`, `token_budget`, `api_timeout_secs`, and
2409 `heartbeat_timeout_secs`. The v0.9.x keys `explorer_model`, `awaiter_model`,
2410 and `review_model` remain accepted as aliases. The `[subagents]
2411 max_concurrent` value overrides
2412 top-level `max_subagents` and is also clamped to `1..=128`. `[subagents]
2413 max_admitted` (aliases: `max_total`, `admission_limit`) is the bounded total
2414 of queued plus running sub-agents; it defaults to `1024`
2415 (`MAX_SUBAGENT_ADMISSION`, `crates/tui/src/config/subagent_limits.rs:21`,
2416 applied at `config.rs:6400`) so high-fanout turns can queue and drain while
2417 runtime launch pressure remains bounded, and is clamped to
2418 `max_concurrent..=1024`. `[subagents]
2419 launch_concurrency` sets how many direct children start at once before the
2420 rest queue for a launch slot; it defaults to the resolved `max_subagents` cap
2421 and is clamped to `1..=max_subagents` (the deprecated
2422 `interactive_max_launch` key is accepted as an alias, with the new key
2423 winning when both are set). `[subagents] token_budget` is an optional
2424 aggregate token ceiling for each root `agent` run and its descendants; unset
2425 or `0` preserves unlimited legacy behavior. `[subagents] api_timeout_secs`
2426 controls the per-step API timeout for sub-agent model calls and is clamped to
2427 `1..=3600`, with `0` or unset preserving the 600 second default; a timed-out
2428 attempt is retried with exponential backoff (up to 5 retries) before the
2429 step interrupts with a preserved checkpoint.
2430 `[subagents] heartbeat_timeout_secs` controls stale running agent cleanup,
2431 defaults to `300`, and is clamped to `30..=3600` while staying above the
2432 resolved API timeout. `[subagents.providers.<provider>]` accepts the same
2433 fanout, depth, budget, and timeout knobs (`enabled`, `max_concurrent`,
2434 `max_admitted`, `launch_concurrency`, `max_depth`, `token_budget`,
2435 `api_timeout_secs`, `heartbeat_timeout_secs`) and inherits the global
2436 `[subagents]` value for any key you omit. Provider keys accept canonical
2437 names such as `deepseek`, `zai`, `openrouter`, `anthropic`, plus convenience
2438 aliases such as `glm` for Z.ai and `deepseek_api` for direct DeepSeek:
2439
2440 ```toml
2441 [subagents]
2442 max_concurrent = 20
2443 launch_concurrency = 20
2444 max_admitted = 200
2445 max_depth = 6
2446
2447 [subagents.providers.deepseek]
2448 max_concurrent = 20
2449 launch_concurrency = 20
2450 max_admitted = 200
2451
2452 [subagents.providers.glm]
2453 max_concurrent = 4
2454 launch_concurrency = 3
2455 max_admitted = 12
2456 max_depth = 2
2457
2458 [subagents.providers.openrouter]
2459 max_concurrent = 5
2460 launch_concurrency = 3
2461 max_admitted = 20
2462 ```
2463
2464 `/config subagents status` prints both global values and the active
2465 provider's resolved profile so rate-limit tuning is visible in the TUI.
2466 `[subagents.models]` accepts lower-case Fleet role keys such as `worker`,
2467 `scout`, `planner`, `reviewer`, `builder`, and `verifier`; legacy type keys
2468 remain accepted during v0.9.x. Values are validated
2469 against the active provider at spawn time; direct DeepSeek requires DeepSeek
2470 IDs, while OpenAI-compatible/custom provider routes pass explicit model IDs
2471 through to that provider. To route a child to a different provider than the
2472 parent session, save a Fleet/AgentProfile with explicit `provider` and
2473 `model` fields (including user-named custom providers such as `lm-studio`)
2474 and call `agent(profile: "...")`; see [SUBAGENTS.md](SUBAGENTS.md).
2475 - `skills_dir` (string, optional): defaults to `~/.codewhale/skills` (each skill is
2476 a directory containing `SKILL.md`). Workspace-local `.agents/skills` or
2477 `./skills` are preferred when present; the runtime also discovers global
2478 agentskills.io-compatible `~/.agents/skills` and the broader Claude-ecosystem
2479 `~/.claude/skills`. First launch installs versioned bundled skills for common
2480 workflows including skill creation, delegation, MCP/plugin scaffolding,
2481 documents, presentations, spreadsheets, PDFs, and Feishu/Lark. Only
2482 Codewhale-owned roots (`<workspace>/.codewhale/skills` and
2483 `~/.codewhale/skills`) are writable install/import targets; compatible harness
2484 roots stay read-only. Bare `/skills` opens the Skills Manager (owned-only,
2485 zero network). See [SKILLS.md](SKILLS.md) for the manager, audit statuses,
2486 provenance markers, and mutation rules, and
2487 [CLAUDE_PLUGIN_COMPAT.md](CLAUDE_PLUGIN_COMPAT.md) for the supported boundary
2488 between portable `SKILL.md` bundles and Claude Code plugin runtimes.
2489 - `[skills].scan_codewhale_only` (bool, default `false`): when `true`, session
2490 skill discovery ignores cross-tool roots such as `.claude/skills`,
2491 `.opencode/skills`, `.cursor/skills`, and `~/.agents/skills`. Codewhale still
2492 scans `<workspace>/.codewhale/skills`, `~/.codewhale/skills`, and any explicit
2493 `skills_dir` override. The Skills Manager can still toggle a local compatible
2494 audit scan independently of this runtime knob — see [SKILLS.md](SKILLS.md).
2495 - `[skills].flat_workspace_root` (bool, default `false`): opt in to the flat
2496 `<workspace>/skills` compatibility root after workspace trust. Without this
2497 opt-in it is an audit candidate only; an explicit `skills_dir` remains an
2498 alternative. `scan_codewhale_only = true` excludes the flat compatibility
2499 root regardless of this flag, unless it is the explicit `skills_dir`.
2500 - `[skills].registry_url` / `[skills].max_install_size_bytes` (optional): used by
2501 `/skills --remote`, `/skills suggest <task>`, `/skills sync`, and `/skill
2502 install|update`. The default manager open path does not contact the registry.
2503 - `[verifier].enabled` (bool, default `false`): enables automatic
2504 claim-of-done verifier preview once that runtime trigger is active. The
2505 manual `run_verifiers` tool is still available when this is false.
2506 - `mcp_config_path` (string, optional): defaults to `~/.codewhale/mcp.json`, with
2507 legacy `~/.deepseek/mcp.json` fallback when the Codewhale path is absent.
2508 Custom paths must be absolute; a relative value falls back to the user-global
2509 path so changing the launch directory cannot silently change the MCP pool.
2510 It is visible in `/config` and can be changed from the TUI. The new path is
2511 used immediately by `/mcp`, but rebuilding the model-visible MCP tool pool
2512 requires restarting the TUI.
2513 - `notes_path` (string, optional): defaults to `~/.codewhale/notes.txt`, with
2514 legacy `~/.deepseek/notes.txt` fallback when the Codewhale path is absent, and
2515 is used by the model-visible `note` tool.
2516 - `[memory].enabled` (bool, optional): defaults to `false`. When `true`,
2517 the TUI loads the user memory file into a `<user_memory>` prompt block,
2518 enables `# foo` quick-capture in the composer, surfaces the `/memory`
2519 slash command, and registers the `remember` tool. The same toggle is
2520 available via `DEEPSEEK_MEMORY=on`.
2521 - `memory_path` (string, optional): anchors the native memory store. The
2522 configured filename is **not** the file that is written. Under the Native
2523 backend (the only backend) the store is re-rooted to
2524 `<parent-of-memory_path>/memory/global/MEMORY.md` — so the default
2525 `~/.codewhale/memory.md` yields `~/.codewhale/memory/global/MEMORY.md`
2526 (plus workspace-scoped files and a rebuildable SQLite FTS5 index). See
2527 [`MEMORY.md`](MEMORY.md) for the full feature surface (`# foo` composer
2528 prefix, `/memory` slash command, `remember` tool, opt-in toggle).
2529 - `snapshots.*` (optional): side-git workspace snapshots for file rollback:
2530 - `[snapshots].enabled` (bool, default `true`)
2531 - `[snapshots].max_age_days` (int, default `7`)
2532 - snapshots live under
2533 `~/.codewhale/snapshots/<project_hash>/<worktree_hash>/.git`, with legacy
2534 `~/.deepseek/snapshots/...` fallback when only the legacy state exists, and
2535 never use the workspace's own `.git` directory
2536 - `context.*` (optional):
2537 - `[context].project_pack` (bool, default `false`): include a deterministic
2538 project context pack (a large pretty-printed directory listing) in the
2539 stable prompt prefix (#4781). Useful for weak tool-calling models; the
2540 model can rebuild the same information with one `File` call.
2541 - The removed seam-manager keys (`enabled`, `verbatim_window_turns`,
2542 `l1_threshold`, `l2_threshold`, `l3_threshold`, `seam_model`) are
2543 ignored if an older config still carries them; they no longer load.
2544 - `compaction.*` (optional, config.toml): how a compaction pass behaves once
2545 it fires. `auto_compact` / `auto_compact_threshold_percent` (settings.toml)
2546 still decide *when* it fires. Both keys are absent by default, and absent
2547 means the built-in behaviour, unchanged:
2548 - `[compaction].summary_instructions` (string, default empty): standing
2549 operator instructions appended to the summarizer prompt as a clearly
2550 delimited "Additional instructions from the operator" section on **every**
2551 pass, manual and automatic — the effort-free counterpart to a one-off
2552 `/compact <focus>`, which still composes after this text. Useful for
2553 "always list exact file paths and line numbers", "always restate open
2554 decisions and their trade-offs", "always write a TL;DR first". Truncated
2555 at 4 000 characters with a warning naming the key; whitespace-only reads
2556 as unset. The summarizer still runs with no system prompt and no tools —
2557 this suffix is the only operator-authored input it sees.
2558 - `[compaction].retained_user_message_tokens` (int, default `20000`, clamped
2559 to `2000`–`200000`; also accepted as `retained_user_message_max_tokens`):
2560 token budget for the recent plain user messages kept **verbatim** in the
2561 replacement history. Raising it keeps more of the user's own earlier
2562 messages instead of only whatever the lossy summary captured; the
2563 last-round survival contract still applies on top. The `/compact` receipt
2564 names the effective budget and whether operator instructions were applied,
2565 so you can tell the knob took effect.
2566 - `retry.*` (optional): retry/backoff settings for API requests:
2567 - `[retry].enabled` (bool, default `true`)
2568 - `[retry].max_retries` (int, default `3`)
2569 - `[retry].initial_delay` (float seconds, default `1.0`)
2570 - `[retry].max_delay` (float seconds, default `60.0`)
2571 - `[retry].exponential_base` (float, default `2.0`)
2572 - `[retry].jitter` (bool, default `true`): randomize each backoff delay
2573 - `[retry].jitter_factor` (float, default `0.1` = ±10%; clamps to `0.0..=1.0`)
2574 - `[retry].respect_retry_after` (bool, default `true`): wait for a server
2575 `Retry-After` header instead of the computed backoff
2576
2577 `[retry]` schedules HTTP-request retries inside the client. The stream-level
2578 budgets that sit above it — how often a turn re-issues a request whose
2579 stream failed to open or died — are the `[stream]` keys below; legacy `tui.stream_max_*` keys remain fallbacks.
2580 - `[notifications]`: notification delivery, attention, categories and audio share one
2581 policy. `quiet = true`, `method = "off"`, `condition = "never"` and disabled
2582 categories suppress both the banner and Codewhale's selected sound.
2583 - `notifications.method`: `auto` (default), `osc9`, `kitty`, `ghostty`, `bel`, `off`.
2584 - `notifications.condition`: `unfocused` (default), `always`, `never`. When absent,
2585 the legacy `tui.notification_condition` remains the fallback. `always` also
2586 bypasses the duration threshold; `unfocused` requires two seconds away.
2587 - `notifications.threshold_secs`: nonnegative integer, default `30`.
2588 - `notifications.include_summary`: boolean, default `false`.
2589 - `notifications.sound`: optional `off`, `whale`, `bell`, `beep`, `file`.
2590 A selected value controls audio across enabled categories. Absent keeps legacy
2591 `completion_sound` and `event_sound` choices; `off` overrides both.
2592 - `notifications.sound_file`: custom local WAV path for `sound = "file"` or legacy
2593 `completion_sound = "file"`.
2594 - `notifications.subagent_completion`: `always`, `final-only` (default), `off`.
2595 Covers all background work that finishes: sub-agents, background shells and
2596 durable tasks. `final-only` sends one notice naming everything that finished
2597 once no agent, workflow or durable task is still running; a running shell
2598 (a dev server, a watcher) never holds it back. `always` sends one per item.
2599 - `notifications.quiet`: boolean, default `false`.
2600 - `notifications.events`: six boolean categories, all enabled by default; see below.
2601 - `notifications.completion_sound`: legacy completion cue, default `off`, with the
2602 same values as `sound`. Used only when `sound` is absent.
2603 - `notifications.event_sound`: legacy `enabled` (default `false`), `events`
2604 (default `["turn-complete", "approval-needed"]`), and `quiet` (default `false`).
2605 `min_interval_ms` (default `2000`) applies to each category's audio in both modes.
2606 - `tui.alternate_screen` (string, optional, default `auto`): which screen an interactive session starts on. `auto` and `always` start on the TUI-owned alternate screen; `never` starts in inline mode — a ratatui viewport the full height of the terminal with no alternate screen, so the shell's scrollback survives the session and stays scrollable after exit. `/fullscreen` and `/inline` switch it in-process; a switch that the terminal refuses rolls back and says why. Inline mode paints the whole transcript inside its viewport — nothing is written into the host scrollback while the session runs.
2607 - `tui.mouse_capture` (bool, optional, default `true` on non-Windows terminals and on Windows Terminal/ConEmu/Cmder when the alternate screen is active; `false` on legacy Windows console and inside JetBrains JediTerm — PyCharm/IDEA/CLion/etc. — where mouse-event escapes leak into the input stream as garbled text, see #878 / #898): enable internal mouse scrolling, transcript selection, right-click context actions, and transcript scrollbar dragging. TUI-owned drag selection copies the intersected cells, removes visual wrap-column line breaks from paragraphs, and keeps selection scoped to the transcript pane; the payload is Markdown source by default, see `tui.selection_copy_markdown` below. Set this to `false` or run with `--no-mouse-capture` for raw terminal selection; set it to `true` or run with `--mouse-capture` to opt in anywhere it's defaulted off. On raw terminal selection, especially on legacy Windows console or when mouse capture is disabled, selection may cross the right workbar and include visual wraps because the terminal, not the TUI, owns the selection.
2608 On Linux, finishing a transcript or composer selection quietly copies text to
2609 PRIMARY, leaving the regular clipboard unchanged. Middle-click inside the
2610 composer pastes PRIMARY at the pointer without submitting it. This uses native
2611 X11 or Wayland data control; compositors must support PRIMARY selection. Over
2612 SSH without a forwarded graphical display, use your terminal's selection/paste
2613 gestures or `--no-mouse-capture`. Explicit Copy still uses the regular clipboard.
2614
2615 - `tui.selection_copy_markdown` (bool, optional, default `true`): copy TUI-owned
2616 transcript selections (drag release, context-menu Copy, and `Cmd+C`/`Ctrl+C`
2617 on an active selection) as Markdown source instead of rendered text. Every
2618 intersected cell serializes through the same canonical projection `Ctrl-Y`
2619 and `/copy` use — user and assistant cells keep their authored Markdown,
2620 other cells keep their full transcript form — partial intersections round out
2621 to whole cells, cells join with blank lines, and a toast names the copied
2622 cell count. Set `false` to copy the rendered text as displayed. Composer
2623 selections and the Linux PRIMARY auto-copy are unchanged; PRIMARY always
2624 carries rendered text.
2625
2626 - `tui.stream_chunk_timeout_secs` (int, optional, default `900`): per-SSE-chunk idle timeout for streamed model responses. Slow local or compatible servers can raise this with `/config stream_chunk_timeout_secs <seconds>` (add `--save` to write canonical `stream.chunk_timeout_secs`); `0` maps to the default and explicit values must be `1..=3600`. The legacy `DEEPSEEK_STREAM_IDLE_TIMEOUT_SECS` env var is still honored when this key is omitted.
2627 - `tui.osc8_links` (bool, optional, default on for macOS/Linux, off for Windows): emit OSC 8 escape sequences around URLs in transcript output so supporting terminals (iTerm2, Terminal.app 13+, Ghostty, Kitty, WezTerm, Alacritty, recent gnome-terminal/konsole) can open them with the terminal's link gesture—usually Cmd-click on macOS and Ctrl-click on Linux/Windows. Terminals without OSC 8 support render the plain label and ignore the escape. The escapes are emitted out-of-band (not inside buffer cells), so column corruption is not a concern; set `false` only for terminals that misrender the OSC 8 terminator itself. Windows legacy consoles default off; opt in with `true`.
2628 - `tui.max_model_steps` (int, optional, default uncapped): optional model-step ceiling for one ordinary turn. Omission or `0` leaves model steps uncapped; explicit positive values are clamped to `1..=100000`. Headless `exec` and Fleet workers also have no implicit model-step ceiling; `exec --max-turns N` and positive worker budgets still apply. At ~80% of an explicit step budget the model gets one soft-landing notice; at exhaustion the turn ends `Failed` with `Maximum model steps reached before completion (limit: N)` after one bounded final-report response when needed. Cumulative wall-clock and per-stream limits remain independent. Active interactive goal turns use `goal.max_steps` instead (default `1000`); see the Goal loop section below.
2629 - `tui.turn_wall_clock_secs` (int, optional, default: no limit): cumulative per-turn wall-clock budget in seconds, measured across every model step of one turn (not per request). Time blocked on a human approval is excluded. Omitted or `0` means no limit; positive values clamp to `30..=86400` (24 hours is the ceiling). When exhausted the turn stops before authorizing another billable request with a message naming the limit and the key to raise.
2630 - `tui.stream_max_resumes` (int, optional, default `3`): how many times one turn re-issues a model request after its stream failed — the request never opened (connect failure or response-header stall), the stream died before any content, the host slept mid-stream, or the network dropped mid-stream. Every one of those paths spends this one budget, and a healthy stream resets it. `0` disables turn-level re-issues (a failed stream then fails the turn); values clamp to `0..=10`.
2631 - `tui.stream_max_transparent_retries` (int, optional, default `2`): in-stream re-requests while nothing has streamed yet. `0` disables them; values clamp to `0..=10`.
2632 - `tui.stream_max_errors` (int, optional, default `5`): recoverable errors tolerated within one stream before it ends. Unlike the two retry counts above, `0` does not switch anything off: like the other finite stream budgets it selects the default. Other values clamp to `1..=50`, so `1` ends the stream on its first recoverable error.
2633 - `tui.stream_open_timeout_secs` (int, optional, default `45`): wait for a streaming request's response headers (connection setup included). A header stall on HTTP/2 retries once over HTTP/1.1 with the same wait. Omitted or `0` falls back to `CODEWHALE_STREAM_OPEN_TIMEOUT_SECS`, then the default; values clamp to `5..=300`.
2634 - `tui.connect_timeout_secs` (int, optional, default `30`): TCP/TLS connect timeout for the model HTTP client. Omitted or `0` uses the default; values clamp to `1..=300`.
2635 - `tui.force_http1` (bool, optional, default `false`): pin the model HTTP client to HTTP/1.1, for provider edges or proxies that mishandle long-lived HTTP/2 streams. `CODEWHALE_FORCE_HTTP1=1` does the same; either one pins.
2636 - `tui.stream_max_content_mb` (int, optional, default `10`) and `tui.stream_max_duration_secs` (int, optional, default `1800`): per-step caps on one stream's accumulated content and wall-clock duration. `0` selects the default; values clamp to `1..=512` MB and `10..=86400` seconds.
2637 - `transcript.prose_measure` (positive integer, optional, default absent = full width): wrap cap, in columns, for prose cells — user messages, assistant answers, and reasoning/thinking blocks — in the live transcript (#5436). Absent (or `0`) spends the full content width, consistent with tool/status cells and the #5322 wide-frame decision; the former 105-column prose rail is gone. Set a positive whole number (e.g. `prose_measure = 120` under `[transcript]`) to restore a bounded reading measure on ultrawide terminals. Narrow terminals always keep their content width — the cap clamps from above only. Tool, diff, and status cells never inherit this cap. Invalid values (negative or non-integer) are rejected at startup with a `transcript.prose_measure` config error. Resolved once per render pass, so the main transcript cache and the full-screen overlay always agree on the effective width.
2638 - `hooks` (optional): lifecycle hooks configuration (see `config.example.toml`).
2639 - `features.*` (optional): feature flag overrides (see below).
2640
2641 ### Stream and transport settings
2642
2643 `[stream]` is the canonical table for model-stream policy and its HTTP clients.
2644 `codewhale config dump` and `codewhale config get stream` show effective values
2645 from the runtime resolver, including environment fallbacks and clamps; this does
2646 not persist defaults. Set or unset individual fields with, for example,
2647 `codewhale config set stream.open_timeout_secs 120` and
2648 `codewhale config unset stream.open_timeout_secs`. Per-run
2649 `--set stream.open_timeout_secs=120` uses the same validation.
2650
2651 | `[stream]` key | Default | Accepted behavior | Legacy `[tui]` fallback |
2652 | --- | --- | --- | --- |
2653 | `open_timeout_secs` | 45 | Positive values clamp to 5–300; 0 or omission falls through to the environment/default | `stream_open_timeout_secs` |
2654 | `chunk_timeout_secs` | 900 | 0 selects default; positive values clamp to 1–3600 | `stream_chunk_timeout_secs` |
2655 | `max_resumes` | 3 | 0 disables whole-request reissues; maximum 10 | `stream_max_resumes` |
2656 | `max_transparent_retries` | 2 | 0 disables retries before any content; maximum 10 | `stream_max_transparent_retries` |
2657 | `max_stream_errors` | 5 | 0 selects default; positive values clamp to 1–50 | `stream_max_errors` |
2658 | `max_duration_secs` | 1800 | Per-stream wall clock; 0 selects default, positive values clamp to 10–86400 | `stream_max_duration_secs` |
2659 | `max_content_mb` | 10 | Per-stream content; 0 selects default, positive values clamp to 1–512 MiB | `stream_max_content_mb` |
2660 | `connect_timeout_secs` | 30 | TCP/TLS setup; 0 selects default, positive values clamp to 1–300 | `connect_timeout_secs` |
2661 | `force_http1` | false | Boolean; a truthy environment pin always enables HTTP/1.1 | `force_http1` |
2662 | `tcp_keepalive_secs` | 30 | Idle time before TCP keepalive probes; 0 disables, positive values clamp to 1–3600 | none |
2663 | `http2_keep_alive_interval_secs` | 15 | PING interval on active HTTP/2 connections; 0 disables, positive values clamp to 1–3600 | none |
2664 | `http2_keep_alive_timeout_secs` | 20 | PING acknowledgement deadline; 0 selects default, positive values clamp to 1–3600 | none |
2665
2666 For each field, an explicit canonical value wins over the legacy `[tui]` field,
2667 including zero or false; an absent canonical field preserves the legacy value.
2668 The existing environment precedence is unchanged: a positive configured header
2669 wait wins over `CODEWHALE_STREAM_OPEN_TIMEOUT_SECS` (then its `DEEPSEEK_` alias);
2670 zero falls through to those variables. An omitted chunk timeout uses
2671 `CODEWHALE_STREAM_IDLE_TIMEOUT_SECS` (then its `DEEPSEEK_` alias), whereas an
2672 explicit zero uses 900. `CODEWHALE_FORCE_HTTP1` (legacy `DEEPSEEK_FORCE_HTTP1`)
2673 is ORed with the selected config flag, so a false config flag cannot defeat a
2674 truthy environment pin. Profile overlays merge canonical fields independently.
2675
2676 HTTP transport values apply to all newly constructed model, catalog and HTTP/1
2677 fallback clients. They do not rebuild an active client, govern MCP/other network
2678 services, or send HTTP/2 PINGs on idle pooled connections. HTTP/2 knobs have no
2679 effect when HTTP/1.1 is pinned. The operating system controls TCP probe details;
2680 these settings do not disable certificate validation. `[retry]` still owns the
2681 HTTP-request backoff schedule independently of stream-level budgets.
2682
2683 ### Workspace notes
2684
2685 `/note` manages a simple notes file in the current workspace at
2686 `.codewhale/notes.md` (legacy `.deepseek/notes.md` is the fallback path when
2687 no `.codewhale/notes.md` exists yet).
2688 Existing `/note <text>` usage still appends a note.
2689 The management forms are:
2690
2691 | Command | Action |
2692 |---|---|
2693 | `/note <text>` | Append a note (legacy shorthand) |
2694 | `/note add <text>` | Append a note explicitly |
2695 | `/note list` | List notes with temporary 1-based numbers |
2696 | `/note show <n>` | Show the full note at number `n` |
2697 | `/note edit <n> <text>` | Replace note `n` with new text |
2698 | `/note remove <n>` | Delete note `n`; `rm` and `delete` are aliases |
2699 | `/note clear` | Empty the workspace notes file |
2700 | `/note path` | Show the resolved workspace notes path |
2701
2702 The numbers shown by `/note list` are not stored in the file; they are derived
2703 from the current order each time notes are read. This keeps the file format
2704 compatible with the existing `---`-separated notes.
2705
2706 ### User memory
2707
2708 User memory is split across one top-level path setting and one opt-in
2709 toggle table:
2710
2711 ```toml
2712 # Anchors the store only — actual writes go to
2713 # ~/.codewhale/memory/global/MEMORY.md (see MEMORY.md).
2714 memory_path = "~/.codewhale/memory.md"
2715
2716 [memory]
2717 enabled = true
2718 ```
2719
2720 Notes:
2721
2722 - `memory_path` stays at the top level beside `notes_path` and
2723 `skills_dir`; it is not nested under `[memory]`.
2724 - The configured path is an **anchor**: its parent directory gains
2725 `memory/global/MEMORY.md`, workspace-scoped files, and `index.db`.
2726 Pointing `memory_path` at the native layout path itself would double-nest
2727 (`…/memory/global/memory/global/MEMORY.md`); keep the legacy-style
2728 anchor filename.
2729 - `DEEPSEEK_MEMORY_PATH` overrides the anchor path from the environment.
2730 - `DEEPSEEK_MEMORY=on` (also `1`, `true`, `yes`, `y`, or `enabled`)
2731 flips the feature on without editing `config.toml`.
2732 - The feature is inert when disabled: no file is injected, `# foo`
2733 falls through to normal message submission, and the model does not
2734 see the `remember` tool.
2735 - See [`MEMORY.md`](MEMORY.md) for examples and the full `/memory`
2736 command surface.
2737
2738 ### Goal loop (`[goal]`)
2739
2740 Operate-mode goals run to their completion gate with no default token, time, or
2741 continuation ceiling (#5052). Token/time budgets, when explicitly supplied,
2742 are telemetry only and do not stop a goal. Users who want a circuit breaker can
2743 opt into one:
2744
2745 ```toml
2746 [goal]
2747 # Optional safety backstop on automatic goal continuation passes.
2748 # Default: 0 (unlimited). Set a positive value to opt into a ceiling.
2749 max_continuations = 100
2750
2751 # Optional cancellable quiet period between successful turns. This is useful
2752 # for coordinator goals that should poll on a cadence instead of keeping one
2753 # provider turn open. Default: 0 (continue immediately).
2754 continuation_delay_seconds = 300
2755
2756 # Per-turn step allowance while a goal is active (#5994). Goal turns get a
2757 # larger but still finite budget than an ordinary interactive turn.
2758 # Default: 1000 (0 or absent resolves to 1000, never unlimited). Range:
2759 # 1..=100,000. This bounds each provider turn, never the number of
2760 # continuation passes.
2761 max_steps = 1000
2762 ```
2763
2764 The effective delay is capped at 86,400 seconds (24 hours); use an automation
2765 for schedules that are less frequent than once per day.
2766
2767 When an explicit backstop fires, the goal pauses with a status message naming
2768 `[goal] max_continuations` and a warning is logged; resume the goal after
2769 inspecting progress, or raise/disable the backstop.
2770
2771 `[goal] max_steps` governs one engine turn at a time: the ordinary interactive
2772 turn has no implicit model-step ceiling. Explicit per-invocation
2773 ceilings — `exec --max-turns N`, child-worker caps — always win over it. At
2774 about 80% of the selected budget the model is told to land; at exhaustion it
2775 gets one bounded final report and the turn classifies as budget-exhausted. An
2776 unfinished goal then pauses with the BudgetLimit reason instead of re-arming
2777 another goal turn — resume it explicitly after reviewing the report. Wall-clock
2778 and stream protections are separate and still apply.
2779
2780 The delay starts only after a successful turn while an explicitly created goal
2781 is still active. `/goal pause`, `/goal done`, `/goal blocked`, `/goal clear`,
2782 Esc, or Ctrl+C cancels a pending continuation before another provider request
2783 starts. Failed turns and policy/route failures never schedule another turn.
2784 Only the numeric cadence is stored in config; no prompt, credential, or secret
2785 is persisted for the loop.
2786
2787 ### Reasoning-only recovery (`[reasoning_only]`)
2788
2789 When a reasoning model (thinking mode) finishes a response with only hidden
2790 reasoning and no answer text or tool call, the engine can automatically
2791 re-request the answer. Configure this behavior with the `[reasoning_only]`
2792 table:
2793
2794 ```toml
2795 [reasoning_only]
2796 # Maximum number of automatic re-requests. Default: 2.
2797 # Set to 0 to disable automatic recovery (fail immediately).
2798 max_reprompts = 10
2799
2800 # Optional custom message sent to the model on each re-request.
2801 # When set, overrides the built-in default message.
2802 # When unset (or commented out), the engine uses:
2803 # "So, what's up ? Keep running !"
2804 reprompt_message = "Allez, répond quelque chose !"
2805 ```
2806
2807 This only applies when the model returns a clean `stop` finish reason with
2808 only thinking content. An output-length stop (`length`/`max_tokens`) is never
2809 retried, and a persistently answerless model still fails honestly after the
2810 configured bound.
2811
2812 To disable the reprompt message entirely (silent retry), set it to an empty
2813 string:
2814
2815 ```toml
2816 [reasoning_only]
2817 reprompt_message = ""
2818 ```
2819
2820 ### Notifications
2821
2822 Notification controls are available in the existing `/config` settings view,
2823 from terminal commands, and through the CLI. All write the same `config.toml`
2824 keys. Terminal changes apply immediately; add `--save` to keep them. CLI writes
2825 apply when the next process loads its configuration.
2826
2827 ```sh
2828 codewhale config set notifications.sound whale
2829 codewhale config set notifications.events.approval-needed false
2830 codewhale config set notifications.quiet true
2831 codewhale config get notifications
2832 codewhale config unset notifications.quiet
2833 ```
2834
2835 ```text
2836 /config notifications sound whale --save
2837 /config notifications condition unfocused --save
2838 /config notifications quiet true
2839 /config notifications status
2840 ```
2841
2842 Nested CLI edits preserve TOML types, unrelated keys and comments. Unset removes
2843 only the selected leaf. `notifications.sound legacy` in the TUI, or CLI unset of
2844 `notifications.sound`, restores previous sound choices. Invalid values are
2845 rejected before file or session changes. In an active TUI profile, saved edits
2846 update its notification table when it owns one; otherwise they update the
2847 inherited root table. The settings detail keeps saved and current values distinct.
2848
2849 ```toml
2850 [notifications]
2851 method = "auto" # auto | osc9 | kitty | ghostty | bel | off
2852 condition = "unfocused" # unfocused | always | never
2853 threshold_secs = 30
2854 include_summary = false
2855 sound = "whale" # optional; sound is opt-in, not enabled by default
2856 quiet = false
2857
2858 [notifications.events]
2859 turn-complete = true
2860 subagent-terminal = true
2861 approval-needed = true
2862 input-needed = true
2863 elevation-needed = true
2864 model-notify = true
2865 ```
2866
2867 `quiet = true` mutes every category without changing saved choices. `method =
2868 "off"` also stops both banner and selected audio. Disabling a category stops its
2869 sound. Attention and duration gates apply before any sink runs. Title animation
2870 completion is silent; only the authorized notification event can request audio.
2871 Successful turn completion is a notification category; failed/cancelled turns do
2872 not create a success notification.
2873
2874 `auto` chooses a recognized terminal protocol or the existing macOS native
2875 fallback; unknown terminals remain unsupported and never invent a bell. `bel`
2876 is an audio-only transport: one selected cue is dispatched, without a second
2877 transport bell. With `sound = "off"`, that transport is silent. `osc9`, `kitty`
2878 and `ghostty` use their terminal notification protocols; tmux passthrough is
2879 preserved. Terminal/OS notification preferences still govern display, attribution
2880 and any sound the host itself adds.
2881
2882 By default the terminal must stay unfocused for two seconds. `condition =
2883 "always"` allows foreground notifications and bypasses the duration threshold;
2884 `"never"` suppresses all delivery. The canonical condition takes precedence over
2885 legacy `[tui].notification_condition`.
2886
2887 The bundled `whale` is a 1.55-second original whale-inspired cue, with no
2888 third-party recording. It remains an opt-in candidate pending listening approval.
2889 WAV playback uses a background worker: macOS `/usr/bin/afplay`, Linux `aplay`
2890 from [ALSA utilities](https://github.com/alsa-project/alsa-utils), or Windows
2891 `PlaySoundW`. A missing player/file or unsupported platform has no fallback bell.
2892 Only one WAV plays at a time. A worker-start receipt is a dispatch attempt, not
2893 proof of audible playback or OS acceptance. The existing macOS `osascript`
2894 banner retains Script Editor attribution; this Core change does not provide a
2895 branded native Apps banner.
2896
2897 #### Previous sound settings
2898
2899 When `notifications.sound` is absent, completion uses a non-off
2900 `completion_sound` selection, and other events use the existing event allow-list.
2901 These are compatibility inputs to the same audio decision, not separate playback
2902 paths. When the global sound is selected, it takes precedence over the previous
2903 completion/event choices. The default remains silent unless a legacy sound or
2904 explicit `bel` transport was already selected.
2905
2906 ```toml
2907 [notifications.event_sound]
2908 enabled = false
2909 events = ["turn-complete", "approval-needed"]
2910 min_interval_ms = 2000
2911 quiet = false
2912 ```
2913
2914 Legacy event cues use one bell for completion, subagent completion, input and
2915 model notices; approval/elevation cues use two. The per-category repeat interval
2916 survives settings refreshes. Old unknown event names are ignored on load; new
2917 CLI/TUI edits require names from the six categories above.
2918
2919 Local approval/input/elevation prompts and error receipts remain available when
2920 external notifications are muted. Action prompts use the selected UI language
2921 and retire when that request settles. Repeated live notices keep their first
2922 expiry; a later routine update does not hide an unresolved warning at completion.
2923 Optional plugin suggestion toasts and contextual tips share one session guidance
2924 budget. `/config contextual_tips off` hides those toasts while preserving
2925 required notices and explicit plugin review requests.
2926
2927 #### What a notification can contain
2928
2929 A desktop notification is a glance surface: on macOS it can appear on the
2930 lock screen, and on every platform it is visible to anyone near the machine.
2931 Codewhale therefore builds notifications from a typed payload with a fixed
2932 per-event disclosure policy rather than from whatever text was on hand:
2933
2934 | Event | Shown | Never shown |
2935 |---|---|---|
2936 | Turn complete | status line (+ elapsed/cost when `include_summary`), preview of the assistant's reply | — |
2937 | Sub-agent finished | status line, agent id, preview of the child's summary line | — |
2938 | Approval needed | the tool name | the tool description, the command, the arguments |
2939 | Input needed | "Answer the question in the terminal to continue" | the question |
2940 | Sandbox elevation needed | the tool name and the denial reason | the command |
2941 | `notify` tool | model-supplied title and body | — |
2942
2943 Every field is capped (80 characters for the status line, 120 for the
2944 identifier, 200 for the preview), stripped of control bytes and escape
2945 sequences, and passed through a redactor that replaces credential-shaped
2946 strings with `[redacted]`, reduces absolute local paths to `…/basename`,
2947 and replaces raw tool JSON with `[details hidden]`. The redactor is
2948 deliberately over-eager: an unbroken 40-character run has no word
2949 structure, so it is redacted even when it is not a secret.
2950
2951 #### macOS: why the banner says "Script Editor"
2952
2953 On macOS terminals that provide no notification escape of their own —
2954 Apple Terminal, the VS Code and JetBrains embedded terminals, plain tmux
2955 without `LC_TERMINAL` — `method = "auto"` falls back to `osascript`'s
2956 `display notification`. That command posts on behalf of the *bundled*
2957 host process, and `/usr/bin/osascript` is unbundled, so macOS attributes
2958 the banner to `com.apple.ScriptEditor2`. That attribution supplies the
2959 Script Editor icon and owns the System Settings → Notifications entry
2960 (alert style, previews, Do Not Disturb). `display notification` has no
2961 icon parameter, so this cannot be fixed from the notification code; it
2962 needs Codewhale to ship a real `.app` bundle. Tracked in
2963 [#4834](https://github.com/codewhale-hq/CodeWhale/issues/4834). In the meantime,
2964 iTerm2, WezTerm, Ghostty, and kitty are matched first and use their own
2965 notification protocols, and `method = "osc9"` / `"bel"` / `"off"` opt out
2966 of the `osascript` path explicitly.
2967
2968 ## Automations in the terminal
2969
2970 Open `/automation`, then choose **New automation** (`n`) or **Edit automation**
2971 (`e`). The form edits the name, multiline prompt, schedule, model, workspace,
2972 and enabled status. `Tab` moves between fields; `Enter` inserts a newline in
2973 the prompt. Schedule presets include daily, weekly, hourly, once, and a custom
2974 RRULE. Time fields accept `HH:MM`; their arrow controls change the time by
2975 15 minutes. Weekly day buttons and model search support mouse and keyboard.
2976
2977 The next-run preview uses the scheduler's local time zone, shown with its UTC
2978 offset. An enabled automation can run after **Save** (`Ctrl+S`); a paused one
2979 has no scheduled next run. **Cancel** (`Esc`) discards the draft. Editing keeps
2980 existing permission settings, custom schedules, and additional workspace
2981 entries unless the corresponding supported field is explicitly changed.
2982
2983 Schedules are evaluated in the machine's local time zone against the wall
2984 clock. A wall time that does not exist on a spring-forward day is skipped and
2985 an ambiguous fall-back time fires once. Occurrences missed while Codewhale was
2986 closed, asleep, or still running the previous occurrence are coalesced: the
2987 next start runs one catch-up occurrence and then continues from the next
2988 future slot, never replaying every missed slot. An occurrence never starts
2989 while an earlier run of the same automation is still queued or running. Each
2990 run is recorded durably with its status, timing and error; a run that needs a
2991 tool approval has no operator to ask and fails once the approval wait expires.
2992
2993 Choosing a concrete model pins both the model and its exact configured
2994 provider, including named custom routes. Later changes to the active provider
2995 do not move that automation's pin. The default-model choice and legacy
2996 definitions without a provider pin keep the runtime's existing default
2997 behavior. New automation records use schema v2 and tasks use v3 so older
2998 runtimes reject records whose provider pins they cannot preserve.
2999
3000 ## Lifecycle Outbox (`[lifecycle_outbox]`)
3001
3002 The lifecycle outbox is an opt-in, machine-readable stream of session,
3003 turn, and sub-agent lifecycle events. With a path configured, Codewhale
3004 appends one JSON line per event to that file — for interactive TUI
3005 sessions *and* headless `codewhale exec` runs — so a supervisor
3006 (terminal multiplexer wrapper, automation harness, alerting setup) can
3007 react to what happened without scraping the screen or installing per-hook
3008 shell commands. Unset or empty `path` = the feature is **off** and
3009 behavior is unchanged.
3010
3011 ```toml
3012 [lifecycle_outbox]
3013 path = "~/.codewhale/notifications/outbox.jsonl" # unset/empty = OFF
3014 webhook_url = "" # optional; POSTs events as JSON when set
3015 webhook_token = "" # optional bearer token for webhook_url
3016 ```
3017
3018 ### Events emitted
3019
3020 | Event | Kind | Fired at |
3021 |---|---|---|
3022 | `turn_start` | `turn.started` | a new turn begins (TUI TurnStarted; `exec` at message dispatch) |
3023 | `turn_end` | `turn.completed` / `turn.failed` / `turn.interrupted` | turn completion, kind projected from the turn status |
3024 | `turn_stalled` | `turn.stalled` | the stall watchdog recovers a wedged turn |
3025 | `subagent_spawn` | `subagent.spawned` | a sub-agent is spawned |
3026 | `subagent_complete` | `subagent.completed` | a sub-agent reaches a terminal state |
3027 | `session_start` | `session.started` | interactive session start |
3028 | `session_end` | `session.ended` | interactive session end |
3029
3030 ### File contract
3031
3032 Each line is a `RuntimeEventEnvelope`:
3033
3034 ```json
3035 {"schema_version": 1, "seq": 3, "event": "turn_start", "kind": "turn.started",
3036 "thread_id": "…", "turn_id": "…", "item_id": null, "timestamp": "…",
3037 "created_at": "…", "payload": {…}}
3038 ```
3039
3040 - `seq` is monotonic per outbox file and recovers from the last written
3041 line when a new process opens the file.
3042 - Lines are written one complete JSON line per append, serialized by an
3043 internal writer task and flushed before the next event; concurrent
3044 sessions writing the same path do not interleave bytes mid-line, but
3045 separate processes each continue from their own recovered `seq`, so seqs
3046 can repeat across processes sharing one file — prefer one file per
3047 process for strict uniqueness.
3048 - Parent directories are created lazily on the first event.
3049 - Payloads are constructed from bounded, pre-redacted fields only — never
3050 raw tool arguments, environment, or full transcript text. Free-form
3051 fields (error messages, previews) are capped at the notification limits
3052 (80 headline / 120 detail / 200 preview characters) and stripped of
3053 control bytes.
3054
3055 ### Webhook delivery
3056
3057 With `webhook_url` set, every event is additionally POSTed as
3058 `{"at": "<ISO 8601 timestamp>", "event": {…}}` with
3059 `Authorization: Bearer <webhook_token>` when a token is configured.
3060 Delivery is best-effort: failures are logged and dropped, never retried
3061 into the agent loop, and a failing webhook never blocks the local file
3062 append.
3063
3064 ## Control Socket (`[control_socket]`)
3065
3066 Per-session control surface for supervised operation: with the feature
3067 enabled, the interactive TUI binds one unix domain socket per *running*
3068 session at `<sessions-dir>/<session-id>/control.sock` (mode `0600`;
3069 `<sessions-dir>` is the same directory the session store uses, typically
3070 `~/.codewhale/sessions`). The socket is removed with the session's
3071 artifact directory, and a stale socket left by a crashed process is taken
3072 over by the next launch. Unset or `enabled = false` = the feature is
3073 **off** (the default) and behavior is unchanged. Unix-only; on other
3074 platforms the key parses but no socket is bound.
3075
3076 ```toml
3077 [control_socket]
3078 enabled = false # default: OFF
3079 ```
3080
3081 The socket speaks newline-framed JSON-RPC, one request per connection:
3082 write one request line, read one response line, close.
3083
3084 ```json
3085 {"id":"1","method":"message","params":{"text":"hello"}}
3086 {"id":"2","method":"interrupt","params":{}}
3087 {"id":"3","method":"status","params":{}}
3088 ```
3089
3090 - `message` — delivers `text` as a structured user message through the
3091 ordinary composer dispatch path; dispatched immediately when idle,
3092 queued when a turn is in flight (the response's `delivery` field says
3093 which).
3094 - `interrupt` — the Esc-shaped cancel of the active turn; `cancelled`
3095 reports whether active work was in flight.
3096 - `status` — answers `turn_state` (`idle` / `in_progress` / `waiting`) and
3097 `goal` (`objective`, `status`, `paused`).
3098
3099 Success responses echo the request id with a `type`-tagged result;
3100 failures carry `error.code` (`invalid_request`, `command_error`,
3101 `timeout`, `server_unavailable`). Requests are bounded at 1 MiB per line
3102 and a handler that does not answer within 5 s is reported as `timeout`.
3103
3104 ## Tool Catalog
3105
3106 Codewhale loads a small core native tool catalog by default and leaves less
3107 common native tools discoverable through ToolSearch. To keep specific native
3108 tools loaded on every request, add them to `[tools].always_load`:
3109
3110 ```toml
3111 [tools]
3112 always_load = ["Git", "notify"]
3113 ```
3114
3115 ### Script tools and overrides
3116
3117 Scripts in `~/.codewhale/tools/` (or `[tools].plugin_dir`) that start with a
3118 `# name:` header become model-visible tools, and `/plugin tools` lists them.
3119 The script reads the tool's JSON input on stdin and writes a JSON
3120 `ToolResult` (`{"content": "...", "success": true}`) on stdout.
3121
3122 ```sh
3123 #!/usr/bin/env sh
3124 # name: word_count
3125 # description: Count words in the given text
3126 # schema: {"type":"object","properties":{"text":{"type":"string"}}}
3127 # approval: required
3128 ```
3129
3130 `# approval:` takes `suggest` (the default) or `required`; either way the
3131 tool follows the session's approval setting. A script cannot approve itself:
3132 `approval: auto` is no longer supported, so such a script gets the default,
3133 and the runtime log (`~/.codewhale/logs/`) and `/plugin tools` name it.
3134
3135 A script cannot replace a built-in tool either. A script whose `# name:` is
3136 already registered is not loaded. `[tools.overrides]` may disable a built-in,
3137 or add a script or command tool under a name of its own:
3138
3139 ```toml
3140 [tools.overrides]
3141 "Web" = { type = "disabled" } # turn a built-in off
3142 "audited_shell" = { type = "script", path = "audit-shell.sh" } # a new tool
3143 "Bash" = { type = "script", path = "audit-shell.sh" } # refused: Bash is built in
3144 ```
3145
3146 A `script` or `command` override keyed by a built-in is refused, and the
3147 built-in stays active. A status line names the key once per session, and the
3148 runtime log records it. To route a
3149 built-in through your own wrapper, disable the built-in and register the
3150 wrapper under a new name. An override keyed by a drop-in script's name still
3151 replaces that script. Relative `path` values resolve against the plugin
3152 directory.
3153
3154 ### `request_user_input` limits
3155
3156 `request_user_input` asks the user a short batch of multiple-choice questions.
3157 Both ceilings are configurable (#5949): raise `user_input_max_questions` when a
3158 research or planning workflow legitimately needs more clarifications, lower it
3159 when interactive triage should stay terse.
3160
3161 ```toml
3162 [tools]
3163 user_input_max_questions = 6 # default 6, clamped to 1..=10
3164 user_input_max_options = 4 # default 4, clamped to 2..=10
3165 ```
3166
3167 The effective values are applied in three places at once: the tool's JSON
3168 schema (`minItems` / `maxItems`), its model-visible description, and the
3169 payload validator. A rejected payload names the key to raise, so the model can
3170 either resize the batch or tell the user which setting to change.
3171
3172 ### User-input wait timeout
3173
3174 Questions from `request_user_input` wait until answered or canceled by default
3175 (#6003). An omitted setting or `0` leaves the wait unbounded; a positive value
3176 cancels the question when that many seconds pass, capped at 86,400 (24 hours).
3177 Headless `exec` runs have no
3178 responder, so `request_user_input` is withheld there by default:
3179 the model reports the tool absent and finishes instead of stalling.
3180
3181 ```toml
3182 [tools]
3183 user_input_timeout_seconds = 300 # opt into 5 minutes; omitted or 0 waits indefinitely; maximum 86400
3184 ```
3185
3186 This key governs question waits only. Approvals have their own clock,
3187 `[approval] timeout_seconds`, which waits indefinitely unless you set it.
3188 Wall-clock and stream protections elsewhere are unaffected.
3189
3190 ## Feature Flags
3191
3192 Feature flags live under the `[features]` table and are merged across profiles.
3193 Defaults are enabled for built-in tooling, so you only need to set entries you
3194 want to force on or off.
3195
3196 ```toml
3197 [features]
3198 shell_tool = true
3199 subagents = true
3200 web_search = true # enables deferred Web; the flag name is retained for config compatibility
3201 apply_patch = true
3202 mcp = true
3203 exec_policy = true
3204 code_mode = true # execute_tools composes MCP/plugin/native calls; false defers it behind tool_search
3205 verify_tool = true # agent-callable `verify` self-critique; false removes it from the tool catalog
3206 vision_model = false # true routes image analysis to the [vision_model] model (see above)
3207 extension_host = false # experimental: run reviewed plugins' native code (see EXTENSIONS.md)
3208 ```
3209
3210 `code_mode` is on by default: `execute_tools` is advertised from the first
3211 request and nested calls go through the same permission gate as direct calls
3212 (see [Tool surface](TOOL_SURFACE.md#code-mode-execute_tools)). Set
3213 `code_mode = false` to defer it behind `tool_search` again.
3214
3215 `extension_host` is experimental and off by default. Turning it on lets reviewed
3216 plugins run their `native` TypeScript/JavaScript tool and slash-command code in
3217 a separate Node (or, opt-in, Bun) process; toggling it in either direction
3218 changes the plugin activation policy, so every plugin is reviewed again after a
3219 restart. Its `[extension_host]` table (`runtime`, `node`, `bun`) and what the
3220 host's sandbox does on each platform are in
3221 [Writing an extension tool](EXTENSIONS.md).
3222
3223 Every flag has a row in [`docs/features.toml`](features.toml), the feature
3224 registry; a test fails when a flag and its row disagree.
3225
3226 You can also override features for a single run:
3227
3228 - `codewhale --enable web_search`
3229 - `codewhale --disable subagents`
3230
3231 Use `codewhale features list` to inspect known flags and their effective state.
3232 The native `/config` view also includes a read-only **Experimental** section
3233 for experimental feature flags. It shows each flag's effective enabled/disabled
3234 state and whether that state comes from the default or a configured override.
3235 Change feature flags in `[features]` or with `--enable` / `--disable`; the
3236 `/config` section is an audit surface, not a stability promise. Goal and
3237 Workflow preview rows may appear there as reserved entries until those workflows
3238 graduate behind real gated flags.
3239
3240 ## Web Search Provider
3241
3242 `web_search` uses keyless Firecrawl by default. Runtime failure or an exhausted
3243 keyless quota degrades visibly through DuckDuckGo and Bing. China deployments
3244 can explicitly select Baidu, Metaso, Volcengine, or a trusted SearXNG endpoint;
3245 Codewhale does not guess geography from locale or model provider.
3246
3247 Configured API providers are attempted first. Runtime failure or an empty
3248 result visibly degrades through DuckDuckGo and then Bing; the structured search
3249 receipt records every hop. Missing configuration and network-policy denials
3250 fail closed without sending the query to another provider.
3251
3252 **Provider-native search.** On routes whose provider offers its own web-search
3253 tool (OpenAI, xAI, Anthropic, DeepSeek, Kimi and others), that search can run
3254 ahead of the configured provider. It is a separate model call on the active
3255 route. `[search] native` decides the order:
3256
3257 - unset (default): native search leads only when no search provider is
3258 configured; a provider chosen in `[search] provider`,
3259 `CODEWHALE_SEARCH_PROVIDER`, or a Tavily key wins;
3260 - `native = true`: native search leads even when a provider is pinned;
3261 - `native = false`: native search is never used.
3262
3263 The native answer is returned whole; oversized tool output spills to a session
3264 artifact the model can page back.
3265
3266 **Recency and locale.** `recency` and `locale` are forwarded where the
3267 backend's API takes them, and the search receipt reports each as honored or
3268 ignored:
3269
3270 | Backend | Recency | Locale |
3271 | --- | --- | --- |
3272 | Firecrawl | `tbs=qdr:d/w/m/y` | `country` from the region (`de-DE` → `DE`); a bare language is ignored |
3273 | Tavily | `time_range` | not sent (Tavily takes country names) |
3274 | SearXNG | `time_range` | `language`, as given |
3275 | Serply | not sent (undocumented) | `hl` language, `gl` country |
3276
3277 Recency is rounded up to the backend's nearest window (day, week, month,
3278 year), so `recency = 10` searches the last month. Other backends ignore both
3279 knobs and say so in the receipt.
3280
3281 For a private/internal search service that serves DuckDuckGo-compatible HTML,
3282 keep `provider = "duckduckgo"` and set `base_url`; Codewhale appends the `q`
3283 query parameter to that endpoint and applies network policy to its host.
3284 Custom endpoints do not fall back to public Bing. `CODEWHALE_SEARCH_BASE_URL`
3285 can override this per process; `DEEPSEEK_SEARCH_BASE_URL` remains accepted as
3286 the legacy alias.
3287
3288 **SearXNG** ([docs](https://docs.searxng.org/dev/search_api.html)) uses the
3289 configured instance's JSON API. Set `provider = "searxng"` and
3290 `base_url = "https://your-searxng.example"`; Codewhale calls
3291 `/search?q=...&format=json`. Codewhale does not use a public SearXNG instance
3292 by default because public instances often disable JSON output or rate-limit API
3293 traffic.
3294
3295 Self-host it as a separate process (Docker is fine); Codewhale never bundles or
3296 manages the search engine itself:
3297
3298 - Enable JSON on the instance (`settings.yml`, `search.formats` must include
3299 `json`) and restart it. An HTML-only instance answers the API with HTTP 403;
3300 Codewhale reports that as a JSON/API-access problem on the SearXNG hop rather
3301 than silently returning no results.
3302 - Bind it to loopback, or to a host and port your network policy allows. The
3303 instance keeps its own engine list, limiter, and limits.
3304 - Point Codewhale at it with `[search] provider = "searxng"` and `base_url`
3305 (required; either the root URL or the `/search` endpoint). No instance ships
3306 as a default, and none is discovered automatically.
3307 - `codewhale doctor --probe-search` sends a transport-only `HEAD` to that
3308 origin — no `q=`, no credentials, no redirects, no audit receipt — so a green
3309 probe proves reachability and network-policy admission, not that JSON is on.
3310
3311 Confirm the JSON API itself before assuming a Codewhale bug:
3312
3313 ```sh
3314 curl -sS "$BASE/search?q=codewhale&format=json" | jq '.results[0] | {title,url,score}'
3315 ```
3316
3317 Codewhale ranks the returned rows by `score`, highest first, and applies
3318 `max_results` to that ranking; rows an instance reports without a usable score
3319 keep their original relative order.
3320
3321 **Metaso** ([metaso.cn](https://metaso.cn)) requires a user-supplied key. Set
3322 `METASO_API_KEY` or `[search] api_key`; Codewhale does not ship a shared key.
3323
3324 **Firecrawl** ([docs](https://docs.firecrawl.dev/sdks/cli)) searches Firecrawl
3325 Cloud without a key using its bounded per-IP daily quota. Set
3326 `FIRECRAWL_API_KEY` or `[search] api_key` for authenticated limits. Codewhale
3327 sends no `Authorization` header in keyless mode.
3328
3329 **Baidu** uses Baidu AI Search at
3330 `https://qianfan.baidubce.com/v2/ai_search/web_search`. Set
3331 `BAIDU_SEARCH_API_KEY` or `[search] api_key`. This is a search-tool backend
3332 only; it does not add a Baidu model provider.
3333
3334 **Sofya** ([sofya.co](https://sofya.co)) returns full extracted page content
3335 rather than snippets. Set `[search] api_key` to your `ay_live_...` key, or the
3336 `SOFYA_API_KEY` env var. This is a search-tool backend only; it does not add a
3337 Sofya model provider.
3338
3339 **Serply** ([serply.io](https://serply.io)) returns Google organic results with
3340 title, URL, and snippet. Set `[search] api_key` to your Serply key, or the
3341 `SERPLY_API_KEY` env var. This is a search-tool backend only; it does not add a
3342 Serply model provider.
3343
3344 **Tavily** ([tavily.com](https://tavily.com)) is selected automatically when a
3345 Tavily key is present and no provider is pinned: `TAVILY_API_KEY` set, or
3346 `[search] api_key` / `CODEWHALE_SEARCH_API_KEY` in the `tvly-` family. Doctor
3347 reports that as `source: tavily key`. Autodetect is runtime-only — Codewhale
3348 never writes `[search] provider` for it, and `TAVILY_API_KEY` is never merged
3349 into `[search] api_key`. An explicit `[search] provider` or
3350 `CODEWHALE_SEARCH_PROVIDER` always wins, so `provider = "firecrawl"` keeps
3351 Firecrawl even with a Tavily key in the environment. Pinned `tavily` accepts
3352 any non-empty `[search] api_key` and is configured by that key or
3353 `TAVILY_API_KEY`; with both empty it fails closed.
3354
3355 ```toml
3356 [search]
3357 provider = "firecrawl" # also duckduckgo | bing | tavily | bocha | metaso | searxng | baidu | volcengine | sofya | serply
3358 # base_url = "https://search.example/" # optional with provider = "duckduckgo"; required with "searxng"
3359 # api_key = "YOUR_KEY" # optional for firecrawl; required by the other API providers
3360 # native = false # provider-native search: unset = only when no provider is configured
3361 ```
3362
3363 ## Local Media Attachments
3364
3365 Use `@path/to/file` in the composer to add local text file or directory context
3366 to the next message. Use `/attach <path>` for local image/video media paths, or
3367 `Ctrl+V` to attach an image from a local clipboard or an explicitly forwarded
3368 X11/Wayland clipboard. SSH terminal paste without a forwarded graphical display
3369 is text-only; use the local terminal's paste command (`Cmd+V` on macOS or
3370 `Ctrl+Shift+V` on Linux/Windows), and use `/attach <path>` for remote image
3371 files. OpenSSH loopback X11 displays are detected automatically. For an
3372 explicitly forwarded Wayland or non-loopback X11 display, set
3373 `CODEWHALE_SSH_CLIPBOARD=graphical`; set it to `terminal` to force terminal
3374 transfer instead of an ambient remote display. DeepSeek's public Chat
3375 Completions API currently accepts text message
3376 content, so media attachments are sent as explicit local path references instead
3377 of native image/video payloads.
3378 Attachment rows appear above the composer before submit; move to the start of
3379 the composer, press `↑` to select an attachment row, then press `Backspace` or
3380 `Delete` to remove it without editing the sample text by hand.
3381
3382 ## Managed Configuration and Requirements
3383
3384 codewhale supports a policy layering model:
3385
3386 1. user config + profile + env overrides
3387 2. managed config (if present)
3388 3. requirements validation (if present)
3389
3390 By default on Unix:
3391 - managed config: `/etc/deepseek/managed_config.toml`
3392 - requirements: `/etc/deepseek/requirements.toml`
3393
3394 Requirements file shape:
3395
3396 ```toml
3397 allowed_approval_policies = ["on-request", "untrusted", "never"]
3398 allowed_sandbox_modes = ["read-only", "workspace-write"]
3399 ```
3400
3401 If configured values violate requirements, startup fails with a descriptive error.
3402
3403 ## Notes On `codewhale doctor`
3404
3405 `codewhale doctor` follows the same config resolution rules as the rest of the
3406 TUI. That means `--config`, `CODEWHALE_CONFIG_PATH`, and the legacy
3407 `DEEPSEEK_CONFIG_PATH` are respected, and MCP/skills
3408 checks use the resolved `mcp_config_path` / `skills_dir` (including env overrides).
3409
3410 To bootstrap missing MCP/skills paths, run `codewhale setup --all`. You can
3411 also run `codewhale setup --skills --local` to create a workspace-local
3412 `./skills` dir.
3413
3414 Both plain `codewhale doctor` and `codewhale doctor --json` are structural and
3415 offline by default. They do not check the release service, hosted provider
3416 APIs, local provider endpoints, or MCP processes, and they do not load a
3417 workspace credential `.env`. Use `--check-updates`,
3418 `--probe-api`, `--probe-local`, or `--probe-mcp` to opt into the corresponding
3419 live boundary; `--probe-local` may start a desktop-managed service such as
3420 Ollama. Only the explicit API/local probe paths may load workspace credential
3421 `.env` values. Live flags conflict with `--json`, so machine-readable doctor
3422 output is always offline. Top-level keys include `version`, `paths`, `secret_backend`,
3423 `config_path`, `config_present`, `workspace`, `api_key.source`,
3424 `api_key.availability`, `base_url`,
3425 `default_text_model`, `mcp`, `skills`, `tools`, `plugins`, `sandbox`,
3426 `platform`, `api_connectivity`, and `capability`. CI consumers should rely on
3427 `api_key.source` (`config_declared`/`env_declared`/`external_auth_declared`/
3428 `secret_store_unprobed`/`secret_store_unavailable`/`oauth_unprobed`/
3429 `external_consent`/`none`/`local_runtime`/`unknown`) and
3430 `api_key.availability`
3431 (`present`/`not_required`/`not_probed`/`unavailable`/`unknown`) rather than parsing the
3432 human-readable `doctor` text. Source is declaration metadata, not proof that a
3433 credential exists or works. Only a non-empty, non-sentinel literal config value
3434 is structurally `present`; no-auth and local routes are `not_required`. Environment,
3435 external-auth, OAuth, consent, and secret-store declarations remain `not_probed`
3436 and cannot make structural Setup or Fleet readiness true. A secret-store sentinel
3437 on a named/custom endpoint that is prohibited from using the shared store is
3438 `secret_store_unavailable`/`unavailable`, while `unknown` remains reserved for
3439 the absence of a supported structural conclusion. Exact and whitespace-wrapped
3440 legacy sentinels are never treated as literal credentials. The structural
3441 loader still honors safe environment routing/model/policy fields, but it never
3442 materializes environment HTTP headers, sandbox API keys, or search API keys;
3443 only an explicit API/local probe switches to the normal credential-loading
3444 path. An opted-in update check also emits only typed generic failures: untrusted
3445 release metadata and transport errors are not echoed.
3446
3447 If configuration loading or validation fails, `doctor --json` returns nonzero
3448 and prints a bounded JSON error envelope with
3449 `status = "error"` and `error.kind = "config_validation"`. It does not emit a
3450 normal route or capability report—or the underlying possibly sensitive error—
3451 for an invalid configuration.
3452
3453 MCP entries are configuration diagnostics unless an explicit MCP command is
3454 run. `mcp.probe_scope` is `configuration`, `mcp.live_health_checked` is false,
3455 and each server separates `checks.configuration` / `checks.command` from
3456 `checks.process_reachable`, `checks.protocol_initialized`, and
3457 `checks.backend_tool_health`. The latter three remain `not_checked` in doctor
3458 output. Run `codewhale mcp validate` to explicitly start enabled servers and
3459 verify protocol initialization/discovery; backend health still requires an
3460 appropriate explicit tool call. Doctor reports only safe structural MCP fields:
3461 URL userinfo/path/query/fragment and raw command arguments, environment values,
3462 header values, and token material are omitted. Provider URLs follow the same
3463 rule and expose only `scheme://host[:explicit-port]`.
3464
3465 The `capability` key contains per-provider capability info derived from
3466 static knowledge (release docs, API guides) rather than live API probes.
3467 Top-level sub-keys: `resolved_provider`, `resolved_model`, `context_window`,
3468 `max_output`, `thinking_supported`, `cache_telemetry_supported`,
3469 and `request_payload_mode`.
3470
3471 Use `capability.context_window` and `capability.max_output` for model-limit
3472 checks in CI scripts; do not treat `capability.max_output` as the per-turn
3473 request budget. Use `capability.thinking_supported` to decide whether to
3474 configure reasoning effort.
3475
3476 ## Setup status, clean, and extension dirs
3477
3478 `codewhale setup` accepts a few flags beyond the existing `--mcp`,
3479 `--skills`, `--local`, `--all`, and `--force`:
3480
3481 - `--status` — print a compact one-screen status (api key, base URL, model,
3482 MCP/skills/tools/plugins counts, sandbox, `.env` presence). Read-only and
3483 network-free; safe to run in CI. If `.env` is missing and `.env.example` is
3484 present in the workspace, the status output points at `cp .env.example .env`.
3485 - `--tools` — scaffold `~/.codewhale/tools/` with a `README.md` describing the
3486 self-describing frontmatter convention (`# name:` / `# description:` /
3487 `# usage:`) and an `example.sh` that follows it. The directory is
3488 intentionally not auto-loaded; wire individual scripts into the agent via
3489 MCP, hooks, or skills.
3490 - `--plugins` — scaffold `~/.codewhale/plugins/` with a `README.md` and an
3491 `example/plugin.toml` plus a namespaced example Skill. Bundles are discovered
3492 read-only, untrusted, and disabled; review them through `/plugin` before
3493 enabling. v0.9.1 activates only declared Skills and MCP servers. See
3494 [PLUGIN_BUNDLES.md](PLUGIN_BUNDLES.md).
3495 - `--all` now scaffolds MCP + skills + tools + plugins together.
3496 - `--clean` — list `~/.codewhale/sessions/checkpoints/latest.json` and
3497 `offline_queue.json` if they exist. Legacy
3498 `~/.deepseek/sessions/checkpoints/` files are not scanned automatically; set
3499 `CODEWHALE_HOME=~/.deepseek` for a one-off legacy cleanup. Pass `--force` to
3500 actually remove matched files. This never touches real session history or the
3501 task queue.
3502
3503 `--status` and `--clean` are mutually exclusive with the scaffold flags.
3504
3505 ## Why the engine strips XML/`[TOOL_CALL]` text
3506
3507 codewhale sends and receives tool calls only over the API tool channel
3508 (structured `tool_use` / `tool_call` items). The streaming loop in
3509 `crates/tui/src/core/engine.rs` recognizes a fixed set of fake-wrapper start
3510 markers — `[TOOL_CALL]`, `<codewhale:tool_call`, `<tool_call`, `<invoke `,
3511 `<function_calls>` — and scrubs them from visible assistant text without ever
3512 turning them into structured tool calls. When a wrapper is stripped, the loop
3513 emits one compact `status` notice per turn so the user can see why their
3514 visible text shrank. Treat any change that re-enables text-based tool
3515 execution as a regression; the protocol-recovery tests in
3516 `crates/tui/tests/integration/protocol_recovery.rs` lock the contract.
3517
3518 ## Model-bound redaction (`[redaction] model_bound`)
3519
3520 Codewhale masks credential-looking values in tool output **before it is sent
3521 to an upstream model** — the "model boundary". A file read by a tool can
3522 contain a configured API key, a bare provider token, or a credential-shaped
3523 opaque string, and the model must not see those bytes. This backstop is
3524 separate from the display/export scrubbers: it decides what the model itself
3525 can quote back, and it is deliberately conservative (`CredentialShaped`
3526 policy, see `crates/config/src/persistence.rs`), so ordinary code and config
3527 stay byte-exact while keys, JWTs, bearer tokens, PEM blocks, and long opaque
3528 runs are masked.
3529
3530 Turning that masking **off** is a security decision, so it is not a plain
3531 boolean:
3532
3533 ```toml
3534 [redaction]
3535 model_bound = "disabled" # "enabled" (default) | "disabled"
3536 ```
3537
3538 Setting `"disabled"` only records a *request*. It takes effect only when all
3539 of these are true:
3540
3541 1. You restart the interactive TUI.
3542 2. The startup gate appears and you press `1`/`Y` on its first stage
3543 ("confirm and disable"). This only advances to a second, final-confirmation
3544 stage - the gate repeats the red warning and asks "are you really sure?".
3545 3. On that second stage you press `1`/`Y` again. The gate is rendered with the
3546 same explicit-key discipline as workspace trust - `Enter` never confirms by
3547 reflex, and `2`/`U` on the second stage steps back.
3548 4. Only that second confirmation persists a receipt to
3549 `~/.codewhale/redaction-state.json` (next to `config.toml`) and rebuilds
3550 the engine with masking off for the rest of this launch and future ones.
3551
3552 The receipt is bound to the config it was made against and is valid only
3553 while that config still requests `"disabled"`. Setting `model_bound` back
3554 to `"enabled"` - or rewriting `config.toml` in any way after the
3555 confirmation - invalidates it, so requesting `"disabled"` again later
3556 asks for a fresh confirmation on the next launch. The receipt checks both the
3557 config contents and modification time; missing, unreadable, or malformed config
3558 and older receipts without this binding keep masking enabled. Legacy-home
3559 installs store the receipt beside their resolved config file.
3560
3561 Until a confirmation exists, the effective mode is always `"enabled"`:
3562
3563 - Choosing `2`/`U` ("keep masking on") leaves the config field untouched, so
3564 the next launch asks again. Edit the field back to `"enabled"` to stop being
3565 asked.
3566 - Non-interactive entry points (`codewhale exec`, hooks, automations, headless
3567 agents) never confirm anything and never apply an unconfirmed request.
3568 - Routing/classification summaries and durable goal-state text keep their own
3569 always-on redaction regardless of this switch; the opt-out exists so the
3570 model can quote file bytes for exact edits, not to relax stored state.
3571
3572 The config value itself is forgiving: `true`/`false`, `"on"`/`"off"`, and
3573 `"enabled"`/`"disabled"` (any casing) all parse, with `false`/`"off"` meaning
3574 `"disabled"`.
3575
3576 A confirmed opt-out still sends your configured API keys to the provider you
3577 are already talking to. Only use it when the model must read and edit files
3578 that contain real credentials.
3579
3580 ### Stored sessions
3581
3582 The masking runs once, when tool output enters the transcript, so saved
3583 session files (`~/.codewhale/sessions/*.json` and their checkpoints) and
3584 Runtime API thread items (their text and tool metadata) do not store a
3585 credential-shaped value from tool output. With the confirmed opt-out above,
3586 tool output is stored as the model saw it.
3587
3588 Not yet masked: large outputs spilled to `~/.codewhale/tool_outputs`, shell
3589 completion evidence artifacts, and tool-call *inputs* (for example a
3590 `curl -H "Authorization: Bearer …"` command line). `doctor` and
3591 `scrub-secrets` do not check those either.
3592
3593 Sessions saved by builds before this change may still hold credentials in
3594 their tool output; nothing rewrites them automatically. `codewhale doctor`
3595 reports them under **Stored Sessions** (it checks the newest 50 files), and
3596 this command finds and masks them across every saved session:
3597
3598 ```sh
3599 codewhale sessions scrub-secrets # report only
3600 codewhale sessions scrub-secrets --apply # rewrite the affected files
3601 ```
3602
3603 `--apply` rewrites each file under the same per-session lock a save takes,
3604 so it never loses a concurrent save. A session that is still open can write
3605 its in-memory copy back on its next save, so close open sessions first, and
3606 rotate any credential that was exposed — masking a stored copy cannot un-leak it.
3607
3607 lines MARKDOWN