| 1 | # Configuration |
| 2 | |
| 3 | codewhale reads configuration from a TOML file plus environment variables. |
| 4 | At process startup it may also load literal built-in-provider credentials from |
| 5 | a workspace-local `.env` file. Use the tracked `.env.example` as the template; |
| 6 | copy it to `.env`, then add only credential values. |
| 7 | |
| 8 | A workspace is not configuration authority. Codewhale therefore ignores |
| 9 | config/profile/home paths, provider/model/base-URL routing, MCP/plugin state, |
| 10 | approval/sandbox/shell posture, executable paths, runtime settings, and every |
| 11 | other non-credential `.env` entry. Variable expansion is rejected so a |
| 12 | repository cannot substitute an ambient secret into a credential value. Use |
| 13 | `config.toml`, CLI flags, or values exported by the launching shell for those |
| 14 | explicit control-plane settings. `.env` is read from a stable regular-file |
| 15 | handle, is capped at 1 MiB, and symbolic links, reparse points, and multiply |
| 16 | linked files are rejected. |
| 17 | |
| 18 | ## Reading and checking configuration from the CLI |
| 19 | |
| 20 | `codewhale config get <key>` reads scalar keys, whole tables such as `tools`, |
| 21 | and nested paths such as `tools.user_input_timeout_seconds`. Displayed tables |
| 22 | and nested values apply the same recursive credential redaction as `config dump`. |
| 23 | |
| 24 | `config set` supports its named scalar keys and the provider, route, and |
| 25 | notification commands. It refuses a key that nothing reads and suggests the |
| 26 | nearest real key (`config set calm_mod on` names `calm_mode`). A settings.toml |
| 27 | key such as `calm_mode` or `tool_collapse` is validated and written to the |
| 28 | user-global settings.toml, not config.toml; `config get` reads it back from |
| 29 | there even when an old config.toml copy (which nothing reads) is still present, |
| 30 | and names that copy so you can `config unset` it. Settings keys have no project |
| 31 | scope, so `--project` refuses them. A config.toml key is checked against the |
| 32 | type its reader expects and stored as a TOML boolean or number where the |
| 33 | reader needs one (`yolo = true`, `max_subagents = 4`); `reasoning_effort` |
| 34 | accepts the same aliases as `/effort` (#6563). Other dotted writes fail before |
| 35 | modifying the file and name the TOML table to edit. For example, set a tools |
| 36 | timeout in the file as: |
| 37 | |
| 38 | ```toml |
| 39 | [tools] |
| 40 | user_input_timeout_seconds = 0 |
| 41 | ``` |
| 42 | |
| 43 | `codewhale config doctor` checks credential presence and endpoint shape, and |
| 44 | warns about each config.toml root key that nothing reads: a settings.toml key |
| 45 | left in config.toml is named as misplaced, and any other unread key gets a |
| 46 | did-you-mean. Keys read by any runtime reader are not reported. The warnings do |
| 47 | not fail the check, and a clean result does not validate every runtime setting's |
| 48 | value (#6083, #6563). |
| 49 | |
| 50 | ## Constitution, project instructions, and repo authority |
| 51 | |
| 52 | Codewhale has several instruction surfaces. They are deliberately separate so a |
| 53 | personal constitution, repo policy, project instructions, and runtime security |
| 54 | controls do not blur together. |
| 55 | |
| 56 | - **Bundled global Constitution** — the compiled base law in the binary. It is |
| 57 | the default floor for every session. |
| 58 | - **User-global constitution** — the normal guided setup output. Manage it with |
| 59 | `/constitution` or `/setup`; Codewhale stores structured data at |
| 60 | `$CODEWHALE_HOME/constitution.json` (default `~/.codewhale/constitution.json`) |
| 61 | and renders it into a separate `<codewhale_user_constitution>` prose block. |
| 62 | This can express preferences and stop conditions, but it does not change |
| 63 | runtime approval policy, sandbox, shell, network, trust, or MCP permissions. |
| 64 | - **Repo-local constitution** — optional project policy in |
| 65 | `.codewhale/constitution.json`, described below. |
| 66 | - **`AGENTS.md`** — cross-agent **project instructions** (prose). This is the |
| 67 | canonical file for "how should an agent work in this repo." Run `/init` to |
| 68 | scaffold one. `CLAUDE.md` and `.claude/instructions.md` are read as |
| 69 | compatibility fallbacks. |
| 70 | - **Memory and handoffs** — recalled state. Useful, but lower authority than |
| 71 | constitutions and project instructions. |
| 72 | |
| 73 | ### Managing the user-global constitution (`/setup` and `/constitution`) |
| 74 | |
| 75 | The bundled **working agreement** is the safe default and no longer adds a |
| 76 | required first-run screen. Customize it later through `/constitution` or the |
| 77 | progressive `/setup` guide. Provider/model readiness, workspace trust, and |
| 78 | runtime posture stay separate from this guidance. |
| 79 | |
| 80 | On the **Constitution** step: |
| 81 | |
| 82 | - **`1`–`6`** tune the guided draft. **`G`** previews it, and **`G`** again |
| 83 | ratifies and saves a fresh structured `constitution.json`. |
| 84 | - **`A`** (shown only when a provider is configured) asks your first configured |
| 85 | model to draft the constitution. Drafting is **not** saving: the draft is |
| 86 | rendered through the same preview and you still press **`G`** to ratify |
| 87 | before anything persists. |
| 88 | - **`K`** keeps your existing loaded constitution unchanged (shown only when a |
| 89 | valid file is already present). |
| 90 | - **`U`** (or `/constitution bundled`) records the bundled/default law. |
| 91 | |
| 92 | `/constitution` (alias `/law`) is the primary management surface once you are |
| 93 | set up. Subcommands: `status` (the default), `preview`, `review`, `repo` (the |
| 94 | repo-local law block), `explain`, `edit`/`guided`, `repair`, `posture`, and |
| 95 | `bundled`. Managing the constitution never changes runtime approval, sandbox, |
| 96 | shell, network, trust, default mode, or MCP authority — those stay in runtime |
| 97 | posture/config. |
| 98 | |
| 99 | Each repo can carry two distinct, complementary files: |
| 100 | |
| 101 | - **`AGENTS.md`** — ordinary project working instructions. |
| 102 | - **`.codewhale/constitution.json`** — Codewhale-specific **repo authority / |
| 103 | prioritization policy**: when local sources conflict, which should Codewhale |
| 104 | trust first, and what to verify before claiming a task is done. `.codewhale/` |
| 105 | lives inside the repo (like `.github/`). Example: |
| 106 | |
| 107 | ```json |
| 108 | { |
| 109 | "schema_version": 1, |
| 110 | "authority": [ |
| 111 | "current user request", |
| 112 | "live code and tests", |
| 113 | "GitHub issue/PR details", |
| 114 | "AGENTS.md", |
| 115 | "memory", |
| 116 | "old handoffs" |
| 117 | ], |
| 118 | "protected_invariants": [ |
| 119 | "do not break old-session transcript replay" |
| 120 | ], |
| 121 | "branch_policy": "PRs target the integration branch, not main", |
| 122 | "verification_policy": { |
| 123 | "before_claiming_done": ["run focused tests", "read changed files back"] |
| 124 | }, |
| 125 | "escalate_when": [ |
| 126 | "a destructive action was not explicitly authorized" |
| 127 | ] |
| 128 | } |
| 129 | ``` |
| 130 | |
| 131 | All fields are optional. When present, the file is rendered into the system |
| 132 | prompt as concise prose in a higher-authority block. Legacy `WHALE.md` files |
| 133 | are ignored and reported as migration-only diagnostics. |
| 134 | |
| 135 | Each `protected_invariants` entry may be either a plain string (advisory |
| 136 | prose, the historical shape) or an object carrying path globs, which is |
| 137 | additionally **mechanically enforced** in the tool gate. See |
| 138 | [Enforced repo-law invariants](#enforced-repo-law-invariants) below. |
| 139 | |
| 140 | This is the **repo-local law** layer in Codewhale's hierarchy: *bundled global |
| 141 | Constitution* → *user-global constitution* (`$CODEWHALE_HOME/constitution.json`, |
| 142 | rendered as prose) → *repo constitution* (`.codewhale/constitution.json`, this |
| 143 | file) → *AGENTS/project instructions* → *memory and handoffs* → *current |
| 144 | request and live evidence for the active turn*. Runtime policy |
| 145 | (permissions/sandbox/cost limits enforced in code) is separate from all of |
| 146 | these prompt layers. The repo constitution gives project decision rules; it |
| 147 | does not replace the bundled Constitution, the user-global constitution, or |
| 148 | the current user request. |
| 149 | |
| 150 | > **`WHALE.md` is deprecated.** It overlapped confusingly with `AGENTS.md`. |
| 151 | > Codewhale no longer reads `WHALE.md` as project or global context. If one is |
| 152 | > present, setup/context diagnostics report it as ignored so you can migrate it. |
| 153 | > Move ordinary instructions to `AGENTS.md` and Codewhale-specific authority |
| 154 | > policy to `.codewhale/constitution.json`. Personal standing guidance belongs |
| 155 | > in `/constitution` / `$CODEWHALE_HOME/constitution.json`. (The global |
| 156 | > Codewhale Constitution shipped in the model prompt is a separate thing and is |
| 157 | > unaffected.) |
| 158 | |
| 159 | ### Enforced repo-law invariants |
| 160 | |
| 161 | By default a `protected_invariants` entry is advisory prose: it is rendered into |
| 162 | the prompt as guidance the agent should honor, but nothing stops a write. An |
| 163 | entry written as an **object with `paths`** is different — it compiles into a |
| 164 | mechanical write hold that the engine's tool gate evaluates before the write |
| 165 | runs. The law becomes mechanism, not just a request. |
| 166 | |
| 167 | An enforced entry has this shape: |
| 168 | |
| 169 | ```json |
| 170 | { |
| 171 | "schema_version": 1, |
| 172 | "protected_invariants": [ |
| 173 | "Keep DeepSeek support first-class.", |
| 174 | { |
| 175 | "text": "The wire format is frozen; protocol changes need a human.", |
| 176 | "paths": ["crates/protocol/**"], |
| 177 | "action": "block" |
| 178 | }, |
| 179 | { |
| 180 | "text": "Release notes need human review.", |
| 181 | "paths": ["CHANGELOG.md"], |
| 182 | "action": "ask" |
| 183 | } |
| 184 | ] |
| 185 | } |
| 186 | ``` |
| 187 | |
| 188 | - `text` — required. The reason surfaced on the hold. An empty `text` is skipped. |
| 189 | - `paths` — workspace-relative globs (globset syntax, e.g. `crates/protocol/**`, |
| 190 | `**/secrets.toml`, `CHANGELOG.md`). An object with no usable `paths` stays |
| 191 | advisory-only despite the object shape. |
| 192 | - `action` — optional, defaults to `ask`. `ask` force-prompts in Ask and |
| 193 | Auto-Review; in Full Access it denies the protected write without opening a |
| 194 | modal. `block` **denies the write outright** in every posture. |
| 195 | |
| 196 | Semantics: |
| 197 | |
| 198 | - **Tighten-only.** The schema has no allow/widen shape, so law can only *add* |
| 199 | holds — a crafted constitution can never grant authority or weaken a gate |
| 200 | above it. |
| 201 | - **Not bypassable by mode.** Like the built-in safety floor, an `ask` hold |
| 202 | force-prompts in Ask and Auto-Review. Full Access never opens approval |
| 203 | modals, so the same hold fails closed as a hard block; `block` always denies. |
| 204 | Mode cannot turn a hold off. |
| 205 | - **Repo-local only.** Only the repo's `.codewhale/constitution.json` |
| 206 | participates. The user-global constitution stays advisory prose and never |
| 207 | reaches this mechanism. |
| 208 | - **Fails safe.** A missing file, parse error, or invalid glob degrades to |
| 209 | fewer or zero rules — never a hold on unprotected paths and never a poisoned |
| 210 | gate. Across matches the strongest action wins, so `block` outranks `ask`. |
| 211 | - **Leaves a receipt.** Every hold emits a `tool.repo_law_decision` tool-audit |
| 212 | event naming the invariant, the matched path, and the source file; the |
| 213 | approval/denial reason names the invariant too. |
| 214 | |
| 215 | **Coverage is deliberately limited.** Holds are evaluated only for the write |
| 216 | tools `write_file`, `edit_file`, `apply_patch`, and `fim_edit`, and only |
| 217 | against the filesystem targets named in their inputs (`path`/`target`/ |
| 218 | `destination`/`file_path`, `changes[].path`, and unified-diff / |
| 219 | `apply_patch`-envelope headers). A shell command that writes a protected path is **not** held by |
| 220 | repo law — those writes are still governed by the ordinary approval, sandbox, |
| 221 | and shell-write gates, not by this mechanism. |
| 222 | |
| 223 | ### Expert full base-prompt override (#3638) |
| 224 | |
| 225 | The global Constitution (the base system prompt, normally compiled in from |
| 226 | `crates/tui/src/prompts/text.rs` as `BASE_PROMPT`) can be replaced per-user |
| 227 | without rebuilding. This is |
| 228 | an expert escape hatch, not the normal `/constitution` guided setup output. |
| 229 | Because this is a prompt trust boundary, it takes **two deliberate steps** — a |
| 230 | file alone is not enough: |
| 231 | |
| 232 | 1. Drop the replacement at `~/.codewhale/prompts/constitution.md` (under |
| 233 | `$CODEWHALE_HOME` when set). |
| 234 | 2. Set the explicit opt-in flag `CODEWHALE_ALLOW_BASE_PROMPT_OVERRIDE=1` |
| 235 | (`true`/`on`/`yes` also accepted). |
| 236 | |
| 237 | If the file exists but the flag is unset, the override is **ignored** (with a |
| 238 | log line pointing to the flag) and the bundled Constitution stays in place. |
| 239 | This is intended for repurposing the TUI beyond software engineering — e.g. |
| 240 | long-form writing or document review — where the engineering-oriented base |
| 241 | prompt is a poor fit. It is loaded once at startup; a **missing or empty file |
| 242 | is a no-op**, so existing installs keep the bundled prompt. |
| 243 | |
| 244 | Scope is deliberately narrow: only the byte-stable **base prompt segment** is |
| 245 | overridable. Mode deltas, the approval policy, the tool taxonomy, Context |
| 246 | Management, and the Compaction Relay are still owned by Codewhale's runtime |
| 247 | assembly, so an override **cannot remove safety-relevant guidance** (sandbox, |
| 248 | approvals) — it only swaps the task/voice framing. To customize ordinary |
| 249 | personal behavior, prefer `/constitution`; to customize per-repo behavior, |
| 250 | prefer `AGENTS.md` + `.codewhale/constitution.json` above. |
| 251 | |
| 252 | ## Where It Looks |
| 253 | |
| 254 | Default config path: |
| 255 | |
| 256 | - `~/.codewhale/config.toml` |
| 257 | - Legacy fallback: `~/.deepseek/config.toml` |
| 258 | |
| 259 | Overrides: |
| 260 | |
| 261 | - CLI: `codewhale --config /path/to/config.toml` |
| 262 | - Env: `CODEWHALE_CONFIG_PATH=/path/to/config.toml` |
| 263 | - Legacy env alias: `DEEPSEEK_CONFIG_PATH=/path/to/config.toml` |
| 264 | |
| 265 | If both are set, `--config` wins. Environment variable overrides are applied after the file is loaded. |
| 266 | |
| 267 | ### TUI editability audit |
| 268 | |
| 269 | Inside the TUI, run `/config audit` to see which documented keys can be changed |
| 270 | from the current session, which ones can also be persisted, and which ones stay |
| 271 | file-only or restart-only. The audit includes current values for the high-impact |
| 272 | runtime controls such as `approval_policy`, `allow_shell`, |
| 273 | `stream_chunk_timeout_secs`, `base_url`, `mcp_config_path`, and the |
| 274 | `[subagents]` concurrency/depth/timeout keys. |
| 275 | |
| 276 | Use the command's "Command / reason" column as the source of truth before |
| 277 | editing by hand. For example, `/config approval_mode on-request --save` writes |
| 278 | top-level `approval_policy = "on-request"`, while provider base URLs are saved |
| 279 | but still require restarting the model client. |
| 280 | |
| 281 | ### User workspace entries |
| 282 | |
| 283 | Interactive Agent sessions expose shell tools by default with approval gating |
| 284 | unless you explicitly disable them. For a shell opt-in that should live in the |
| 285 | user's global config for noninteractive or durable-task profiles rather than in |
| 286 | the repository, add a workspace-scoped entry: |
| 287 | |
| 288 | ```toml |
| 289 | [workspace.'/absolute/path/to/project'] |
| 290 | allow_shell = true |
| 291 | ``` |
| 292 | |
| 293 | The entry applies only when the launched workspace path matches the table key. |
| 294 | The legacy `[projects."/absolute/path/to/project"]` table is also accepted for |
| 295 | this user-owned override. |
| 296 | |
| 297 | In interactive mode, the per-project overlay |
| 298 | `<workspace>/.codewhale/config.toml` is applied after this user entry. A |
| 299 | project-level `allow_shell = false` can still tighten the session; project-level |
| 300 | `allow_shell = true` is ignored. |
| 301 | |
| 302 | ### Per-project overlay (#485) |
| 303 | |
| 304 | When the TUI starts in a workspace that contains a regular-file |
| 305 | `<workspace>/.codewhale/config.toml`, the safe values declared in that file are |
| 306 | merged on top of the global config. Legacy |
| 307 | `<workspace>/.deepseek/config.toml` files are still read when the Codewhale path |
| 308 | is absent. Symlinked project config files are rejected. This lets a repo suggest |
| 309 | a model or tighten local safety posture without touching the user's |
| 310 | `~/.codewhale/config.toml`. Pass `--no-project-config` to skip the overlay for |
| 311 | one launch. |
| 312 | |
| 313 | Supported keys in the project overlay (top-level fields only): |
| 314 | |
| 315 | | Key | Effect | |
| 316 | |---|---| |
| 317 | | `model` | override `default_text_model` | |
| 318 | | `reasoning_effort` | force `"high"` / `"max"` for a complex repo | |
| 319 | | `approval_policy` | only values that tighten the user's current permission posture | |
| 320 | | `sandbox_mode` | only values that tighten the user's current sandbox posture | |
| 321 | | `max_subagents` | clamp sub-agent concurrency for a constrained repo (clamped to 1..=128) | |
| 322 | | `allow_shell` | `false` can disable shell access; `true` is ignored | |
| 323 | |
| 324 | The overlay is intentionally narrow — it covers the fields a repo |
| 325 | maintainer is most likely to want to standardize across contributors. |
| 326 | Credential, endpoint, provider-selection, MCP config, hooks, skills, |
| 327 | retry, hotbar bindings, and `instructions = [...]` settings stay user-global. |
| 328 | If a repo-local config declares `api_key`, `base_url`, `providers`, `provider`, |
| 329 | `mcp_config_path`, `notes_path`, `hotbar`, `allow_shell = true`, or `instructions`, |
| 330 | Codewhale ignores that key and keeps the user's global setting. |
| 331 | |
| 332 | The consolidated `codewhale` runtime uses one config file for DeepSeek auth |
| 333 | and model defaults. `codewhale auth set --provider deepseek` saves |
| 334 | the key to `~/.codewhale/config.toml` (migrating legacy `~/.deepseek/config.toml` |
| 335 | on first launch when needed), and `codewhale --model deepseek-v4-flash` is |
| 336 | forwarded to the TUI as `CODEWHALE_MODEL`. The dispatcher no longer writes the |
| 337 | `DEEPSEEK_*` twins of these variables; a `DEEPSEEK_*` value you set yourself is |
| 338 | still read as a legacy alias when the `CODEWHALE_*` one is unset. |
| 339 | |
| 340 | `codewhale login` signs in to the Codewhale account — it is the same browser |
| 341 | device flow as `codewhale account login`, not a provider-key command. Provider |
| 342 | credentials are configured exclusively through `codewhale auth set |
| 343 | --provider <provider>`. |
| 344 | |
| 345 | That provider credential is distinct from the optional managed-product |
| 346 | account. `codewhale account login` starts the Codewhale browser device flow; |
| 347 | `codewhale account status` and `codewhale account logout` inspect or remove the |
| 348 | session for the selected `--profile`. Account sessions prefer the OS |
| 349 | credential manager and fall back automatically to the private `0600` |
| 350 | Codewhale secrets file when no credential manager is available (headless |
| 351 | hosts, SSH, containers). |
| 352 | `codewhale account keys list|set|remove` manages the |
| 353 | signed-in account's BYOK vault without displaying secret values. The older |
| 354 | `codewhale cloud ...` spelling remains a command alias. |
| 355 | |
| 356 | ### Portable config bundles |
| 357 | |
| 358 | `codewhale config export --portable [--project] [--out FILE]` writes a |
| 359 | portable, secret-free bundle of your configuration: sorted TOML with |
| 360 | credential and machine-specific keys (API keys, base URLs, socket paths) |
| 361 | dropped, never a redacted placeholder in their place. Without `--out` the |
| 362 | bundle goes to stdout. Typed tables, arrays, numbers, booleans, and datetimes |
| 363 | remain typed. Machine-bound authority is deliberately non-portable: project |
| 364 | trust overlays, credential readers, auto-running hooks, executable LSP |
| 365 | definitions, and local path bindings are omitted rather than copied to a new |
| 366 | host. So is trust posture (`yolo`, `allow_shell`, `approval_policy`, |
| 367 | `sandbox_mode`, `sandbox_network_access`, and the sandbox read-denylist |
| 368 | paths) and local endpoints or executables (`[lifecycle_outbox]`, |
| 369 | `[extension_host]`, `[control_socket]`). URLs that carry a credential — |
| 370 | userinfo, a token-named query parameter, or a Slack/Discord/Teams webhook |
| 371 | path — are treated as credentials. |
| 372 | |
| 373 | `codewhale config import <FILE|HTTPS_URL|-> [--dry-run] [--yes] [--project]` |
| 374 | applies a bundle. The envelope is strict (`schema_version = 1`, kind |
| 375 | `codewhale.portable-config`; unknown fields fail). Import prints a |
| 376 | deterministic plan — added / changed / skipped / conflicting / rejected — |
| 377 | then asks for consent unless `--yes` is given; headless use requires it. |
| 378 | Credential-shaped entries are rejected by key name and by value shape; |
| 379 | rejections name the field, never a value. Remote bundles come from HTTPS |
| 380 | only (loopback http excepted) with a 5 MiB cap. Application backs up the |
| 381 | target document to `<config>.bundle-backup-<timestamp>-<random>`, rolls back on any |
| 382 | failure, and re-importing an applied bundle changes nothing. |
| 383 | |
| 384 | Import also rejects the non-portable authority classes omitted by export, |
| 385 | including nested/camel/dotted credential keys and cookie headers. This keeps a |
| 386 | hand-authored or remote bundle from reintroducing machine trust, local |
| 387 | credential access, or automatically executable commands that a local export |
| 388 | would refuse to carry. Structured tables are deep-merged: a portable model or |
| 389 | preference update does not erase target-local provider credentials, endpoints, |
| 390 | hooks, or executable definitions that were deliberately omitted from the |
| 391 | bundle. Arrays and scalar values still replace the corresponding portable |
| 392 | value. |
| 393 | |
| 394 | Sections map to scope: `[project]` entries land only in the workspace |
| 395 | document (`--project`, which must target an actual workspace config), |
| 396 | `[global]` only in the user-global one; `preferences`, `profiles`, and |
| 397 | `plugins` apply at either scope. A global bundle operation refuses a workspace |
| 398 | document just as a project operation refuses the user-global document. |
| 399 | |
| 400 | ### Credential read precedence (#5197) |
| 401 | |
| 402 | Credential reads are **folder-independent by default**: every layer below is |
| 403 | user-global or process-scoped, so a key saved in one repo resolves identically |
| 404 | in every other repo. Repo-local config never carries credential material — the |
| 405 | project overlay above reads only its allowlisted keys and ignores `api_key`, |
| 406 | and a credential write aimed at a workspace-scoped config is rescoped to the |
| 407 | user-global `~/.codewhale/config.toml` (#5045, #5193). |
| 408 | |
| 409 | For the active provider, the runtime resolves the API key in this exact order |
| 410 | (first match wins): |
| 411 | |
| 412 | 1. **Route-specific auth contract.** Routes whose `auth_mode` disables API |
| 413 | keys stop here with no credential. OAuth routes use their explicitly |
| 414 | granted token: the official `openai-codex` plan route accepts only |
| 415 | Codewhale-owned credentials from `codewhale auth chatgpt`, bound to its |
| 416 | issued registration and verified account. Imported Codex CLI and process |
| 417 | tokens cannot authorize this route; |
| 418 | `[providers.xai] auth_mode = "oauth"` reads Codewhale's own xAI |
| 419 | device-login store (or a consent-granted Grok CLI file). |
| 420 | 2. **Explicit CLI key.** `--api-key` forwarded with its source marker wins |
| 421 | over every saved slot. |
| 422 | 3. **Config file `api_key`.** The `[providers.<name>] api_key` table slot for |
| 423 | the active provider. `deepseek-CN` also reads `[providers.deepseek]`. An |
| 424 | older top-level `api_key` is read as `[providers.deepseek] api_key` (see |
| 425 | [Legacy top-level `base_url` and `api_key`](#legacy-top-level-base_url-and-api_key)). |
| 426 | File-owned keys stay bound to their file-owned endpoint: when the |
| 427 | environment replaces the route's base URL with a custom host, the saved |
| 428 | key is not sent there. |
| 429 | 4. **`api_key_env` binding.** `[providers.<name>] api_key_env = "VAR"` reads |
| 430 | the named environment variable. For custom providers an unset or empty |
| 431 | binding is a loud error, not a silent fallback (#5104). |
| 432 | 5. **Secret store.** The durable per-provider slot written by |
| 433 | `codewhale auth set` (file-backed under `~/.codewhale/secrets/` by |
| 434 | default; the OS keyring only when explicitly selected). Skipped for named |
| 435 | custom routes, self-hosted providers, custom endpoints other than an |
| 436 | explicitly authenticated loopback, and routes whose `auth_mode` needs no |
| 437 | key. |
| 438 | 6. **Ambient environment.** The provider's own variable |
| 439 | (`DEEPSEEK_API_KEY`, `OPENROUTER_API_KEY`, `MOONSHOT_API_KEY`, …). |
| 440 | Ambient keys are only ever sent to the provider's official endpoint and |
| 441 | are skipped under the same conditions as the secret store. |
| 442 | 7. **Keyless fallback.** Self-hosted providers and loopback endpoints may run |
| 443 | with no credential; every other route fails with provider-specific setup |
| 444 | guidance. |
| 445 | |
| 446 | Legacy compatibility: `~/.deepseek/config.toml` is migrated into |
| 447 | `~/.codewhale/config.toml` on first launch, `DEEPSEEK_*` environment |
| 448 | variables remain accepted aliases for the `CODEWHALE_*` forms, and |
| 449 | `DEEPSEEK_SECRET_BACKEND` is the legacy alias for `CODEWHALE_SECRET_BACKEND`. |
| 450 | |
| 451 | Run `codewhale auth status` to inspect the active provider's config |
| 452 | file, OS keyring backend, environment variable, winning source, and last-four |
| 453 | label without printing the key itself. The command only probes the active |
| 454 | provider's keyring entry. |
| 455 | |
| 456 | For hosted, generic OpenAI-compatible, self-hosted, OpenAI Responses, or native |
| 457 | Anthropic providers, set `provider = "<id>"` or pass |
| 458 | `codewhale --provider <id>`. The canonical provider IDs are `deepseek`, |
| 459 | `nvidia-nim`, `openai`, `atlascloud`, `wanjie-ark`, `volcengine`, |
| 460 | `openrouter`, `orcarouter`, `xiaomi-mimo`, `novita`, `fireworks`, |
| 461 | `siliconflow`, `arcee`, `siliconflow-CN`, `moonshot`, `sglang`, `vllm`, |
| 462 | `ollama`, `ollama-cloud`, `huggingface`, `modelscope`, `together`, `qianfan`, `openai-codex`, |
| 463 | `anthropic`, `openmodel`, `zai`, `stepfun`, `minimax`, `deepinfra`, |
| 464 | `sakana`, `longcat`, `opencode-go`, `opencode-zen`, `meta`, `xai`, |
| 465 | `mistral`, `telecomjs`, `modelstudio-token-plan`, `google`, |
| 466 | `edenai`, `concentrate`, `codewhale`, and `custom` (a user-defined OpenAI-compatible endpoint via |
| 467 | `[providers.<name>]`). |
| 468 | For the provider-by-provider registry, including wire protocol, auth variables, |
| 469 | default base URLs, model IDs, and capability metadata, see |
| 470 | [PROVIDERS.md](PROVIDERS.md). |
| 471 | The facade saves provider credentials to the shared user config and forwards |
| 472 | the resolved key, base URL, provider, and model to the TUI process. Use |
| 473 | `codewhale auth set --provider nvidia-nim --api-key "YOUR_NVIDIA_API_KEY"` or |
| 474 | `codewhale auth set --provider openai --api-key "YOUR_OPENAI_COMPATIBLE_API_KEY"` or |
| 475 | `codewhale auth set --provider atlascloud --api-key "YOUR_ATLASCLOUD_API_KEY"` or |
| 476 | `codewhale auth set --provider wanjie-ark --api-key "YOUR_WANJIE_API_KEY"` or |
| 477 | `codewhale auth set --provider xiaomi-mimo --api-key "YOUR_XIAOMI_KEY"` or |
| 478 | `codewhale auth set --provider fireworks --api-key "YOUR_FIREWORKS_API_KEY"` or |
| 479 | `codewhale auth set --provider siliconflow --api-key "YOUR_SILICONFLOW_API_KEY"` or |
| 480 | `codewhale auth set --provider arcee --api-key "YOUR_ARCEE_API_KEY"` or the |
| 481 | matching provider ID from [PROVIDERS.md](PROVIDERS.md) to save provider keys |
| 482 | through the facade. The generic `openai` provider defaults |
| 483 | to `https://api.openai.com/v1`, accepts `OPENAI_BASE_URL`, and defaults to |
| 484 | `gpt-5.6`. A custom OpenAI-compatible gateway can still select its own model |
| 485 | explicitly. `atlascloud` defaults to |
| 486 | `https://api.atlascloud.ai/v1`, accepts `ATLASCLOUD_BASE_URL`, and uses |
| 487 | `deepseek-ai/deepseek-v4-flash` as its default model. `wanjie-ark` targets |
| 488 | Wanjie Ark's OpenAI-compatible endpoint at |
| 489 | `https://maas-openapi.wanjiedata.com/api/v1`, defaults to `deepseek-reasoner`, |
| 490 | and passes model IDs through unchanged because Wanjie model access is |
| 491 | account-scoped. SGLang, vLLM, and Ollama are |
| 492 | self-hosted and can run without an API key by default. Ollama defaults to |
| 493 | `http://localhost:11434/v1` and sends model tags such as `codewhale-coder:1.3b` |
| 494 | or `qwen2.5-coder:7b` unchanged. Self-hosted providers and loopback custom |
| 495 | URLs (`localhost`, `127.0.0.1`, `[::1]`, `0.0.0.0`) do not read the secret store |
| 496 | unless API-key auth is explicitly requested; use an env var or config-file key |
| 497 | when a local server does require bearer auth. |
| 498 | Ollama Cloud is the separate hosted `ollama-cloud` provider. It defaults to |
| 499 | `https://ollama.com/v1` and `gpt-oss:120b`; save its key with |
| 500 | `codewhale auth set --provider ollama-cloud`. Ambient auth reads |
| 501 | `OLLAMA_CLOUD_API_KEY` first, then Ollama's official `OLLAMA_API_KEY`. |
| 502 | SiliconFlow defaults to `https://api.siliconflow.com/v1`, accepts |
| 503 | `SILICONFLOW_BASE_URL`, and uses `deepseek-ai/DeepSeek-V4-Pro` by default. |
| 504 | `provider = "siliconflow-CN"` selects the China regional default |
| 505 | `https://api.siliconflow.cn/v1` with the `[providers.siliconflow_cn]` table and |
| 506 | `SILICONFLOW_API_KEY` credential slot. |
| 507 | Arcee AI defaults to `https://api.arcee.ai/api/v1`, accepts `ARCEE_BASE_URL`, |
| 508 | and uses `trinity-large-thinking` by default for Codewhale agent work. |
| 509 | `trinity-large-preview` is also listed as a direct Arcee API model; OpenRouter's |
| 510 | `arcee-ai/trinity-large-thinking` remains the OpenRouter namespaced form, while |
| 511 | the direct Arcee provider uses the bare `trinity-large-thinking` ID. Direct |
| 512 | Arcee large-model API calls are tracked as 256K-context BF16 serving; Thinking |
| 513 | is reasoning-capable, while Preview is not marked as a thinking model. |
| 514 | |
| 515 | ### OpenRouter vendor pinning |
| 516 | |
| 517 | OpenRouter serves each model through several upstream vendors, and Codewhale |
| 518 | can pin requests to a vendor with `[providers.openrouter] vendor` (#6007): |
| 519 | |
| 520 | ```toml |
| 521 | provider = "openrouter" |
| 522 | [providers.openrouter] |
| 523 | model = "deepseek/deepseek-v4-pro" |
| 524 | vendor = "deepinfra" # copy the vendor slug from the model's OpenRouter page |
| 525 | ``` |
| 526 | |
| 527 | This sends `"provider": {"order": ["deepinfra"], "allow_fallbacks": false}` |
| 528 | on OpenRouter requests. A base slug can match multiple endpoint variants; |
| 529 | copy a full slug such as `deepinfra/turbo` to select one variant. An unavailable |
| 530 | pin fails at OpenRouter. Codewhale's separate `fallback_providers` setting can |
| 531 | still switch the whole route after a recoverable error. |
| 532 | |
| 533 | The pin applies across OpenRouter models, including auxiliary requests on that |
| 534 | route. Set `vendor = ""` to clear it. Reload config or restart to apply edits; |
| 535 | requests already in flight keep their captured route. Other providers do not |
| 536 | inherit the pin. `/preview-request` shows the primary request's routing fields. |
| 537 | |
| 538 | Model strings still pass through verbatim: `:floor` sorts by price, `:nitro` |
| 539 | sorts by throughput, and `@preset/my-team-preset` references an account preset. |
| 540 | See OpenRouter's [provider routing](https://openrouter.ai/docs/guides/routing/provider-selection) |
| 541 | and [presets](https://openrouter.ai/docs/guides/features/presets) documentation. |
| 542 | Codewhale does not fetch per-vendor endpoint prices or availability; pinned |
| 543 | usage reports a routing-dependent unknown cost instead of a catalog estimate. |
| 544 | |
| 545 | ### Custom OpenAI-Compatible Gateways |
| 546 | |
| 547 | For a single third-party service that implements the OpenAI Chat Completions |
| 548 | API, the simplest setup is the built-in `openai` provider name pointed at the |
| 549 | gateway: |
| 550 | |
| 551 | ```toml |
| 552 | provider = "openai" |
| 553 | default_text_model = "your-model-id" |
| 554 | |
| 555 | [providers.openai] |
| 556 | api_key = "YOUR_OPENAI_COMPATIBLE_API_KEY" |
| 557 | base_url = "https://your-gateway.example/v1" |
| 558 | ``` |
| 559 | |
| 560 | Put the endpoint under `[providers.openai]`, not the legacy top-level |
| 561 | `base_url`, so the OpenAI-compatible provider receives it. `default_text_model` |
| 562 | is the model ID sent to the gateway; `[providers.openai].model` can be used as |
| 563 | the OpenAI-provider-specific override. |
| 564 | |
| 565 | If you keep several OpenAI-compatible gateways, or need a stable name for an |
| 566 | AgentProfile provider pin, define a user-named custom provider table: |
| 567 | |
| 568 | ```toml |
| 569 | provider = "lm-studio" |
| 570 | |
| 571 | [providers.lm-studio] |
| 572 | kind = "openai-compatible" |
| 573 | base_url = "http://127.0.0.1:1234/v1" |
| 574 | api_key = "lm-studio" |
| 575 | model = "qwen-2.5-7b" |
| 576 | ``` |
| 577 | |
| 578 | Custom provider names may be selected with `provider = "<name>"`, |
| 579 | `--provider <name>`, or an AgentProfile `provider = "<name>"` when the matching |
| 580 | `[providers.<name>]` table exists. |
| 581 | |
| 582 | StepFun has a first-class provider entry, so keep Coding Plan credentials and |
| 583 | base URL scoped to `[providers.stepfun]`: |
| 584 | |
| 585 | ```toml |
| 586 | provider = "stepfun" |
| 587 | |
| 588 | [providers.stepfun] |
| 589 | api_key = "YOUR_STEPFUN_API_KEY" |
| 590 | base_url = "https://api.stepfun.ai/step_plan/v1" |
| 591 | model = "step-3.7-flash" |
| 592 | ``` |
| 593 | |
| 594 | `/provider` setup asks which StepFun billing route the key belongs to — |
| 595 | pay-as-you-go (`https://api.stepfun.ai/v1`) or a Step Plan subscription |
| 596 | (`https://api.stepfun.ai/step_plan/v1`) — and validates the key against the |
| 597 | endpoint you pick before saving it. The answer is written to |
| 598 | `[providers.stepfun].base_url` and nowhere else. If that key already holds a |
| 599 | base URL Codewhale does not recognize as one of those two routes, the question |
| 600 | is skipped and your value is left untouched. |
| 601 | |
| 602 | Alibaba Bailian / Model Studio DashScope Qwen routes use the same OpenAI |
| 603 | provider shape: |
| 604 | |
| 605 | ```toml |
| 606 | provider = "openai" |
| 607 | |
| 608 | [providers.openai] |
| 609 | api_key = "YOUR_DASHSCOPE_API_KEY" |
| 610 | base_url = "https://dashscope-intl.aliyuncs.com/compatible-mode/v1" |
| 611 | model = "qwen-plus" |
| 612 | context_window = 1000000 |
| 613 | ``` |
| 614 | |
| 615 | Use the regional DashScope `compatible-mode/v1` base URL that matches the |
| 616 | region of your API key. Codewhale keeps `qwen-plus` scoped to the `openai` |
| 617 | provider route and does not infer a different provider from the model prefix. |
| 618 | The same rule applies to all provider-prefixed model strings: a prefix such as |
| 619 | `deepseek-ai/...` or `deepseek/...` is a provider-owned wire ID under the |
| 620 | selected provider, not an automatic switch to the DeepSeek provider. |
| 621 | Set `context_window` to the gateway/model's real total context window when it |
| 622 | differs from Codewhale's static model metadata. See |
| 623 | [Context length (context window)](#context-length-context-window) for the full |
| 624 | resolution order and for how to check which value is in effect. |
| 625 | |
| 626 | If the gateway accepts `POST /chat/completions` but rejects |
| 627 | `/v1/chat/completions`, set a provider-local `path_suffix`: |
| 628 | |
| 629 | ```toml |
| 630 | [providers.openai] |
| 631 | base_url = "https://your-gateway.example/v1" |
| 632 | path_suffix = "/chat/completions" |
| 633 | ``` |
| 634 | |
| 635 | The suffix applies only to chat-completion requests. Model listing and |
| 636 | DeepSeek beta paths keep their built-in routing so a generic gateway override |
| 637 | does not accidentally rewrite `/models` or `/beta/completions`. |
| 638 | |
| 639 | For private gateways with broken or intercepted certificates, use |
| 640 | `SSL_CERT_FILE` with a trusted CA bundle. The legacy provider-table key |
| 641 | `insecure_skip_tls_verify = true` is still parsed so `codewhale doctor` can |
| 642 | report stale configs, but provider clients reject it instead of disabling TLS |
| 643 | certificate verification. |
| 644 | |
| 645 | Local HTTP endpoints such as Ollama, SGLang, and vLLM are allowed by default |
| 646 | when they use localhost or loopback addresses. For a non-local `http://` |
| 647 | gateway, launch with `DEEPSEEK_ALLOW_INSECURE_HTTP=1` only on a trusted network: |
| 648 | |
| 649 | ```bash |
| 650 | DEEPSEEK_ALLOW_INSECURE_HTTP=1 codewhale |
| 651 | ``` |
| 652 | |
| 653 | Third-party OpenAI-compatible gateways that need extra request headers can set |
| 654 | `http_headers = { "X-Model-Provider-Id" = "your-model-provider" }` at the top |
| 655 | level or under a provider table such as `[providers.deepseek]`. When configured, |
| 656 | codewhale sends those custom headers on model API requests. The equivalent |
| 657 | environment override is `DEEPSEEK_HTTP_HEADERS`, using comma-separated |
| 658 | `name=value` pairs such as |
| 659 | `X-Model-Provider-Id=your-model-provider,X-Gateway-Route=dev`. `Authorization` |
| 660 | and `Content-Type` are managed by the client and are not overridden by this |
| 661 | setting. |
| 662 | |
| 663 | ### Vision Model |
| 664 | |
| 665 | Codewhale's chat provider and `image_analyze` tool are configured separately. |
| 666 | The main chat path remains the selected text/tool provider; image analysis runs |
| 667 | through `[vision_model]` when the `vision_model` feature is enabled. |
| 668 | |
| 669 | Xiaomi's current image-understanding docs include `mimo-v2.5` for image input. |
| 670 | To use MiMo for `image_analyze`, configure the vision model explicitly: |
| 671 | |
| 672 | ```toml |
| 673 | [features] |
| 674 | vision_model = true |
| 675 | |
| 676 | [vision_model] |
| 677 | model = "mimo-v2.5" |
| 678 | api_key = "YOUR_XIAOMI_KEY" |
| 679 | base_url = "https://api.xiaomimimo.com/v1" |
| 680 | ``` |
| 681 | |
| 682 | The example above uses Xiaomi MiMo's pay-as-you-go OpenAI-compatible endpoint. |
| 683 | If you are using a Token Plan key (`tp-...`) for `[vision_model]`, you must set |
| 684 | `base_url` explicitly because this generic OpenAI-compatible block does not |
| 685 | auto-select MiMo endpoints. Use |
| 686 | `https://token-plan-sgp.xiaomimimo.com/v1` for Singapore accounts, |
| 687 | `https://token-plan-cn.xiaomimimo.com/v1` for China-region accounts, or |
| 688 | `https://token-plan-ams.xiaomimimo.com/v1` for Europe/Amsterdam accounts. |
| 689 | |
| 690 | ### Auto Model Routing (`[auto]`, `[auto.router]`) |
| 691 | |
| 692 | With `model = "auto"`, each turn runs on your **declared default model** unless |
| 693 | you have opted into something else. Auto never guesses a cheaper or stronger |
| 694 | model from how a request is worded; the old keyword-and-length heuristic was |
| 695 | removed (`auto_route_declared_fallback` in `crates/tui/src/model_routing.rs`). |
| 696 | Two optional layers change that: |
| 697 | |
| 698 | - an `[auto.router]` classifier, which you write down, that picks a model per |
| 699 | turn; and |
| 700 | - `[auto] cost_saving`, which prefers the active provider's fast sibling. |
| 701 | |
| 702 | With neither set, Auto is local and free: the turn uses the default model and |
| 703 | no classifier call is made. |
| 704 | |
| 705 | **There is no default classifier.** With `[auto.router]` unset, no classifier |
| 706 | call happens, whatever keys you hold. Holding a DeepSeek key used to elect |
| 707 | `deepseek-v4-flash` automatically. That was removed because it spent tokens on |
| 708 | a route the user never chose and privileged one provider over the rest |
| 709 | (`AutoRouterConfig` in `crates/tui/src/config.rs`). Electing a network |
| 710 | classifier is now something you write down. |
| 711 | |
| 712 | Point the classifier at any configured provider with `[auto.router]`: |
| 713 | |
| 714 | ```toml |
| 715 | [auto.router] |
| 716 | provider = "zai" |
| 717 | model = "glm-5-turbo" |
| 718 | thinking = "off" # optional; defaults to off |
| 719 | timeout_secs = 4 # optional; default 4, 0 = default, capped at 300 |
| 720 | ``` |
| 721 | |
| 722 | A classifier call happens only when `[auto.router]` names both `provider` and |
| 723 | `model` *and* that provider has a key: |
| 724 | `router_available = router_configured && has_api_key_for(...)` in |
| 725 | `ModelInventory::from_config` (`crates/tui/src/model_inventory.rs`). If either |
| 726 | condition fails, or the classifier call errors or times out, the local |
| 727 | fallback decides: the default model, or the fast sibling under `cost_saving`. |
| 728 | The turn's route receipt (Turn Inspector, Ctrl+Alt+O or `/turn inspect`, "Model |
| 729 | route + tokens/cost") records which path was taken, and a |
| 730 | router you configured that cannot run or fails (missing key, HTTP error, |
| 731 | timeout, invalid answer) is shown as `Auto router: failing — …` rather than |
| 732 | silently ignored. |
| 733 | |
| 734 | #### Set up model routing |
| 735 | |
| 736 | `/router` (also `/model router`) opens one Router setup view with these |
| 737 | presets. Each writes only `[auto.router]`; none is ever chosen for you. |
| 738 | |
| 739 | | Preset | What it writes | Cost and privacy | |
| 740 | | --- | --- | --- | |
| 741 | | `/router jev` | Jev, TypeSafe's decision model, over OpenRouter (`typesafe/jev-1.13`) or TypeSafe direct, whichever key you have | About $0.00002 per turn ($0.042 per million input tokens, output free). Your latest request and up to six recent context lines go to OpenRouter → TypeSafe (or TypeSafe). | |
| 742 | | `/router fast` | The active provider's runnable fast tier with thinking off | Your existing key; the classifier sees the same request text. | |
| 743 | | `/router off` | Removes `[auto.router]` | No router call; Auto turns use the default model, or the fast tier while `[auto] cost_saving = true` (Off leaves that setting alone). | |
| 744 | | `/router custom` | Nothing; prints the TOML to edit | — | |
| 745 | |
| 746 | Choosing Jev or Fast makes **one test call** with a fixed sample request and |
| 747 | shows the tier it picked, the probabilities and confidence, the latency and the |
| 748 | provider-reported cost. `Enter` (or `/router save <preset>`) then saves through |
| 749 | the normal config writer; `Esc` discards it. TypeSafe paused new signups on |
| 750 | 2026-09-22, so OpenRouter is the default route for new users. A TypeSafe key is |
| 751 | read from `TYPESAFE_API_KEY`, the `typesafe` secret-store entry, or |
| 752 | `[providers.typesafe] api_key` / `api_key_env`. |
| 753 | |
| 754 | #### Decision routers (`kind = "decision"`) |
| 755 | |
| 756 | A decision router asks a non-generative decision model typed Choice questions |
| 757 | for the active provider's `fast` and `strong` tiers and thinking level, then |
| 758 | reads calibrated probabilities. No prose is parsed. |
| 759 | |
| 760 | ```toml |
| 761 | [auto.router] |
| 762 | kind = "decision" # default "chat" |
| 763 | provider = "openrouter" # or "typesafe" |
| 764 | model = "typesafe/jev-1.13" # "~typesafe/jev-latest" also works; TypeSafe direct: "jev-latest" |
| 765 | timeout_secs = 2 |
| 766 | min_confidence = 0.5 # default 0.5, clamped to 0..1 |
| 767 | ``` |
| 768 | |
| 769 | - The router is called only when the active provider has a runnable strong/fast |
| 770 | pair; otherwise there is no call and no spend. |
| 771 | - An answer below `min_confidence` takes the local fallback. Under |
| 772 | `[auto] cost_saving`, a `strong` answer also needs a probability of at least |
| 773 | 0.75, or the turn stays on the fast tier. |
| 774 | - An unknown `kind`, or a decision `provider` other than `openrouter` / |
| 775 | `typesafe`, leaves the router unconfigured and shown as failing. |
| 776 | - `[auto.router] thinking` is ignored for decision routers. Both routes settle tokens through |
| 777 | the originating session's routed-usage ledger. TypeSafe is a named Custom route |
| 778 | with unknown billing; its price is never borrowed from the active chat provider. |
| 779 | Provider-reported cost remains verbatim on the decision receipt. |
| 780 | - OpenRouter uses `POST /api/alpha/decisions`; TypeSafe uses `POST /v1/systemone`. |
| 781 | The shared client validates Choice, Noul and Score against the offered questions |
| 782 | and bounds responses to 256 KiB. Malformed answers fail open with usage retained. |
| 783 | |
| 784 | #### Shadow Decision Gate (experimental, off by default) |
| 785 | |
| 786 | The Superfast Decision Gate asks a decision model three typed questions about |
| 787 | the latest user message — does it need a tool, can it be answered from the |
| 788 | conversation, and what is its intent — and logs a conservative recommendation. |
| 789 | It is shadow-only: it never changes routing, never skips or delays the model |
| 790 | call, and fails open on any error, timeout or malformed answer. It uses the same |
| 791 | System One client as the decision router above; there is no separate HTTP |
| 792 | client. It is configured from the environment and reads it when a turn starts: |
| 793 | |
| 794 | ```sh |
| 795 | SUPERFAST_ENABLED=1 # off unless set |
| 796 | SUPERFAST_PROVIDER=typesafe # or openrouter; required when enabled |
| 797 | SUPERFAST_BASE_URL=http://localhost:8000/v1 # optional TypeSafe-route base, e.g. self-hosted |
| 798 | SUPERFAST_MODEL=jev-latest # default jev-latest / ~typesafe/jev-latest |
| 799 | SUPERFAST_TIMEOUT_MS=150 # 1..=10000, default 150 |
| 800 | ``` |
| 801 | |
| 802 | - Enabling the gate never picks an endpoint by itself: without |
| 803 | `SUPERFAST_PROVIDER` nothing is sent and a warning is logged. |
| 804 | - The key comes from the same place the decision router reads it. The TypeSafe |
| 805 | route always authenticates, so a self-hosted server that ignores auth still |
| 806 | needs a placeholder `TYPESAFE_API_KEY`. |
| 807 | - Only the latest user message is sent, truncated to 4,000 characters and |
| 808 | redacted of configured secrets. The log (target `superfast`) carries the |
| 809 | route, failure class and latency, never prompt text. |
| 810 | - The gate retains the originating turn's accounting owner and cancellation. |
| 811 | Its tokens settle through the shared ledger; missing usage or a cancelled/timed |
| 812 | out request after dispatch creates an explicit coverage gap. Unknown TypeSafe |
| 813 | pricing is recorded as unpriced rather than free. |
| 814 | - Provider-reported cost survives rejected answers and late responses in bounded |
| 815 | receipts on the originating turn or session. These are diagnostic evidence; |
| 816 | unknown pricing never becomes an authoritative dollar total. Incomplete token |
| 817 | counters record a coverage gap while preserving the reported raw evidence. |
| 818 | |
| 819 | The wire contracts are documented in [TypeSafe's OpenAPI schema](https://api.typesafe.ai/openapi.json) |
| 820 | and [OpenRouter's Decisions examples](https://openrouter.ai/blog/insights/what-is-jev/). |
| 821 | |
| 822 | The Decision Gate concept and reference implementation are by Andrea Bruno, |
| 823 | released under CC BY 4.0: |
| 824 | [harness-superfast](https://github.com/Andrea-Bruno/harness-superfast). |
| 825 | |
| 826 | Two `[auto]` keys shape routing (`AutoConfig` in `crates/tui/src/config.rs`): |
| 827 | |
| 828 | ```toml |
| 829 | [auto] |
| 830 | cost_saving = false # default false |
| 831 | cross_provider = false # default false |
| 832 | ``` |
| 833 | |
| 834 | - **`cost_saving`** (default `false`). Without a classifier, Auto pins the |
| 835 | active provider's validated fast sibling instead of the default model. A |
| 836 | provider with no runnable fast sibling stays on the default. With a |
| 837 | classifier, the classifier is told to prefer the fast tier for routine or |
| 838 | ambiguous work and to pick the strong tier only for clearly agentic, |
| 839 | multi-step, architecture, security or debugging work. Cost-saving never |
| 840 | switches provider just to save money. |
| 841 | - **`cross_provider`** (default `false`). Auto stays on the provider the session |
| 842 | is configured to use. The classifier is only shown that provider's models, |
| 843 | and the fallback never leaves it. Setting `cross_provider = true` lets the |
| 844 | classifier choose among every runnable provider. There is no interactive |
| 845 | toggle for `cross_provider`; it has to be set in config. A decision router |
| 846 | always chooses within the active provider. |
| 847 | |
| 848 | To bootstrap MCP and skills directories at their resolved paths, run `codewhale setup`. |
| 849 | To only scaffold MCP, run `codewhale mcp init`. |
| 850 | |
| 851 | Note: `setup`, `doctor`, `mcp`, `features`, `sessions`, `resume`/`fork`, `exec`, |
| 852 | `review`, and `eval` are all available from the installed `codewhale` command. |
| 853 | The consolidated dispatcher also provides `auth`, `config`, `model`, `thread`, `sandbox`, |
| 854 | `app-server`, `mcp-server`, `completions`, `login`/`logout`, `account`, |
| 855 | `metrics`, `update`, `lane`, `workflow`, and `web`. Plain prompts enter the |
| 856 | in-process TUI runtime. Release installers expose the same bytes as `codew`. |
| 857 | |
| 858 | ### Startup Update Checks |
| 859 | |
| 860 | By default, the TUI starts a background check for the latest stable Codewhale |
| 861 | release and shows a short toast only when a newer release is available and the |
| 862 | official release assets are complete. The check never blocks startup, never |
| 863 | blocks a turn, and fails silently when offline. |
| 864 | |
| 865 | Disable the startup check entirely for air-gapped, corporate-proxy, or managed |
| 866 | desktop environments: |
| 867 | |
| 868 | ```toml |
| 869 | [update] |
| 870 | check_for_updates = false |
| 871 | ``` |
| 872 | |
| 873 | #### Throttling |
| 874 | |
| 875 | The answer is cached in `~/.codewhale/update-check.json` and reused for |
| 876 | `check_interval_hours` (default `1`). Only the *network request* is throttled — |
| 877 | the notice still appears on every launch while an update is outstanding. Set `0` |
| 878 | to check on every launch. |
| 879 | |
| 880 | ```toml |
| 881 | [update] |
| 882 | check_interval_hours = 1 |
| 883 | ``` |
| 884 | |
| 885 | A failed check is not cached, so an outage does not suppress the notice until |
| 886 | the interval elapses. |
| 887 | |
| 888 | #### Automatic suppression |
| 889 | |
| 890 | Checks are skipped, without contacting the network, when any of these is set to |
| 891 | a non-falsey value: |
| 892 | |
| 893 | | Variable | Why | |
| 894 | | --- | --- | |
| 895 | | `CODEWHALE_NO_UPDATE_CHECK` | Explicit opt-out. | |
| 896 | | `NO_UPDATE_NOTIFIER` | The cross-CLI convention, honored for compatibility. | |
| 897 | | `CI`, `CONTINUOUS_INTEGRATION`, `GITHUB_ACTIONS`, `GITLAB_CI`, `BUILDKITE`, `CIRCLECI`, `JENKINS_URL`, `TEAMCITY_VERSION`, `TF_BUILD` | Automated build; nobody is at the terminal. | |
| 898 | |
| 899 | Values of `""`, `0`, `false`, `no`, and `off` do not count as set, so a |
| 900 | `CI=false` export does not disable checks for ordinary users. |
| 901 | |
| 902 | #### Which update command is offered |
| 903 | |
| 904 | Codewhale never installs anything on its own — it only tells you an update |
| 905 | exists. The command it names depends on how the running binary was installed, |
| 906 | detected from its path: |
| 907 | |
| 908 | | Install | Command offered | |
| 909 | | --- | --- | |
| 910 | | GitHub release binary (including Termux) | `codewhale update` | |
| 911 | | npm (`node_modules` on the path) | `npm install -g codewhale@latest` | |
| 912 | | Homebrew (`Cellar` / `linuxbrew` prefix) | `brew upgrade codewhale` | |
| 913 | | `cargo install` (`~/.cargo/bin`) | `cargo install codewhale-cli --locked --force` | |
| 914 | |
| 915 | For package-managed installs the notice also warns against `codewhale update`: |
| 916 | replacing a binary Homebrew or npm owns leaves the manager describing a version |
| 917 | that is no longer on disk, and the next upgrade silently reverts you. |
| 918 | |
| 919 | Override the detection with `CODEWHALE_INSTALL_METHOD=npm|homebrew|cargo|binary` |
| 920 | if you relocated the binary somewhere the path heuristics cannot read. |
| 921 | |
| 922 | To redirect the startup check, set `update_uri` to an internal endpoint that |
| 923 | returns GitHub-compatible latest-release JSON. Minimal mirror metadata with a |
| 924 | `tag_name` field is accepted; if `assets` are present, Codewhale requires the |
| 925 | same uploaded asset set as the official release before showing the toast. |
| 926 | |
| 927 | ```toml |
| 928 | [update] |
| 929 | check_for_updates = true |
| 930 | update_uri = "https://internal.mirror.example/codewhale/releases/latest" |
| 931 | ``` |
| 932 | |
| 933 | When `update_uri` is not set, startup checks honor release mirror environment |
| 934 | variables such as `CODEWHALE_RELEASE_BASE_URL` before falling back to the |
| 935 | official GitHub API endpoint. If a configured `update_uri` cannot be fetched or |
| 936 | parsed and a release mirror env var is set, the TUI falls back to that mirror |
| 937 | instead of failing startup. |
| 938 | |
| 939 | ## Workshop output budgets |
| 940 | |
| 941 | By default an oversized tool result uses bounded spillover: a result |
| 942 | under the byte threshold stays inline, a larger one gets a head/tail |
| 943 | preview plus a session artifact the model can read back. There is no |
| 944 | synthesis sub-agent and no per-call `raw = true` escape. |
| 945 | |
| 946 | `[workshop] large_output_threshold_tokens` and |
| 947 | `[workshop.per_tool_thresholds]` (exact model-visible tool names such as |
| 948 | `bash`) only take effect when the process opts in to adaptive evidence |
| 949 | routing with `CODEWHALE_ADAPTIVE_OUTPUT_ROUTING=1`; they set where a result |
| 950 | stops being inline and becomes handle-only evidence. |
| 951 | |
| 952 | Two optional byte ceilings (#5367) apply either way. They raise the |
| 953 | model-visible floor and never lower it: |
| 954 | |
| 955 | - `read_result_max_bytes` — cap for a single `read` / `read_file` |
| 956 | result. Absent keeps the compile-time defaults (100000 bytes for |
| 957 | `read`, which has no line cap; 16KiB / 500 lines for `read_file`). |
| 958 | For `read` this is the middle of a three-layer budget: the model's |
| 959 | own per-call `max_bytes` (hard maximum 500000) raises the budget for |
| 960 | one call, this setting raises the floor for the whole process, and |
| 961 | either way 2MiB is the absolute ceiling. Highest wins; neither layer |
| 962 | can lower a budget the other granted. |
| 963 | - `tool_result_max_bytes` — cap for a generic tool result after |
| 964 | spillover. Absent keeps the 12K-character compact floor (48K on |
| 965 | windows ≥500K tokens). Hard cap is 2MiB. |
| 966 | |
| 967 | ## Context length (context window) |
| 968 | |
| 969 | Also called context size, context limit, max context, or window. This is the |
| 970 | total token window Codewhale budgets against, and it drives the header/footer |
| 971 | context percent, the auto-compaction trigger, context-pressure checks, and the |
| 972 | request output cap. If Codewhale compacts at 128K on a model you know serves a |
| 973 | 1M window, this is the setting to change (#5134). |
| 974 | |
| 975 | **See what is in effect, and where the value came from.** Every one of these |
| 976 | prints the resolved window *and* its source: |
| 977 | |
| 978 | - `/status` — a `Context window:` row with the percent and token counts, and a |
| 979 | `Window source:` row naming the provenance and the exact key that overrides |
| 980 | it. |
| 981 | - `/config` → Provider — `Context window` (your override, or `(not set)`) and |
| 982 | `Effective context window` (`1048576 tokens · configured`). Typing |
| 983 | `context length` in the `/config` filter jumps straight to them. |
| 984 | - `/context report` — `Window: 1048576 tokens (12.4% used, ...; source: configured)`. |
| 985 | - `/context json` — machine-readable `context_window_tokens` and |
| 986 | `context_window_source`. |
| 987 | |
| 988 | **Change it** with the provider-table key `context_window`: |
| 989 | |
| 990 | ```toml |
| 991 | [providers.moonshot] |
| 992 | context_window = 1048576 |
| 993 | ``` |
| 994 | |
| 995 | or from the CLI: |
| 996 | |
| 997 | ```bash |
| 998 | codewhale config set providers.moonshot.context_window 1048576 |
| 999 | codewhale config unset providers.moonshot.context_window # back to automatic |
| 1000 | ``` |
| 1001 | |
| 1002 | Use the table for the provider you are actually on (`providers.openai`, |
| 1003 | `providers.deepseek`, `providers.moonshot`, …); `/status` names it for you. The |
| 1004 | value is a positive token count for the route's *total* window. |
| 1005 | |
| 1006 | When one gateway fronts models with heterogeneous windows, scope the override |
| 1007 | to an exact wire model id with `[providers.<name>.model_context_windows]`: |
| 1008 | |
| 1009 | ```toml |
| 1010 | [providers.command_code] |
| 1011 | context_window = 204800 |
| 1012 | |
| 1013 | [providers.command_code.model_context_windows] |
| 1014 | "MiniMaxAI/MiniMax-M2.5" = 204800 |
| 1015 | "google/gemini-3.1-flash-lite" = 1000000 |
| 1016 | ``` |
| 1017 | |
| 1018 | Keys are the exact wire model ids the route sends (dotted and `org/model` |
| 1019 | spellings both work as TOML keys when quoted); each value must be a positive |
| 1020 | token count. A matching entry beats the provider-level `context_window` for |
| 1021 | that model only — every other model on the provider still resolves against |
| 1022 | `context_window` and the rungs below. From the CLI: |
| 1023 | |
| 1024 | ```bash |
| 1025 | codewhale config set 'providers.command_code.model_context_windows."MiniMaxAI/MiniMax-M2.5"' 204800 |
| 1026 | codewhale config unset 'providers.command_code.model_context_windows."MiniMaxAI/MiniMax-M2.5"' |
| 1027 | ``` |
| 1028 | |
| 1029 | ### How the effective window is resolved |
| 1030 | |
| 1031 | First match wins, and the source label each surface prints is exactly this |
| 1032 | rung: |
| 1033 | |
| 1034 | 1. `configured (per-model)` — a `[providers.<name>.model_context_windows]` |
| 1035 | entry keyed by the route's exact wire model id. A hard override for that |
| 1036 | model only; it never rewrites another model's window. |
| 1037 | 2. `configured` — `[providers.<name>] context_window` in `config.toml`. A hard |
| 1038 | override for every model on the provider: nothing below it can raise or |
| 1039 | lower the result. Read-time aliases: |
| 1040 | `contextWindow`, `context_window_tokens`, `contextWindowTokens`, |
| 1041 | `context_length`, `contextLength`. |
| 1042 | 3. `provider-reported` — route-scoped 1M metadata a provider actually reported |
| 1043 | for the Kimi Code `k3` route, when it was observed within the last 24 hours. |
| 1044 | 4. `static Kimi Code safe floor` — 262,144 tokens for Kimi Code memberships, |
| 1045 | because 1M access is plan-gated (Allegretto and above). |
| 1046 | 5. `catalog` — the bundled route catalog (hand-curated offerings first, then |
| 1047 | the bundled Models.dev rows). The official `openai-codex` roster supplies |
| 1048 | account-specific model choices, not context-window metadata. Refresh it |
| 1049 | with `codewhale models --update --provider openai-codex`; Codewhale does |
| 1050 | not read `$CODEX_HOME` for this route. |
| 1051 | 6. `model-name hint` — an `_Nk` suffix parsed from the model name itself |
| 1052 | (`qwen3-32b-256k` → 256,000), vendor-agnostic. A naming convention the |
| 1053 | serving engine may not honor is not a fact about the route, so this rung |
| 1054 | sits *below* the catalog: any catalog row for the same id beats it (#5441). |
| 1055 | 7. `fallback` — the static per-provider capability table: 200,000 for |
| 1056 | Anthropic-wire routes, 128,000 for `openai-codex`, 8,192 for Ollama, |
| 1057 | otherwise Codewhale's static per-model metadata, and finally 128,000 when |
| 1058 | the model is unknown. |
| 1059 | |
| 1060 | ### What "(unverified)" means |
| 1061 | |
| 1062 | The `model-name hint` and `fallback` rungs still drive real budgets — the |
| 1063 | compaction trigger, the context meter, and the output reservation all use the |
| 1064 | number — but they are guesses, not capabilities anyone checked. Every surface |
| 1065 | that renders one of these windows appends `(unverified)` to its source label |
| 1066 | (the status line, the context-pressure message, `/status`, `/config`, and the |
| 1067 | model picker chip), so a window you did not configure and no provider reported |
| 1068 | can never read as a verified limit (#5239, #5441). The `context_window` and |
| 1069 | `model_context_windows` provider-table keys above are the fix: a configured |
| 1070 | window is a hard override and renders as `configured` (or |
| 1071 | `configured (per-model)`) with no marker. |
| 1072 | |
| 1073 | Output ceilings follow the same rule (#5440): an Anthropic-family model the |
| 1074 | catalog does not describe keeps the 64K Messages floor as its clamp, labeled |
| 1075 | `unverified` (or an "assumed floor") instead of `documented`. The official |
| 1076 | ChatGPT plan preview does not accept `max_output_tokens`; Codewhale sends no |
| 1077 | output-token cap on that route. A local request budget is not an upstream |
| 1078 | limit or a guarantee about the completed response length. |
| 1079 | |
| 1080 | There is no environment variable for the context window; the provider-table |
| 1081 | `context_window` and per-model `model_context_windows` keys are the user |
| 1082 | knobs. They are the right ones to set when a gateway or self-hosted runtime |
| 1083 | serves a window Codewhale's catalog does not model — per-model when only some |
| 1084 | of a provider's routes differ, provider-wide when they all do. Codewhale will |
| 1085 | not invent a window it cannot justify — it falls back to a conservative value, |
| 1086 | labels it `fallback`, and marks it `(unverified)` at every surface that shows |
| 1087 | it. |
| 1088 | |
| 1089 | ### Adjacent knobs |
| 1090 | |
| 1091 | - `auto_compact_threshold_percent` (settings.toml; also accepted as |
| 1092 | `auto_compact_threshold`; `10`–`100`, default `80`): the share of the full |
| 1093 | route context window at which auto-compaction fires, clamped so it can never |
| 1094 | cross the spendable input ceiling after output reservation and headroom. |
| 1095 | Editable from `/config`. Raising the window without touching this raises the |
| 1096 | absolute compaction point along with it. |
| 1097 | - `auto_compact` (settings.toml, on/off): turns automatic compaction off |
| 1098 | entirely; `/compact` and Ctrl+L stay available. |
| 1099 | - `[compaction] summary_instructions` and |
| 1100 | `[compaction] retained_user_message_tokens` (config.toml): standing |
| 1101 | summarizer instructions and the verbatim user-message retention budget. See |
| 1102 | the `compaction.*` entry in the config-key reference below. |
| 1103 | - `CODEWHALE_MAX_OUTPUT_TOKENS` (environment variable; legacy alias |
| 1104 | `DEEPSEEK_MAX_OUTPUT_TOKENS`): overrides the requested output cap. Without an |
| 1105 | override, Codewhale starts at the safe `65536` request cap and intersects it |
| 1106 | with any smaller documented model or route ceiling; a catalog `max_output` |
| 1107 | such as DeepSeek V4's 384K remains a capability ceiling, not the amount every |
| 1108 | response requests. Explicit overrides are preserved within the resolved |
| 1109 | route context window and any route output ceiling, and preflight/emergency |
| 1110 | budgeting reserves the same effective value that can reach the wire. A |
| 1111 | separately documented route input ceiling also clamps preflight and |
| 1112 | compaction even when the total context window is larger. A blank canonical |
| 1113 | variable falls through to a nonblank legacy value; a nonblank invalid or |
| 1114 | zero canonical value is authoritative and falls back to the safe automatic |
| 1115 | default instead of activating a stale legacy setting. There is no |
| 1116 | `max_output_tokens` key in `config.toml`. |
| 1117 | |
| 1118 | Before compaction replaces conversation history, Codewhale durably saves the |
| 1119 | original messages to the session's `artifacts/context-transfer-<id>.json` and, |
| 1120 | when a model summary is produced, its handoff to the matching `.md` file. |
| 1121 | These use the existing session artifact store and persistence redaction. |
| 1122 | A failed write aborts compaction without replacing context. Pruning-only passes |
| 1123 | save the original messages without making an extra model call. Pressure metadata |
| 1124 | shows estimated input tokens and the configured trigger; it is an estimate, not |
| 1125 | an exact promise about a provider's remaining context. |
| 1126 | |
| 1127 | Compaction history is available through `codewhale metrics` (or `--json`) and |
| 1128 | `audit.log` in the Codewhale home. Completed passes record their trigger, |
| 1129 | summary/pruning path, message and estimated-token counts, effective threshold, |
| 1130 | and summarizer token usage. Automatic refusals are recorded once per turn with |
| 1131 | their reason. These are local diagnostics, not provider invoice totals; earlier |
| 1132 | artifacts are not retroactively counted. A text-mode `exec` that attempts |
| 1133 | compaction saves its owning session at turn completion so its recovery artifacts |
| 1134 | remain discoverable. A killed process may leave artifacts without that final |
| 1135 | session snapshot; the audit writer reports I/O failures instead of inventing data. |
| 1136 | |
| 1137 | See [Settings File](#settings-file-persistent-ui-preferences) for the |
| 1138 | compaction settings and [Token Quantities and |
| 1139 | Drivers](#token-quantities-and-drivers) for what each displayed token number |
| 1140 | actually measures. |
| 1141 | |
| 1142 | ## Profiles |
| 1143 | |
| 1144 | You can define multiple profiles in the same file: |
| 1145 | |
| 1146 | ```toml |
| 1147 | default_text_model = "deepseek-flash" |
| 1148 | |
| 1149 | [providers.deepseek] |
| 1150 | api_key = "PERSONAL_KEY" |
| 1151 | |
| 1152 | [profiles.work.providers.deepseek] |
| 1153 | api_key = "WORK_KEY" |
| 1154 | base_url = "https://api.deepseek.com/beta" |
| 1155 | |
| 1156 | [profiles.nvidia-nim] |
| 1157 | provider = "nvidia-nim" |
| 1158 | default_text_model = "deepseek-ai/deepseek-v4-pro" |
| 1159 | |
| 1160 | [profiles.nvidia-nim.providers.nvidia_nim] |
| 1161 | api_key = "NVIDIA_KEY" |
| 1162 | base_url = "https://integrate.api.nvidia.com/v1" |
| 1163 | |
| 1164 | [profiles.fireworks] |
| 1165 | provider = "fireworks" |
| 1166 | default_text_model = "accounts/fireworks/models/deepseek-v4-pro" |
| 1167 | |
| 1168 | [profiles.siliconflow] |
| 1169 | provider = "siliconflow" |
| 1170 | default_text_model = "deepseek-ai/DeepSeek-V4-Pro" |
| 1171 | |
| 1172 | [profiles.siliconflow.providers.siliconflow] |
| 1173 | base_url = "https://api.siliconflow.com/v1" |
| 1174 | |
| 1175 | [profiles.openai-compatible] |
| 1176 | provider = "openai" |
| 1177 | |
| 1178 | [profiles.openai-compatible.providers.openai] |
| 1179 | base_url = "https://openai-compatible.example/v4" |
| 1180 | model = "glm-5" |
| 1181 | |
| 1182 | [profiles.atlascloud] |
| 1183 | provider = "atlascloud" |
| 1184 | |
| 1185 | [profiles.atlascloud.providers.atlascloud] |
| 1186 | base_url = "https://api.atlascloud.ai/v1" |
| 1187 | model = "deepseek-ai/deepseek-v4-flash" |
| 1188 | |
| 1189 | [profiles.sglang] |
| 1190 | provider = "sglang" |
| 1191 | |
| 1192 | [profiles.sglang.providers.sglang] |
| 1193 | base_url = "http://localhost:30000/v1" |
| 1194 | model = "deepseek-ai/DeepSeek-V4-Pro" |
| 1195 | |
| 1196 | [profiles.vllm] |
| 1197 | provider = "vllm" |
| 1198 | |
| 1199 | [profiles.vllm.providers.vllm] |
| 1200 | base_url = "http://localhost:8000/v1" |
| 1201 | model = "deepseek-ai/DeepSeek-V4-Pro" |
| 1202 | |
| 1203 | [profiles.ollama] |
| 1204 | provider = "ollama" |
| 1205 | |
| 1206 | [profiles.ollama.providers.ollama] |
| 1207 | base_url = "http://localhost:11434/v1" |
| 1208 | model = "qwen2.5-coder:7b" # any tag `ollama list` shows |
| 1209 | |
| 1210 | [profiles.ollama-cloud] |
| 1211 | provider = "ollama-cloud" |
| 1212 | |
| 1213 | [profiles.ollama-cloud.providers.ollama_cloud] |
| 1214 | base_url = "https://ollama.com/v1" |
| 1215 | model = "gpt-oss:120b" |
| 1216 | ``` |
| 1217 | |
| 1218 | Select a profile with: |
| 1219 | |
| 1220 | - CLI: `codewhale --profile work` |
| 1221 | - Env: `DEEPSEEK_PROFILE=work` |
| 1222 | |
| 1223 | If a profile is selected but missing, codewhale exits with an error listing available profiles. |
| 1224 | |
| 1225 | ## Legacy top-level `base_url` and `api_key` |
| 1226 | |
| 1227 | Older releases kept DeepSeek's endpoint and key at the top of `config.toml`, |
| 1228 | and each reader decided for itself which other routes inherited them. They |
| 1229 | now live in provider tables only. An older file keeps working unchanged: |
| 1230 | every load reads the top-level keys as if they were already in their table, |
| 1231 | by one rule, per file and per `[profiles.<name>]`: |
| 1232 | |
| 1233 | 1. `provider = "custom"` with no `[providers.custom]` table: the endpoint, key |
| 1234 | and a copy of the model become `[providers.custom]`. |
| 1235 | 2. An endpoint on another vendor's official host (for example |
| 1236 | `integrate.api.nvidia.com`, `xiaomimimo.com`, `openrouter.ai`, or the |
| 1237 | ChatGPT Codex endpoint) belongs to that vendor's table, so a DeepSeek route |
| 1238 | never sends requests there. |
| 1239 | 3. Anything else belongs to `[providers.deepseek]`, which DeepSeek-CN also |
| 1240 | reads for its endpoint and key. |
| 1241 | |
| 1242 | The key goes with DeepSeek (or the literal custom route). It follows the |
| 1243 | endpoint to another vendor only when the same table sets `provider` to that |
| 1244 | vendor, the host is that vendor's own, and the vendor's table has no key. When |
| 1245 | no `provider` is set, a NIM host still selects `nvidia-nim` and |
| 1246 | `api.deepseeki.com` still selects `deepseek-cn`. A `[vision_model]` without a |
| 1247 | key of its own keeps using the top-level key. A profile's top-level |
| 1248 | `base_url` now overrides the base file's `[providers.deepseek] base_url` |
| 1249 | (before, the base table silently won). |
| 1250 | |
| 1251 | When a top-level value and its table disagree, the table's `base_url` and the |
| 1252 | top-level `api_key` are used, which is what the runtime sent before. |
| 1253 | |
| 1254 | Loading never rewrites the file. The next save Codewhale makes (for example |
| 1255 | `codewhale config set`, `/config ... --save`, or `auth set`) moves the keys, |
| 1256 | keeps comments, writes a one-time credential-free copy of the old file to |
| 1257 | `config.toml.pre-migrate.bak`, and prints one line saying what moved. A |
| 1258 | disagreeing pair is never resolved in the file on its own; both values stay |
| 1259 | until you choose, and `codewhale config doctor` / `codewhale doctor` report |
| 1260 | it (sources only, never values). `codewhale config get|set|unset base_url` |
| 1261 | (and `api_key`) address the active provider's table. |
| 1262 | |
| 1263 | To tidy the file yourself: |
| 1264 | |
| 1265 | ```bash |
| 1266 | codewhale config migrate --dry-run # show what would move |
| 1267 | codewhale config migrate # move it (backup first) |
| 1268 | codewhale config migrate --prefer top-level # resolve a conflict: keep the top-level value |
| 1269 | codewhale config migrate --prefer table # resolve a conflict: keep the table value |
| 1270 | ``` |
| 1271 | |
| 1272 | ## Environment Variables |
| 1273 | |
| 1274 | Most runtime environment variables override config values. API-key variables are |
| 1275 | fallbacks after saved config and keyring credentials. |
| 1276 | |
| 1277 | The three user-facing slots — provider, model, base URL — expose `CODEWHALE_*` |
| 1278 | aliases. When both forms are set the `CODEWHALE_*` value wins; the |
| 1279 | `DEEPSEEK_*` form is kept for older shells: |
| 1280 | |
| 1281 | - `CODEWHALE_PROVIDER` (preferred) / `DEEPSEEK_PROVIDER` (legacy alias) — |
| 1282 | `deepseek|deepseek-anthropic|nvidia-nim|openai|atlascloud|wanjie-ark|volcengine|openrouter|xiaomi-mimo|novita|fireworks|siliconflow|arcee|siliconflow-CN|moonshot|sglang|vllm|ollama|ollama-cloud|huggingface|modelscope|together|qianfan|openai-codex|anthropic|openmodel|zai|stepfun|minimax|deepinfra|mistral` |
| 1283 | - `CODEWHALE_MODEL` (preferred) / `DEEPSEEK_MODEL` (legacy alias) — default model for the active provider |
| 1284 | - `CODEWHALE_BASE_URL` (preferred) / `DEEPSEEK_BASE_URL` (legacy alias) — base URL for the active provider |
| 1285 | |
| 1286 | `CODEWHALE_BASE_URL` applies to the **active** route only. A request pinned to |
| 1287 | another provider — a subagent or fleet child, a routed tool, the per-turn |
| 1288 | auto-router, a picker preview — resolves its endpoint from that provider's own |
| 1289 | `[providers.<table>]`, then its provider-scoped variable (`MOONSHOT_BASE_URL`, |
| 1290 | `OPENAI_BASE_URL`, …), then that provider's default. It never inherits the |
| 1291 | active session's host, and a custom route with no configured `base_url` fails |
| 1292 | closed on a loopback placeholder rather than borrowing another provider's |
| 1293 | endpoint. The environment writes the value into the active identity's own |
| 1294 | table: DeepSeek-CN falls back to `[providers.deepseek]` for its endpoint and |
| 1295 | key, but never to a DeepSeek value the environment addressed to DeepSeek alone. |
| 1296 | A managed-config overlay that supplies or reselects the effective |
| 1297 | route's endpoint takes the generic override away from every route. |
| 1298 | |
| 1299 | Remaining variables: |
| 1300 | |
| 1301 | - `DEEPSEEK_API_KEY` |
| 1302 | - `DEEPSEEK_ANTHROPIC_BASE_URL` |
| 1303 | - `DEEPSEEK_HTTP_HEADERS` (custom model request headers, comma-separated `name=value` pairs) |
| 1304 | - `DEEPSEEK_DEFAULT_TEXT_MODEL` (extra legacy alias of `DEEPSEEK_MODEL`) |
| 1305 | - `DEEPSEEK_STREAM_IDLE_TIMEOUT_SECS` (stream idle timeout in seconds; default `900`, clamped to `1..=3600`) |
| 1306 | - `DEEPSEEK_STREAM_OPEN_TIMEOUT_SECS` (connection setup + response-header wait in seconds; default `45`, clamped to `5..=300`; distinct from the per-chunk idle timeout; `stream.open_timeout_secs` (legacy `tui.stream_open_timeout_secs` fallback) wins when positive) |
| 1307 | - `CODEWHALE_CACHE_MAXIMAL` (`1`/`true`/`on`/`yes`) — cache-maximal context mode (#528). When on, the Repo Working Set block materializes the **full current contents** of the top active files into the system prompt each turn (deterministic order, byte-bounded), instead of only listing their paths. The block stays byte-stable while those files are unchanged so DeepSeek's KV prefix cache keeps hitting; editing a file cache-misses from its block onward. Off by default (path list only). Byte caps default to 24 KB per file / 96 KB total. |
| 1308 | - `NVIDIA_API_KEY` or `NVIDIA_NIM_API_KEY` (when provider is `nvidia-nim`) |
| 1309 | - `NVIDIA_NIM_BASE_URL`, `NIM_BASE_URL`, or `NVIDIA_BASE_URL` |
| 1310 | - `NVIDIA_NIM_MODEL` |
| 1311 | - `OPENAI_API_KEY` |
| 1312 | - `OPENAI_BASE_URL` |
| 1313 | - `OPENAI_MODEL` |
| 1314 | - `ATLASCLOUD_API_KEY` |
| 1315 | - `ATLASCLOUD_BASE_URL` |
| 1316 | - `ATLASCLOUD_MODEL` |
| 1317 | - `WANJIE_ARK_API_KEY`, `WANJIE_API_KEY`, or `WANJIE_MAAS_API_KEY` |
| 1318 | - `WANJIE_ARK_BASE_URL`, `WANJIE_BASE_URL`, or `WANJIE_MAAS_BASE_URL` |
| 1319 | - `WANJIE_ARK_MODEL`, `WANJIE_MODEL`, or `WANJIE_MAAS_MODEL` |
| 1320 | - `VOLCENGINE_API_KEY`, `VOLCENGINE_ARK_API_KEY`, or `ARK_API_KEY` |
| 1321 | - `VOLCENGINE_BASE_URL`, `VOLCENGINE_ARK_BASE_URL`, or `ARK_BASE_URL` |
| 1322 | - `VOLCENGINE_MODEL` or `VOLCENGINE_ARK_MODEL` |
| 1323 | - `OPENROUTER_API_KEY` |
| 1324 | - `OPENROUTER_BASE_URL` |
| 1325 | - `OPENROUTER_MODEL` |
| 1326 | - `XIAOMI_MIMO_TOKEN_PLAN_API_KEY`, `MIMO_TOKEN_PLAN_API_KEY`, `XIAOMI_MIMO_API_KEY`, `XIAOMI_API_KEY`, or `MIMO_API_KEY` |
| 1327 | - `XIAOMI_MIMO_BASE_URL` or `MIMO_BASE_URL` |
| 1328 | - `XIAOMI_MIMO_MODEL` or `MIMO_MODEL` |
| 1329 | - `XIAOMI_MIMO_MODE` or `MIMO_MODE` (`token-plan-sgp`, `token-plan-cn`, |
| 1330 | `token-plan-ams`, or `pay-as-you-go`) |
| 1331 | - `NOVITA_API_KEY` |
| 1332 | - `NOVITA_BASE_URL` |
| 1333 | - `NOVITA_MODEL` |
| 1334 | - `FIREWORKS_API_KEY` |
| 1335 | - `FIREWORKS_BASE_URL` |
| 1336 | - `FIREWORKS_MODEL` |
| 1337 | - `HUGGINGFACE_API_KEY` or `HF_TOKEN` (`HF_TOKEN` is a fallback alias accepted when provider is `huggingface`) |
| 1338 | - `MODELSCOPE_API_KEY` |
| 1339 | - `HUGGINGFACE_BASE_URL` or `HF_BASE_URL` |
| 1340 | - `HUGGINGFACE_MODEL` or `HF_MODEL` |
| 1341 | - `SILICONFLOW_API_KEY` |
| 1342 | - `SILICONFLOW_BASE_URL` |
| 1343 | - `SILICONFLOW_MODEL` |
| 1344 | - `ARCEE_API_KEY` |
| 1345 | - `ARCEE_BASE_URL` |
| 1346 | - `ARCEE_MODEL` |
| 1347 | - `TOGETHER_API_KEY` |
| 1348 | - `TOGETHER_BASE_URL` |
| 1349 | - `TOGETHER_MODEL` |
| 1350 | - `QIANFAN_API_KEY` or `BAIDU_QIANFAN_API_KEY` |
| 1351 | - `QIANFAN_BASE_URL` or `BAIDU_QIANFAN_BASE_URL` |
| 1352 | - `QIANFAN_MODEL` or `BAIDU_QIANFAN_MODEL` |
| 1353 | - `OPENAI_CODEX_ACCESS_TOKEN` or `CODEX_ACCESS_TOKEN` (legacy process tokens; ignored by the official ChatGPT plan route) |
| 1354 | - `OPENAI_CODEX_BASE_URL` or `CODEX_BASE_URL` |
| 1355 | - `OPENAI_CODEX_MODEL` or `CODEX_MODEL` |
| 1356 | - `OPENAI_CODEX_ACCOUNT_ID` or `CODEX_ACCOUNT_ID` (legacy identity hints; ignored by the official ChatGPT plan route) |
| 1357 | - `CODEWHALE_CHATGPT_NEW_ACCOUNT=1` explicitly replaces the single selected ChatGPT registration after a validated browser sign-in |
| 1358 | - `ANTHROPIC_API_KEY` |
| 1359 | - `ANTHROPIC_BASE_URL` |
| 1360 | - `ANTHROPIC_MODEL` |
| 1361 | - `ZAI_API_KEY` or `Z_AI_API_KEY` |
| 1362 | - `ZAI_BASE_URL` or `Z_AI_BASE_URL` |
| 1363 | - `ZAI_MODEL` or `Z_AI_MODEL` |
| 1364 | - `STEPFUN_API_KEY` or `STEP_API_KEY` |
| 1365 | - `STEPFUN_BASE_URL` or `STEP_BASE_URL` |
| 1366 | - `STEPFUN_MODEL` or `STEP_MODEL` |
| 1367 | - `MINIMAX_API_KEY` |
| 1368 | - `MINIMAX_BASE_URL` |
| 1369 | - `MINIMAX_MODEL` |
| 1370 | - `DEEPINFRA_API_KEY` or `DEEPINFRA_TOKEN` |
| 1371 | - `DEEPINFRA_BASE_URL` |
| 1372 | - `DEEPINFRA_MODEL` |
| 1373 | - `MISTRAL_API_KEY` |
| 1374 | - `MISTRAL_BASE_URL` |
| 1375 | - `MISTRAL_MODEL` |
| 1376 | - `MOONSHOT_API_KEY` or `KIMI_API_KEY` |
| 1377 | - `MOONSHOT_BASE_URL` or `KIMI_BASE_URL` |
| 1378 | - `MOONSHOT_MODEL`, `KIMI_MODEL_NAME`, or `KIMI_MODEL` |
| 1379 | - `SGLANG_BASE_URL` |
| 1380 | - `SGLANG_MODEL` |
| 1381 | - `SGLANG_API_KEY` (optional; many localhost SGLang servers do not require auth) |
| 1382 | - `VLLM_BASE_URL` |
| 1383 | - `VLLM_MODEL` |
| 1384 | - `VLLM_API_KEY` (optional; many localhost vLLM servers do not require auth) |
| 1385 | - `OLLAMA_BASE_URL` |
| 1386 | - `OLLAMA_MODEL` |
| 1387 | - `OLLAMA_API_KEY` (optional; many localhost Ollama servers do not require auth) |
| 1388 | - `OLLAMA_CLOUD_BASE_URL` |
| 1389 | - `OLLAMA_CLOUD_MODEL` |
| 1390 | - `OLLAMA_CLOUD_API_KEY` (preferred Cloud key; `OLLAMA_API_KEY` is the official fallback) |
| 1391 | For every product-level `CODEWHALE_*` variable below, the matching legacy |
| 1392 | `DEEPSEEK_*` name is still read as a compatibility fallback; when both are set, |
| 1393 | the `CODEWHALE_*` value wins. |
| 1394 | |
| 1395 | - `CODEWHALE_LOG_LEVEL` or `RUST_LOG` (`info`/`debug`/`trace` enables lightweight verbose logs) |
| 1396 | - `CODEWHALE_SKILLS_DIR` |
| 1397 | - `CODEWHALE_MCP_CONFIG` |
| 1398 | - `CODEWHALE_NOTES_PATH` |
| 1399 | - `CODEWHALE_MEMORY` (`1|on|true|yes|y|enabled` turns user memory on) |
| 1400 | - `CODEWHALE_MEMORY_PATH` |
| 1401 | - `CODEWHALE_TELEMETRY` / `DEEPSEEK_TELEMETRY` (legacy alias) — anonymous usage |
| 1402 | counting is on by default in the current 0.9.12 source, with a disclosure |
| 1403 | naming Codewhale and PostHog and an easy durable opt-out. Prior explicit |
| 1404 | declines remain off. Accepts `0|1|true|false|yes|no|on|off|enabled| |
| 1405 | disabled`. An explicit "off" is a **floor**: it beats `--telemetry true` and |
| 1406 | `telemetry = true` in config, and a value this list cannot read also resolves |
| 1407 | to off, because a typo in a kill switch must never resolve to "on". See |
| 1408 | [`TELEMETRY.md`](TELEMETRY.md). |
| 1409 | - `CODEWHALE_TELEMETRY_ENDPOINT` / `DEEPSEEK_TELEMETRY_ENDPOINT` (legacy alias) |
| 1410 | — `https://`, or plain `http://` only for loopback. Overrides the config file. |
| 1411 | Unset selects the shipped default, |
| 1412 | `https://telemetry.codewhale.net/v1/telemetry`; setting it to the **empty |
| 1413 | string** routes batches to a local dry-run file and contacts nobody. Either |
| 1414 | way it only decides where a session sends — it cannot override an opt-out. |
| 1415 | - `CODEWHALE_ALLOW_SHELL` (`1`/`true` enables) |
| 1416 | - `CODEWHALE_APPROVAL_POLICY` (`on-request|untrusted|never`) |
| 1417 | - `CODEWHALE_SANDBOX_MODE` (`read-only|workspace-write|danger-full-access|external-sandbox`) |
| 1418 | - `CODEWHALE_NO_NEW_PRIVS` (`0`/`false`/`no`/`off`/`disabled` opts out) — Linux only. The |
| 1419 | TUI process sets the kernel's irreversible no-new-privileges flag at startup |
| 1420 | as defense-in-depth, which blocks `sudo`/`su`/setuid helpers for Codewhale's |
| 1421 | whole process tree. The flag is already skipped when the startup sandbox |
| 1422 | mode resolves to `danger-full-access` (#5723), so this variable is the |
| 1423 | explicit override in the remaining cases: set it to a falsey value before |
| 1424 | launching if you administer through Codewhale as a wheel-group user under a |
| 1425 | narrower posture and need escalation to work (#5413), or set a truthy value |
| 1426 | to force the flag on even under `danger-full-access`. Unset, the posture |
| 1427 | decides; the other startup hardening (no ptrace, no core dumps) always |
| 1428 | stays on. |
| 1429 | - `CODEWHALE_MANAGED_CONFIG_PATH` |
| 1430 | - `CODEWHALE_REQUIREMENTS_PATH` |
| 1431 | - `CODEWHALE_MAX_SUBAGENTS` (clamped to `1..=128`) |
| 1432 | - `CODEWHALE_TASKS_DIR` (runtime task queue/artifact storage, default |
| 1433 | `~/.codewhale/tasks`, with legacy `~/.deepseek/tasks` fallback when only the |
| 1434 | legacy directory exists) |
| 1435 | - `CODEWHALE_RUNTIME_DIR` (override the runtime thread store root). Interactive |
| 1436 | sessions default to `$CODEWHALE_HOME/sessions/<session-id>/runtime` so each |
| 1437 | Codewhale process owns its own store (#5630). The store is single-owner: a |
| 1438 | second process on the **same** root fails at startup. Set this variable to |
| 1439 | share one store across processes, or when the runtime API server should use a |
| 1440 | stable non-session path. Unset, the API/server path remains |
| 1441 | `$CODEWHALE_HOME/tasks/runtime`. Legacy alias: `DEEPSEEK_RUNTIME_DIR`. |
| 1442 | - `CODEWHALE_ALLOW_INSECURE_HTTP` (`1`/`true` allows non-local `http://` base URLs; default is reject) |
| 1443 | - `CODEWHALE_FORCE_HTTP1` (`1|true|yes|on` pins the HTTP client to HTTP/1.1, disabling HTTP/2; useful on Windows or behind proxies that mishandle long-lived H2 streams) |
| 1444 | - `CODEWHALE_HOME` (override the base data directory; defaults to `~/.codewhale`). |
| 1445 | If you previously exported `DEEPSEEK_HOME`, rename it to `CODEWHALE_HOME`; |
| 1446 | the old env var is not used for new Codewhale state paths. |
| 1447 | - `CODEWHALE_RELEASE_BASE_URL` (release asset mirror used by `codewhale update` |
| 1448 | and by TUI startup update checks when `[update].update_uri` is not set, or as |
| 1449 | a fallback when that configured URI cannot be fetched) |
| 1450 | - `CODEWHALE_AUTOMATIONS_DIR` (override the automations storage directory; uses |
| 1451 | `~/.codewhale/automations` by default, with legacy `~/.deepseek/automations` |
| 1452 | fallback when only the legacy directory exists) |
| 1453 | - `NO_ANIMATIONS` (`1|true|yes|on` forces `low_motion = true` and |
| 1454 | `fancy_animations = false` at startup, regardless of the saved |
| 1455 | settings; see [`docs/ACCESSIBILITY.md`](./ACCESSIBILITY.md)). |
| 1456 | - `SSL_CERT_FILE` — corporate-proxy / TLS-inspecting MITM users |
| 1457 | point this at a PEM bundle (or single DER cert) and the cert(s) |
| 1458 | get added alongside the platform's system trust store. Failures |
| 1459 | log a warning and continue — the existing system roots still |
| 1460 | apply. |
| 1461 | |
| 1462 | ### Instruction sources (`instructions = [...]`, #454) |
| 1463 | |
| 1464 | Add a list of additional system-prompt sources that get |
| 1465 | concatenated, in declared order, alongside the auto-loaded |
| 1466 | `AGENTS.md`: |
| 1467 | |
| 1468 | ```toml |
| 1469 | instructions = [ |
| 1470 | "./AGENTS.md", |
| 1471 | "~/.codewhale/global.md", |
| 1472 | "~/team/agents-shared.md", |
| 1473 | ] |
| 1474 | ``` |
| 1475 | |
| 1476 | Rules: |
| 1477 | |
| 1478 | - Paths run through `expand_path` so `~` and env vars work. |
| 1479 | - Each file is capped at 100 KiB; oversized files are |
| 1480 | truncated with a `[…elided]` marker rather than skipped. |
| 1481 | - Missing files are skipped with a tracing warning so a stale |
| 1482 | entry doesn't fail the launch. |
| 1483 | - Only user-owned config, profiles, and managed config may set this array. |
| 1484 | Project config (`<workspace>/.codewhale/config.toml`, or legacy |
| 1485 | `<workspace>/.deepseek/config.toml`) ignores `instructions` so a cloned repo |
| 1486 | cannot choose arbitrary local files to place into the prompt. |
| 1487 | |
| 1488 | ### Hooks |
| 1489 | |
| 1490 | Hooks are a **TUI runtime feature**. They fire from the interactive TUI and |
| 1491 | the engine turn loop it drives; `codewhale exec`, the CLI subcommands, the |
| 1492 | app-server / ACP surfaces, and the `workflow` tool do not fire them. |
| 1493 | |
| 1494 | [`docs/HOOKS.md`](HOOKS.md) is the authoritative reference for all eleven hook |
| 1495 | events — their firing points, environment variables, stdin payloads, timeout |
| 1496 | and background semantics, and which three of them can steer Codewhale. The |
| 1497 | sections below cover the configuration surface and the steering contracts in |
| 1498 | more depth. |
| 1499 | |
| 1500 | Two contract points worth reading there before writing a hook: |
| 1501 | |
| 1502 | - `background = true` means **submitted and never awaited**. The hook still |
| 1503 | gets the documented stdin payload and the same timeout, but it has no exit |
| 1504 | code and cannot steer. |
| 1505 | - A condition that references context its event never carries (an `exit_code` |
| 1506 | condition outside `tool_call_after` / `on_error`, a `mode` condition on |
| 1507 | `shell_env`, a tool condition on a non-tool event) is **rejected at load**, |
| 1508 | logged, and shown in `/hooks list`. It does not silently never match. |
| 1509 | Rejection is per entry, so a broken hook never drops another one that merely |
| 1510 | shares its `name` or is likewise unnamed. |
| 1511 | |
| 1512 | ### `/hooks` listing |
| 1513 | |
| 1514 | Run `/hooks` (or `/hooks list`) inside the TUI to see every |
| 1515 | configured lifecycle hook grouped by event, including each |
| 1516 | hook's name, command preview, effective timeout, and condition. When |
| 1517 | `[hooks].default_timeout_secs` is set it replaces every per-hook |
| 1518 | `timeout_secs`, and the listing shows that effective value and names the |
| 1519 | override rather than echoing the per-hook number. A |
| 1520 | `default_timeout_secs = 0` is rejected at load — it would expire every hook |
| 1521 | in the config immediately — so the override is ignored, per-hook |
| 1522 | `timeout_secs` applies, the listing shows that per-hook value with no |
| 1523 | override provenance, and the rejection appears under `configuration |
| 1524 | problems`. The |
| 1525 | `[hooks].enabled` flag's state is shown at the top so it's |
| 1526 | obvious when hooks are globally suppressed, and any entry rejected |
| 1527 | at load is listed under `configuration problems` with the reason. |
| 1528 | Hooks are configured under `[[hooks.hooks]]` entries — see |
| 1529 | [`docs/HOOKS.md`](HOOKS.md) for the full schema. |
| 1530 | |
| 1531 | ### Mutable `message_submit` hooks |
| 1532 | |
| 1533 | `message_submit` hooks run before a submitted message is added to |
| 1534 | history or sent to the model. Unlike observer-only lifecycle hooks, |
| 1535 | non-background `message_submit` hooks can replace or block the |
| 1536 | submitted text. |
| 1537 | |
| 1538 | ```toml |
| 1539 | [[hooks.hooks]] |
| 1540 | event = "message_submit" |
| 1541 | command = "~/.codewhale/hooks/inject-context.sh" |
| 1542 | timeout_secs = 2 |
| 1543 | continue_on_error = true |
| 1544 | ``` |
| 1545 | |
| 1546 | The hook receives JSON on stdin: |
| 1547 | |
| 1548 | ```json |
| 1549 | { |
| 1550 | "event": "message_submit", |
| 1551 | "text": "original user text", |
| 1552 | "text_bytes": 18, |
| 1553 | "text_original_bytes": 18, |
| 1554 | "text_truncated": false, |
| 1555 | "session_id": "sess_12345678", |
| 1556 | "workspace": "/path/to/workspace", |
| 1557 | "mode": "agent", |
| 1558 | "model": "deepseek-chat", |
| 1559 | "total_tokens": 1234 |
| 1560 | } |
| 1561 | ``` |
| 1562 | |
| 1563 | The entire serialized document is capped at 32 KiB. Codewhale retains the |
| 1564 | largest UTF-8-safe `text` prefix that fits after JSON escaping and bounded |
| 1565 | metadata, and the three `text_*` fields make truncation explicit. Immediate |
| 1566 | messages, restored queue entries, merged steers, and prior-hook replacements |
| 1567 | all cross this same serialization boundary. |
| 1568 | |
| 1569 | If the hook exits `0` and prints JSON with a non-empty string `text` field, |
| 1570 | that value replaces the submitted text: |
| 1571 | |
| 1572 | ```json |
| 1573 | { "text": "replacement user text" } |
| 1574 | ``` |
| 1575 | |
| 1576 | Exit `0` with empty stdout, or stdout JSON without `text`, leaves |
| 1577 | the current text unchanged. A JSON `text` field must not be empty; |
| 1578 | `{"text":""}` is treated as invalid stdout and ignored. Exit `2` |
| 1579 | blocks the submission before the turn starts; a structured `reason` field can |
| 1580 | provide the bounded, redacted status message shown in the TUI. Raw stdout, |
| 1581 | stderr, and process-error text are not copied into denial receipts. |
| 1582 | Other non-zero exits follow the hook's `continue_on_error` setting. |
| 1583 | Timeouts and spawn failures are also surfaced as transient TUI status |
| 1584 | messages when `continue_on_error = true` lets submission continue. |
| 1585 | |
| 1586 | Multiple `message_submit` hooks run in config order, and each hook |
| 1587 | receives the text produced by the previous hook. Hooks marked |
| 1588 | `background = true` are observer-only and cannot transform or block |
| 1589 | the message — they still receive the same stdin payload and the same |
| 1590 | environment, they are simply never awaited. Existing environment |
| 1591 | variables remain available. |
| 1592 | `shell_env` hooks keep their existing `KEY=VALUE` stdout contract; |
| 1593 | JSON stdout contracts exist for `message_submit` (above) and |
| 1594 | `tool_call_before` (below). |
| 1595 | |
| 1596 | ### `tool_call_before` decision hooks |
| 1597 | |
| 1598 | `tool_call_before` hooks run before each tool call executes. In |
| 1599 | addition to the legacy hard deny (exit code `2`, which always wins |
| 1600 | regardless of stdout), a foreground hook may print a JSON decision on |
| 1601 | stdout with exit code `0`: |
| 1602 | |
| 1603 | ```json |
| 1604 | { |
| 1605 | "decision": "allow" | "deny" | "ask", |
| 1606 | "reason": "human-readable explanation (used for deny)", |
| 1607 | "updatedInput": { "command": "ls -la" }, |
| 1608 | "additionalContext": "text appended to the tool result for the model" |
| 1609 | } |
| 1610 | ``` |
| 1611 | |
| 1612 | All fields are optional. Empty stdout, non-JSON stdout, and JSON |
| 1613 | without a `decision` field behave exactly as before (allow). An |
| 1614 | unrecognized `decision` string logs a fixed warning without echoing the |
| 1615 | untrusted value and is treated as allow. |
| 1616 | |
| 1617 | - `deny` blocks the tool; the model receives a permission-denied tool |
| 1618 | result containing `reason`. |
| 1619 | - `ask` forces the interactive approval prompt in Ask and Auto-Review even for |
| 1620 | tools that would otherwise auto-run. Full Access does not open tool-approval |
| 1621 | prompts, so hook `ask` does not downgrade that posture. |
| 1622 | - `updatedInput` must be a JSON object; it replaces the tool input |
| 1623 | before execution. When several hooks supply it, the last hook wins. |
| 1624 | - `additionalContext` is appended to the tool result sent back to the |
| 1625 | model as `[hook context] ...`. Multiple hooks' contexts are |
| 1626 | concatenated. |
| 1627 | |
| 1628 | When multiple hooks match, precedence is deny > ask > allow. Hooks |
| 1629 | marked `background = true` cannot steer tool calls — they are |
| 1630 | submitted and never awaited, so they have no verdict to contribute. |
| 1631 | |
| 1632 | A foreground hook that produced no verdict at all — it hit its timeout, the |
| 1633 | process could not be started, or a strict process exited non-zero without an |
| 1634 | explicit JSON decision — is not treated as permission. If |
| 1635 | *that* hook is configured with `continue_on_error = false`, the outcome |
| 1636 | denies the tool call and the denial names the hook and a bounded reason. |
| 1637 | Strictness is read off the hooks that actually matched this call, so a |
| 1638 | strict gate scoped to another tool cannot deny it, and a lenient hook's |
| 1639 | timeout does not deny merely because a strict hook exists elsewhere in |
| 1640 | config. Under the default `continue_on_error = true` the outcome is |
| 1641 | logged and the call proceeds. |
| 1642 | |
| 1643 | `reason` and `additionalContext` are capped (2 000 characters per field, |
| 1644 | 8 000 for the concatenated context of one call) and stripped of control |
| 1645 | characters before they reach the TUI or the model. |
| 1646 | |
| 1647 | Example deny hook: |
| 1648 | |
| 1649 | ```toml |
| 1650 | [[hooks.hooks]] |
| 1651 | event = "tool_call_before" |
| 1652 | command = '''echo '{"decision":"deny","reason":"blocked by project policy"}' ''' |
| 1653 | condition = { type = "tool_name", name = "exec_shell" } |
| 1654 | ``` |
| 1655 | |
| 1656 | Example ask hook (force approval for every MCP tool): |
| 1657 | |
| 1658 | ```toml |
| 1659 | [[hooks.hooks]] |
| 1660 | event = "tool_call_before" |
| 1661 | command = '''echo '{"decision":"ask"}' ''' |
| 1662 | condition = { type = "tool_name", name = "mcp__*" } |
| 1663 | ``` |
| 1664 | |
| 1665 | Example input rewrite: |
| 1666 | |
| 1667 | ```toml |
| 1668 | [[hooks.hooks]] |
| 1669 | event = "tool_call_before" |
| 1670 | command = "~/.codewhale/hooks/clamp-shell-timeout.sh" |
| 1671 | condition = { type = "tool_name", name = "exec_shell" } |
| 1672 | ``` |
| 1673 | |
| 1674 | where the script reads the hook context, then prints |
| 1675 | `{"updatedInput": {...}}` with the adjusted arguments. |
| 1676 | |
| 1677 | `tool_name` conditions support `*` globs: `mcp__*` matches every MCP |
| 1678 | tool (e.g. `mcp__github__create_issue`) but not built-ins like |
| 1679 | `read_file`; exact names keep matching exactly. Other regex |
| 1680 | metacharacters in the pattern are matched literally. |
| 1681 | |
| 1682 | ### Project-local hooks |
| 1683 | |
| 1684 | Repositories can ship policy in `<workspace>/.codewhale/hooks.toml`, |
| 1685 | using the same shape as the `[hooks]` table (top-level fields plus |
| 1686 | `[[hooks]]` entries). Project hooks are executable shell |
| 1687 | configuration, so Codewhale only loads them after the workspace has |
| 1688 | been trusted in user-owned config through the trust prompt or a |
| 1689 | `[projects."<workspace>"] trust_level = "trusted"` entry. Session |
| 1690 | `/trust on` mode does not enable repo-supplied hooks by itself, and |
| 1691 | repo-local legacy markers such as `.deepseek/trusted` do not enable |
| 1692 | project hooks. Once trusted, project hooks are appended after global |
| 1693 | hooks from `config.toml`, so they run last and, for `updatedInput`, |
| 1694 | win ties. A malformed trusted project file logs a warning and startup |
| 1695 | falls back to global hooks only. |
| 1696 | |
| 1697 | ```toml |
| 1698 | # .codewhale/hooks.toml |
| 1699 | [[hooks]] |
| 1700 | event = "tool_call_before" |
| 1701 | command = '''echo '{"decision":"deny","reason":"no shell in this repo"}' ''' |
| 1702 | condition = { type = "tool_name", name = "exec_shell" } |
| 1703 | ``` |
| 1704 | |
| 1705 | ### Turn-end observer hooks |
| 1706 | |
| 1707 | `turn_end` hooks observe the end of each model turn after post-turn |
| 1708 | state, usage totals, cost accounting, notifications, receipts, and |
| 1709 | queue recovery have been updated. They receive JSON on stdin and are |
| 1710 | observer-only: stdout is ignored, failures are logged as warnings, and |
| 1711 | the hook cannot block user input, mutate the transcript, or change the |
| 1712 | next queued follow-up. |
| 1713 | |
| 1714 | Observer-only UI events share one 32-entry queue and two persistent workers; |
| 1715 | the terminal loop uses non-blocking submission and does not create a thread per |
| 1716 | event. A full queue or unavailable dispatcher drops that observer event and is |
| 1717 | kept as an event-specific error toast, independent of ordinary agent/turn |
| 1718 | status text. |
| 1719 | |
| 1720 | ```toml |
| 1721 | [[hooks.hooks]] |
| 1722 | event = "turn_end" |
| 1723 | command = "~/.codewhale/hooks/turn-audit.sh" |
| 1724 | timeout_secs = 2 |
| 1725 | continue_on_error = true |
| 1726 | ``` |
| 1727 | |
| 1728 | The payload includes common hook metadata plus post-turn accounting: |
| 1729 | |
| 1730 | ```json |
| 1731 | { |
| 1732 | "event": "turn_end", |
| 1733 | "session_id": "sess_12345678", |
| 1734 | "workspace": "/path/to/workspace", |
| 1735 | "mode": "agent", |
| 1736 | "created_at": "2026-07-12T10:30:00+00:00", |
| 1737 | "model_backed": true, |
| 1738 | "provider": "deepseek", |
| 1739 | "model": "deepseek-chat", |
| 1740 | "billing_surface": null, |
| 1741 | "turn_id": "turn_12345678", |
| 1742 | "status": "completed", |
| 1743 | "error": null, |
| 1744 | "duration_ms": 1834, |
| 1745 | "usage": { |
| 1746 | "input_tokens": 1200, |
| 1747 | "output_tokens": 180, |
| 1748 | "prompt_cache_hit_tokens": 900, |
| 1749 | "prompt_cache_miss_tokens": 300, |
| 1750 | "prompt_cache_write_tokens": 0, |
| 1751 | "reasoning_tokens": null, |
| 1752 | "reasoning_replay_tokens": null |
| 1753 | }, |
| 1754 | "totals": { |
| 1755 | "session_tokens": 1380, |
| 1756 | "conversation_tokens": 1380, |
| 1757 | "input_tokens": 1200, |
| 1758 | "output_tokens": 180 |
| 1759 | }, |
| 1760 | "tool_count": 2, |
| 1761 | "queued_message_count": 1, |
| 1762 | "stop_hook_active": false |
| 1763 | } |
| 1764 | ``` |
| 1765 | |
| 1766 | `created_at` anchors time-window pricing; `provider` and `model` identify the |
| 1767 | effective route used for model-backed turns. `billing_surface` is an optional, |
| 1768 | non-secret classification derived from the endpoint that actually served the |
| 1769 | turn. Recognized StepFun routes emit `stepfun-payg` or `stepfun-plan`; the raw |
| 1770 | base URL is never written to hook or runtime records. Runtime `TurnRecord` |
| 1771 | exports call the same field `effective_billing_surface`, which `scorecard` |
| 1772 | accepts as an alias. This keeps subscription quota separate from token-priced |
| 1773 | usage. Unrecognized and custom endpoints remain `null` and unpriced. |
| 1774 | |
| 1775 | Shell-only lifecycle completions set `model_backed` to `false` and may report a |
| 1776 | `null` provider; offline scorecards exclude those records from model token and |
| 1777 | cost totals. Completion-only shell, manual-compaction, and purge events that do |
| 1778 | not have a matching `TurnStarted` retain the observer notification with a |
| 1779 | synthetic `lifecycle_<uuid>` turn id and the time the completion was observed. |
| 1780 | |
| 1781 | For `interrupted` or `failed` turns, `status` reflects that terminal |
| 1782 | state and `error` carries the engine error string when one is available. |
| 1783 | `stop_hook_active` is reserved for future re-entry protection and is |
| 1784 | currently always `false`. |
| 1785 | |
| 1786 | ### Sub-agent lifecycle hooks |
| 1787 | |
| 1788 | `subagent_spawn` and `subagent_complete` hooks observe sub-agent lifecycle |
| 1789 | events. They receive bounded JSON metadata on stdin and are observer-only: |
| 1790 | hook failures are logged as warnings and do not block sub-agent scheduling, |
| 1791 | change prompts, or change results. For these observer events, |
| 1792 | `continue_on_error` has no effect: later matching hooks still run even when an |
| 1793 | earlier hook exits non-zero. |
| 1794 | |
| 1795 | ```toml |
| 1796 | [[hooks.hooks]] |
| 1797 | event = "subagent_complete" |
| 1798 | command = "~/.codewhale/hooks/subagent-audit.sh" |
| 1799 | timeout_secs = 2 |
| 1800 | continue_on_error = true |
| 1801 | ``` |
| 1802 | |
| 1803 | `subagent_spawn` receives: |
| 1804 | |
| 1805 | ```json |
| 1806 | { |
| 1807 | "event": "subagent_spawn", |
| 1808 | "agent_id": "agent_12345678", |
| 1809 | "session_id": "sess_12345678", |
| 1810 | "workspace": "/path/to/workspace", |
| 1811 | "mode": "agent", |
| 1812 | "model": "deepseek-chat", |
| 1813 | "total_tokens": 1234, |
| 1814 | "prompt_preview": "bounded prompt preview", |
| 1815 | "prompt_truncated": false |
| 1816 | } |
| 1817 | ``` |
| 1818 | |
| 1819 | `subagent_complete` receives the same common fields plus terminal metadata: |
| 1820 | |
| 1821 | ```json |
| 1822 | { |
| 1823 | "event": "subagent_complete", |
| 1824 | "agent_id": "agent_12345678", |
| 1825 | "session_id": "sess_12345678", |
| 1826 | "workspace": "/path/to/workspace", |
| 1827 | "mode": "agent", |
| 1828 | "model": "deepseek-chat", |
| 1829 | "total_tokens": 1234, |
| 1830 | "status": "completed", |
| 1831 | "result_preview": "bounded result preview", |
| 1832 | "result_truncated": false |
| 1833 | } |
| 1834 | ``` |
| 1835 | |
| 1836 | Previews are capped before delivery so lifecycle hooks do not receive full |
| 1837 | sub-agent prompts, transcripts, or unbounded results. Use the transcript handle |
| 1838 | returned by `agent` when full sub-agent details are needed. |
| 1839 | |
| 1840 | ### Running-turn input |
| 1841 | |
| 1842 | Composer shortcuts keep the same role throughout a session: |
| 1843 | |
| 1844 | - **Enter** sends when idle and queues a next-turn follow-up while busy. The |
| 1845 | behavior does not change before versus after the provider's first token. |
| 1846 | - With an empty composer and queued follow-ups visible, **Enter** sends the |
| 1847 | oldest queued follow-up into the active turn now. |
| 1848 | - **Ctrl+Enter** (or **Cmd+Enter** when the terminal forwards it) explicitly |
| 1849 | steers the active turn. It sends normally when idle. |
| 1850 | - By default, **Shift+Enter**, **Alt+Enter**, and **Ctrl+J** insert a newline. |
| 1851 | - Set `composer_multiline_mode = true` to make **Enter** insert a newline and |
| 1852 | **Shift+Enter** send instead. **Alt+Enter**, **Ctrl+J**, and supported |
| 1853 | **Ctrl+Enter** / **Cmd+Enter** behavior is unchanged. |
| 1854 | - **Ctrl+G** and **Ctrl+S** only stash drafts; they never send or steer. |
| 1855 | |
| 1856 | ### Composer stash (`/stash`, Ctrl+G / Ctrl+S) |
| 1857 | |
| 1858 | Press **Ctrl+G** in the composer to park the current draft to |
| 1859 | `~/.codewhale/composer_stash.jsonl`. `/stash list` shows parked |
| 1860 | drafts with one-line previews and timestamps; `/stash pop` |
| 1861 | restores the most recently parked draft (LIFO); `/stash clear` |
| 1862 | wipes the file. Capped at 200 entries; multiline drafts round-trip intact. |
| 1863 | **Ctrl+S** remains an alias in terminals that forward it; Cursor and VS Code |
| 1864 | reserve Ctrl+S for Save, so Ctrl+G is the portable default. |
| 1865 | |
| 1866 | ## Settings File (Persistent UI Preferences) |
| 1867 | |
| 1868 | codewhale also stores user preferences in: |
| 1869 | |
| 1870 | - `~/.codewhale/settings.toml` on new installs |
| 1871 | - `~/.deepseek/settings.toml` or the legacy platform config-dir |
| 1872 | `deepseek/settings.toml` when an existing settings file is present |
| 1873 | |
| 1874 | Notable settings include `auto_compact`, which uses a model-aware default-on |
| 1875 | policy for known context windows up to the 1M-token V4 class. Automatic |
| 1876 | compaction runs before the active model limit and carries the compacted summary |
| 1877 | forward into the next request. The trigger defaults to |
| 1878 | `auto_compact_threshold_percent = 80`. Users who prefer manual continuity can |
| 1879 | persist `auto_compact = false`; manual `/compact` / Ctrl+L remains available. |
| 1880 | You can inspect or update these from the TUI with `/settings` and `/config` |
| 1881 | (interactive editor). |
| 1882 | |
| 1883 | Common settings keys: |
| 1884 | |
| 1885 | - `theme` (`system`, `terminal`, `underwater`, `underwater-retro`, |
| 1886 | `shoreline`, `shoreline-light`, `dark`, `light`, `grayscale`, |
| 1887 | `catppuccin-mocha`, `tokyo-night`, `dracula`, `gruvbox-dark`, `claude`, |
| 1888 | `matrix`, `solarized-light`, `uwu`; default `underwater`): `underwater` is |
| 1889 | the dark navy fresh-install default, `shoreline` is its warm charcoal |
| 1890 | alternative and `shoreline-light` the light variant, `system` follows terminal |
| 1891 | background detection, `dark`/`light` use the Codewhale Whale pair, |
| 1892 | `terminal` inherits the host terminal, `grayscale` is the low-opinion |
| 1893 | black/white theme, and the named community presets apply across the TUI. |
| 1894 | Aliases such as `whale`, `mono`, `black-white`, `tokyonight`, and `gruvbox` |
| 1895 | are accepted. In Whale, cobalt blue owns action/focus, seafoam owns live |
| 1896 | work, Signal Gold owns human decisions and the whale, coral owns warnings, |
| 1897 | rose owns danger, violet owns Operate, and green remains completed/verified. |
| 1898 | Text labels, markers, and motion policy carry the same states when color is |
| 1899 | unavailable; color is never the only cue. |
| 1900 | User-authored overlays live only at `~/.codewhale/themes/<name>.json` (or |
| 1901 | `$CODEWHALE_HOME/themes/<name>.json`) and are selected with |
| 1902 | `/theme custom:<name>`. The filename is a bounded slug, symlinks and files |
| 1903 | over 64 KiB are refused, colors must be `#RRGGBB`, and unknown fields fail |
| 1904 | validation. `/theme schema` prints the embedded JSON Schema and `/theme path` |
| 1905 | shows the exact directory. An overlay names one compiled `base` theme and |
| 1906 | changes only listed semantic colors; it cannot include or read another file. |
| 1907 | Open `/theme` to browse valid overlays, preview them live, and keep the |
| 1908 | active `custom:<name>` selector when the picker is opened without moving. |
| 1909 | - `auto_compact` (on/off, model-aware default on for known context windows |
| 1910 | unless explicitly configured) |
| 1911 | - `auto_compact_threshold_percent` (10-100, default `80`): pre-send |
| 1912 | auto-compaction threshold used only when `auto_compact` is enabled. |
| 1913 | - `paste_burst_detection` (on/off, default on): fallback rapid-key paste |
| 1914 | detection for terminals that do not emit bracketed-paste events. This is |
| 1915 | independent of terminal bracketed-paste mode. |
| 1916 | - `work_surface_placement` (`bottom`, `top`, `left`, `right`, or `off`; |
| 1917 | default `bottom`): places the workbar — Tasks / To-do / Workers — under the |
| 1918 | composer (the default bottom workbar), above the transcript, in a side |
| 1919 | workbar, or hides it entirely (`off`). Side choices fall back to the top |
| 1920 | layout on narrow terminals without changing the saved preference. Set it |
| 1921 | live with `/config work_surface_placement right --save` (or `left` / `top` / |
| 1922 | `bottom` / `off`). |
| 1923 | - `rail_panel` (`tasks`, `agents`, `background`, `files`, `notepad`, |
| 1924 | `context`, `git`, `price`; default `tasks`, alias key `rail`): which panel |
| 1925 | the workbar shows. Panel selection is orthogonal to placement. `tasks` is |
| 1926 | the full live work list (to-dos, then sub-agents); `agents` narrows to the |
| 1927 | sub-agent rows; `background` lists background shells and automations; |
| 1928 | `files` lists touched files; `notepad` shows the workspace notes; `context` |
| 1929 | is a read-only session-facts list; `git` shows branch status; `price` |
| 1930 | shows cost. In every panel except `context`, rows are selectable and |
| 1931 | clickable and open their detail surface. `Alt+!`/`Alt+@`/`Alt+#`/`Alt+$` |
| 1932 | switch panels live. |
| 1933 | - `work_surface_top_height` (2–16) and `work_surface_side_width` (26–80): |
| 1934 | ceilings for the top strip's height and the side workbar's width. Both are |
| 1935 | normally persisted by dragging the divider rather than edited by hand; the |
| 1936 | strip still auto-fits its content below the ceiling. |
| 1937 | - `focus_texture` (`off`, `scrim`, or `grain`; default `off`): focus-context |
| 1938 | texture for modal views. `scrim` dims the already-rendered background |
| 1939 | outside the focused modal toward the theme surface; `grain` sprinkles |
| 1940 | sparse dots over blank cells there. The texture is static (no time |
| 1941 | component, so it is unaffected by `low_motion`), never writes over a cell |
| 1942 | that carries text, and preserves the 4.5:1 body-text contrast floor |
| 1943 | wherever both colors are resolvable. It is skipped entirely on frames |
| 1944 | below the ambient-life minimum size and when the focused modal already |
| 1945 | covers 90% or more of the frame. Set it live with |
| 1946 | `/config focus_texture scrim --save`. |
| 1947 | - `mention_menu_limit` (integer, default `128`): maximum number of |
| 1948 | `@`-mention popup candidates retained before the composer renders the |
| 1949 | visible window. The visible rows still depend on terminal height. |
| 1950 | - `mention_walk_depth` (integer, default `10`): maximum workspace depth for |
| 1951 | `@`-mention completion walks. Set to `0` for unlimited depth in deeply |
| 1952 | nested workspaces; keep the default in very large repos unless needed. |
| 1953 | - `mention_menu_behavior` (`fuzzy`, `browser`; default `fuzzy`): controls how |
| 1954 | `@`-mention completions are populated. `fuzzy` searches the workspace and |
| 1955 | applies mention frecency. `browser` lists only the immediate children of the |
| 1956 | currently typed directory segment in deterministic alphabetical order. |
| 1957 | - `show_thinking` (on/off) |
| 1958 | - `thinking_default_expanded` (on/off, default off): renders thinking blocks |
| 1959 | expanded initially when `show_thinking` is enabled. Space still toggles the |
| 1960 | selected block, and it decides only the blocks you have not touched: a block |
| 1961 | you expanded or collapsed yourself keeps that state if you change this |
| 1962 | setting later. This is useful in SSH/tmux environments where the Space |
| 1963 | binding may be intercepted. |
| 1964 | - `thinking_preview_lines` (integer, default `2`): how many body rows a |
| 1965 | **collapsed** completed thought still shows. `0` is header-only; `10` is |
| 1966 | the older dump. Live streaming preview is unchanged. Expand a block with |
| 1967 | Space, or set `thinking_default_expanded` to open every block. |
| 1968 | - `help_expand_groups` (on/off, default off): start Help/shortcuts with every |
| 1969 | group expanded. Default folds the long tail (Grok-style); type-to-filter |
| 1970 | still unfolds matches. |
| 1971 | - `pin_last_prompt` (on/off, default on): pin the last user prompt at the top |
| 1972 | of the transcript viewport after it scrolls off. |
| 1973 | - `show_tool_details` (on/off) |
| 1974 | - `inline_diffs` (`full`, `summary`, or `off`; default `full`): controls the |
| 1975 | inline presentation of successful structured File mutations. `full` shows a |
| 1976 | bounded red/green diff and semantic statistics, `summary` keeps only the |
| 1977 | statistics, and `off` keeps the calm changed-file outcome. All three retain |
| 1978 | the exact applied change in the selected File receipt's Alt/Option+V detail. |
| 1979 | Failure and cancellation never render a successful diff. Save the choice |
| 1980 | with `/config inline_diffs <mode> --save`. |
| 1981 | - `locale` (`auto`, `en`, `ja`, `zh-Hans`, `zh-Hant`, `pt-BR`, `es-419`, `vi`, |
| 1982 | `ko`; default `auto`): UI chrome locale. `auto` checks `LC_ALL`, |
| 1983 | `LC_MESSAGES`, then `LANG`; unsupported locale selections resolve to English. |
| 1984 | Every shipped pack holds full `en.json` parity, so no string falls back |
| 1985 | to English. The runtime also exposes the resolved locale in the system |
| 1986 | prompt as the fallback natural language for V4 reasoning and replies when the |
| 1987 | latest user message is ambiguous. Clear user language still takes priority; |
| 1988 | Chinese turns should produce Chinese `reasoning_content` and Chinese final |
| 1989 | replies even when the resolved locale is English. |
| 1990 | - `background_color` (`#RRGGBB`, `RRGGBB`, or `default`): optional main TUI |
| 1991 | background color applied to the root, header, transcript, and footer |
| 1992 | surfaces while preserving panel contrast. |
| 1993 | - `cost_currency` (`usd`, `cny`; default `usd`): currency used by the footer, |
| 1994 | context panel, `/cost`, `/tokens`, and long-turn notification summaries. The |
| 1995 | aliases `rmb` and `yuan` normalize to `cny`. |
| 1996 | - `default_mode` (`agent`, `plan`, or `operate`; legacy values are accepted for migration but are not live mode vocabulary) |
| 1997 | - `launch_screen` (legacy, migration-only): this historical `on`/`off` value |
| 1998 | is still accepted when reading an existing settings file, but it no longer |
| 1999 | changes behavior and is omitted from new saves. A fresh interactive launch |
| 2000 | always opens Tideline Startup; only an explicit resume or an explicit |
| 2001 | initial prompt enters the live session directly. |
| 2002 | - `sidebar_focus` (legacy, migration-only): the classic right sidebar this key |
| 2003 | configured was removed in the 0.9.4 rail unification. The key is still read |
| 2004 | once so old settings carry forward, then folds into the live keys: |
| 2005 | `pinned`/`work`/`plan`/`todos` become `rail_panel = "pinned"`, |
| 2006 | `agents`/`subagents` become `rail_panel = "agents"`, `context`/`session` |
| 2007 | become `rail_panel = "context"`, `tasks`/`auto` (the old default) become the |
| 2008 | `tasks` panel, `sessions` enables `sessions_rail`, and `hidden` turns the |
| 2009 | workbar off via `work_surface_placement = "off"`. An explicit `rail_panel` |
| 2010 | in the file always wins over the migrated value. Configure the workbar with |
| 2011 | `rail_panel` and `work_surface_placement`, not this key. |
| 2012 | - `sessions_rail` (`on`/`off`; default `off`): show the persistent Sessions |
| 2013 | list in the workbar. Rows list this workspace's recent |
| 2014 | non-archived sessions, newest first, with the active one marked; activating a |
| 2015 | row opens the session picker preselected on it (`/sessions open <id>`), so |
| 2016 | resume keeps its single implementation. Rows are projected from cached |
| 2017 | session metadata — the list never reads a transcript per frame, and never |
| 2018 | contacts a provider. |
| 2019 | - `session_auto_resume` (`on`/`off`; default `off`): reattach to this |
| 2020 | workspace's most recent session when Codewhale starts. Off by default so |
| 2021 | plain `codewhale` keeps starting fresh. `--resume`, `--continue`, and |
| 2022 | `--fresh` always take precedence. When it is on, startup still refuses to |
| 2023 | resume a session that is archived, fails to load, or is recorded against a |
| 2024 | different workspace; each of those falls back to a fresh transcript and says |
| 2025 | which session was skipped and why. It applies to the interactive launch only |
| 2026 | — `codewhale "<prompt>"` and `codewhale exec` are never silently prefixed |
| 2027 | with a prior conversation. |
| 2028 | - `max_input_history` (number of submitted input history entries; cleared |
| 2029 | drafts are also kept locally for composer history search). Note the spelling: |
| 2030 | the serde field on disk is `max_input_history` |
| 2031 | (`crates/tui/src/settings.rs:426`, default 100). `max_history` is the key |
| 2032 | name accepted by `/config set` and `settings.set()` (`settings.rs:1388`), not |
| 2033 | a settings.toml key — writing `max_history` into the file is silently |
| 2034 | ignored. |
| 2035 | - `default_model` (model name override) |
| 2036 | |
| 2037 | `/task digest` (alias `/tasks digest`) renders the canonical Work Graph |
| 2038 | operations and four-state To-do list as plain text, running work first. It |
| 2039 | reads the same snapshots as the styled Work surface and owns no parallel |
| 2040 | progress state. |
| 2041 | |
| 2042 | Plan and Work are the everyday visible modes in the UI; Operate is an explicit |
| 2043 | preview entry while its Workflow control surface is still being built. Switch |
| 2044 | between them with `/mode`. For compatibility, older settings files with |
| 2045 | `default_mode = "normal"` still load as `agent`. |
| 2046 | |
| 2047 | Localization scope is tracked in [LOCALIZATION.md](LOCALIZATION.md). The v0.7.6 |
| 2048 | core pack covers high-visibility TUI chrome only; provider/tool schemas, |
| 2049 | personality prompts, and full documentation remain English unless explicitly |
| 2050 | translated later. |
| 2051 | |
| 2052 | Readability semantics: |
| 2053 | |
| 2054 | - Selection uses a unified style across transcript, composer menus, and modals. |
| 2055 | - Footer hints use a dedicated semantic role (`FOOTER_HINT`) so hint text stays readable across themes. |
| 2056 | |
| 2057 | ### Token Quantities and Drivers |
| 2058 | |
| 2059 | DeepSeek V4 prefix caching makes token labels matter. These quantities are kept |
| 2060 | separate: |
| 2061 | |
| 2062 | | Quantity | Meaning | Allowed to drive | |
| 2063 | |---|---|---| |
| 2064 | | Active request input estimate | Conservative estimate of the next request's live system prompt and transcript payload. | Header/footer context percent, auto-compaction trigger, opt-in Flash seam trigger, and emergency overflow preflight. | |
| 2065 | | Reserved response headroom | The effective request cap plus `1024` safety tokens on every route. Normal no-override requests start at `65536`; a smaller route/provider ceiling narrows that value, and an explicit output override raises it only within the resolved route window and output ceiling. The identical cap reaches the wire and drives preflight; reasoning effort does not add a second hidden reservation. A separately published route input ceiling independently clamps the spendable input budget. | Emergency overflow budget checks only. | |
| 2066 | | Cumulative API usage | Provider-reported input plus output tokens summed across completed API calls; multi-tool turns may count the same stable prefix more than once. | Session usage and approximate cost telemetry only. | |
| 2067 | | Prompt cache hit/miss | Provider cache telemetry for the most recent call when available. | Cache-hit display and cost estimation only; never compaction or seam triggers. | |
| 2068 | | Context percent | Active request input estimate divided by the model context window. | Display only; it mirrors the active-input basis used by context safeguards. | |
| 2069 | | Cost estimate | Approximate spend from provider usage and configured DeepSeek rates. | Display only. | |
| 2070 | |
| 2071 | For known context-window models, including 1M-class V4 models, replacement |
| 2072 | compaction is enabled by default unless the user explicitly configures |
| 2073 | `auto_compact = false`. It fires at the active model's compaction threshold and |
| 2074 | replaces old history with recent user context followed by one ordinary |
| 2075 | checkpoint message. The standing system prompt remains unchanged. Unknown model |
| 2076 | ids remain opt-in. |
| 2077 | |
| 2078 | ### Command Migration Notes |
| 2079 | |
| 2080 | If you are upgrading from older releases: |
| 2081 | |
| 2082 | - Old: `/codewhale` |
| 2083 | New: `/links` (aliases: `/dashboard`, `/api`) |
| 2084 | - Old: `/set model deepseek-reasoner` |
| 2085 | New: `/config` and edit the `model` row to `deepseek-v4-pro` or `deepseek-v4-flash` |
| 2086 | - Old: visible `Normal` mode or `default_mode = "normal"` |
| 2087 | New: use `Agent` / `default_mode = "agent"`; legacy `normal` still maps to `agent` |
| 2088 | - Old: discover `/set` in slash UX/help |
| 2089 | New: use `/config` for editing and `/settings` for read-only inspection |
| 2090 | |
| 2091 | ## Key Reference |
| 2092 | |
| 2093 | ### Kimi Code membership model IDs |
| 2094 | |
| 2095 | The exact `https://api.kimi.com/coding/v1` endpoint accepts `k3`, `k3-256k`, |
| 2096 | `kimi-for-coding`, and `kimi-for-coding-highspeed`. Use `k3-256k` for a fixed |
| 2097 | 262,144-token K3 window; use bare `k3` with `context_window = 1048576` only |
| 2098 | when the membership plan includes the 1M entitlement. Both K3 ids use the same |
| 2099 | reasoning contract, and all four membership ids omit generic sampling fields. |
| 2100 | |
| 2101 | ### Core keys (used by the TUI/engine) |
| 2102 | |
| 2103 | - `provider` (string, optional): `deepseek` (default), `deepseek-anthropic`, `nvidia-nim`, `openai`, `atlascloud`, `wanjie-ark`, `volcengine`, `openrouter`, `xiaomi-mimo`, `novita`, `fireworks`, `siliconflow`, `arcee`, `siliconflow-CN`, `moonshot`, `sglang`, `vllm`, `ollama`, `ollama-cloud`, `huggingface`, `modelscope`, `together`, `qianfan`, `openai-codex`, `anthropic`, `openmodel`, `zai`, `stepfun`, `minimax`, `deepinfra`, `sakana`, `longcat`, `opencode-go`, `meta`, `mistral`, `telecomjs`, `xai`, `orcarouter`, `modelstudio-token-plan`, `google`, `edenai`, or `custom`. Legacy `deepseek-cn` configs are still accepted as an alias for `deepseek`; DeepSeek uses the same official host [`https://api.deepseek.com`](https://api-docs.deepseek.com/) worldwide. `deepseek-anthropic` targets DeepSeek's Anthropic Messages-compatible endpoint at `https://api.deepseek.com/anthropic` using `DEEPSEEK_API_KEY`; `nvidia-nim` targets NVIDIA's NIM-hosted DeepSeek endpoints through `https://integrate.api.nvidia.com/v1`; `openai` targets a generic OpenAI-compatible endpoint, defaulting to `https://api.openai.com/v1`; `atlascloud` targets AtlasCloud's OpenAI-compatible endpoint at `https://api.atlascloud.ai/v1`; `wanjie-ark` targets Wanjie Ark's OpenAI-compatible endpoint at `https://maas-openapi.wanjiedata.com/api/v1`; `volcengine` targets Volcengine Ark's OpenAI-compatible coding endpoint at `https://ark.cn-beijing.volces.com/api/coding/v3`; `openrouter` targets `https://openrouter.ai/api/v1`; `xiaomi-mimo` targets Xiaomi MiMo's OpenAI-compatible endpoint, using `https://token-plan-sgp.xiaomimimo.com/v1` by default for Token Plan keys (`tp-...`) and `https://api.xiaomimimo.com/v1` for pay-as-you-go keys. For Token Plan accounts outside the Singapore default, set `base_url` explicitly or use `mode = "token-plan-cn"` for China and `mode = "token-plan-ams"` for Europe/Amsterdam; `novita` targets `https://api.novita.ai/openai/v1`; `fireworks` targets `https://api.fireworks.ai/inference/v1`; `siliconflow` targets SiliconFlow, defaulting to `https://api.siliconflow.com/v1`; `arcee` targets Arcee AI's OpenAI-compatible endpoint at `https://api.arcee.ai/api/v1`; `siliconflow-CN` targets the SiliconFlow China regional endpoint through `[providers.siliconflow_cn]`; `moonshot` targets Moonshot/Kimi, defaulting to `https://api.moonshot.ai/v1`; `sglang` targets a self-hosted OpenAI-compatible endpoint, defaulting to `http://localhost:30000/v1`; `vllm` targets a self-hosted vLLM OpenAI-compatible endpoint, defaulting to `http://localhost:8000/v1`; `ollama` targets Ollama's OpenAI-compatible endpoint, defaulting to `http://localhost:11434/v1`; `huggingface` targets Hugging Face Inference Providers at `https://router.huggingface.co/v1`; `modelscope` targets ModelScope's OpenAI-compatible inference API at `https://api-inference.modelscope.cn/v1`; `together` targets Together AI at `https://api.together.xyz/v1`; `qianfan` targets Baidu Qianfan at `https://api.baiduqianfan.ai/v1`; `openai-codex` targets official ChatGPT plan inference with Codewhale-owned OAuth at `https://api.openai.com/v1`; `anthropic` targets Claude's native Messages API; `openmodel` targets OpenModel's Anthropic-compatible Messages API at `https://api.openmodel.ai`; `zai` targets Z.ai at `https://api.z.ai/api/coding/paas/v4`; `stepfun` targets StepFun at `https://api.stepfun.ai/v1`; `minimax` targets MiniMax at `https://api.minimax.io/v1`; `deepinfra` targets DeepInfra at `https://api.deepinfra.com/v1/openai`; `sakana` targets Sakana AI Fugu at `https://api.sakana.ai/v1`; `longcat` targets Meituan LongCat at `https://api.longcat.chat/openai/v1`; `opencode-go` targets the subscription-backed OpenCode Go model-aware route (Chat Completions, Responses, or Messages according to the documented model) at `https://opencode.ai/zen/go/v1`; `meta` targets Meta Model API; `mistral` targets Mistral AI's OpenAI-compatible endpoint at `https://api.mistral.ai/v1`; `telecomjs` targets TelecomJS TokenHub at `https://aigw.telecomjs.com/v1`; and `xai` targets xAI's API-key or OAuth route. |
| 2104 | - `opencode-zen` (string provider value): selects the model-aware OpenCode Zen gateway through `[providers.opencode_zen]`. The default base URL is `https://opencode.ai/zen/v1`, the default model is `gpt-5.6`, and credentials come from `api_key`, `OPENCODE_ZEN_API_KEY`, or fallback `OPENCODE_API_KEY`—never ChatGPT/Codex OAuth. `OPENCODE_ZEN_BASE_URL` and `OPENCODE_ZEN_MODEL` are accepted. The selected model's wire comes from the curated Zen snapshot, then from the AI SDK package its Models.dev `opencode` row names (refresh with `codewhale models --update`): GPT, Grok, and Muse Spark use Responses; Claude and most Qwen rows use Anthropic Messages; DeepSeek, MiniMax, GLM, Kimi, `qwen3.8-max`, and the free rows use Chat Completions. Gemini, Models.dev rows marked `deprecated`, and models no loaded catalog lists fail closed because Codewhale has no proven supported wire contract for them. See the exact current model groups in [`PROVIDERS.md`](PROVIDERS.md#opencode-zen-protocol-catalog). |
| 2105 | - `minimax-anthropic` (string provider value): selects MiniMax's Anthropic-compatible Messages route through `[providers.minimax_anthropic]`. The default Base URL is `https://api.minimax.io/anthropic`; set `https://api.minimaxi.com/anthropic` for China. Keep the `/anthropic` suffix because Codewhale appends `/v1/messages`. The route uses `MINIMAX_API_KEY` and defaults to `MiniMax-M3`; `MiniMax-M2.7` is also registered. Official M3 input modalities are text, image, and video, with adaptive or disabled thinking. M2.7 is text-only and always keeps thinking enabled. |
| 2106 | - `api_key` (string, required for hosted providers): must be non-empty for DeepSeek/hosted providers (or set the provider API key env var). Self-hosted SGLang, vLLM, and local `ollama` can omit it. `ollama-cloud` requires a key saved for that provider or supplied by `OLLAMA_CLOUD_API_KEY`, then `OLLAMA_API_KEY`. |
| 2107 | - `auth_mode` (string, optional provider-table key): selects a provider-specific authentication contract. Kimi Code membership uses `auth_mode = "api_key"` (or omit the field), a key created in the [Kimi Code console](https://www.kimi.com/code/console), `base_url = "https://api.kimi.com/coding/v1"`, and bare `model = "k3"` for K3. Codewhale gives that route a safe 262,144-token baseline; set `context_window = 1048576` only when the Kimi Code plan includes 1M access (Allegretto and above). `k3[1m]` is a Claude Code-only convention, not an API model ID, and Codewhale rejects it instead of silently changing the wire model or assuming an entitlement. `model = "kimi-for-coding"` remains the valid K2.7 compatibility route available to all Kimi Code members. Legacy `auth_mode = "kimi_oauth"` fails closed with API-key guidance and never probes, reads, refreshes, or rewrites `kimi_cli`/`kimi_code_cli` credential files. First-class OAuth requires Codewhale's own vendor-registered client identity and remains tracked in #4417. |
| 2108 | - `base_url` (string, optional, `[providers.<name>]` key; see [Legacy top-level `base_url` and `api_key`](#legacy-top-level-base_url-and-api_key) for the older top-level spelling): defaults to `https://api.deepseek.com/beta` for DeepSeek's OpenAI-compatible Chat Completions API, including legacy `provider = "deepseek-cn"` configs. Other defaults are `https://api.deepseek.com/anthropic` for `deepseek-anthropic`, `https://integrate.api.nvidia.com/v1` for `nvidia-nim`, `https://api.openai.com/v1` for `openai`, `https://api.atlascloud.ai/v1` for `atlascloud`, `https://maas-openapi.wanjiedata.com/api/v1` for `wanjie-ark`, `https://ark.cn-beijing.volces.com/api/coding/v3` for `volcengine`, `https://openrouter.ai/api/v1` for `openrouter`, `https://token-plan-sgp.xiaomimimo.com/v1` for `xiaomi-mimo` when the API key starts with `tp-...` and `https://api.xiaomimimo.com/v1` otherwise, `https://api.novita.ai/openai/v1` for `novita`, `https://api.fireworks.ai/inference/v1` for `fireworks`, `https://api.siliconflow.com/v1` for `siliconflow`, `https://api.siliconflow.cn/v1` for `siliconflow-CN`, `https://api.arcee.ai/api/v1` for `arcee`, `https://api.moonshot.ai/v1` for `moonshot`, `https://api.minimax.io/v1` for `minimax`, `https://api.openmodel.ai` for `openmodel`, `https://api.z.ai/api/coding/paas/v4` for `zai`, `https://api.stepfun.ai/v1` for `stepfun`, `https://api.deepinfra.com/v1/openai` for `deepinfra`, `https://api.sakana.ai/v1` for `sakana`, `https://router.huggingface.co/v1` for `huggingface`, `https://api-inference.modelscope.cn/v1` for `modelscope`, `https://api.together.xyz/v1` for `together`, `https://api.baiduqianfan.ai/v1` for `qianfan`, `https://api.openai.com/v1` for `openai-codex`, `https://api.anthropic.com` for `anthropic`, `https://api.mistral.ai/v1` for `mistral`, `http://localhost:30000/v1` for `sglang`, `http://localhost:8000/v1` for `vllm`, `http://localhost:11434/v1` for `ollama`, and `https://ollama.com/v1` for `ollama-cloud`. Set `base_url = "https://token-plan-cn.xiaomimimo.com/v1"` for China-region Xiaomi MiMo Token Plan accounts or `base_url = "https://token-plan-ams.xiaomimimo.com/v1"` for Europe/Amsterdam accounts. Mistral-specific reasoning fields and polymorphic replay are enabled only on the documented first-party HTTPS `/v1` hosts; a custom Mistral base URL keeps generic Chat semantics. Set `https://api.deepseek.com` or `https://api.deepseek.com/v1` explicitly to opt out of DeepSeek beta features. |
| 2109 | - `ollama-cloud` route: select `provider = "ollama-cloud"`, configure `[providers.ollama_cloud]` when overriding the default `https://ollama.com/v1` / `gpt-oss:120b` tuple, and save a key from [Ollama account settings](https://ollama.com/settings/keys) with `codewhale auth set --provider ollama-cloud`. Ambient precedence is `OLLAMA_CLOUD_API_KEY`, then `OLLAMA_API_KEY`; arbitrary Ollama model IDs pass through unchanged. |
| 2110 | - Legacy Ollama Cloud migration: a released `provider = "ollama"` config whose normalized `[providers.ollama].base_url` is exactly `https://ollama.com/v1` is upgraded to the `ollama-cloud` runtime identity in memory. Only that exact tuple may read its old `ollama` provider table and secret slot. The config and secrets are never rewritten, and neighboring paths, HTTP downgrades, lookalike hosts, or an explicit `ollama-cloud` selection never consume the fallback. |
| 2111 | - `telecomjs` base URL and catalog: `[providers.telecomjs]` defaults to `https://aigw.telecomjs.com/v1`; `TELECOMJS_BASE_URL` overrides it. With `TELECOMJS_API_KEY`, `/models` refreshes a key-scoped catalog without mixing rows into another provider. |
| 2112 | - `edenai` gateway: select `provider = "edenai"`; `[providers.edenai]` defaults to `https://api.edenai.run/v3` and `deepseek/deepseek-v4-pro`. `EDENAI_API_KEY`, `EDENAI_BASE_URL`, and `EDENAI_MODEL` are accepted. Use `EDENAI_BASE_URL = "https://api.eu.edenai.run/v3"` for Eden AI's documented EU endpoint; the default `deepseek/deepseek-v4-pro` is only listed on the global catalog, so pair the EU endpoint with an EU-listed model such as `qwen/deepseek-v4-pro` via `EDENAI_MODEL` or `model`. The provider refreshes Eden AI's `/models` catalog, but leaves model-specific reasoning controls untouched because the gateway spans multiple model families. |
| 2113 | - `codewhale` (Codewhale API): select `provider = "codewhale"`; `[providers.codewhale]` defaults to `https://api.codewhale.net/v1` and `deepseek/deepseek-v4-pro`. The credential is a Codewhale account API key (`cwc_key_…`) with the `models:infer` scope, read from `CODEWHALE_API_KEY` or the `codewhale` secret-store slot; `codewhale account api-keys create --name <name> --use` mints one and saves it locally. `CODEWHALE_API_BASE` overrides the origin and must be HTTPS except on loopback. The model catalog is the account's own authenticated `GET /v1/models`: ids are `provider/model` and each row states its wire (`chat-completions` → `/v1/chat/completions`, `anthropic-messages` → `/v1/messages`, `responses` → `/v1/responses`). Connect the underlying provider keys with `codewhale account keys set <provider>`. |
| 2114 | - `concentrate` gateway: select `provider = "concentrate"`; `[providers.concentrate]` defaults to `https://api.concentrate.ai/v1` and `deepseek-v4-pro` over the OpenAI Responses wire. `CONCENTRATE_API_KEY`, `CONCENTRATE_BASE_URL`, and `CONCENTRATE_MODEL` are accepted. Model ids pass through verbatim (`gpt-5.6-sol`, `openai/gpt-5.6-sol`, or `concentrate/auto` for the gateway router). BYOK only; see [PROVIDERS.md](PROVIDERS.md#concentrate-notes). |
| 2115 | - `mistral` model and reasoning contract: `[providers.mistral]` defaults to `mistral-code-latest`; `MISTRAL_MODEL` overrides it and the generic `CODEWHALE_MODEL` override wins when both are set. The current picker also lists `mistral-medium-latest`, `mistral-small-latest`, and `mistral-large-latest`. On exact first-party HTTPS `/v1` routes, Medium and Small accept only `reasoning_effort = "none" | "high"` and replay polymorphic thinking blocks. Deprecated native Magistral IDs may still be configured explicitly, remain always-reasoning, and never receive the adjustable effort field. |
| 2116 | - `context_window` (integer, optional provider-table key): override the total context window for the active `[providers.<name>]` route when an OpenAI-compatible gateway, hosted model alias, or self-hosted runtime has a different limit than Codewhale's static model table. For example, `[providers.openai] context_window = 1000000` lets an OpenAI-compatible DashScope/Qwen route budget against a 1M-token window instead of the conservative fallback. For Kimi Code K3, keep `model = "k3"` and set `[providers.moonshot] context_window = 1048576` only when the membership plan includes 1M access; otherwise omit it to retain the 262,144-token safe baseline. The value must be greater than 0 and affects prompt context notes, compaction thresholds, context-pressure checks, and request output caps. Full resolution order, and how to see which rung produced the current window: [Context length (context window)](#context-length-context-window). |
| 2117 | - `path_suffix` (string, optional provider-table key): override the chat-completions path for OpenAI-compatible gateways that do not serve `/v1/chat/completions`. For example, `[providers.openai] path_suffix = "/chat/completions"` sends chat requests to the unversioned base URL plus `/chat/completions`; `models` and `beta/*` requests keep their normal routing. |
| 2118 | - `reasoning_stream_style` (string, optional provider-table key): override how streaming reasoning is separated from answer text for the active provider route. Use `separate_field` for `reasoning_content` / `reasoning` deltas, `inline_tags` for gateways that stream `<think>...</think>` inside `delta.content`, or `none` to render incoming content exactly as answer text. When unset, every Chat Completions route uses `separate_field` (Mistral's first-party route uses its typed thinking blocks); set `none` only for a gateway that streams its answer inside `reasoning_content`. |
| 2119 | - `[providers.<name>.auth]` (table, optional): provider-scoped auth source metadata. `source = "command"` stores a command argv plus optional `timeout_ms`; `source = "secret"` stores a `secret_id`. This slice lets provider readiness, `/provider`, and doctor JSON report the auth source class without exposing command argv output or secret values; executing commands and resolving external secret material is handled by the follow-up resolver work. |
| 2120 | - `insecure_skip_tls_verify` (bool, optional provider-table key): legacy compatibility key, disabled by default. When true on the active provider table, provider clients reject the configuration instead of skipping TLS certificate verification. Use `SSL_CERT_FILE` for corporate or private CA bundles; `codewhale doctor` reports stale uses of this setting. |
| 2121 | - `default_text_model` (string, optional): defaults to `deepseek-flash` for DeepSeek and `deepseek-anthropic`, `gpt-5.6` for OpenAI, `grok-4.6` for xAI, `deepseek-ai/deepseek-v4-pro` for NVIDIA NIM, `deepseek-ai/deepseek-v4-flash` for AtlasCloud, `deepseek-reasoner` for Wanjie Ark, `DeepSeek-V4-Pro` for Volcengine Ark, `deepseek/deepseek-v4-pro` for OpenRouter and Novita, `mimo-v2.5-pro` for Xiaomi MiMo, `accounts/fireworks/models/deepseek-v4-pro` for Fireworks, `deepseek-ai/DeepSeek-V4-Pro` for SiliconFlow and DeepInfra, `trinity-large-thinking` for Arcee AI, `kimi-k2.7-code` for Moonshot, `MiniMax-M3` for MiniMax, `GLM-5.3` for Z.ai, `step-3.7-flash` for StepFun, `ernie-4.0-turbo-8k` for Qianfan, `fugu` for Sakana AI, `deepseek-ai/DeepSeek-V4-Pro` for SGLang/vLLM, no fixed id for local Ollama (Codewhale adopts a chat-capable tag from the live local catalog — coder or tool-capable tags first, then the largest context — and never an embedding or reranker model), and `gpt-oss:120b` for Ollama Cloud. Hugging Face and Together AI both default to `deepseek-ai/DeepSeek-V4-Pro`; `openai-codex` defaults to `gpt-5.6`; `anthropic` defaults to `claude-sonnet-4-6`; `openmodel` defaults to `deepseek-v4-flash`. Current public DeepSeek IDs include `deepseek-v4-pro` and `deepseek-flash` (V4.1 Flash, shipped as the unversioned id), both with 1M context windows, 384K max output, and thinking mode enabled by default. DeepSeek's live pricing/model page now labels the Pro backend `DeepSeek-V4-Pro-0813`; the callable API ID remains `deepseek-v4-pro`, so Codewhale does not send the backend label or the Claude Code-specific `deepseek-v4-pro[1m]` selector. DeepSeek retires `deepseek-chat` and `deepseek-reasoner` on July 24, 2026; direct first-party routes migrate both to `deepseek-v4-flash`, with omitted reasoning settings preserving their former non-thinking (`off`) and thinking (`high`) intent. Explicit `reasoning_effort` wins, and provider-owned ids on Wanjie Ark, aggregators, self-hosted runtimes, and custom endpoints are not globally rewritten. SiliconFlow retains its own mapping: `deepseek-reasoner` and `deepseek-r1` select its Pro model while `deepseek-chat` and `deepseek-v3` select Flash. Provider-specific mappings translate `deepseek-v4-pro` / `deepseek-v4-flash` to each provider's model ID where supported. OpenRouter also recognizes recent large IDs such as `arcee-ai/trinity-large-thinking`, `minimax/minimax-m3`, `minimax/minimax-m2.7`, `xiaomi/mimo-v2.5-pro`, `qwen/qwen3.6-flash`, `qwen/qwen3.6-35b-a3b`, `qwen/qwen3.6-max-preview`, `qwen/qwen3.6-27b`, `qwen/qwen3.6-plus`, `qwen/qwen3.7-max`, `google/gemma-4-31b-it`, `moonshotai/kimi-k2.7-code`, `moonshotai/kimi-k2.6`, `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free`, and `nvidia/nemotron-3-ultra-550b-a55b`; direct Arcee uses bare IDs such as `trinity-large-thinking` and `trinity-large-preview`; direct Moonshot recognizes `kimi-k3`, `kimi-k2.7-code`, and `kimi-k2.6`. The exact Kimi Code endpoint recognizes bare `k3` for K3 and `kimi-for-coding` for K2.7; those membership IDs are distinct from the direct Moonshot IDs and are never rewritten across routes. Direct MiniMax recognizes `MiniMax-M3` and the documented M2.x chat model IDs; direct Z.ai recognizes `GLM-5.3` (the default), `GLM-5.2`, `GLM-5.1`, and `GLM-5-Turbo`, and OpenRouter recognizes the matching `z-ai/glm-5.1`, `z-ai/glm-5.2`, `z-ai/glm-5.3`, and `z-ai/glm-5-turbo` IDs — `GLM-5.3` has been live on the Z.ai Coding Plan since 2026-08-13; it inherits its catalog metadata from `GLM-5.2` until Z.ai publishes distinct 5.3 numbers and carries no price, and an explicit `GLM-5.2` selection keeps its own id; direct Sakana recognizes `fugu` and `fugu-ultra-20260615`; direct Xiaomi MiMo recognizes chat IDs `mimo-v2.5-pro`, `mimo-v2.5-pro-ultraspeed`, and `mimo-v2.5`, while TTS IDs are selected through `codewhale speech` / `tts`. Generic `openai`, `atlascloud`, `wanjie-ark`, `xiaomi-mimo`, `arcee`, `moonshot`, `minimax`, `openmodel`, `zai`, `stepfun`, `qianfan`, `sakana`, local Ollama, and Ollama Cloud model IDs are passed through unchanged after known aliases are normalized. OpenRouter and SiliconFlow provider configs with a custom `base_url` also preserve explicit model values, which lets OpenAI-compatible gateways accept bare model IDs. Use `/models` or `codewhale models` to discover live IDs from your configured endpoint. `CODEWHALE_MODEL` overrides this for a single process; `DEEPSEEK_MODEL` is the legacy alias. |
| 2122 | - TelecomJS uses `deepseek-v4-pro` only as a conservative pre-refresh fallback. Once its key-scoped `/models` catalog is available, the picker uses those live rows; Codewhale omits unsupported reasoning request fields on this route. |
| 2123 | - `reasoning_effort` (string, optional): `auto`, `off`, `minimal`, `low`, `medium`, `high`, `xhigh`, `ultra` (`ultracode` is an alias of `ultra`), or `max`; defaults to the configured UI tier. DeepSeek Platform receives top-level `thinking` / `reasoning_effort` fields. Ollama Cloud's OpenAI-compatible Chat Completions route preserves its documented `none` / `low` / `medium` / `high` / `max` ladder (`off` is sent as `none`; `xhigh` and `ultracode` normalize to `max`). Direct xAI `grok-4.7` and `grok-4.6` on exact `https://api.x.ai/v1` receive top-level `reasoning_effort = "low" | "medium" | "high" | "xhigh"` (the ladder comes from the bundled catalog row; `grok-4.5` maps `xhigh` to `high`, and rows without a documented effort such as `grok-4.3` get no field); Grok reasoning cannot be disabled, so `off` normalizes to `high`, `max`/`ultracode` to `xhigh`, and `auto` leaves the field omitted so xAI's documented default `high` applies. A custom xAI-compatible `base_url` does not inherit that dialect. Direct Moonshot `kimi-k3` on exact `https://api.moonshot.ai/v1` is always-thinking and receives only top-level `reasoning_effort = "low" | "high" | "max"`; `off` normalizes to `low`, and `medium` to `high`. Kimi Code membership `k3` on exact `https://api.kimi.com/coding/v1` instead receives nested `thinking.effort`, and its `off` setting also normalizes to enabled `low`. Normal dispatched `auto` uses Codewhale's auto-reasoning selector and sends a concrete route-normalized tier; only an omitted reasoning setting leaves the provider default in control. Neighboring gateways and model/endpoint combinations retain the generic Moonshot contract. OpenAI Codex normalizes stale `off` to `low` and sends `max` / `ultracode` as Responses `xhigh`. Z.ai receives documented `thinking` controls and treats enabled thinking as the GLM coding high/max lane. NVIDIA NIM receives equivalent settings through `chat_template_kwargs`. |
| 2124 | - `verbosity` (string, optional): `normal` or `concise`. `normal` keeps the |
| 2125 | default conversational prompt. `concise` appends a prompt discipline block |
| 2126 | for direct, low-chatter output; CLI noninteractive commands (`exec` and |
| 2127 | `eval`) default to `concise` unless config/env/CLI overrides it. |
| 2128 | Override per process with `CODEWHALE_VERBOSITY` or the legacy |
| 2129 | `DEEPSEEK_VERBOSITY` alias. |
| 2130 | - `telemetry` (bool, optional): anonymous usage counting, **`true` by default |
| 2131 | in the current 0.9.12 source**. Notice version `5` names Codewhale and PostHog |
| 2132 | and describes the opt-out policy; no acceptance is invented for a default |
| 2133 | user. Existing explicit declines remain off. An explicit `false` here is the |
| 2134 | durable *opt-out*: it deletes the random install id, truncates buffered and |
| 2135 | dry-run events, and leaves a tombstone reasserted while the setting is false. |
| 2136 | It is a floor: `--telemetry true` and `CODEWHALE_TELEMETRY=1` lose to it. |
| 2137 | Use `/settings` or `codewhale config set telemetry true` to turn counting |
| 2138 | back on explicitly for new sessions through the existing privacy transition. `CODEWHALE_TELEMETRY` |
| 2139 | (legacy alias `DEEPSEEK_TELEMETRY`) and `--telemetry false` provide a run-scoped |
| 2140 | kill switch that stops collection and delivery without erasing the owner's |
| 2141 | state. A repo-local `.codewhale/config.toml` cannot set this preference. |
| 2142 | `codewhale config telemetry` shows the disclosure; `codewhale config get |
| 2143 | telemetry` reports preference and privacy status. Full schema and opt-out |
| 2144 | behavior: [`TELEMETRY.md`](TELEMETRY.md). |
| 2145 | - `telemetry_endpoint` (string, optional): where batches are POSTed. Leaving it |
| 2146 | unset selects the shipped default, |
| 2147 | **`https://telemetry.codewhale.net/v1/telemetry`** — the first-party ingest |
| 2148 | service described in [`TELEMETRY.md`](TELEMETRY.md), whose source is in |
| 2149 | `telemetry-ingest/`. This key decides only *where* a permitted session sends; |
| 2150 | it cannot override an opt-out. Setting it |
| 2151 | to the **empty string** is how you stay enabled and contact nobody: each batch |
| 2152 | is then written to `$CODEWHALE_HOME/telemetry/dryrun.jsonl` and no HTTP client |
| 2153 | is constructed at all, so you can read exactly what would have been sent. Any |
| 2154 | other value replaces the default outright. `https://` is required; plain |
| 2155 | `http://` is accepted only for loopback hosts, and there is no environment |
| 2156 | variable that overrides that refusal. A rejected endpoint turns telemetry off |
| 2157 | for the run rather than falling back to plaintext or to the default. Override |
| 2158 | per process with `CODEWHALE_TELEMETRY_ENDPOINT` (legacy alias |
| 2159 | `DEEPSEEK_TELEMETRY_ENDPOINT`), where an empty value means the same "contact |
| 2160 | nobody". A repo-local `.codewhale/config.toml` cannot set it. |
| 2161 | - `allow_shell` (bool, optional): in interactive TUI Agent sessions, omitting |
| 2162 | this keeps shell tools available with approval prompts; setting it to `false` |
| 2163 | hides shell tools. Headless, durable-task, and other noninteractive profiles |
| 2164 | keep the conservative omitted-field default and require `allow_shell = true` |
| 2165 | to expose shell. Plan mode always hides shell; Full Access enables shell and |
| 2166 | auto-approval. |
| 2167 | - `approval_policy` (string, optional): `on-request`, `untrusted`, or `never`. Runtime `approval_mode` editing in `/config` also accepts `on-request` and `untrusted` aliases. |
| 2168 | - `[approval] default_selection` (string, optional): which option an approval |
| 2169 | card highlights when it first appears — `deny` (default) or `allow_once`. |
| 2170 | `deny` means a reflexive Enter on a card you have not read refuses the call. |
| 2171 | Set `allow_once` to restore the pre-v0.9.6 Enter-to-approve muscle memory |
| 2172 | (#5293). It moves the highlight only: which calls are prompted for is still |
| 2173 | `approval_policy` plus the rules in `permissions.toml`. |
| 2174 | |
| 2175 | ```toml |
| 2176 | [approval] |
| 2177 | default_selection = "allow_once" |
| 2178 | ``` |
| 2179 | - `[approval] timeout_seconds` (integer, optional): bound how long an |
| 2180 | approval may wait — the TUI card and the Runtime API approvals the desktop |
| 2181 | app and web use alike. When the window elapses the call is refused |
| 2182 | (fail-closed) and recorded as a timeout, not as the operator's denial. |
| 2183 | Omitted or `0` waits indefinitely, which is the default everywhere: no |
| 2184 | approval is ever denied on your behalf unless you set this. Values above |
| 2185 | 24h clamp with a warning (#6101). |
| 2186 | |
| 2187 | ```toml |
| 2188 | [approval] |
| 2189 | timeout_seconds = 300 |
| 2190 | ``` |
| 2191 | - `sandbox_mode` (string, optional): `read-only`, `workspace-write`, `danger-full-access`, `external-sandbox`. |
| 2192 | Platform support is not identical. macOS uses Seatbelt when its runtime |
| 2193 | probe succeeds. Linux uses bubblewrap only when `prefer_bwrap = true` and |
| 2194 | `/usr/bin/bwrap` is executable; without that opt-in it reports no OS command |
| 2195 | sandbox. Windows does not currently advertise an OS sandbox; its planned helper contract starts |
| 2196 | with process-tree containment only and must not be described as read-only |
| 2197 | filesystem isolation, workspace-write enforcement, network blocking, |
| 2198 | registry isolation, or AppContainer isolation until those are implemented. |
| 2199 | - The cross-layer relationship between mode admission, hooks, registered tool |
| 2200 | requirements, typed rules, auto-review, repo law, human approval, and the |
| 2201 | execution sandbox is defined in |
| 2202 | [Authorization order](AUTHORIZATION_ORDER.md). |
| 2203 | - **Read deny-list.** Every sandbox posture — `read-only` included — grants |
| 2204 | read access to the whole filesystem; the postures differ in what they may |
| 2205 | *write* and whether they may reach the network. The read deny-list narrows |
| 2206 | that: |
| 2207 | - `sandbox_read_denylist_defaults` (bool, default `true`): apply the built-in |
| 2208 | credential-store set — `~/.ssh`, `~/.gnupg`, cloud credential directories |
| 2209 | (`~/.aws`, `~/.config/gcloud`, `~/.azure`, `~/.kube`, …), `~/.netrc`, |
| 2210 | `~/.npmrc`, `~/.git-credentials`, macOS keychains, browser profiles, |
| 2211 | Codewhale's own secret stores, and `.env` files (but not `.env.example` |
| 2212 | and friends). Ordinary source, `Cargo.toml`, `~/.gitconfig`, `~/.cargo`, |
| 2213 | and `~/.npm` stay readable so builds and tests still work. Set `false` to |
| 2214 | restore the pre-0.9.12 full-disk-read behavior. |
| 2215 | - `sandbox_denied_read_paths` (list of paths): additional denied subpaths. |
| 2216 | `~` expands. These can never be exempted. |
| 2217 | - `sandbox_read_denylist_exempt` (list of paths): subtract a path from the |
| 2218 | *built-in defaults* when a project genuinely needs it. Deny wins over |
| 2219 | allow: this never reopens anything in `sandbox_denied_read_paths`. |
| 2220 | |
| 2221 | Exemption granularity is **whole-rule**, not per-file. An exempt path |
| 2222 | removes a built-in rule only when the rule's own path is at or below it, |
| 2223 | so exempting `~/.ssh/config` does nothing: the `~/.ssh` rule still denies |
| 2224 | it, because `~/.ssh` does not lie within `~/.ssh/config`. To reopen that |
| 2225 | one file you must exempt `~/.ssh` itself — which also reopens the private |
| 2226 | keys next to it. That is the documented tradeoff: there is no shipped way |
| 2227 | to narrow a built-in rule to "everything except one file"; copy what you |
| 2228 | need out of the denied tree instead. The one name-shaped rule, `.env` |
| 2229 | files, is exempted by *name* rather than by path: any exempt entry whose |
| 2230 | file name is exactly `.env` — bare `.env`, `~/.env`, |
| 2231 | `some/project/.env` — disables the entire `.env` filename rule, i.e. |
| 2232 | every `.env` and `.env.<name>` on disk rather than one project's. |
| 2233 | (`.env.example` and friends are never denied, so they need no exemption.) |
| 2234 | |
| 2235 | Enforced at two points: sandboxed shell commands (Seatbelt last-match-wins |
| 2236 | `deny file-read*` rules; bubblewrap masks each path) and Codewhale's own |
| 2237 | in-process tools, which the OS sandbox never wraps — `read_file` / `read` / |
| 2238 | `read_media` for contents, `list_dir` / `file_search` for *enumeration* |
| 2239 | (listing a denied directory, or searching one, is refused just as Seatbelt |
| 2240 | blocks its readdir; a name search rooted above a denied tree skips entries |
| 2241 | inside it). A refused read is always an explicit error, never an empty |
| 2242 | result, and the error names the path as the caller spelled it rather than a |
| 2243 | symlink target's real location. |
| 2244 | |
| 2245 | **This is defense-in-depth, not a security boundary.** It does not stop a |
| 2246 | hardlink to a denied file, a secret already copied into the workspace, an |
| 2247 | indirect read (`ssh-agent`, `security find-generic-password`, `aws sts …`), |
| 2248 | reads under `danger-full-access` shell commands, reads by MCP servers or |
| 2249 | other unwrapped child processes, or exfiltration of anything that *was* |
| 2250 | read. Keep least-privilege credentials and short-lived tokens doing the real |
| 2251 | work. |
| 2252 | - `permissions.toml` (sibling file, optional): typed permission rule records |
| 2253 | loaded next to `config.toml`, for example `~/.codewhale/permissions.toml`. |
| 2254 | This active user file is the only permission-rule source today; project |
| 2255 | config overlays do not load a project-local `permissions.toml`. A rule's |
| 2256 | optional `workspace` field is its repository scope, not a second source. |
| 2257 | Manually authored `[[rules]]` entries accept `tool`, optional `command` or |
| 2258 | `path`, optional absolute `workspace`, optional `command_exact = true`, and |
| 2259 | optional `action = "deny" | "ask" | "allow"`; omitted `action` defaults to |
| 2260 | `"ask"`. `workspace` limits a rule to that repository, while |
| 2261 | `command_exact = true` changes a command rule from the historical |
| 2262 | arity-aware prefix match to a complete-command match. `deny` blocks matching |
| 2263 | invocations before mode-based |
| 2264 | approval handling, `allow` skips approval for matching invocations, and |
| 2265 | `ask` forces approval only in modes that can prompt. Outside the TUI |
| 2266 | auto-approve path, a matching `ask` rule under `approval_policy = "never"` |
| 2267 | is rejected because no prompt can be shown. In Full Access / auto-approval sessions, |
| 2268 | `ask` rules do not downgrade the session into prompting or blocking; explicit |
| 2269 | `deny` rules still block according to the current execution-policy logic. |
| 2270 | |
| 2271 | In a supported approval card, press `S` to allow the request once and append |
| 2272 | exact `action = "ask"` rules to this file. For eligible safe requests, choose |
| 2273 | **Always allow this exact rule in this repo** (shortcut `P`) to append an |
| 2274 | `action = "allow"` rule with the current absolute `workspace` scope. |
| 2275 | Remembered shell grants set `command_exact = true`, so later commands with |
| 2276 | extra arguments do not inherit the grant. File and patch grants retain the |
| 2277 | exact workspace-relative paths produced by the existing validation path. |
| 2278 | Supported saves are intentionally narrow: |
| 2279 | `exec_shell` stores the exact approved command string; `write_file` and |
| 2280 | `edit_file` store the exact workspace-relative file path; `apply_patch` |
| 2281 | stores one exact workspace-relative `path` rule per validated touched file |
| 2282 | from apply-patch preflight. Existing exec command matching remains |
| 2283 | arity-aware for manually authored prefix rules; approval-card allow grants |
| 2284 | use complete-command matching. File paths are normalized to the same |
| 2285 | workspace-relative form used by runtime matching. |
| 2286 | |
| 2287 | `read_file` rules can still be authored manually when you want future reads |
| 2288 | of a specific path to ask, allow, or deny, but the approval UI does not save |
| 2289 | `read_file` rules. Commands classified as requiring approval or dangerous, |
| 2290 | critical approval cards, and repo-law prompts cannot save allow grants and |
| 2291 | continue to require review. |
| 2292 | |
| 2293 | `/permissions` (or `/permissions list`) is the narrow rule-management |
| 2294 | surface. It lists each numbered rule with the active user-file source, its |
| 2295 | exact effective matcher (tool-wide, command prefix, exact command, or exact |
| 2296 | normalized path), global or repository scope, and whether that scope |
| 2297 | applies in the current workspace. `/config ask-rules` remains a compatibility |
| 2298 | entry to the same list. |
| 2299 | |
| 2300 | Deletion is review-gated: `/permissions remove <number>` only previews the |
| 2301 | selected rule and prints a confirmation command. That command carries an |
| 2302 | opaque token bound to the exact file bytes and rule index; if another writer |
| 2303 | changes `permissions.toml`, confirmation fails instead of deleting a rule |
| 2304 | that moved into the old position. Confirmed removal and approval-card appends |
| 2305 | share the adjacent `permissions.toml.lock`, preserve unrelated TOML comments |
| 2306 | and formatting, and atomically replace the file. The running TUI reloads the |
| 2307 | user ruleset without clearing session-only approvals. |
| 2308 | |
| 2309 | This editor intentionally does not create or rewrite rules, persist deny |
| 2310 | choices from approval cards, expand globs, or create broad |
| 2311 | directory/recursive rules. Author those supported exact/prefix records |
| 2312 | manually when needed. |
| 2313 | - `[[hotbar]]` (array of tables, optional): user-owned 1-8 slot bindings for |
| 2314 | the TUI hotbar. Each entry has `slot`, `action`, and optional `label`. |
| 2315 | The Hotbar is hidden by default (#3807): omitting `hotbar` and setting |
| 2316 | `hotbar = []` both mean no hotbar. `/hotbar on` writes the eight default |
| 2317 | slots (`slash.workflow`, `slash.goal`, `slash.auto`, `mode.plan`, |
| 2318 | `mode.agent`, `mode.operate`, `palette.open`, `sidebar.toggle`) as explicit |
| 2319 | `[[hotbar]]` tables. When one or more `[[hotbar]]` tables are present, that |
| 2320 | list is the whole bar; missing slots stay empty. Invalid slots outside `1..=8` are skipped with a warning, duplicate |
| 2321 | slots use the later entry, and unknown action IDs are kept so the UI can show |
| 2322 | a disabled/unknown cell instead of silently deleting user config. Trusted |
| 2323 | user config, profiles, and managed config replace the whole list; project |
| 2324 | overlays cannot change hotbar bindings. Setup or wizard flows that persist |
| 2325 | hotbar bindings write this same schema to the resolved `~/.codewhale/config.toml` |
| 2326 | path, preserving legacy `~/.deepseek/config.toml` only when that fallback file |
| 2327 | is already the active config. |
| 2328 | |
| 2329 | ```toml |
| 2330 | [[hotbar]] |
| 2331 | slot = 1 |
| 2332 | action = "mode.plan" |
| 2333 | label = "Plan" |
| 2334 | |
| 2335 | [[hotbar]] |
| 2336 | slot = 2 |
| 2337 | action = "session.compact" |
| 2338 | ``` |
| 2339 | - `[auto_review]` (table, optional): tool-call review policy — a deterministic floor plus a model guardian tier. |
| 2340 | This layer sits on top of the existing permission posture; it can hold or block a |
| 2341 | tool call, but it is not an auto-push, auto-merge, or hosted review service. |
| 2342 | Block rules are checked first, then the built-in safety floor, then allow |
| 2343 | rules. In Ask, a safety hold opens approval; in Auto-Review, Full Access, |
| 2344 | or a non-interactive `never` posture it fails closed as a hard block. The |
| 2345 | safety floor still covers publish-like actions and destructive |
| 2346 | background/headless actions even if an allow rule matches. |
| 2347 | |
| 2348 | ```toml |
| 2349 | [[auto_review.allow]] |
| 2350 | id = "read-only-inspection" |
| 2351 | action_kind = "read" |
| 2352 | reason = "Read-only inspection is safe to run automatically." |
| 2353 | |
| 2354 | [[auto_review.block]] |
| 2355 | id = "no-release-publish" |
| 2356 | action_kind = "publish" |
| 2357 | reason = "Release and publish actions require maintainer review." |
| 2358 | ``` |
| 2359 | |
| 2360 | Rule matchers are exact `tool` and/or `action_kind`. At least one matcher is |
| 2361 | required. `action_kind` accepts the six decision-relevant kinds `read`, |
| 2362 | `write`, `shell`, `external`, `publish`, and `destructive`. Invalid names |
| 2363 | fail config validation instead of silently broadening into another policy |
| 2364 | class. In block rules, the old names remain conservative compatibility |
| 2365 | aliases: `network`, `git`, `mcp_action`, `browser`, and `unknown` map to |
| 2366 | `external`; `secret` maps to `destructive`; and `mcp_read` maps to `read`. |
| 2367 | Retired narrow kinds in allow rules fail validation rather than widening to |
| 2368 | a broader class. The retired `text_contains` matcher likewise fails |
| 2369 | validation instead of silently broadening an old intent-dependent rule. |
| 2370 | Fallback holds in interactive Auto-Review escalate to one stateless guardian |
| 2371 | request. The request contains the exact held call and deterministic |
| 2372 | observations as separate JSON fields. Conversation history, skill |
| 2373 | instructions, attached file contents, and other expanded model context are |
| 2374 | excluded. The guardian does not infer user intent or compute an authorization |
| 2375 | score. It exposes no tools and returns a risk level, allow/deny, and a |
| 2376 | rationale. High or critical risk cannot auto-run even if the model says |
| 2377 | allow. An oversized exact call is denied rather than truncated. Exactly one |
| 2378 | reviewer request is made; incomplete or malformed output, timeout, |
| 2379 | cancellation, provider failure, or an empty rationale all fail closed. The |
| 2380 | deterministic floor is never model-reviewed, and headless adapters use the |
| 2381 | deterministic-only tier. The pinned Codex, Kimi, and DeepSeek source |
| 2382 | boundaries are linked from [Permission Posture](MODES.md#permission-posture). |
| 2383 | Reviewer outcomes emit `tool.auto_review` audit events with |
| 2384 | `gate = "guardian"`. |
| 2385 | |
| 2386 | Auto-review decisions emit `tool.auto_review` audit events with |
| 2387 | `gate = "deterministic"` when tool |
| 2388 | audit logging is enabled. Future PreToolUse/PostToolUse hooks can add |
| 2389 | observer input around this layer, but the configured auto-review policy is |
| 2390 | evaluated before a tool call is allowed to proceed. |
| 2391 | - `managed_config_path` (string, optional): managed config file loaded after user/env config. |
| 2392 | - `requirements_path` (string, optional): requirements file used to enforce allowed approval/sandbox values. |
| 2393 | - `max_subagents` (int, optional): defaults to `64` and is clamped to `1..=128`. |
| 2394 | - `subagents.*` (optional compatibility table): manual per-role model pins |
| 2395 | for direct and Workflow `agent` starts. An explicit saved profile wins, |
| 2396 | then a manual role pin, then a unique saved role pin. Conflicting tool |
| 2397 | `model` or `model_strength` choices are refused before admission. Unpinned |
| 2398 | roles allow task model/strength choices before inherited defaults. |
| 2399 | `[subagents.roles.<role>] model = "provider/model"` folds into the existing override |
| 2400 | map and wins over `[subagents.models]`, then the convenience keys. Structured |
| 2401 | canonical role keys win over legacy aliases. Only this structured syntax |
| 2402 | separates the explicit provider from the model suffix; unknown providers |
| 2403 | fail before admission. Bare structured model ids inherit the active provider. |
| 2404 | Legacy scalar/map values preserve namespaced provider-owned ids unchanged. |
| 2405 | Supported convenience keys are |
| 2406 | `default_model`, `worker_model`, `scout_model`, `planner_model`, |
| 2407 | `reviewer_model`, `custom_model`, `max_concurrent`, `max_admitted`, |
| 2408 | `launch_concurrency`, `token_budget`, `api_timeout_secs`, and |
| 2409 | `heartbeat_timeout_secs`. The v0.9.x keys `explorer_model`, `awaiter_model`, |
| 2410 | and `review_model` remain accepted as aliases. The `[subagents] |
| 2411 | max_concurrent` value overrides |
| 2412 | top-level `max_subagents` and is also clamped to `1..=128`. `[subagents] |
| 2413 | max_admitted` (aliases: `max_total`, `admission_limit`) is the bounded total |
| 2414 | of queued plus running sub-agents; it defaults to `1024` |
| 2415 | (`MAX_SUBAGENT_ADMISSION`, `crates/tui/src/config/subagent_limits.rs:21`, |
| 2416 | applied at `config.rs:6400`) so high-fanout turns can queue and drain while |
| 2417 | runtime launch pressure remains bounded, and is clamped to |
| 2418 | `max_concurrent..=1024`. `[subagents] |
| 2419 | launch_concurrency` sets how many direct children start at once before the |
| 2420 | rest queue for a launch slot; it defaults to the resolved `max_subagents` cap |
| 2421 | and is clamped to `1..=max_subagents` (the deprecated |
| 2422 | `interactive_max_launch` key is accepted as an alias, with the new key |
| 2423 | winning when both are set). `[subagents] token_budget` is an optional |
| 2424 | aggregate token ceiling for each root `agent` run and its descendants; unset |
| 2425 | or `0` preserves unlimited legacy behavior. `[subagents] api_timeout_secs` |
| 2426 | controls the per-step API timeout for sub-agent model calls and is clamped to |
| 2427 | `1..=3600`, with `0` or unset preserving the 600 second default; a timed-out |
| 2428 | attempt is retried with exponential backoff (up to 5 retries) before the |
| 2429 | step interrupts with a preserved checkpoint. |
| 2430 | `[subagents] heartbeat_timeout_secs` controls stale running agent cleanup, |
| 2431 | defaults to `300`, and is clamped to `30..=3600` while staying above the |
| 2432 | resolved API timeout. `[subagents.providers.<provider>]` accepts the same |
| 2433 | fanout, depth, budget, and timeout knobs (`enabled`, `max_concurrent`, |
| 2434 | `max_admitted`, `launch_concurrency`, `max_depth`, `token_budget`, |
| 2435 | `api_timeout_secs`, `heartbeat_timeout_secs`) and inherits the global |
| 2436 | `[subagents]` value for any key you omit. Provider keys accept canonical |
| 2437 | names such as `deepseek`, `zai`, `openrouter`, `anthropic`, plus convenience |
| 2438 | aliases such as `glm` for Z.ai and `deepseek_api` for direct DeepSeek: |
| 2439 | |
| 2440 | ```toml |
| 2441 | [subagents] |
| 2442 | max_concurrent = 20 |
| 2443 | launch_concurrency = 20 |
| 2444 | max_admitted = 200 |
| 2445 | max_depth = 6 |
| 2446 | |
| 2447 | [subagents.providers.deepseek] |
| 2448 | max_concurrent = 20 |
| 2449 | launch_concurrency = 20 |
| 2450 | max_admitted = 200 |
| 2451 | |
| 2452 | [subagents.providers.glm] |
| 2453 | max_concurrent = 4 |
| 2454 | launch_concurrency = 3 |
| 2455 | max_admitted = 12 |
| 2456 | max_depth = 2 |
| 2457 | |
| 2458 | [subagents.providers.openrouter] |
| 2459 | max_concurrent = 5 |
| 2460 | launch_concurrency = 3 |
| 2461 | max_admitted = 20 |
| 2462 | ``` |
| 2463 | |
| 2464 | `/config subagents status` prints both global values and the active |
| 2465 | provider's resolved profile so rate-limit tuning is visible in the TUI. |
| 2466 | `[subagents.models]` accepts lower-case Fleet role keys such as `worker`, |
| 2467 | `scout`, `planner`, `reviewer`, `builder`, and `verifier`; legacy type keys |
| 2468 | remain accepted during v0.9.x. Values are validated |
| 2469 | against the active provider at spawn time; direct DeepSeek requires DeepSeek |
| 2470 | IDs, while OpenAI-compatible/custom provider routes pass explicit model IDs |
| 2471 | through to that provider. To route a child to a different provider than the |
| 2472 | parent session, save a Fleet/AgentProfile with explicit `provider` and |
| 2473 | `model` fields (including user-named custom providers such as `lm-studio`) |
| 2474 | and call `agent(profile: "...")`; see [SUBAGENTS.md](SUBAGENTS.md). |
| 2475 | - `skills_dir` (string, optional): defaults to `~/.codewhale/skills` (each skill is |
| 2476 | a directory containing `SKILL.md`). Workspace-local `.agents/skills` or |
| 2477 | `./skills` are preferred when present; the runtime also discovers global |
| 2478 | agentskills.io-compatible `~/.agents/skills` and the broader Claude-ecosystem |
| 2479 | `~/.claude/skills`. First launch installs versioned bundled skills for common |
| 2480 | workflows including skill creation, delegation, MCP/plugin scaffolding, |
| 2481 | documents, presentations, spreadsheets, PDFs, and Feishu/Lark. Only |
| 2482 | Codewhale-owned roots (`<workspace>/.codewhale/skills` and |
| 2483 | `~/.codewhale/skills`) are writable install/import targets; compatible harness |
| 2484 | roots stay read-only. Bare `/skills` opens the Skills Manager (owned-only, |
| 2485 | zero network). See [SKILLS.md](SKILLS.md) for the manager, audit statuses, |
| 2486 | provenance markers, and mutation rules, and |
| 2487 | [CLAUDE_PLUGIN_COMPAT.md](CLAUDE_PLUGIN_COMPAT.md) for the supported boundary |
| 2488 | between portable `SKILL.md` bundles and Claude Code plugin runtimes. |
| 2489 | - `[skills].scan_codewhale_only` (bool, default `false`): when `true`, session |
| 2490 | skill discovery ignores cross-tool roots such as `.claude/skills`, |
| 2491 | `.opencode/skills`, `.cursor/skills`, and `~/.agents/skills`. Codewhale still |
| 2492 | scans `<workspace>/.codewhale/skills`, `~/.codewhale/skills`, and any explicit |
| 2493 | `skills_dir` override. The Skills Manager can still toggle a local compatible |
| 2494 | audit scan independently of this runtime knob — see [SKILLS.md](SKILLS.md). |
| 2495 | - `[skills].flat_workspace_root` (bool, default `false`): opt in to the flat |
| 2496 | `<workspace>/skills` compatibility root after workspace trust. Without this |
| 2497 | opt-in it is an audit candidate only; an explicit `skills_dir` remains an |
| 2498 | alternative. `scan_codewhale_only = true` excludes the flat compatibility |
| 2499 | root regardless of this flag, unless it is the explicit `skills_dir`. |
| 2500 | - `[skills].registry_url` / `[skills].max_install_size_bytes` (optional): used by |
| 2501 | `/skills --remote`, `/skills suggest <task>`, `/skills sync`, and `/skill |
| 2502 | install|update`. The default manager open path does not contact the registry. |
| 2503 | - `[verifier].enabled` (bool, default `false`): enables automatic |
| 2504 | claim-of-done verifier preview once that runtime trigger is active. The |
| 2505 | manual `run_verifiers` tool is still available when this is false. |
| 2506 | - `mcp_config_path` (string, optional): defaults to `~/.codewhale/mcp.json`, with |
| 2507 | legacy `~/.deepseek/mcp.json` fallback when the Codewhale path is absent. |
| 2508 | Custom paths must be absolute; a relative value falls back to the user-global |
| 2509 | path so changing the launch directory cannot silently change the MCP pool. |
| 2510 | It is visible in `/config` and can be changed from the TUI. The new path is |
| 2511 | used immediately by `/mcp`, but rebuilding the model-visible MCP tool pool |
| 2512 | requires restarting the TUI. |
| 2513 | - `notes_path` (string, optional): defaults to `~/.codewhale/notes.txt`, with |
| 2514 | legacy `~/.deepseek/notes.txt` fallback when the Codewhale path is absent, and |
| 2515 | is used by the model-visible `note` tool. |
| 2516 | - `[memory].enabled` (bool, optional): defaults to `false`. When `true`, |
| 2517 | the TUI loads the user memory file into a `<user_memory>` prompt block, |
| 2518 | enables `# foo` quick-capture in the composer, surfaces the `/memory` |
| 2519 | slash command, and registers the `remember` tool. The same toggle is |
| 2520 | available via `DEEPSEEK_MEMORY=on`. |
| 2521 | - `memory_path` (string, optional): anchors the native memory store. The |
| 2522 | configured filename is **not** the file that is written. Under the Native |
| 2523 | backend (the only backend) the store is re-rooted to |
| 2524 | `<parent-of-memory_path>/memory/global/MEMORY.md` — so the default |
| 2525 | `~/.codewhale/memory.md` yields `~/.codewhale/memory/global/MEMORY.md` |
| 2526 | (plus workspace-scoped files and a rebuildable SQLite FTS5 index). See |
| 2527 | [`MEMORY.md`](MEMORY.md) for the full feature surface (`# foo` composer |
| 2528 | prefix, `/memory` slash command, `remember` tool, opt-in toggle). |
| 2529 | - `snapshots.*` (optional): side-git workspace snapshots for file rollback: |
| 2530 | - `[snapshots].enabled` (bool, default `true`) |
| 2531 | - `[snapshots].max_age_days` (int, default `7`) |
| 2532 | - snapshots live under |
| 2533 | `~/.codewhale/snapshots/<project_hash>/<worktree_hash>/.git`, with legacy |
| 2534 | `~/.deepseek/snapshots/...` fallback when only the legacy state exists, and |
| 2535 | never use the workspace's own `.git` directory |
| 2536 | - `context.*` (optional): |
| 2537 | - `[context].project_pack` (bool, default `false`): include a deterministic |
| 2538 | project context pack (a large pretty-printed directory listing) in the |
| 2539 | stable prompt prefix (#4781). Useful for weak tool-calling models; the |
| 2540 | model can rebuild the same information with one `File` call. |
| 2541 | - The removed seam-manager keys (`enabled`, `verbatim_window_turns`, |
| 2542 | `l1_threshold`, `l2_threshold`, `l3_threshold`, `seam_model`) are |
| 2543 | ignored if an older config still carries them; they no longer load. |
| 2544 | - `compaction.*` (optional, config.toml): how a compaction pass behaves once |
| 2545 | it fires. `auto_compact` / `auto_compact_threshold_percent` (settings.toml) |
| 2546 | still decide *when* it fires. Both keys are absent by default, and absent |
| 2547 | means the built-in behaviour, unchanged: |
| 2548 | - `[compaction].summary_instructions` (string, default empty): standing |
| 2549 | operator instructions appended to the summarizer prompt as a clearly |
| 2550 | delimited "Additional instructions from the operator" section on **every** |
| 2551 | pass, manual and automatic — the effort-free counterpart to a one-off |
| 2552 | `/compact <focus>`, which still composes after this text. Useful for |
| 2553 | "always list exact file paths and line numbers", "always restate open |
| 2554 | decisions and their trade-offs", "always write a TL;DR first". Truncated |
| 2555 | at 4 000 characters with a warning naming the key; whitespace-only reads |
| 2556 | as unset. The summarizer still runs with no system prompt and no tools — |
| 2557 | this suffix is the only operator-authored input it sees. |
| 2558 | - `[compaction].retained_user_message_tokens` (int, default `20000`, clamped |
| 2559 | to `2000`–`200000`; also accepted as `retained_user_message_max_tokens`): |
| 2560 | token budget for the recent plain user messages kept **verbatim** in the |
| 2561 | replacement history. Raising it keeps more of the user's own earlier |
| 2562 | messages instead of only whatever the lossy summary captured; the |
| 2563 | last-round survival contract still applies on top. The `/compact` receipt |
| 2564 | names the effective budget and whether operator instructions were applied, |
| 2565 | so you can tell the knob took effect. |
| 2566 | - `retry.*` (optional): retry/backoff settings for API requests: |
| 2567 | - `[retry].enabled` (bool, default `true`) |
| 2568 | - `[retry].max_retries` (int, default `3`) |
| 2569 | - `[retry].initial_delay` (float seconds, default `1.0`) |
| 2570 | - `[retry].max_delay` (float seconds, default `60.0`) |
| 2571 | - `[retry].exponential_base` (float, default `2.0`) |
| 2572 | - `[retry].jitter` (bool, default `true`): randomize each backoff delay |
| 2573 | - `[retry].jitter_factor` (float, default `0.1` = ±10%; clamps to `0.0..=1.0`) |
| 2574 | - `[retry].respect_retry_after` (bool, default `true`): wait for a server |
| 2575 | `Retry-After` header instead of the computed backoff |
| 2576 | |
| 2577 | `[retry]` schedules HTTP-request retries inside the client. The stream-level |
| 2578 | budgets that sit above it — how often a turn re-issues a request whose |
| 2579 | stream failed to open or died — are the `[stream]` keys below; legacy `tui.stream_max_*` keys remain fallbacks. |
| 2580 | - `[notifications]`: notification delivery, attention, categories and audio share one |
| 2581 | policy. `quiet = true`, `method = "off"`, `condition = "never"` and disabled |
| 2582 | categories suppress both the banner and Codewhale's selected sound. |
| 2583 | - `notifications.method`: `auto` (default), `osc9`, `kitty`, `ghostty`, `bel`, `off`. |
| 2584 | - `notifications.condition`: `unfocused` (default), `always`, `never`. When absent, |
| 2585 | the legacy `tui.notification_condition` remains the fallback. `always` also |
| 2586 | bypasses the duration threshold; `unfocused` requires two seconds away. |
| 2587 | - `notifications.threshold_secs`: nonnegative integer, default `30`. |
| 2588 | - `notifications.include_summary`: boolean, default `false`. |
| 2589 | - `notifications.sound`: optional `off`, `whale`, `bell`, `beep`, `file`. |
| 2590 | A selected value controls audio across enabled categories. Absent keeps legacy |
| 2591 | `completion_sound` and `event_sound` choices; `off` overrides both. |
| 2592 | - `notifications.sound_file`: custom local WAV path for `sound = "file"` or legacy |
| 2593 | `completion_sound = "file"`. |
| 2594 | - `notifications.subagent_completion`: `always`, `final-only` (default), `off`. |
| 2595 | Covers all background work that finishes: sub-agents, background shells and |
| 2596 | durable tasks. `final-only` sends one notice naming everything that finished |
| 2597 | once no agent, workflow or durable task is still running; a running shell |
| 2598 | (a dev server, a watcher) never holds it back. `always` sends one per item. |
| 2599 | - `notifications.quiet`: boolean, default `false`. |
| 2600 | - `notifications.events`: six boolean categories, all enabled by default; see below. |
| 2601 | - `notifications.completion_sound`: legacy completion cue, default `off`, with the |
| 2602 | same values as `sound`. Used only when `sound` is absent. |
| 2603 | - `notifications.event_sound`: legacy `enabled` (default `false`), `events` |
| 2604 | (default `["turn-complete", "approval-needed"]`), and `quiet` (default `false`). |
| 2605 | `min_interval_ms` (default `2000`) applies to each category's audio in both modes. |
| 2606 | - `tui.alternate_screen` (string, optional, default `auto`): which screen an interactive session starts on. `auto` and `always` start on the TUI-owned alternate screen; `never` starts in inline mode — a ratatui viewport the full height of the terminal with no alternate screen, so the shell's scrollback survives the session and stays scrollable after exit. `/fullscreen` and `/inline` switch it in-process; a switch that the terminal refuses rolls back and says why. Inline mode paints the whole transcript inside its viewport — nothing is written into the host scrollback while the session runs. |
| 2607 | - `tui.mouse_capture` (bool, optional, default `true` on non-Windows terminals and on Windows Terminal/ConEmu/Cmder when the alternate screen is active; `false` on legacy Windows console and inside JetBrains JediTerm — PyCharm/IDEA/CLion/etc. — where mouse-event escapes leak into the input stream as garbled text, see #878 / #898): enable internal mouse scrolling, transcript selection, right-click context actions, and transcript scrollbar dragging. TUI-owned drag selection copies the intersected cells, removes visual wrap-column line breaks from paragraphs, and keeps selection scoped to the transcript pane; the payload is Markdown source by default, see `tui.selection_copy_markdown` below. Set this to `false` or run with `--no-mouse-capture` for raw terminal selection; set it to `true` or run with `--mouse-capture` to opt in anywhere it's defaulted off. On raw terminal selection, especially on legacy Windows console or when mouse capture is disabled, selection may cross the right workbar and include visual wraps because the terminal, not the TUI, owns the selection. |
| 2608 | On Linux, finishing a transcript or composer selection quietly copies text to |
| 2609 | PRIMARY, leaving the regular clipboard unchanged. Middle-click inside the |
| 2610 | composer pastes PRIMARY at the pointer without submitting it. This uses native |
| 2611 | X11 or Wayland data control; compositors must support PRIMARY selection. Over |
| 2612 | SSH without a forwarded graphical display, use your terminal's selection/paste |
| 2613 | gestures or `--no-mouse-capture`. Explicit Copy still uses the regular clipboard. |
| 2614 | |
| 2615 | - `tui.selection_copy_markdown` (bool, optional, default `true`): copy TUI-owned |
| 2616 | transcript selections (drag release, context-menu Copy, and `Cmd+C`/`Ctrl+C` |
| 2617 | on an active selection) as Markdown source instead of rendered text. Every |
| 2618 | intersected cell serializes through the same canonical projection `Ctrl-Y` |
| 2619 | and `/copy` use — user and assistant cells keep their authored Markdown, |
| 2620 | other cells keep their full transcript form — partial intersections round out |
| 2621 | to whole cells, cells join with blank lines, and a toast names the copied |
| 2622 | cell count. Set `false` to copy the rendered text as displayed. Composer |
| 2623 | selections and the Linux PRIMARY auto-copy are unchanged; PRIMARY always |
| 2624 | carries rendered text. |
| 2625 | |
| 2626 | - `tui.stream_chunk_timeout_secs` (int, optional, default `900`): per-SSE-chunk idle timeout for streamed model responses. Slow local or compatible servers can raise this with `/config stream_chunk_timeout_secs <seconds>` (add `--save` to write canonical `stream.chunk_timeout_secs`); `0` maps to the default and explicit values must be `1..=3600`. The legacy `DEEPSEEK_STREAM_IDLE_TIMEOUT_SECS` env var is still honored when this key is omitted. |
| 2627 | - `tui.osc8_links` (bool, optional, default on for macOS/Linux, off for Windows): emit OSC 8 escape sequences around URLs in transcript output so supporting terminals (iTerm2, Terminal.app 13+, Ghostty, Kitty, WezTerm, Alacritty, recent gnome-terminal/konsole) can open them with the terminal's link gesture—usually Cmd-click on macOS and Ctrl-click on Linux/Windows. Terminals without OSC 8 support render the plain label and ignore the escape. The escapes are emitted out-of-band (not inside buffer cells), so column corruption is not a concern; set `false` only for terminals that misrender the OSC 8 terminator itself. Windows legacy consoles default off; opt in with `true`. |
| 2628 | - `tui.max_model_steps` (int, optional, default uncapped): optional model-step ceiling for one ordinary turn. Omission or `0` leaves model steps uncapped; explicit positive values are clamped to `1..=100000`. Headless `exec` and Fleet workers also have no implicit model-step ceiling; `exec --max-turns N` and positive worker budgets still apply. At ~80% of an explicit step budget the model gets one soft-landing notice; at exhaustion the turn ends `Failed` with `Maximum model steps reached before completion (limit: N)` after one bounded final-report response when needed. Cumulative wall-clock and per-stream limits remain independent. Active interactive goal turns use `goal.max_steps` instead (default `1000`); see the Goal loop section below. |
| 2629 | - `tui.turn_wall_clock_secs` (int, optional, default: no limit): cumulative per-turn wall-clock budget in seconds, measured across every model step of one turn (not per request). Time blocked on a human approval is excluded. Omitted or `0` means no limit; positive values clamp to `30..=86400` (24 hours is the ceiling). When exhausted the turn stops before authorizing another billable request with a message naming the limit and the key to raise. |
| 2630 | - `tui.stream_max_resumes` (int, optional, default `3`): how many times one turn re-issues a model request after its stream failed — the request never opened (connect failure or response-header stall), the stream died before any content, the host slept mid-stream, or the network dropped mid-stream. Every one of those paths spends this one budget, and a healthy stream resets it. `0` disables turn-level re-issues (a failed stream then fails the turn); values clamp to `0..=10`. |
| 2631 | - `tui.stream_max_transparent_retries` (int, optional, default `2`): in-stream re-requests while nothing has streamed yet. `0` disables them; values clamp to `0..=10`. |
| 2632 | - `tui.stream_max_errors` (int, optional, default `5`): recoverable errors tolerated within one stream before it ends. Unlike the two retry counts above, `0` does not switch anything off: like the other finite stream budgets it selects the default. Other values clamp to `1..=50`, so `1` ends the stream on its first recoverable error. |
| 2633 | - `tui.stream_open_timeout_secs` (int, optional, default `45`): wait for a streaming request's response headers (connection setup included). A header stall on HTTP/2 retries once over HTTP/1.1 with the same wait. Omitted or `0` falls back to `CODEWHALE_STREAM_OPEN_TIMEOUT_SECS`, then the default; values clamp to `5..=300`. |
| 2634 | - `tui.connect_timeout_secs` (int, optional, default `30`): TCP/TLS connect timeout for the model HTTP client. Omitted or `0` uses the default; values clamp to `1..=300`. |
| 2635 | - `tui.force_http1` (bool, optional, default `false`): pin the model HTTP client to HTTP/1.1, for provider edges or proxies that mishandle long-lived HTTP/2 streams. `CODEWHALE_FORCE_HTTP1=1` does the same; either one pins. |
| 2636 | - `tui.stream_max_content_mb` (int, optional, default `10`) and `tui.stream_max_duration_secs` (int, optional, default `1800`): per-step caps on one stream's accumulated content and wall-clock duration. `0` selects the default; values clamp to `1..=512` MB and `10..=86400` seconds. |
| 2637 | - `transcript.prose_measure` (positive integer, optional, default absent = full width): wrap cap, in columns, for prose cells — user messages, assistant answers, and reasoning/thinking blocks — in the live transcript (#5436). Absent (or `0`) spends the full content width, consistent with tool/status cells and the #5322 wide-frame decision; the former 105-column prose rail is gone. Set a positive whole number (e.g. `prose_measure = 120` under `[transcript]`) to restore a bounded reading measure on ultrawide terminals. Narrow terminals always keep their content width — the cap clamps from above only. Tool, diff, and status cells never inherit this cap. Invalid values (negative or non-integer) are rejected at startup with a `transcript.prose_measure` config error. Resolved once per render pass, so the main transcript cache and the full-screen overlay always agree on the effective width. |
| 2638 | - `hooks` (optional): lifecycle hooks configuration (see `config.example.toml`). |
| 2639 | - `features.*` (optional): feature flag overrides (see below). |
| 2640 | |
| 2641 | ### Stream and transport settings |
| 2642 | |
| 2643 | `[stream]` is the canonical table for model-stream policy and its HTTP clients. |
| 2644 | `codewhale config dump` and `codewhale config get stream` show effective values |
| 2645 | from the runtime resolver, including environment fallbacks and clamps; this does |
| 2646 | not persist defaults. Set or unset individual fields with, for example, |
| 2647 | `codewhale config set stream.open_timeout_secs 120` and |
| 2648 | `codewhale config unset stream.open_timeout_secs`. Per-run |
| 2649 | `--set stream.open_timeout_secs=120` uses the same validation. |
| 2650 | |
| 2651 | | `[stream]` key | Default | Accepted behavior | Legacy `[tui]` fallback | |
| 2652 | | --- | --- | --- | --- | |
| 2653 | | `open_timeout_secs` | 45 | Positive values clamp to 5–300; 0 or omission falls through to the environment/default | `stream_open_timeout_secs` | |
| 2654 | | `chunk_timeout_secs` | 900 | 0 selects default; positive values clamp to 1–3600 | `stream_chunk_timeout_secs` | |
| 2655 | | `max_resumes` | 3 | 0 disables whole-request reissues; maximum 10 | `stream_max_resumes` | |
| 2656 | | `max_transparent_retries` | 2 | 0 disables retries before any content; maximum 10 | `stream_max_transparent_retries` | |
| 2657 | | `max_stream_errors` | 5 | 0 selects default; positive values clamp to 1–50 | `stream_max_errors` | |
| 2658 | | `max_duration_secs` | 1800 | Per-stream wall clock; 0 selects default, positive values clamp to 10–86400 | `stream_max_duration_secs` | |
| 2659 | | `max_content_mb` | 10 | Per-stream content; 0 selects default, positive values clamp to 1–512 MiB | `stream_max_content_mb` | |
| 2660 | | `connect_timeout_secs` | 30 | TCP/TLS setup; 0 selects default, positive values clamp to 1–300 | `connect_timeout_secs` | |
| 2661 | | `force_http1` | false | Boolean; a truthy environment pin always enables HTTP/1.1 | `force_http1` | |
| 2662 | | `tcp_keepalive_secs` | 30 | Idle time before TCP keepalive probes; 0 disables, positive values clamp to 1–3600 | none | |
| 2663 | | `http2_keep_alive_interval_secs` | 15 | PING interval on active HTTP/2 connections; 0 disables, positive values clamp to 1–3600 | none | |
| 2664 | | `http2_keep_alive_timeout_secs` | 20 | PING acknowledgement deadline; 0 selects default, positive values clamp to 1–3600 | none | |
| 2665 | |
| 2666 | For each field, an explicit canonical value wins over the legacy `[tui]` field, |
| 2667 | including zero or false; an absent canonical field preserves the legacy value. |
| 2668 | The existing environment precedence is unchanged: a positive configured header |
| 2669 | wait wins over `CODEWHALE_STREAM_OPEN_TIMEOUT_SECS` (then its `DEEPSEEK_` alias); |
| 2670 | zero falls through to those variables. An omitted chunk timeout uses |
| 2671 | `CODEWHALE_STREAM_IDLE_TIMEOUT_SECS` (then its `DEEPSEEK_` alias), whereas an |
| 2672 | explicit zero uses 900. `CODEWHALE_FORCE_HTTP1` (legacy `DEEPSEEK_FORCE_HTTP1`) |
| 2673 | is ORed with the selected config flag, so a false config flag cannot defeat a |
| 2674 | truthy environment pin. Profile overlays merge canonical fields independently. |
| 2675 | |
| 2676 | HTTP transport values apply to all newly constructed model, catalog and HTTP/1 |
| 2677 | fallback clients. They do not rebuild an active client, govern MCP/other network |
| 2678 | services, or send HTTP/2 PINGs on idle pooled connections. HTTP/2 knobs have no |
| 2679 | effect when HTTP/1.1 is pinned. The operating system controls TCP probe details; |
| 2680 | these settings do not disable certificate validation. `[retry]` still owns the |
| 2681 | HTTP-request backoff schedule independently of stream-level budgets. |
| 2682 | |
| 2683 | ### Workspace notes |
| 2684 | |
| 2685 | `/note` manages a simple notes file in the current workspace at |
| 2686 | `.codewhale/notes.md` (legacy `.deepseek/notes.md` is the fallback path when |
| 2687 | no `.codewhale/notes.md` exists yet). |
| 2688 | Existing `/note <text>` usage still appends a note. |
| 2689 | The management forms are: |
| 2690 | |
| 2691 | | Command | Action | |
| 2692 | |---|---| |
| 2693 | | `/note <text>` | Append a note (legacy shorthand) | |
| 2694 | | `/note add <text>` | Append a note explicitly | |
| 2695 | | `/note list` | List notes with temporary 1-based numbers | |
| 2696 | | `/note show <n>` | Show the full note at number `n` | |
| 2697 | | `/note edit <n> <text>` | Replace note `n` with new text | |
| 2698 | | `/note remove <n>` | Delete note `n`; `rm` and `delete` are aliases | |
| 2699 | | `/note clear` | Empty the workspace notes file | |
| 2700 | | `/note path` | Show the resolved workspace notes path | |
| 2701 | |
| 2702 | The numbers shown by `/note list` are not stored in the file; they are derived |
| 2703 | from the current order each time notes are read. This keeps the file format |
| 2704 | compatible with the existing `---`-separated notes. |
| 2705 | |
| 2706 | ### User memory |
| 2707 | |
| 2708 | User memory is split across one top-level path setting and one opt-in |
| 2709 | toggle table: |
| 2710 | |
| 2711 | ```toml |
| 2712 | # Anchors the store only — actual writes go to |
| 2713 | # ~/.codewhale/memory/global/MEMORY.md (see MEMORY.md). |
| 2714 | memory_path = "~/.codewhale/memory.md" |
| 2715 | |
| 2716 | [memory] |
| 2717 | enabled = true |
| 2718 | ``` |
| 2719 | |
| 2720 | Notes: |
| 2721 | |
| 2722 | - `memory_path` stays at the top level beside `notes_path` and |
| 2723 | `skills_dir`; it is not nested under `[memory]`. |
| 2724 | - The configured path is an **anchor**: its parent directory gains |
| 2725 | `memory/global/MEMORY.md`, workspace-scoped files, and `index.db`. |
| 2726 | Pointing `memory_path` at the native layout path itself would double-nest |
| 2727 | (`…/memory/global/memory/global/MEMORY.md`); keep the legacy-style |
| 2728 | anchor filename. |
| 2729 | - `DEEPSEEK_MEMORY_PATH` overrides the anchor path from the environment. |
| 2730 | - `DEEPSEEK_MEMORY=on` (also `1`, `true`, `yes`, `y`, or `enabled`) |
| 2731 | flips the feature on without editing `config.toml`. |
| 2732 | - The feature is inert when disabled: no file is injected, `# foo` |
| 2733 | falls through to normal message submission, and the model does not |
| 2734 | see the `remember` tool. |
| 2735 | - See [`MEMORY.md`](MEMORY.md) for examples and the full `/memory` |
| 2736 | command surface. |
| 2737 | |
| 2738 | ### Goal loop (`[goal]`) |
| 2739 | |
| 2740 | Operate-mode goals run to their completion gate with no default token, time, or |
| 2741 | continuation ceiling (#5052). Token/time budgets, when explicitly supplied, |
| 2742 | are telemetry only and do not stop a goal. Users who want a circuit breaker can |
| 2743 | opt into one: |
| 2744 | |
| 2745 | ```toml |
| 2746 | [goal] |
| 2747 | # Optional safety backstop on automatic goal continuation passes. |
| 2748 | # Default: 0 (unlimited). Set a positive value to opt into a ceiling. |
| 2749 | max_continuations = 100 |
| 2750 | |
| 2751 | # Optional cancellable quiet period between successful turns. This is useful |
| 2752 | # for coordinator goals that should poll on a cadence instead of keeping one |
| 2753 | # provider turn open. Default: 0 (continue immediately). |
| 2754 | continuation_delay_seconds = 300 |
| 2755 | |
| 2756 | # Per-turn step allowance while a goal is active (#5994). Goal turns get a |
| 2757 | # larger but still finite budget than an ordinary interactive turn. |
| 2758 | # Default: 1000 (0 or absent resolves to 1000, never unlimited). Range: |
| 2759 | # 1..=100,000. This bounds each provider turn, never the number of |
| 2760 | # continuation passes. |
| 2761 | max_steps = 1000 |
| 2762 | ``` |
| 2763 | |
| 2764 | The effective delay is capped at 86,400 seconds (24 hours); use an automation |
| 2765 | for schedules that are less frequent than once per day. |
| 2766 | |
| 2767 | When an explicit backstop fires, the goal pauses with a status message naming |
| 2768 | `[goal] max_continuations` and a warning is logged; resume the goal after |
| 2769 | inspecting progress, or raise/disable the backstop. |
| 2770 | |
| 2771 | `[goal] max_steps` governs one engine turn at a time: the ordinary interactive |
| 2772 | turn has no implicit model-step ceiling. Explicit per-invocation |
| 2773 | ceilings — `exec --max-turns N`, child-worker caps — always win over it. At |
| 2774 | about 80% of the selected budget the model is told to land; at exhaustion it |
| 2775 | gets one bounded final report and the turn classifies as budget-exhausted. An |
| 2776 | unfinished goal then pauses with the BudgetLimit reason instead of re-arming |
| 2777 | another goal turn — resume it explicitly after reviewing the report. Wall-clock |
| 2778 | and stream protections are separate and still apply. |
| 2779 | |
| 2780 | The delay starts only after a successful turn while an explicitly created goal |
| 2781 | is still active. `/goal pause`, `/goal done`, `/goal blocked`, `/goal clear`, |
| 2782 | Esc, or Ctrl+C cancels a pending continuation before another provider request |
| 2783 | starts. Failed turns and policy/route failures never schedule another turn. |
| 2784 | Only the numeric cadence is stored in config; no prompt, credential, or secret |
| 2785 | is persisted for the loop. |
| 2786 | |
| 2787 | ### Reasoning-only recovery (`[reasoning_only]`) |
| 2788 | |
| 2789 | When a reasoning model (thinking mode) finishes a response with only hidden |
| 2790 | reasoning and no answer text or tool call, the engine can automatically |
| 2791 | re-request the answer. Configure this behavior with the `[reasoning_only]` |
| 2792 | table: |
| 2793 | |
| 2794 | ```toml |
| 2795 | [reasoning_only] |
| 2796 | # Maximum number of automatic re-requests. Default: 2. |
| 2797 | # Set to 0 to disable automatic recovery (fail immediately). |
| 2798 | max_reprompts = 10 |
| 2799 | |
| 2800 | # Optional custom message sent to the model on each re-request. |
| 2801 | # When set, overrides the built-in default message. |
| 2802 | # When unset (or commented out), the engine uses: |
| 2803 | # "So, what's up ? Keep running !" |
| 2804 | reprompt_message = "Allez, répond quelque chose !" |
| 2805 | ``` |
| 2806 | |
| 2807 | This only applies when the model returns a clean `stop` finish reason with |
| 2808 | only thinking content. An output-length stop (`length`/`max_tokens`) is never |
| 2809 | retried, and a persistently answerless model still fails honestly after the |
| 2810 | configured bound. |
| 2811 | |
| 2812 | To disable the reprompt message entirely (silent retry), set it to an empty |
| 2813 | string: |
| 2814 | |
| 2815 | ```toml |
| 2816 | [reasoning_only] |
| 2817 | reprompt_message = "" |
| 2818 | ``` |
| 2819 | |
| 2820 | ### Notifications |
| 2821 | |
| 2822 | Notification controls are available in the existing `/config` settings view, |
| 2823 | from terminal commands, and through the CLI. All write the same `config.toml` |
| 2824 | keys. Terminal changes apply immediately; add `--save` to keep them. CLI writes |
| 2825 | apply when the next process loads its configuration. |
| 2826 | |
| 2827 | ```sh |
| 2828 | codewhale config set notifications.sound whale |
| 2829 | codewhale config set notifications.events.approval-needed false |
| 2830 | codewhale config set notifications.quiet true |
| 2831 | codewhale config get notifications |
| 2832 | codewhale config unset notifications.quiet |
| 2833 | ``` |
| 2834 | |
| 2835 | ```text |
| 2836 | /config notifications sound whale --save |
| 2837 | /config notifications condition unfocused --save |
| 2838 | /config notifications quiet true |
| 2839 | /config notifications status |
| 2840 | ``` |
| 2841 | |
| 2842 | Nested CLI edits preserve TOML types, unrelated keys and comments. Unset removes |
| 2843 | only the selected leaf. `notifications.sound legacy` in the TUI, or CLI unset of |
| 2844 | `notifications.sound`, restores previous sound choices. Invalid values are |
| 2845 | rejected before file or session changes. In an active TUI profile, saved edits |
| 2846 | update its notification table when it owns one; otherwise they update the |
| 2847 | inherited root table. The settings detail keeps saved and current values distinct. |
| 2848 | |
| 2849 | ```toml |
| 2850 | [notifications] |
| 2851 | method = "auto" # auto | osc9 | kitty | ghostty | bel | off |
| 2852 | condition = "unfocused" # unfocused | always | never |
| 2853 | threshold_secs = 30 |
| 2854 | include_summary = false |
| 2855 | sound = "whale" # optional; sound is opt-in, not enabled by default |
| 2856 | quiet = false |
| 2857 | |
| 2858 | [notifications.events] |
| 2859 | turn-complete = true |
| 2860 | subagent-terminal = true |
| 2861 | approval-needed = true |
| 2862 | input-needed = true |
| 2863 | elevation-needed = true |
| 2864 | model-notify = true |
| 2865 | ``` |
| 2866 | |
| 2867 | `quiet = true` mutes every category without changing saved choices. `method = |
| 2868 | "off"` also stops both banner and selected audio. Disabling a category stops its |
| 2869 | sound. Attention and duration gates apply before any sink runs. Title animation |
| 2870 | completion is silent; only the authorized notification event can request audio. |
| 2871 | Successful turn completion is a notification category; failed/cancelled turns do |
| 2872 | not create a success notification. |
| 2873 | |
| 2874 | `auto` chooses a recognized terminal protocol or the existing macOS native |
| 2875 | fallback; unknown terminals remain unsupported and never invent a bell. `bel` |
| 2876 | is an audio-only transport: one selected cue is dispatched, without a second |
| 2877 | transport bell. With `sound = "off"`, that transport is silent. `osc9`, `kitty` |
| 2878 | and `ghostty` use their terminal notification protocols; tmux passthrough is |
| 2879 | preserved. Terminal/OS notification preferences still govern display, attribution |
| 2880 | and any sound the host itself adds. |
| 2881 | |
| 2882 | By default the terminal must stay unfocused for two seconds. `condition = |
| 2883 | "always"` allows foreground notifications and bypasses the duration threshold; |
| 2884 | `"never"` suppresses all delivery. The canonical condition takes precedence over |
| 2885 | legacy `[tui].notification_condition`. |
| 2886 | |
| 2887 | The bundled `whale` is a 1.55-second original whale-inspired cue, with no |
| 2888 | third-party recording. It remains an opt-in candidate pending listening approval. |
| 2889 | WAV playback uses a background worker: macOS `/usr/bin/afplay`, Linux `aplay` |
| 2890 | from [ALSA utilities](https://github.com/alsa-project/alsa-utils), or Windows |
| 2891 | `PlaySoundW`. A missing player/file or unsupported platform has no fallback bell. |
| 2892 | Only one WAV plays at a time. A worker-start receipt is a dispatch attempt, not |
| 2893 | proof of audible playback or OS acceptance. The existing macOS `osascript` |
| 2894 | banner retains Script Editor attribution; this Core change does not provide a |
| 2895 | branded native Apps banner. |
| 2896 | |
| 2897 | #### Previous sound settings |
| 2898 | |
| 2899 | When `notifications.sound` is absent, completion uses a non-off |
| 2900 | `completion_sound` selection, and other events use the existing event allow-list. |
| 2901 | These are compatibility inputs to the same audio decision, not separate playback |
| 2902 | paths. When the global sound is selected, it takes precedence over the previous |
| 2903 | completion/event choices. The default remains silent unless a legacy sound or |
| 2904 | explicit `bel` transport was already selected. |
| 2905 | |
| 2906 | ```toml |
| 2907 | [notifications.event_sound] |
| 2908 | enabled = false |
| 2909 | events = ["turn-complete", "approval-needed"] |
| 2910 | min_interval_ms = 2000 |
| 2911 | quiet = false |
| 2912 | ``` |
| 2913 | |
| 2914 | Legacy event cues use one bell for completion, subagent completion, input and |
| 2915 | model notices; approval/elevation cues use two. The per-category repeat interval |
| 2916 | survives settings refreshes. Old unknown event names are ignored on load; new |
| 2917 | CLI/TUI edits require names from the six categories above. |
| 2918 | |
| 2919 | Local approval/input/elevation prompts and error receipts remain available when |
| 2920 | external notifications are muted. Action prompts use the selected UI language |
| 2921 | and retire when that request settles. Repeated live notices keep their first |
| 2922 | expiry; a later routine update does not hide an unresolved warning at completion. |
| 2923 | Optional plugin suggestion toasts and contextual tips share one session guidance |
| 2924 | budget. `/config contextual_tips off` hides those toasts while preserving |
| 2925 | required notices and explicit plugin review requests. |
| 2926 | |
| 2927 | #### What a notification can contain |
| 2928 | |
| 2929 | A desktop notification is a glance surface: on macOS it can appear on the |
| 2930 | lock screen, and on every platform it is visible to anyone near the machine. |
| 2931 | Codewhale therefore builds notifications from a typed payload with a fixed |
| 2932 | per-event disclosure policy rather than from whatever text was on hand: |
| 2933 | |
| 2934 | | Event | Shown | Never shown | |
| 2935 | |---|---|---| |
| 2936 | | Turn complete | status line (+ elapsed/cost when `include_summary`), preview of the assistant's reply | — | |
| 2937 | | Sub-agent finished | status line, agent id, preview of the child's summary line | — | |
| 2938 | | Approval needed | the tool name | the tool description, the command, the arguments | |
| 2939 | | Input needed | "Answer the question in the terminal to continue" | the question | |
| 2940 | | Sandbox elevation needed | the tool name and the denial reason | the command | |
| 2941 | | `notify` tool | model-supplied title and body | — | |
| 2942 | |
| 2943 | Every field is capped (80 characters for the status line, 120 for the |
| 2944 | identifier, 200 for the preview), stripped of control bytes and escape |
| 2945 | sequences, and passed through a redactor that replaces credential-shaped |
| 2946 | strings with `[redacted]`, reduces absolute local paths to `…/basename`, |
| 2947 | and replaces raw tool JSON with `[details hidden]`. The redactor is |
| 2948 | deliberately over-eager: an unbroken 40-character run has no word |
| 2949 | structure, so it is redacted even when it is not a secret. |
| 2950 | |
| 2951 | #### macOS: why the banner says "Script Editor" |
| 2952 | |
| 2953 | On macOS terminals that provide no notification escape of their own — |
| 2954 | Apple Terminal, the VS Code and JetBrains embedded terminals, plain tmux |
| 2955 | without `LC_TERMINAL` — `method = "auto"` falls back to `osascript`'s |
| 2956 | `display notification`. That command posts on behalf of the *bundled* |
| 2957 | host process, and `/usr/bin/osascript` is unbundled, so macOS attributes |
| 2958 | the banner to `com.apple.ScriptEditor2`. That attribution supplies the |
| 2959 | Script Editor icon and owns the System Settings → Notifications entry |
| 2960 | (alert style, previews, Do Not Disturb). `display notification` has no |
| 2961 | icon parameter, so this cannot be fixed from the notification code; it |
| 2962 | needs Codewhale to ship a real `.app` bundle. Tracked in |
| 2963 | [#4834](https://github.com/codewhale-hq/CodeWhale/issues/4834). In the meantime, |
| 2964 | iTerm2, WezTerm, Ghostty, and kitty are matched first and use their own |
| 2965 | notification protocols, and `method = "osc9"` / `"bel"` / `"off"` opt out |
| 2966 | of the `osascript` path explicitly. |
| 2967 | |
| 2968 | ## Automations in the terminal |
| 2969 | |
| 2970 | Open `/automation`, then choose **New automation** (`n`) or **Edit automation** |
| 2971 | (`e`). The form edits the name, multiline prompt, schedule, model, workspace, |
| 2972 | and enabled status. `Tab` moves between fields; `Enter` inserts a newline in |
| 2973 | the prompt. Schedule presets include daily, weekly, hourly, once, and a custom |
| 2974 | RRULE. Time fields accept `HH:MM`; their arrow controls change the time by |
| 2975 | 15 minutes. Weekly day buttons and model search support mouse and keyboard. |
| 2976 | |
| 2977 | The next-run preview uses the scheduler's local time zone, shown with its UTC |
| 2978 | offset. An enabled automation can run after **Save** (`Ctrl+S`); a paused one |
| 2979 | has no scheduled next run. **Cancel** (`Esc`) discards the draft. Editing keeps |
| 2980 | existing permission settings, custom schedules, and additional workspace |
| 2981 | entries unless the corresponding supported field is explicitly changed. |
| 2982 | |
| 2983 | Schedules are evaluated in the machine's local time zone against the wall |
| 2984 | clock. A wall time that does not exist on a spring-forward day is skipped and |
| 2985 | an ambiguous fall-back time fires once. Occurrences missed while Codewhale was |
| 2986 | closed, asleep, or still running the previous occurrence are coalesced: the |
| 2987 | next start runs one catch-up occurrence and then continues from the next |
| 2988 | future slot, never replaying every missed slot. An occurrence never starts |
| 2989 | while an earlier run of the same automation is still queued or running. Each |
| 2990 | run is recorded durably with its status, timing and error; a run that needs a |
| 2991 | tool approval has no operator to ask and fails once the approval wait expires. |
| 2992 | |
| 2993 | Choosing a concrete model pins both the model and its exact configured |
| 2994 | provider, including named custom routes. Later changes to the active provider |
| 2995 | do not move that automation's pin. The default-model choice and legacy |
| 2996 | definitions without a provider pin keep the runtime's existing default |
| 2997 | behavior. New automation records use schema v2 and tasks use v3 so older |
| 2998 | runtimes reject records whose provider pins they cannot preserve. |
| 2999 | |
| 3000 | ## Lifecycle Outbox (`[lifecycle_outbox]`) |
| 3001 | |
| 3002 | The lifecycle outbox is an opt-in, machine-readable stream of session, |
| 3003 | turn, and sub-agent lifecycle events. With a path configured, Codewhale |
| 3004 | appends one JSON line per event to that file — for interactive TUI |
| 3005 | sessions *and* headless `codewhale exec` runs — so a supervisor |
| 3006 | (terminal multiplexer wrapper, automation harness, alerting setup) can |
| 3007 | react to what happened without scraping the screen or installing per-hook |
| 3008 | shell commands. Unset or empty `path` = the feature is **off** and |
| 3009 | behavior is unchanged. |
| 3010 | |
| 3011 | ```toml |
| 3012 | [lifecycle_outbox] |
| 3013 | path = "~/.codewhale/notifications/outbox.jsonl" # unset/empty = OFF |
| 3014 | webhook_url = "" # optional; POSTs events as JSON when set |
| 3015 | webhook_token = "" # optional bearer token for webhook_url |
| 3016 | ``` |
| 3017 | |
| 3018 | ### Events emitted |
| 3019 | |
| 3020 | | Event | Kind | Fired at | |
| 3021 | |---|---|---| |
| 3022 | | `turn_start` | `turn.started` | a new turn begins (TUI TurnStarted; `exec` at message dispatch) | |
| 3023 | | `turn_end` | `turn.completed` / `turn.failed` / `turn.interrupted` | turn completion, kind projected from the turn status | |
| 3024 | | `turn_stalled` | `turn.stalled` | the stall watchdog recovers a wedged turn | |
| 3025 | | `subagent_spawn` | `subagent.spawned` | a sub-agent is spawned | |
| 3026 | | `subagent_complete` | `subagent.completed` | a sub-agent reaches a terminal state | |
| 3027 | | `session_start` | `session.started` | interactive session start | |
| 3028 | | `session_end` | `session.ended` | interactive session end | |
| 3029 | |
| 3030 | ### File contract |
| 3031 | |
| 3032 | Each line is a `RuntimeEventEnvelope`: |
| 3033 | |
| 3034 | ```json |
| 3035 | {"schema_version": 1, "seq": 3, "event": "turn_start", "kind": "turn.started", |
| 3036 | "thread_id": "…", "turn_id": "…", "item_id": null, "timestamp": "…", |
| 3037 | "created_at": "…", "payload": {…}} |
| 3038 | ``` |
| 3039 | |
| 3040 | - `seq` is monotonic per outbox file and recovers from the last written |
| 3041 | line when a new process opens the file. |
| 3042 | - Lines are written one complete JSON line per append, serialized by an |
| 3043 | internal writer task and flushed before the next event; concurrent |
| 3044 | sessions writing the same path do not interleave bytes mid-line, but |
| 3045 | separate processes each continue from their own recovered `seq`, so seqs |
| 3046 | can repeat across processes sharing one file — prefer one file per |
| 3047 | process for strict uniqueness. |
| 3048 | - Parent directories are created lazily on the first event. |
| 3049 | - Payloads are constructed from bounded, pre-redacted fields only — never |
| 3050 | raw tool arguments, environment, or full transcript text. Free-form |
| 3051 | fields (error messages, previews) are capped at the notification limits |
| 3052 | (80 headline / 120 detail / 200 preview characters) and stripped of |
| 3053 | control bytes. |
| 3054 | |
| 3055 | ### Webhook delivery |
| 3056 | |
| 3057 | With `webhook_url` set, every event is additionally POSTed as |
| 3058 | `{"at": "<ISO 8601 timestamp>", "event": {…}}` with |
| 3059 | `Authorization: Bearer <webhook_token>` when a token is configured. |
| 3060 | Delivery is best-effort: failures are logged and dropped, never retried |
| 3061 | into the agent loop, and a failing webhook never blocks the local file |
| 3062 | append. |
| 3063 | |
| 3064 | ## Control Socket (`[control_socket]`) |
| 3065 | |
| 3066 | Per-session control surface for supervised operation: with the feature |
| 3067 | enabled, the interactive TUI binds one unix domain socket per *running* |
| 3068 | session at `<sessions-dir>/<session-id>/control.sock` (mode `0600`; |
| 3069 | `<sessions-dir>` is the same directory the session store uses, typically |
| 3070 | `~/.codewhale/sessions`). The socket is removed with the session's |
| 3071 | artifact directory, and a stale socket left by a crashed process is taken |
| 3072 | over by the next launch. Unset or `enabled = false` = the feature is |
| 3073 | **off** (the default) and behavior is unchanged. Unix-only; on other |
| 3074 | platforms the key parses but no socket is bound. |
| 3075 | |
| 3076 | ```toml |
| 3077 | [control_socket] |
| 3078 | enabled = false # default: OFF |
| 3079 | ``` |
| 3080 | |
| 3081 | The socket speaks newline-framed JSON-RPC, one request per connection: |
| 3082 | write one request line, read one response line, close. |
| 3083 | |
| 3084 | ```json |
| 3085 | {"id":"1","method":"message","params":{"text":"hello"}} |
| 3086 | {"id":"2","method":"interrupt","params":{}} |
| 3087 | {"id":"3","method":"status","params":{}} |
| 3088 | ``` |
| 3089 | |
| 3090 | - `message` — delivers `text` as a structured user message through the |
| 3091 | ordinary composer dispatch path; dispatched immediately when idle, |
| 3092 | queued when a turn is in flight (the response's `delivery` field says |
| 3093 | which). |
| 3094 | - `interrupt` — the Esc-shaped cancel of the active turn; `cancelled` |
| 3095 | reports whether active work was in flight. |
| 3096 | - `status` — answers `turn_state` (`idle` / `in_progress` / `waiting`) and |
| 3097 | `goal` (`objective`, `status`, `paused`). |
| 3098 | |
| 3099 | Success responses echo the request id with a `type`-tagged result; |
| 3100 | failures carry `error.code` (`invalid_request`, `command_error`, |
| 3101 | `timeout`, `server_unavailable`). Requests are bounded at 1 MiB per line |
| 3102 | and a handler that does not answer within 5 s is reported as `timeout`. |
| 3103 | |
| 3104 | ## Tool Catalog |
| 3105 | |
| 3106 | Codewhale loads a small core native tool catalog by default and leaves less |
| 3107 | common native tools discoverable through ToolSearch. To keep specific native |
| 3108 | tools loaded on every request, add them to `[tools].always_load`: |
| 3109 | |
| 3110 | ```toml |
| 3111 | [tools] |
| 3112 | always_load = ["Git", "notify"] |
| 3113 | ``` |
| 3114 | |
| 3115 | ### Script tools and overrides |
| 3116 | |
| 3117 | Scripts in `~/.codewhale/tools/` (or `[tools].plugin_dir`) that start with a |
| 3118 | `# name:` header become model-visible tools, and `/plugin tools` lists them. |
| 3119 | The script reads the tool's JSON input on stdin and writes a JSON |
| 3120 | `ToolResult` (`{"content": "...", "success": true}`) on stdout. |
| 3121 | |
| 3122 | ```sh |
| 3123 | #!/usr/bin/env sh |
| 3124 | # name: word_count |
| 3125 | # description: Count words in the given text |
| 3126 | # schema: {"type":"object","properties":{"text":{"type":"string"}}} |
| 3127 | # approval: required |
| 3128 | ``` |
| 3129 | |
| 3130 | `# approval:` takes `suggest` (the default) or `required`; either way the |
| 3131 | tool follows the session's approval setting. A script cannot approve itself: |
| 3132 | `approval: auto` is no longer supported, so such a script gets the default, |
| 3133 | and the runtime log (`~/.codewhale/logs/`) and `/plugin tools` name it. |
| 3134 | |
| 3135 | A script cannot replace a built-in tool either. A script whose `# name:` is |
| 3136 | already registered is not loaded. `[tools.overrides]` may disable a built-in, |
| 3137 | or add a script or command tool under a name of its own: |
| 3138 | |
| 3139 | ```toml |
| 3140 | [tools.overrides] |
| 3141 | "Web" = { type = "disabled" } # turn a built-in off |
| 3142 | "audited_shell" = { type = "script", path = "audit-shell.sh" } # a new tool |
| 3143 | "Bash" = { type = "script", path = "audit-shell.sh" } # refused: Bash is built in |
| 3144 | ``` |
| 3145 | |
| 3146 | A `script` or `command` override keyed by a built-in is refused, and the |
| 3147 | built-in stays active. A status line names the key once per session, and the |
| 3148 | runtime log records it. To route a |
| 3149 | built-in through your own wrapper, disable the built-in and register the |
| 3150 | wrapper under a new name. An override keyed by a drop-in script's name still |
| 3151 | replaces that script. Relative `path` values resolve against the plugin |
| 3152 | directory. |
| 3153 | |
| 3154 | ### `request_user_input` limits |
| 3155 | |
| 3156 | `request_user_input` asks the user a short batch of multiple-choice questions. |
| 3157 | Both ceilings are configurable (#5949): raise `user_input_max_questions` when a |
| 3158 | research or planning workflow legitimately needs more clarifications, lower it |
| 3159 | when interactive triage should stay terse. |
| 3160 | |
| 3161 | ```toml |
| 3162 | [tools] |
| 3163 | user_input_max_questions = 6 # default 6, clamped to 1..=10 |
| 3164 | user_input_max_options = 4 # default 4, clamped to 2..=10 |
| 3165 | ``` |
| 3166 | |
| 3167 | The effective values are applied in three places at once: the tool's JSON |
| 3168 | schema (`minItems` / `maxItems`), its model-visible description, and the |
| 3169 | payload validator. A rejected payload names the key to raise, so the model can |
| 3170 | either resize the batch or tell the user which setting to change. |
| 3171 | |
| 3172 | ### User-input wait timeout |
| 3173 | |
| 3174 | Questions from `request_user_input` wait until answered or canceled by default |
| 3175 | (#6003). An omitted setting or `0` leaves the wait unbounded; a positive value |
| 3176 | cancels the question when that many seconds pass, capped at 86,400 (24 hours). |
| 3177 | Headless `exec` runs have no |
| 3178 | responder, so `request_user_input` is withheld there by default: |
| 3179 | the model reports the tool absent and finishes instead of stalling. |
| 3180 | |
| 3181 | ```toml |
| 3182 | [tools] |
| 3183 | user_input_timeout_seconds = 300 # opt into 5 minutes; omitted or 0 waits indefinitely; maximum 86400 |
| 3184 | ``` |
| 3185 | |
| 3186 | This key governs question waits only. Approvals have their own clock, |
| 3187 | `[approval] timeout_seconds`, which waits indefinitely unless you set it. |
| 3188 | Wall-clock and stream protections elsewhere are unaffected. |
| 3189 | |
| 3190 | ## Feature Flags |
| 3191 | |
| 3192 | Feature flags live under the `[features]` table and are merged across profiles. |
| 3193 | Defaults are enabled for built-in tooling, so you only need to set entries you |
| 3194 | want to force on or off. |
| 3195 | |
| 3196 | ```toml |
| 3197 | [features] |
| 3198 | shell_tool = true |
| 3199 | subagents = true |
| 3200 | web_search = true # enables deferred Web; the flag name is retained for config compatibility |
| 3201 | apply_patch = true |
| 3202 | mcp = true |
| 3203 | exec_policy = true |
| 3204 | code_mode = true # execute_tools composes MCP/plugin/native calls; false defers it behind tool_search |
| 3205 | verify_tool = true # agent-callable `verify` self-critique; false removes it from the tool catalog |
| 3206 | vision_model = false # true routes image analysis to the [vision_model] model (see above) |
| 3207 | extension_host = false # experimental: run reviewed plugins' native code (see EXTENSIONS.md) |
| 3208 | ``` |
| 3209 | |
| 3210 | `code_mode` is on by default: `execute_tools` is advertised from the first |
| 3211 | request and nested calls go through the same permission gate as direct calls |
| 3212 | (see [Tool surface](TOOL_SURFACE.md#code-mode-execute_tools)). Set |
| 3213 | `code_mode = false` to defer it behind `tool_search` again. |
| 3214 | |
| 3215 | `extension_host` is experimental and off by default. Turning it on lets reviewed |
| 3216 | plugins run their `native` TypeScript/JavaScript tool and slash-command code in |
| 3217 | a separate Node (or, opt-in, Bun) process; toggling it in either direction |
| 3218 | changes the plugin activation policy, so every plugin is reviewed again after a |
| 3219 | restart. Its `[extension_host]` table (`runtime`, `node`, `bun`) and what the |
| 3220 | host's sandbox does on each platform are in |
| 3221 | [Writing an extension tool](EXTENSIONS.md). |
| 3222 | |
| 3223 | Every flag has a row in [`docs/features.toml`](features.toml), the feature |
| 3224 | registry; a test fails when a flag and its row disagree. |
| 3225 | |
| 3226 | You can also override features for a single run: |
| 3227 | |
| 3228 | - `codewhale --enable web_search` |
| 3229 | - `codewhale --disable subagents` |
| 3230 | |
| 3231 | Use `codewhale features list` to inspect known flags and their effective state. |
| 3232 | The native `/config` view also includes a read-only **Experimental** section |
| 3233 | for experimental feature flags. It shows each flag's effective enabled/disabled |
| 3234 | state and whether that state comes from the default or a configured override. |
| 3235 | Change feature flags in `[features]` or with `--enable` / `--disable`; the |
| 3236 | `/config` section is an audit surface, not a stability promise. Goal and |
| 3237 | Workflow preview rows may appear there as reserved entries until those workflows |
| 3238 | graduate behind real gated flags. |
| 3239 | |
| 3240 | ## Web Search Provider |
| 3241 | |
| 3242 | `web_search` uses keyless Firecrawl by default. Runtime failure or an exhausted |
| 3243 | keyless quota degrades visibly through DuckDuckGo and Bing. China deployments |
| 3244 | can explicitly select Baidu, Metaso, Volcengine, or a trusted SearXNG endpoint; |
| 3245 | Codewhale does not guess geography from locale or model provider. |
| 3246 | |
| 3247 | Configured API providers are attempted first. Runtime failure or an empty |
| 3248 | result visibly degrades through DuckDuckGo and then Bing; the structured search |
| 3249 | receipt records every hop. Missing configuration and network-policy denials |
| 3250 | fail closed without sending the query to another provider. |
| 3251 | |
| 3252 | **Provider-native search.** On routes whose provider offers its own web-search |
| 3253 | tool (OpenAI, xAI, Anthropic, DeepSeek, Kimi and others), that search can run |
| 3254 | ahead of the configured provider. It is a separate model call on the active |
| 3255 | route. `[search] native` decides the order: |
| 3256 | |
| 3257 | - unset (default): native search leads only when no search provider is |
| 3258 | configured; a provider chosen in `[search] provider`, |
| 3259 | `CODEWHALE_SEARCH_PROVIDER`, or a Tavily key wins; |
| 3260 | - `native = true`: native search leads even when a provider is pinned; |
| 3261 | - `native = false`: native search is never used. |
| 3262 | |
| 3263 | The native answer is returned whole; oversized tool output spills to a session |
| 3264 | artifact the model can page back. |
| 3265 | |
| 3266 | **Recency and locale.** `recency` and `locale` are forwarded where the |
| 3267 | backend's API takes them, and the search receipt reports each as honored or |
| 3268 | ignored: |
| 3269 | |
| 3270 | | Backend | Recency | Locale | |
| 3271 | | --- | --- | --- | |
| 3272 | | Firecrawl | `tbs=qdr:d/w/m/y` | `country` from the region (`de-DE` → `DE`); a bare language is ignored | |
| 3273 | | Tavily | `time_range` | not sent (Tavily takes country names) | |
| 3274 | | SearXNG | `time_range` | `language`, as given | |
| 3275 | | Serply | not sent (undocumented) | `hl` language, `gl` country | |
| 3276 | |
| 3277 | Recency is rounded up to the backend's nearest window (day, week, month, |
| 3278 | year), so `recency = 10` searches the last month. Other backends ignore both |
| 3279 | knobs and say so in the receipt. |
| 3280 | |
| 3281 | For a private/internal search service that serves DuckDuckGo-compatible HTML, |
| 3282 | keep `provider = "duckduckgo"` and set `base_url`; Codewhale appends the `q` |
| 3283 | query parameter to that endpoint and applies network policy to its host. |
| 3284 | Custom endpoints do not fall back to public Bing. `CODEWHALE_SEARCH_BASE_URL` |
| 3285 | can override this per process; `DEEPSEEK_SEARCH_BASE_URL` remains accepted as |
| 3286 | the legacy alias. |
| 3287 | |
| 3288 | **SearXNG** ([docs](https://docs.searxng.org/dev/search_api.html)) uses the |
| 3289 | configured instance's JSON API. Set `provider = "searxng"` and |
| 3290 | `base_url = "https://your-searxng.example"`; Codewhale calls |
| 3291 | `/search?q=...&format=json`. Codewhale does not use a public SearXNG instance |
| 3292 | by default because public instances often disable JSON output or rate-limit API |
| 3293 | traffic. |
| 3294 | |
| 3295 | Self-host it as a separate process (Docker is fine); Codewhale never bundles or |
| 3296 | manages the search engine itself: |
| 3297 | |
| 3298 | - Enable JSON on the instance (`settings.yml`, `search.formats` must include |
| 3299 | `json`) and restart it. An HTML-only instance answers the API with HTTP 403; |
| 3300 | Codewhale reports that as a JSON/API-access problem on the SearXNG hop rather |
| 3301 | than silently returning no results. |
| 3302 | - Bind it to loopback, or to a host and port your network policy allows. The |
| 3303 | instance keeps its own engine list, limiter, and limits. |
| 3304 | - Point Codewhale at it with `[search] provider = "searxng"` and `base_url` |
| 3305 | (required; either the root URL or the `/search` endpoint). No instance ships |
| 3306 | as a default, and none is discovered automatically. |
| 3307 | - `codewhale doctor --probe-search` sends a transport-only `HEAD` to that |
| 3308 | origin — no `q=`, no credentials, no redirects, no audit receipt — so a green |
| 3309 | probe proves reachability and network-policy admission, not that JSON is on. |
| 3310 | |
| 3311 | Confirm the JSON API itself before assuming a Codewhale bug: |
| 3312 | |
| 3313 | ```sh |
| 3314 | curl -sS "$BASE/search?q=codewhale&format=json" | jq '.results[0] | {title,url,score}' |
| 3315 | ``` |
| 3316 | |
| 3317 | Codewhale ranks the returned rows by `score`, highest first, and applies |
| 3318 | `max_results` to that ranking; rows an instance reports without a usable score |
| 3319 | keep their original relative order. |
| 3320 | |
| 3321 | **Metaso** ([metaso.cn](https://metaso.cn)) requires a user-supplied key. Set |
| 3322 | `METASO_API_KEY` or `[search] api_key`; Codewhale does not ship a shared key. |
| 3323 | |
| 3324 | **Firecrawl** ([docs](https://docs.firecrawl.dev/sdks/cli)) searches Firecrawl |
| 3325 | Cloud without a key using its bounded per-IP daily quota. Set |
| 3326 | `FIRECRAWL_API_KEY` or `[search] api_key` for authenticated limits. Codewhale |
| 3327 | sends no `Authorization` header in keyless mode. |
| 3328 | |
| 3329 | **Baidu** uses Baidu AI Search at |
| 3330 | `https://qianfan.baidubce.com/v2/ai_search/web_search`. Set |
| 3331 | `BAIDU_SEARCH_API_KEY` or `[search] api_key`. This is a search-tool backend |
| 3332 | only; it does not add a Baidu model provider. |
| 3333 | |
| 3334 | **Sofya** ([sofya.co](https://sofya.co)) returns full extracted page content |
| 3335 | rather than snippets. Set `[search] api_key` to your `ay_live_...` key, or the |
| 3336 | `SOFYA_API_KEY` env var. This is a search-tool backend only; it does not add a |
| 3337 | Sofya model provider. |
| 3338 | |
| 3339 | **Serply** ([serply.io](https://serply.io)) returns Google organic results with |
| 3340 | title, URL, and snippet. Set `[search] api_key` to your Serply key, or the |
| 3341 | `SERPLY_API_KEY` env var. This is a search-tool backend only; it does not add a |
| 3342 | Serply model provider. |
| 3343 | |
| 3344 | **Tavily** ([tavily.com](https://tavily.com)) is selected automatically when a |
| 3345 | Tavily key is present and no provider is pinned: `TAVILY_API_KEY` set, or |
| 3346 | `[search] api_key` / `CODEWHALE_SEARCH_API_KEY` in the `tvly-` family. Doctor |
| 3347 | reports that as `source: tavily key`. Autodetect is runtime-only — Codewhale |
| 3348 | never writes `[search] provider` for it, and `TAVILY_API_KEY` is never merged |
| 3349 | into `[search] api_key`. An explicit `[search] provider` or |
| 3350 | `CODEWHALE_SEARCH_PROVIDER` always wins, so `provider = "firecrawl"` keeps |
| 3351 | Firecrawl even with a Tavily key in the environment. Pinned `tavily` accepts |
| 3352 | any non-empty `[search] api_key` and is configured by that key or |
| 3353 | `TAVILY_API_KEY`; with both empty it fails closed. |
| 3354 | |
| 3355 | ```toml |
| 3356 | [search] |
| 3357 | provider = "firecrawl" # also duckduckgo | bing | tavily | bocha | metaso | searxng | baidu | volcengine | sofya | serply |
| 3358 | # base_url = "https://search.example/" # optional with provider = "duckduckgo"; required with "searxng" |
| 3359 | # api_key = "YOUR_KEY" # optional for firecrawl; required by the other API providers |
| 3360 | # native = false # provider-native search: unset = only when no provider is configured |
| 3361 | ``` |
| 3362 | |
| 3363 | ## Local Media Attachments |
| 3364 | |
| 3365 | Use `@path/to/file` in the composer to add local text file or directory context |
| 3366 | to the next message. Use `/attach <path>` for local image/video media paths, or |
| 3367 | `Ctrl+V` to attach an image from a local clipboard or an explicitly forwarded |
| 3368 | X11/Wayland clipboard. SSH terminal paste without a forwarded graphical display |
| 3369 | is text-only; use the local terminal's paste command (`Cmd+V` on macOS or |
| 3370 | `Ctrl+Shift+V` on Linux/Windows), and use `/attach <path>` for remote image |
| 3371 | files. OpenSSH loopback X11 displays are detected automatically. For an |
| 3372 | explicitly forwarded Wayland or non-loopback X11 display, set |
| 3373 | `CODEWHALE_SSH_CLIPBOARD=graphical`; set it to `terminal` to force terminal |
| 3374 | transfer instead of an ambient remote display. DeepSeek's public Chat |
| 3375 | Completions API currently accepts text message |
| 3376 | content, so media attachments are sent as explicit local path references instead |
| 3377 | of native image/video payloads. |
| 3378 | Attachment rows appear above the composer before submit; move to the start of |
| 3379 | the composer, press `↑` to select an attachment row, then press `Backspace` or |
| 3380 | `Delete` to remove it without editing the sample text by hand. |
| 3381 | |
| 3382 | ## Managed Configuration and Requirements |
| 3383 | |
| 3384 | codewhale supports a policy layering model: |
| 3385 | |
| 3386 | 1. user config + profile + env overrides |
| 3387 | 2. managed config (if present) |
| 3388 | 3. requirements validation (if present) |
| 3389 | |
| 3390 | By default on Unix: |
| 3391 | - managed config: `/etc/deepseek/managed_config.toml` |
| 3392 | - requirements: `/etc/deepseek/requirements.toml` |
| 3393 | |
| 3394 | Requirements file shape: |
| 3395 | |
| 3396 | ```toml |
| 3397 | allowed_approval_policies = ["on-request", "untrusted", "never"] |
| 3398 | allowed_sandbox_modes = ["read-only", "workspace-write"] |
| 3399 | ``` |
| 3400 | |
| 3401 | If configured values violate requirements, startup fails with a descriptive error. |
| 3402 | |
| 3403 | ## Notes On `codewhale doctor` |
| 3404 | |
| 3405 | `codewhale doctor` follows the same config resolution rules as the rest of the |
| 3406 | TUI. That means `--config`, `CODEWHALE_CONFIG_PATH`, and the legacy |
| 3407 | `DEEPSEEK_CONFIG_PATH` are respected, and MCP/skills |
| 3408 | checks use the resolved `mcp_config_path` / `skills_dir` (including env overrides). |
| 3409 | |
| 3410 | To bootstrap missing MCP/skills paths, run `codewhale setup --all`. You can |
| 3411 | also run `codewhale setup --skills --local` to create a workspace-local |
| 3412 | `./skills` dir. |
| 3413 | |
| 3414 | Both plain `codewhale doctor` and `codewhale doctor --json` are structural and |
| 3415 | offline by default. They do not check the release service, hosted provider |
| 3416 | APIs, local provider endpoints, or MCP processes, and they do not load a |
| 3417 | workspace credential `.env`. Use `--check-updates`, |
| 3418 | `--probe-api`, `--probe-local`, or `--probe-mcp` to opt into the corresponding |
| 3419 | live boundary; `--probe-local` may start a desktop-managed service such as |
| 3420 | Ollama. Only the explicit API/local probe paths may load workspace credential |
| 3421 | `.env` values. Live flags conflict with `--json`, so machine-readable doctor |
| 3422 | output is always offline. Top-level keys include `version`, `paths`, `secret_backend`, |
| 3423 | `config_path`, `config_present`, `workspace`, `api_key.source`, |
| 3424 | `api_key.availability`, `base_url`, |
| 3425 | `default_text_model`, `mcp`, `skills`, `tools`, `plugins`, `sandbox`, |
| 3426 | `platform`, `api_connectivity`, and `capability`. CI consumers should rely on |
| 3427 | `api_key.source` (`config_declared`/`env_declared`/`external_auth_declared`/ |
| 3428 | `secret_store_unprobed`/`secret_store_unavailable`/`oauth_unprobed`/ |
| 3429 | `external_consent`/`none`/`local_runtime`/`unknown`) and |
| 3430 | `api_key.availability` |
| 3431 | (`present`/`not_required`/`not_probed`/`unavailable`/`unknown`) rather than parsing the |
| 3432 | human-readable `doctor` text. Source is declaration metadata, not proof that a |
| 3433 | credential exists or works. Only a non-empty, non-sentinel literal config value |
| 3434 | is structurally `present`; no-auth and local routes are `not_required`. Environment, |
| 3435 | external-auth, OAuth, consent, and secret-store declarations remain `not_probed` |
| 3436 | and cannot make structural Setup or Fleet readiness true. A secret-store sentinel |
| 3437 | on a named/custom endpoint that is prohibited from using the shared store is |
| 3438 | `secret_store_unavailable`/`unavailable`, while `unknown` remains reserved for |
| 3439 | the absence of a supported structural conclusion. Exact and whitespace-wrapped |
| 3440 | legacy sentinels are never treated as literal credentials. The structural |
| 3441 | loader still honors safe environment routing/model/policy fields, but it never |
| 3442 | materializes environment HTTP headers, sandbox API keys, or search API keys; |
| 3443 | only an explicit API/local probe switches to the normal credential-loading |
| 3444 | path. An opted-in update check also emits only typed generic failures: untrusted |
| 3445 | release metadata and transport errors are not echoed. |
| 3446 | |
| 3447 | If configuration loading or validation fails, `doctor --json` returns nonzero |
| 3448 | and prints a bounded JSON error envelope with |
| 3449 | `status = "error"` and `error.kind = "config_validation"`. It does not emit a |
| 3450 | normal route or capability report—or the underlying possibly sensitive error— |
| 3451 | for an invalid configuration. |
| 3452 | |
| 3453 | MCP entries are configuration diagnostics unless an explicit MCP command is |
| 3454 | run. `mcp.probe_scope` is `configuration`, `mcp.live_health_checked` is false, |
| 3455 | and each server separates `checks.configuration` / `checks.command` from |
| 3456 | `checks.process_reachable`, `checks.protocol_initialized`, and |
| 3457 | `checks.backend_tool_health`. The latter three remain `not_checked` in doctor |
| 3458 | output. Run `codewhale mcp validate` to explicitly start enabled servers and |
| 3459 | verify protocol initialization/discovery; backend health still requires an |
| 3460 | appropriate explicit tool call. Doctor reports only safe structural MCP fields: |
| 3461 | URL userinfo/path/query/fragment and raw command arguments, environment values, |
| 3462 | header values, and token material are omitted. Provider URLs follow the same |
| 3463 | rule and expose only `scheme://host[:explicit-port]`. |
| 3464 | |
| 3465 | The `capability` key contains per-provider capability info derived from |
| 3466 | static knowledge (release docs, API guides) rather than live API probes. |
| 3467 | Top-level sub-keys: `resolved_provider`, `resolved_model`, `context_window`, |
| 3468 | `max_output`, `thinking_supported`, `cache_telemetry_supported`, |
| 3469 | and `request_payload_mode`. |
| 3470 | |
| 3471 | Use `capability.context_window` and `capability.max_output` for model-limit |
| 3472 | checks in CI scripts; do not treat `capability.max_output` as the per-turn |
| 3473 | request budget. Use `capability.thinking_supported` to decide whether to |
| 3474 | configure reasoning effort. |
| 3475 | |
| 3476 | ## Setup status, clean, and extension dirs |
| 3477 | |
| 3478 | `codewhale setup` accepts a few flags beyond the existing `--mcp`, |
| 3479 | `--skills`, `--local`, `--all`, and `--force`: |
| 3480 | |
| 3481 | - `--status` — print a compact one-screen status (api key, base URL, model, |
| 3482 | MCP/skills/tools/plugins counts, sandbox, `.env` presence). Read-only and |
| 3483 | network-free; safe to run in CI. If `.env` is missing and `.env.example` is |
| 3484 | present in the workspace, the status output points at `cp .env.example .env`. |
| 3485 | - `--tools` — scaffold `~/.codewhale/tools/` with a `README.md` describing the |
| 3486 | self-describing frontmatter convention (`# name:` / `# description:` / |
| 3487 | `# usage:`) and an `example.sh` that follows it. The directory is |
| 3488 | intentionally not auto-loaded; wire individual scripts into the agent via |
| 3489 | MCP, hooks, or skills. |
| 3490 | - `--plugins` — scaffold `~/.codewhale/plugins/` with a `README.md` and an |
| 3491 | `example/plugin.toml` plus a namespaced example Skill. Bundles are discovered |
| 3492 | read-only, untrusted, and disabled; review them through `/plugin` before |
| 3493 | enabling. v0.9.1 activates only declared Skills and MCP servers. See |
| 3494 | [PLUGIN_BUNDLES.md](PLUGIN_BUNDLES.md). |
| 3495 | - `--all` now scaffolds MCP + skills + tools + plugins together. |
| 3496 | - `--clean` — list `~/.codewhale/sessions/checkpoints/latest.json` and |
| 3497 | `offline_queue.json` if they exist. Legacy |
| 3498 | `~/.deepseek/sessions/checkpoints/` files are not scanned automatically; set |
| 3499 | `CODEWHALE_HOME=~/.deepseek` for a one-off legacy cleanup. Pass `--force` to |
| 3500 | actually remove matched files. This never touches real session history or the |
| 3501 | task queue. |
| 3502 | |
| 3503 | `--status` and `--clean` are mutually exclusive with the scaffold flags. |
| 3504 | |
| 3505 | ## Why the engine strips XML/`[TOOL_CALL]` text |
| 3506 | |
| 3507 | codewhale sends and receives tool calls only over the API tool channel |
| 3508 | (structured `tool_use` / `tool_call` items). The streaming loop in |
| 3509 | `crates/tui/src/core/engine.rs` recognizes a fixed set of fake-wrapper start |
| 3510 | markers — `[TOOL_CALL]`, `<codewhale:tool_call`, `<tool_call`, `<invoke `, |
| 3511 | `<function_calls>` — and scrubs them from visible assistant text without ever |
| 3512 | turning them into structured tool calls. When a wrapper is stripped, the loop |
| 3513 | emits one compact `status` notice per turn so the user can see why their |
| 3514 | visible text shrank. Treat any change that re-enables text-based tool |
| 3515 | execution as a regression; the protocol-recovery tests in |
| 3516 | `crates/tui/tests/integration/protocol_recovery.rs` lock the contract. |
| 3517 | |
| 3518 | ## Model-bound redaction (`[redaction] model_bound`) |
| 3519 | |
| 3520 | Codewhale masks credential-looking values in tool output **before it is sent |
| 3521 | to an upstream model** — the "model boundary". A file read by a tool can |
| 3522 | contain a configured API key, a bare provider token, or a credential-shaped |
| 3523 | opaque string, and the model must not see those bytes. This backstop is |
| 3524 | separate from the display/export scrubbers: it decides what the model itself |
| 3525 | can quote back, and it is deliberately conservative (`CredentialShaped` |
| 3526 | policy, see `crates/config/src/persistence.rs`), so ordinary code and config |
| 3527 | stay byte-exact while keys, JWTs, bearer tokens, PEM blocks, and long opaque |
| 3528 | runs are masked. |
| 3529 | |
| 3530 | Turning that masking **off** is a security decision, so it is not a plain |
| 3531 | boolean: |
| 3532 | |
| 3533 | ```toml |
| 3534 | [redaction] |
| 3535 | model_bound = "disabled" # "enabled" (default) | "disabled" |
| 3536 | ``` |
| 3537 | |
| 3538 | Setting `"disabled"` only records a *request*. It takes effect only when all |
| 3539 | of these are true: |
| 3540 | |
| 3541 | 1. You restart the interactive TUI. |
| 3542 | 2. The startup gate appears and you press `1`/`Y` on its first stage |
| 3543 | ("confirm and disable"). This only advances to a second, final-confirmation |
| 3544 | stage - the gate repeats the red warning and asks "are you really sure?". |
| 3545 | 3. On that second stage you press `1`/`Y` again. The gate is rendered with the |
| 3546 | same explicit-key discipline as workspace trust - `Enter` never confirms by |
| 3547 | reflex, and `2`/`U` on the second stage steps back. |
| 3548 | 4. Only that second confirmation persists a receipt to |
| 3549 | `~/.codewhale/redaction-state.json` (next to `config.toml`) and rebuilds |
| 3550 | the engine with masking off for the rest of this launch and future ones. |
| 3551 | |
| 3552 | The receipt is bound to the config it was made against and is valid only |
| 3553 | while that config still requests `"disabled"`. Setting `model_bound` back |
| 3554 | to `"enabled"` - or rewriting `config.toml` in any way after the |
| 3555 | confirmation - invalidates it, so requesting `"disabled"` again later |
| 3556 | asks for a fresh confirmation on the next launch. The receipt checks both the |
| 3557 | config contents and modification time; missing, unreadable, or malformed config |
| 3558 | and older receipts without this binding keep masking enabled. Legacy-home |
| 3559 | installs store the receipt beside their resolved config file. |
| 3560 | |
| 3561 | Until a confirmation exists, the effective mode is always `"enabled"`: |
| 3562 | |
| 3563 | - Choosing `2`/`U` ("keep masking on") leaves the config field untouched, so |
| 3564 | the next launch asks again. Edit the field back to `"enabled"` to stop being |
| 3565 | asked. |
| 3566 | - Non-interactive entry points (`codewhale exec`, hooks, automations, headless |
| 3567 | agents) never confirm anything and never apply an unconfirmed request. |
| 3568 | - Routing/classification summaries and durable goal-state text keep their own |
| 3569 | always-on redaction regardless of this switch; the opt-out exists so the |
| 3570 | model can quote file bytes for exact edits, not to relax stored state. |
| 3571 | |
| 3572 | The config value itself is forgiving: `true`/`false`, `"on"`/`"off"`, and |
| 3573 | `"enabled"`/`"disabled"` (any casing) all parse, with `false`/`"off"` meaning |
| 3574 | `"disabled"`. |
| 3575 | |
| 3576 | A confirmed opt-out still sends your configured API keys to the provider you |
| 3577 | are already talking to. Only use it when the model must read and edit files |
| 3578 | that contain real credentials. |
| 3579 | |
| 3580 | ### Stored sessions |
| 3581 | |
| 3582 | The masking runs once, when tool output enters the transcript, so saved |
| 3583 | session files (`~/.codewhale/sessions/*.json` and their checkpoints) and |
| 3584 | Runtime API thread items (their text and tool metadata) do not store a |
| 3585 | credential-shaped value from tool output. With the confirmed opt-out above, |
| 3586 | tool output is stored as the model saw it. |
| 3587 | |
| 3588 | Not yet masked: large outputs spilled to `~/.codewhale/tool_outputs`, shell |
| 3589 | completion evidence artifacts, and tool-call *inputs* (for example a |
| 3590 | `curl -H "Authorization: Bearer …"` command line). `doctor` and |
| 3591 | `scrub-secrets` do not check those either. |
| 3592 | |
| 3593 | Sessions saved by builds before this change may still hold credentials in |
| 3594 | their tool output; nothing rewrites them automatically. `codewhale doctor` |
| 3595 | reports them under **Stored Sessions** (it checks the newest 50 files), and |
| 3596 | this command finds and masks them across every saved session: |
| 3597 | |
| 3598 | ```sh |
| 3599 | codewhale sessions scrub-secrets # report only |
| 3600 | codewhale sessions scrub-secrets --apply # rewrite the affected files |
| 3601 | ``` |
| 3602 | |
| 3603 | `--apply` rewrites each file under the same per-session lock a save takes, |
| 3604 | so it never loses a concurrent save. A session that is still open can write |
| 3605 | its in-memory copy back on its next save, so close open sessions first, and |
| 3606 | rotate any credential that was exposed — masking a stored copy cannot un-leak it. |
| 3607 |