返回 CodeWhale
RUNTIME_API.md
根目录 / docs / RUNTIME_API.md
1 # Codewhale Engine Runtime API & Integration Contract
2
3 > 阅读简体中文版:[zh_hans/RUNTIME_API.md](zh_hans/RUNTIME_API.md)。
4
5 `codewhale app-server` exposes the [Codewhale Engine](ARCHITECTURE.md) through
6 its canonical local Runtime API and control plane.
7 Local SDKs, mobile/remote-control clients, and editor integrations talk to it
8 instead of screen-scraping terminal output. It serves the full HTTP/SSE runtime
9 API (`/v1/*`), a JSON-RPC control transport over stdio, and the phone-friendly
10 mobile page. `codewhale doctor --json` provides machine-readable health, and
11 `codewhale serve --acp` speaks the Agent Client Protocol over stdio for editors
12 such as Zed.
13
14 `codewhale serve --http` / `serve --mobile` remain as **compatibility aliases**
15 for `codewhale app-server --http` / `--mobile`; both launch the identical
16 server. New integrations should target `app-server`.
17
18 `codewhale exec` is the separate one-shot headless worker path (stream-json,
19 fleet worker subprocess, CI primitive). It is not part of this API, but it
20 shares the same runtime, provider/model resolution, permission profiles, and
21 event vocabulary.
22
23 This document is the stable integration contract for native workbench
24 applications (and other local supervisors) that embed the Codewhale Engine.
25
26 ## Architecture
27
28 ```
29 local supervisor / SDK / automation harness
30 │
31 ├─ codewhale app-server --http → HTTP/SSE runtime API (/v1/*) [canonical]
32 ├─ codewhale app-server --mobile → runtime API + mobile control page
33 ├─ codewhale app-server --stdio → JSON-RPC control transport over stdio
34 ├─ codewhale app-server --socket → same JSON-RPC over a unix domain socket (desktop daemon)
35 ├─ codewhale doctor --json → machine-readable health & capability
36 ├─ codewhale serve --acp → ACP stdio agent for editors such as Zed
37 ├─ codewhale serve --mcp → MCP stdio server
38 ├─ codewhale serve --http/--mobile → legacy aliases for `app-server --http/--mobile`
39 └─ codewhale exec [args] → one-shot headless worker (stream-json)
40 ```
41
42 The engine runs as a local-only process. All APIs bind to `localhost` by
43 default. No hosted relay, no provider-token custody, no secret leakage.
44
45 For the read-only record of what a thread or turn did, see
46 [`docs/RECEIPTS.md`](RECEIPTS.md): `codewhale receipts` on the CLI and the
47 `/receipt` routes under **Threads** below.
48
49 ## Runtime API entrypoints
50
51 | Entry | Transport | Use |
52 |---|---|---|
53 | `codewhale web [--port 7878]` | HTTP/SSE on `127.0.0.1:7878` + embedded client | First-class loopback-only browser client; opens the default browser |
54 | `codewhale app-server --http` | HTTP/SSE on `127.0.0.1:7878` | Full `/v1/*` runtime API (canonical) |
55 | `codewhale app-server --mobile` | HTTP/SSE on loopback + `/mobile` | Runtime API + local mobile control page |
56 | `codewhale app-server --stdio` | JSON-RPC 2.0 over stdio | Local SDK / control probe (no listener) |
57 | `codewhale app-server --socket [--socket-path P]` | JSON-RPC 2.0 over a `0600` unix domain socket | Desktop daemon: multi-client, peer-uid checked, `daemon/attach` claim handshake (macOS/Linux; Windows named pipe reserved, not implemented) |
58 | `codewhale app-server` | HTTP on `127.0.0.1:8787` | Legacy in-process app-server (`/healthz`, `/thread`, `/app`, `/prompt`, `/jobs`); `/prompt` and `/thread` messages execute real turns via the runtime bridge. There is no direct `/tool` route: tools run only inside Engine turns, under the Engine's tool catalog and approval posture. This legacy server does not surface approvals: its bridge forwards only text deltas and the turn's completion, and it has no decision route, so an approval-gated call waits unanswered. Drive approval-gated work through the Runtime API (`/v1/threads/*` events and `POST /v1/approvals/{approval_id}`) |
59 | `codewhale serve --http` / `--mobile` | same server as `app-server --http`/`--mobile` | Compatibility aliases |
60
61 `app-server --http` and `--mobile` launch the same mature runtime API server
62 historically reached through `serve --http` — no routes or behavior changed, so
63 every endpoint documented below is identical across both entrypoints. The
64 runtime API token is read from `--auth-token`, then `CODEWHALE_RUNTIME_TOKEN`,
65 then `DEEPSEEK_RUNTIME_TOKEN`; use `--insecure-no-auth` only with a loopback
66 bind. The `serve` compatibility aliases keep their `--insecure` flag.
67 The legacy in-process `codewhale app-server` also requires an explicit
68 `--auth-token` or `CODEWHALE_APP_SERVER_TOKEN` before binding a non-loopback
69 host; its generated one-time `cwapp_*` token is loopback-only.
70
71 ### Workspace file suggestions
72
73 `GET /v1/workspace/files/search?query=runtime&limit=20` returns
74 `{"paths":["src/runtime.rs"]}` through the existing authenticated `/v1/*`
75 router. It searches only the server's configured workspace, not a thread's
76 workspace or the process's current directory. No workspace/path override is
77 accepted. The response contains workspace-relative file paths with `/`
78 separators, never file contents, absolute paths, or directories.
79
80 - `query` is a literal partial filename/path, without an `@` prefix, at most
81 256 UTF-8 bytes. Missing, empty, or whitespace-only queries return an empty
82 list without walking the filesystem. No match also returns an empty list.
83 - `limit` defaults to 20; accepted values are 1–100. Invalid limits, oversized
84 queries, and unknown query parameters return HTTP 400.
85 - Matching reuses TUI fuzzy `@file` discovery/ranking: case-insensitive path
86 prefix matches first, then substring matches, alphabetically within each
87 group. This is not glob, subsequence, content, or semantic search, and does
88 not apply the TUI's personal frecency boosts.
89 - Discovery shares the composer's ignore policy, including `.ignore` and
90 `.deepseekignore`, always-discoverable AI directories, and the bounded
91 hidden/gitignored local-reference fallback. The special `.agents`, `.claude`,
92 `.cursor`, and `.deepseek` walks intentionally bypass ignore rules, as in
93 the TUI. Ignore files are not confidentiality boundaries.
94 - Directory symlinks are not traversed. Files are canonicalized and filtered
95 for containment in the workspace before applying the result limit; external
96 and broken file symlinks are omitted. In-workspace file symlinks may appear
97 by their relative names. Suggestions are a filesystem snapshot, not
98 authorization to read a file later; consumers must revalidate when opening it.
99
100 Discovery runs off the async executor, with the shared default depth of 10,
101 at most 20,000 candidates, and a cooperative two-second discovery budget.
102 Results are best-effort, not an exhaustive listing; a slow filesystem operation
103 can finish after that budget. Each request scans anew; there is no new index or
104 cache. This read-only endpoint does not alter sessions or the pinned model
105 prompt/tool prefix.
106
107 ### Workspace files and session artifacts
108
109 Native clients (the GPUI desktop's Files and Preview modules) browse and edit
110 the server's configured workspace through three authenticated routes. They
111 read and write the workspace directly; there is no second file store, cache
112 or index, and no path override: the workspace root is the only root.
113
114 - `GET /v1/workspace/files?path=<dir>&limit=<1-2000>` lists one directory.
115 `path` is workspace-relative with `/` separators; empty or `.` is the root.
116 Each entry carries `name`, `path`, `kind` (`file`, `directory`, `symlink`,
117 `other`), and for files `size` and `modified` (RFC 3339). Directories sort
118 first, then names case-insensitively. `limit` defaults to 200; `truncated`
119 reports a cut. `.git` is never listed or served, and symlinks are listed by
120 name only: they are never followed, so `path=<link>` returns 403.
121 - `GET /v1/workspace/files/read?path=<file>&offset=<bytes>&limit=<1-4194304>`
122 returns one byte window of a regular file with `size`, `revision` (the
123 SHA-256 hex of the **whole** file, not of the window), `modified`,
124 `offset`, `bytes`, `truncated`, `encoding` and `content`. Text windows are
125 `utf-8`; a window with a NUL byte, invalid UTF-8, or a split multi-byte
126 character is `base64`. `limit` defaults to 256 KiB. Files above 16 MiB are
127 refused with 413; a directory is 400; a link is 403; a missing file is 404.
128 - `PUT /v1/workspace/files` with `{"path", "content", "encoding"?,
129 "expected_revision"?}` writes one file atomically through the same confined
130 opener Fleet artifacts use. `encoding` is `utf-8` (default) or `base64`;
131 bodies above 4 MiB are 413. Creating a new file requires **no**
132 `expected_revision` (and creates missing parent directories inside the
133 workspace); overwriting requires the `revision` from the read that the
134 edit was based on, and a stale or missing one is 409 with the current
135 revision in the error message so the client can re-read and merge. This is
136 optimistic concurrency, not a lock: two writers racing between the check and
137 the write can still interleave. The response carries `path`, `size`,
138 `revision`, `created` and `written_at`; 201 for a new file, 200 otherwise.
139 Writes through a link, into `.git`, or to a directory are refused.
140
141 Every path is validated before any filesystem access: absolute paths,
142 backslashes, `.` or `..` components are 400, and each directory on the way is
143 opened without following links (`O_NOFOLLOW` per component on Unix, reparse
144 point checks on Windows). These routes use the runtime bearer token like every
145 other `/v1/*` route; they do not consult the model's tool permission posture,
146 because the caller is the authenticated operator, not the model.
147
148 Session artifacts are the oversized tool outputs a session recorded as
149 `ArtifactRecord`s (`crates/tui/src/artifacts.rs`), stored under
150 `sessions/<id>/artifacts/`:
151
152 - `GET /v1/sessions/{id}/artifacts` lists the records a saved session carries:
153 `id`, `kind`, `tool_call_id`, `tool_name`, `created_at`, `byte_size`,
154 `preview` and the session-relative `path`.
155 - `GET /v1/sessions/{id}/artifacts/{artifact_id}?offset=&limit=` reads one
156 artifact with the same window, `revision` and `encoding` contract as the
157 workspace file read. A record whose stored path is absolute or leaves the
158 session directory is 403; a record whose file is gone is 404.
159
160 These routes only serve what a SavedSession indexes (plus immutable image
161 evidence). A runtime turn's spills are read through the turn route below. An
162 unbound runtime thread's engine has no SavedSession index at all, so for its
163 spills the turn record is the only way in.
164
165 #### Turn artifacts
166
167 A turn records what it produced as typed references on its items and on the
168 turn itself. Nothing is scanned to build them. Each fact is recorded where the
169 bytes were written:
170 - file tools and `apply_patch` report `size`/`sha256` in `mutation.files[]`;
171 - spills report `artifact_digest`;
172 - tool media reports `sha256`.
173
174 The workspace-level half comes from the turn's own restore points: the
175 `pre_turn` and `post_turn` receipts in `TurnRecord.workspace_snapshots` (see
176 "Workspace restore points" below), diffed by tree id in the existing side
177 repo. No second store, snapshot or event is involved.
178
179 A reference (`TurnArtifactRef`) carries:
180
181 | Field | Meaning |
182 | --- | --- |
183 | `id` | Stable within the turn. A file's id is `file_` plus the first 32 hex digits of SHA-256(path). A spill's is `art_<call>`. Media's is `art_image_<sha256>`. |
184 | `kind` | `file`, `tool_output` or `media`. |
185 | `path` | For `file`: workspace-relative with `/` separators. For `tool_output` and `media`: session-relative (`artifacts/...`). |
186 | `change` | For `file` only: `created`, `updated`, `deleted` or `renamed`. A rename also has `previous_path`. |
187 | `size` | Byte size. Absent when the file was deleted. |
188 | `revision` | SHA-256 hex of the whole content. This is the value `GET /v1/workspace/files/read` reports as `revision`, and file-revert's `expected_hash` is `sha256:` + `revision`. Absent when the file was deleted, or for a delta blob over 16 MiB. |
189 | `content_type` | For `media`: the exact media type. |
190 | `session_id` | For `tool_output` and `media`: the artifact session that owns the bytes. |
191 | `item_id`, `tool_call_id`, `tool_name` | The tool call that wrote it. Absent for a change seen only in the workspace delta. |
192 | `source` | `tool_mutation`, `tool_output_spill`, `tool_media`, or `workspace_changed_during_turn`. |
193 | `restore_snapshot_id` | A restore point `POST /v1/threads/{id}/file-revert` accepts for this path on this thread: the `tree_id` of a receipt recorded on this turn's `workspace_snapshots`, so the thread owns it whether or not it is bound to a saved session. For a tool write it is the call's `tool` receipt; for a delta change it is the turn's `pre_turn` receipt. Absent when the turn recorded no such receipt. |
194 | `recorded_at` | When the reference was recorded. |
195
196 Where references appear:
197 - **Items.** `TurnItemRecord.artifacts` lists what one tool call produced. It
198 is set when the call completes, succeeded or failed, so `item.completed` and
199 `item.failed` carry it live.
200 - **Legacy projection.** `artifact_refs` is derived from `artifacts`. It holds
201 only the workspace-relative paths of files that still exist: never spills,
202 media or deleted files.
203 - **Turns.** `TurnRecord.artifacts` is the turn aggregate. It is computed by
204 one merge and ordered most recent first.
205 - Spill and media refs are always kept.
206 - File refs compose in item order: created then deleted drops the file,
207 created then updated stays `created`, and a rename folds its origin.
208 - Once the workspace delta settles, it is authoritative for the net change,
209 `size` and `revision` of every path the snapshots can see. It adds files
210 no tool receipt named, such as shell and sub-agent writes. It drops an item
211 path the snapshots track but that ended the turn unchanged.
212 - The aggregate is capped at 1000 refs, and `workspace.truncated` /
213 `workspace.omitted` report the cut.
214
215 `TurnRecord.workspace` follows the delta's lifecycle. It is `null` while the
216 turn runs. `turn.completed` carries it with one of these states:
217 - `pending`: the turn recorded both a `pre_turn` and a `post_turn` receipt,
218 and their diff is still running. When it finishes, the runtime publishes
219 `turn.artifacts`
220 (`{turn_id, workspace, artifacts}`). That event may arrive after the next
221 turn's `turn.started`, so key it by `turn_id`.
222 - `settled`: the delta is merged. `pre_turn_snapshot_id` and
223 `post_turn_snapshot_id` are the tree ids of the pair.
224 - `unavailable`: no delta will come. `reason` says why:
225 - `snapshots_disabled`, `workspace_too_large`, `too_many_files`,
226 `unsafe_location`, `snapshot_failed`: the snapshot gates. A turn with a
227 `pre_turn` receipt but no `post_turn` one keeps `pre_turn_snapshot_id`.
228 - `not_captured`: the turn recorded no restore point: a compaction or purge
229 operation, or a turn whose engine ended before taking one.
230 - `runtime_restarted`: the process stopped before the delta settled. This is
231 reconciled at startup and never recomputed.
232 - `delta_failed`: the diff itself failed.
233
234 `artifacts` still holds what the tool receipts recorded. `turn.artifacts` is
235 published for every settlement outcome, so a client waiting on `pending`
236 always hears back.
237
238 What the delta means, and what it cannot see:
239 - **It is a workspace diff, not attribution.** A delta change is everything
240 that changed in the workspace while the turn ran. That includes an editor,
241 another thread, or a background job writing the same workspace at the same
242 time. `source: workspace_changed_during_turn` says exactly that.
243 - **Excluded paths are invisible to snapshots.** This covers the built-in
244 excludes (for example `node_modules/`, `target/`, `dist/`, `build/`,
245 `.next/`, and binary and media extensions) and the workspace's `.gitignore`. A file tool's write to
246 such a path is still reported from its receipt. A shell command's write to
247 one is not reported at all.
248 - **Shell writes have no per-call record.** A shell command's writes are
249 never attributed to its item, only to the turn, and only when snapshots are
250 enabled.
251
252 Routes:
253
254 - `GET /v1/threads/{id}/turns/{turn_id}/artifacts` returns
255 `{thread_id, turn_id, workspace, artifacts}` from the runtime store's turn
256 record. While the turn runs, `artifacts` is merged from its items on the fly
257 and `workspace` is `null`. An unknown turn, or a turn of another thread, is
258 404.
259 - `GET /v1/threads/{id}/turns/{turn_id}/artifacts/{artifact_id}?offset=&limit=&revision=`
260 reads one reference.
261 - It uses the workspace file read's window contract (`size`, `revision`,
262 `offset`, `bytes`, `truncated`, `encoding`, `content`) and adds:
263 - `artifact`: the reference;
264 - `source`: `workspace`, `snapshot` or `session_artifact`;
265 - `current`: whether the workspace still holds these bytes, or `null`
266 when that is not a question for this reference.
267 - `revision` selects an intermediate revision one of the turn's items
268 recorded. The default is the reference's own revision.
269 - A `file` is served from the workspace when it still holds the recorded
270 revision (`current: true`). Otherwise it comes from the turn's post-turn
271 snapshot (`current: false`).
272 - A `tool_output` or `media` reference is read under the session artifact
273 root the writer used. The same confinement, image-manifest and integrity
274 checks apply as for the session route.
275
276 | Status | When |
277 | --- | --- |
278 | 404 | Unknown thread, turn or artifact id, or a `revision` this turn never recorded. |
279 | 409 | A file's recorded revision is in neither the workspace nor the snapshot store (snapshots are pruned after 50 per workspace or 7 days), or a session artifact's bytes no longer hash to the recorded revision. The message names the current revision. |
280 | 410 | The turn deleted the file (restore it with `file-revert` and `restore_snapshot_id`), or a session artifact's bytes were pruned. |
281 | 413 | Content over 16 MiB. |
282 | 403 | A symlink, or a reference that leaves its root. |
283
284 Fleet receipt artifacts keep their own route
285 (`GET /v1/fleet/runs/{run_id}/receipts/{task_id}/evidence`).
286
287 ### Runtime and account identity
288
289 `GET /v1/runtime/info` reports `codewhale_version` plus the full 40-character
290 `codewhale_commit` embedded by the shared CLI/TUI build. A source archive that
291 cannot provide an exact commit reports `unknown`, allowing compatibility
292 clients to fail closed rather than accepting an ambiguous binary pair.
293
294 The same response advertises `capabilities.account_session: true` and
295 `capabilities.turn_operation_idempotency: true`. A client must require the
296 latter before relying on `operation_key`; do not infer support from a 2xx turn
297 response because an older tolerant reader may ignore an unknown request field.
298 `capabilities.turn_operation_lookup: true` separately advertises the read-only
299 operation lookup below; clients must require it before relying on GET-based
300 recovery of a lost turn response.
301 The response also includes a token-free account receipt:
302
303 ```json
304 {
305 "account": {
306 "schema_version": 1,
307 "state": "authenticated",
308 "api_base": "https://api.codewhale.net",
309 "account_id": "acct_...",
310 "session_id": "session_...",
311 "scopes": [],
312 "expires_at": "2026-08-01T20:00:00Z"
313 }
314 }
315 ```
316
317 The Runtime reads this receipt from the exact profile- and API-origin-scoped
318 secure record written by `codewhale account login`; it does not run a second
319 login flow. States are `signed_out`, `authenticated`, `offline_cached`,
320 `expired`, or `revoked`. Scopes are copied only from explicit stored session
321 grants and are never inferred from account identity. Access/refresh tokens,
322 email, provider profile, and provider credentials are never returned.
323 `account_id` and `session_id` are included only for a request authorized with
324 the Runtime token (or an explicitly insecure loopback server); the public
325 bootstrap response remains usable but reports `signed_out`. Signed-out local
326 Work remains supported and never allocates cloud compute implicitly.
327
328 The `--stdio` control transport is newline-delimited JSON-RPC 2.0. Probe it
329 without spending model tokens:
330
331 ```bash
332 printf '%s\n' \
333 '{"jsonrpc":"2.0","id":1,"method":"healthz"}' \
334 '{"jsonrpc":"2.0","id":2,"method":"capabilities"}' \
335 '{"jsonrpc":"2.0","id":3,"method":"shutdown"}' \
336 | codewhale app-server --stdio
337 ```
338
339 `capabilities` returns the advertised method families (`thread/*`, `app/*`,
340 `prompt/*`) and the full method list; `thread/capabilities`,
341 `app/capabilities`, and `prompt/capabilities` scope it per family. The method
342 set is pinned by a drift test in `crates/app-server/src/lib.rs`, so SDK and
343 local integration clients can rely on it not changing silently.
344
345 ### Daemon socket: `codewhale app-server --socket`
346
347 The desktop shell (DESKTOP-APP-BRIEF §2) attaches to a long-lived daemon over
348 a unix domain socket. The wire is the `--stdio` transport verbatim — the same
349 newline-delimited JSON-RPC 2.0 methods, dispatched by the same code — with one
350 handshake in front of it.
351
352 > Note (2026-09-14): the Tauri desktop shell named above is retiring under the
353 > 2026-09-14 product-client transition, and the DESKTOP-APP-BRIEF reference is
354 > a dangling pointer (that brief does not exist in this repo). The GPUI client
355 > in the private `codehwhale-gpui` repo is the successor daemon consumer over
356 > this HTTP runtime API; the socket protocol described here is unchanged.
357
358 **Endpoint.** `--socket-path` if given; else `$CODEWHALE_HOME/run/daemon.sock`
359 when `CODEWHALE_HOME` is set (an explicit home is an isolation boundary); else
360 `$XDG_RUNTIME_DIR/codewhale/daemon.sock`; else
361 `~/Library/Application Support/codewhale/daemon.sock` on macOS or
362 `~/.codewhale/run/daemon.sock` elsewhere. The directory is created `0700`, the
363 socket is `0600`, and every accepted peer must present the daemon's own uid.
364 On start, a socket file nobody answers on is removed; a live one makes the new
365 daemon exit with `a live listener already answers on <path>; refusing to
366 replace it`; a non-socket file at the path is never touched. On Windows `--socket` fails with a typed
367 `UnsupportedPlatform` error naming the reserved pipe `\\.\pipe\codewhale-daemon`
368 — there is no silent TCP fallback. The daemon prints
369 `codewhale daemon: listening on <path>` to stderr once it is accepting.
370
371 **Handshake.** The first request on a connection must be `daemon/attach`
372 (`healthz` is also allowed beforehand, so a shell can probe liveness). Every
373 other method is refused with `-32010 attach_required` until then.
374
375 ```json
376 {"jsonrpc":"2.0","id":1,"method":"daemon/attach","params":{
377 "client":{"name":"codewhale-desktop","version":"1.2.3","pid":4242},
378 "mode":"claim",
379 "expect_daemon_version":"0.9.11"}}
380 ```
381
382 `mode` is `"claim"` (this client spawned the daemon and manages its lifetime)
383 or `"attach"` (default: a guest that found a healthy daemon). A claim while
384 another connection owns the daemon fails with `-32011 daemon_already_claimed`
385 (`data.owner` names the holder) and the client should retry with `attach`.
386 `expect_daemon_version`, when present, must equal the daemon's crate version
387 or the attach fails with `-32013 daemon_version_skew` (the bundle-skew guard).
388 The reply reports the granted `role` (`owner` / `attached`), the daemon's
389 `pid`, `version`, `socket_path`, and `uptime_ms`, the current `owner`, and the
390 live `connections` count. A second `daemon/attach` on an attached connection
391 is `-32014 already_attached`.
392
393 **Capabilities.** On this transport `capabilities.methods` is the pinned stdio
394 set plus `daemon/attach` (second entry, after `healthz`); `transport` reads
395 `unix-socket`. `shutdown` is advertised to every connection because the method
396 exists, but only the owner may call it (below).
397
398 **Ownership.** Only the owner may `shutdown`; a guest's `shutdown` is refused
399 with `-32012 not_daemon_owner` and does not interrupt anyone's turn. When the
400 owner disconnects the slot frees, so a relaunched shell re-claims the daemon it
401 left running. The owner's `shutdown` stops the listener, closes every
402 connection, and removes the socket file. Journal replay from a client's
403 last-seen `seq` is not part of this transport yet.
404
405 ### Interrupting a turn
406
407 `thread/message` streams until the turn reaches a terminal state, which can
408 take minutes. The read loop keeps polling stdin while a turn streams, so a
409 client can send:
410
411 ```json
412 {"jsonrpc":"2.0","id":9,"method":"thread/interrupt","params":{"thread_id":"thr_..."}}
413 ```
414
415 and the runtime is asked to interrupt that turn
416 (`POST /v1/threads/{id}/turns/{turn_id}/interrupt`). The reply carries
417 `interrupted: false` when no turn is streaming for that thread — this is not
418 an error, just nothing to stop. The interrupted `thread/message` then fails
419 with a `turn interrupted` error, and its reply is written before the
420 interrupt's own reply, since the turn owns the writer until it unwinds.
421
422 `shutdown` sent during a live turn also interrupts first: it needs the same
423 bridge that the turn holds, so without that it would wait for the very turn
424 it was meant to stop. Other requests that arrive mid-turn are queued and run
425 in order once the turn finishes.
426
427 ### Running a prompt
428
429 `prompt/request` and `prompt/run` (byte-identical aliases) and the legacy
430 HTTP `POST /prompt` all execute a **real turn** on the runtime, through the
431 same bridge `thread/message` uses. There is no local fallback: nothing else
432 in the app-server can produce model output, so a prompt either runs or fails.
433
434 - `params.prompt` is required and must be non-empty (`-32602` otherwise).
435 - `params.thread_id` is optional. With one, the prompt runs on that thread and
436 its history. Without one, the runtime gets a fresh thread for that single
437 turn; the mapping is dropped when the turn ends, so a one-shot prompt is not
438 addressable by `thread/interrupt`. Use `thread/message` when you need to be
439 able to interrupt.
440 - `params.model` selects the model only when the call is the one that creates
441 the runtime thread; an existing thread keeps the model it was created with.
442 - The response carries what the model actually said: `output` is the
443 concatenated `agent_message` text, `model` is the model the runtime reports
444 for the thread that ran it, and `events` are the real
445 `response_start`/`response_delta`/`response_end` frames. Over stdio the same
446 frames are also streamed to stdout while the turn runs, exactly as for
447 `thread/message`.
448 - If the runtime cannot be reached, the call fails with `-32005`
449 (`runtime_unavailable`) on stdio, or HTTP `503` with
450 `{"error":{"code":"runtime_unavailable", ...}}` on `POST /prompt`. Failures
451 are never shaped like a successful `PromptResponse`.
452
453 `POST /thread` with a `Message` body behaves the same way — it runs the turn
454 and replies `status: "completed"` with the streamed frames in `events` — where
455 it previously replied `accepted` without doing anything.
456
457 ### Thread ids and restarts
458
459 `thread/message`, `thread/request` messages, and HTTP `POST /thread` messages
460 take a thread id from `thread/create` (or `thread/fork`). An id that was never
461 created fails with `-32004` (`thread_not_found`) on stdio, or HTTP `404` on
462 `/thread`, before any runtime thread is started. `/prompt`, `prompt/request`,
463 and `prompt/run` are different: their optional `thread_id` is any key the
464 caller chooses, and a new key starts a new conversation.
465
466 The authenticated canonical owner keeps the full saved conversation graph,
467 its selected branch and the thread/session binding. Compatibility controls use
468 that owner; the old SQLite history remains a protected, read-only import source.
469 Import compares the complete source graph and selected leaf before publishing
470 the canonical alias. A failed alias publication retains the source and the
471 actual canonical result so recovery can report what completed.
472
473 Existing validated legacy goals are imported into the owner goal store. Active
474 goals are paused during import; historical data never starts a provider call.
475 Source goal fields participate in the same protected source comparison.
476
477 Saved-session forks preserve the complete journal, including inactive branches,
478 and copy a validated local session-goal sidecar into the new saved session.
479 An active local goal is copied as paused; the source remains unchanged.
480 That sidecar is separate from the public Runtime thread goal. A native Runtime
481 thread fork does not automatically inherit the public thread goal.
482
483 `thread/create`, `thread/start`, `thread/resume` and `thread/fork` carry a
484 client-generated `operation_key`. Capture one key for each user intent before
485 sending it, retain it after an uncertain response, and reuse it when recovering
486 that same intent. Create carries the key in `metadata.operation_key`; Start,
487 Resume and Fork carry it in `operation_key`. Two intentional forks use different
488 keys. Recovery checks
489 the original operation in the existing owner store; it never creates another
490 thread merely because the response was lost or the source later grew.
491
492 `POST /v1/thread-history/operations/lookup` is read-only. Its closed request
493 contains `version: 1`, `operation_key`, `expected_data_dir`,
494 `expected_execution_scope` and `workspace`; the response is `absent`, `pending`
495 or `committed`. A pending or committed response includes the retained receipt
496 and its exact action/source `association`.
497
498 `POST /v1/thread-history/operations/recover` accepts that lookup request in
499 `operation` and the expected `association`. It explicitly finishes an already
500 prepared target under the same owner, after checking the saved document, full
501 graph, workspace, checkpoint and action/source identity. It does not rebuild
502 the original intent from a source that may already have changed. An unprepared
503 target stays pending; a changed or unverifiable target refuses completion.
504 Repeat recovery with the same key to observe the same committed result.
505
506 The selected workspace comes from the acknowledged owner or the explicitly
507 admitted request. History and historical receipts cannot supply permissions,
508 credentials, endpoints or a different owner. A missing, changed, busy or
509 incompatible bound store is an explicit failure. A missing canonical target
510 does not start an empty replacement conversation.
511
512 `codewhale thread resume` and `codewhale thread fork` perform the durable
513 owner control and print its committed thread, session and operation receipt.
514 They do not launch an interactive frontend.
515
516 Global `--workspace` (also `--cd`), `--profile` and `--config` select an explicit
517 control scope. Relative paths are captured before attachment; the client
518 authenticates the owner, then admits that scope against the same owner receipt
519 and its captured worker setting. An incompatible profile or config, missing
520 scope facts, or a changed owner fails explicitly. Without these options, the
521 acknowledged owner supplies the workspace. Thread listing remains store-wide.
522
523 For a fresh `thread resume` or `thread fork`, global `--provider`, `--model`,
524 `--approval-policy` and `--sandbox-mode` supply proposals to the existing owner
525 decoder and permission checks. The corresponding `--set` keys are `provider`,
526 `model`, `default_text_model`, `approval_policy` and `sandbox_mode`. Credentials
527 and endpoints stay with the owner: `--api-key`, `--base-url` and other per-run
528 settings are refused for these controls. Configure and authenticate the owning
529 Runtime before using them. With a retained `--operation-key`, newly supplied
530 model, provider, policy or sandbox proposals are refused; recovery observes
531 the original admitted intent.
532
533 Interactive `codewhale resume` and `codewhale fork` use the same canonical
534 history operation. An inactive local owner is shut down and joined before the
535 existing TUI acquires the saved-session lease and store. A live owner or an
536 uncertain handoff refuses attachment. `--operation-key <KEY>` recovers the
537 original outcome, including completing a verified prepared target; an
538 unprepared or unverifiable result preserves uncertainty.
539
540 ### Changing config
541
542 `app/config/set` and `app/config/unset` write the change to the config file
543 before replying: the `--config` path, or the default `config.toml` used by the
544 canonical owner when no `--config` is given. They change user settings that
545 outlive the app-server, not just this session. The change is applied to the
546 file as it is on disk, so edits saved by other processes are kept. If the file
547 cannot be read, parsed, or written, the reply is `ok: false` and nothing
548 changes; a file that does not parse is not rewritten, so fix it by hand (or
549 `app/config/reload` after fixing it). Over HTTP `/app`, a rejected key or value
550 is `400` and a read or write failure is `500`.
551
552 ### Answering a clarification question
553
554 When a headless turn calls `request_user_input`, the runtime emits a
555 `user_input.required` event carrying a `request_id`. Reply on the runtime API:
556
557 ```
558 POST /v1/user-input/{thread_id}/{request_id}
559 ```
560
561 The app-server control transport cannot accept that reply.
562 `app/request` with `SubmitUserInput` returns `ok: false` and
563 `error: "user_input_reply_unsupported"`. This is a property of the transport,
564 not an omission: while a turn is streaming, the stdio loop executes only
565 `thread/interrupt` and queues everything else, so an answer sent there would
566 wait on the very turn that is waiting for it.
567
568 ## SDK contract
569
570 The app-server exists so an external SDK can answer — without scraping TUI
571 output — *what route ran, which provider/model/reasoning/permission profile was
572 effective, what events happened, how many tokens were used, and how the run
573 finished.* The durable Thread/Turn/Item data model already carries most of
574 this; the table maps each integration need to where a local client reads it.
575
576 | Integration need | Where it comes from | Status |
577 |---|---|---|
578 | Route / effective model / billing surface | `TurnRecord` + thread `model`; per-run `--provider`/`--model` overrides | available |
579 | Permission / sandbox / approval profile | thread `auto_approve`, sandbox + approval policy; `TurnRecord.permission_posture` + `TurnRecord.mode` for how *that* run was governed (the thread's own `mode` may have been switched since) | available |
580 | Run / thread / turn IDs | `thread_id`, `turn_id`, SSE event envelope | available |
581 | Event stream | `GET /v1/threads/{id}/events` (replay + live SSE) | available |
582 | Turn status / terminal classification | `TurnRecord.status` + error summary | available |
583 | Token usage | `TurnRecord.usage`; aggregate via `GET /v1/usage` | available |
584 | Action receipt (files, commands, web/MCP calls, agents, approvals and who decided, failures) | `GET /v1/threads/{id}/receipt`, `GET /v1/threads/{id}/turns/{turn_id}/receipt` | available ([RECEIPTS.md](RECEIPTS.md)) |
585
586 For one-shot/headless automation, prefer `codewhale exec` with explicit
587 `--provider <id> --model <id>` so a failure identifies the exact provider/model
588 pair. Use `app-server` when a local integration needs to start, resume, steer,
589 or interrupt turns, list models/capabilities, follow the event stream, or read
590 usage. Both paths share the same runtime, so route-effective model resolution
591 and the event vocabulary match.
592
593 ### Release smoke
594
595 `scripts/release/app-server-smoke.sh` is the committed pre-release check:
596
597 ```bash
598 scripts/release/app-server-smoke.sh # stdio health/capabilities probe (no tokens)
599 scripts/release/app-server-smoke.sh --matrix # + print the configured provider/model matrix
600 scripts/release/app-server-smoke.sh --matrix --real # + exec a cheap sentinel per provider
601 ```
602
603 The stdio probe runs against a throwaway config, so it never reads real keys.
604 The matrix discovers configured providers from `codewhale auth list`, skips
605 unconfigured providers, and maps a provider to a cheap sentinel model only when
606 it has a built-in cheap default. That built-in set is deliberately conservative
607 (currently `deepseek`, `zai`, `moonshot`, and `openai`); every other provider —
608 including `arcee`, `openrouter`, `xiaomi-mimo`, and `openai-codex` — is left
609 unmapped on purpose and must be given a model per run via `SMOKE_MODEL_<SLUG>`
610 rather than a guessed default (#3205). Any configured-but-unmapped provider
611 fails loudly in `--real` mode. `auth list` reports presence flags only and exec
612 output is passed through a redactor, so secrets are never printed. The parser is
613 covered by `scripts/release/app-server-smoke.test.sh` against a fake `codewhale`
614 binary.
615
616 ## ACP stdio adapter: `codewhale serve --acp`
617
618 ACP JSON-RPC over newline-delimited stdio is a transport over the existing
619 RuntimeThreadManager and Engine. It has no provider/tool round loop, executable
620 registry, prompt composer, or conversation writer of its own. The server loads
621 the selected config/profile and plugin discovery once; each prompt uses a
622 canonical thread, Core turn, event timeline, approval waiter, and full Engine
623 session snapshot.
624
625 The editor surface supports `initialize`, `session/new`, `session/list`,
626 `session/load` (including durable ID prefixes), `session/prompt`, `session/cancel`,
627 model discovery/selection, and the advertised mode/model config options.
628 `session/new` creates a durable bare UUID and empty checkpoint without calling a
629 provider. At most 64 idle transport bindings are retained; eviction keeps the
630 durable conversation. Resume deduplicates against the existing thread binding.
631 Full histories, tool pairs, signatures, media and partial-effect receipts are
632 saved under the same checkpoint guard and session write lease used by HTTP.
633 ACP display text may be shortened; the saved history is the full Core snapshot.
634
635 The trusted local ACP profile narrows Core to file/search/git/patch and admitted
636 foreground shell tools. Shell requires both editor terminal support and operator
637 `allow_shell`; a requested external sandbox that is unavailable removes shell.
638 MCP, dynamic tools, task/PTY/background shell, interpreter, subagent and RLM
639 lifecycles are unavailable on this transport. Built-in overrides remove the
640 whole compatibility alias family. Final dispatch rechecks the profile, so a
641 fabricated alias, hook rewrite or omitted catalog entry cannot bypass it. Plan
642 remains read-only under Full Access. Full Access and the ordinary approval
643 posture are server-owned read-only options; the editor cannot widen them.
644
645 Tool updates begin `pending`. `in_progress` means Core reached final dispatch;
646 `completed`/`failed` and typed image blocks come from the actual Core result.
647 `session/request_permission` correlates a private JSON-RPC ID to the same
648 Runtime-minted pending approval and Core execution ID. Only an exact live
649 `allow-once` response can release that waiter; wrong IDs are ignored, invalid
650 options deny, and cancellation withdraws it. ACP grants neither remembered
651 permissions nor Native capabilities. Core typed rules, strict hooks, repo law,
652 Headless Auto-Review and hard floors remain in force, including the absence of a
653 workspace-write carve-out; a guardian-only decision refuses on this host.
654
655 Replay uses the existing bounded event reader and sequence deduplication. A
656 replay gap, closed owner or terminal store fault is an explicit error, never a
657 re-execution. Input is serviced between replayed events; each transport write has a 30-second
658 deadline. Cancel, EOF and writer
659 failure interrupt only this transport's claimed Core turn; settlement uses its
660 actual terminal receipt and preserves any completed effects. An unconfirmed
661 cancellation does not fabricate success. `stopReason` is `end_turn`, `cancelled`,
662 or typed `max_turn_requests`; Core failure remains an error. Prompts use at most
663 50 model steps (a lower configured limit still wins) and Core's bounded final
664 report response; this profile does not dispatch autonomous goal continuations.
665
666 ACP presently claims its own exclusive canonical Runtime owner. If another
667 process holds that store, startup refuses. A saved conversation bound to another
668 Runtime store also refuses; authenticated cross-process owner attachment remains
669 unqualified and ACP does not copy a live conversation into a random store. This
670 does not expose every `/v1/*` steering, job or control method to an ACP editor.
671 Use `codewhale app-server --http` for that full runtime API.
672
673
674 ## Capability endpoint: `codewhale doctor --json`
675
676 Returns a JSON object describing the current installation's readiness state.
677 Suitable for health-check polling from a macOS workbench. This command is
678 strictly structural and offline: it does not load workspace credential
679 `.env` files, inspect credential environment values, open secret/OAuth files,
680 probe an OS keyring, contact providers, or start MCP processes.
681
682 ```bash
683 codewhale doctor --json
684 ```
685
686 ### Response schema (key fields)
687
688 | Field | Type | Description |
689 |---|---|---|
690 | `version` | string | Installed version (e.g. `"0.8.9"`) |
691 | `config_path` | string | Resolved config file path |
692 | `config_present` | bool | Whether the config file exists |
693 | `paths` | object | Canonical config, settings, state, sessions, logs, automations, and secrets paths |
694 | `secret_backend` | object | Metadata-only file-store shape, or literal `unknown` / `not_probed` for system and unsupported backends |
695 | `workspace` | string | Default workspace directory |
696 | `legacy_state.primary_root` | string | Primary Codewhale state root inspected for known state paths |
697 | `legacy_state.legacy_root` | string | Legacy `.deepseek` state root inspected for known state paths |
698 | `legacy_state.needs_attention` | bool | Whether known `~/.deepseek` state paths need review or the read-only session recovery diagnostic found missing destination filenames / could not complete |
699 | `legacy_state.legacy_only_count` | number | Count of known state paths present only under the legacy root |
700 | `legacy_state.dual_present_count` | number | Count of known state paths present under both primary and legacy roots |
701 | `legacy_state.entries` | array | Per-path migration status: `{name, primary_present, legacy_present, status}` |
702 | `legacy_state.session_recovery.status` | string | `isolated`, `no_legacy_sessions`, `migration_pending`, `migration_incomplete`, `migration_complete`, or `scan_failed` |
703 | `legacy_state.session_recovery.read_only` | bool | Always true; doctor never invokes session migration or modifies either session directory |
704 | `legacy_state.session_recovery.chat_contents_read` | bool | Always false; comparison is based only on top-level `.json` filenames and filesystem metadata |
705 | `legacy_state.session_recovery.checkpoint_internals_scanned` | bool | Always false; `sessions/checkpoints/` and all other directories are skipped |
706 | `legacy_state.session_recovery.recoverable_files` | array | Bounded sample of up to 100 missing destination filenames with source and destination paths; no chat payloads |
707 | `legacy_state.session_recovery.recoverable_file_count` | number | Total missing destination filename count, including entries beyond the bounded sample |
708 | `legacy_state.session_recovery.recoverable_files_truncated` | bool | Whether more than 100 recoverable filenames were found |
709 | `legacy_state.session_recovery.recovery_command` | string or null | `codewhale sessions` when additive automatic recovery is available; null for isolated, complete, empty, or failed scans |
710 | `api_key.source` | string | Structural source state: `config_declared`, `env_declared`, `external_auth_declared`, `secret_store_unprobed`, `secret_store_unavailable`, `oauth_unprobed`, `external_consent`, `none`, `local_runtime`, or `unknown`; declarations are not availability proof |
711 | `api_key.availability` | string | Literal `present`, `not_required`, `not_probed`, `unavailable`, or `unknown`; only `present` and `not_required` certify structural Setup/Fleet credential readiness |
712 | `base_url` | string | Provider URL authority only (`scheme://host[:explicit-port]`); userinfo, path, query, and fragment are omitted |
713 | `default_text_model` | string | Default model |
714 | `memory.enabled` | bool | Whether the memory feature is on |
715 | `memory.path` | string | Path to memory file |
716 | `memory.file_present` | bool | Whether memory file exists |
717 | `mcp.config_path` | string | MCP config file path |
718 | `mcp.present` | bool | Whether MCP config exists |
719 | `mcp.probe_scope` | string | `configuration`; doctor does not start MCP servers |
720 | `mcp.live_health_checked` | bool | Always false for doctor JSON |
721 | `mcp.servers` | array | Per-server structural result and counts plus separate `checks`; URL userinfo/path/query/fragment and command argv, environment, header, and token values are never emitted, and all live stages are `not_checked` |
722 | `skills.selected` | string | Resolved skills directory |
723 | `skills.global.path` / `.present` / `.count` | — | Codewhale global skills dir (`~/.codewhale/skills`, with legacy `~/.deepseek/skills` support) |
724 | `skills.agents.path` / `.present` / `.count` | — | Workspace `.agents/skills/` dir |
725 | `skills.agents_global.path` / `.present` / `.count` | — | agentskills.io global skills dir (`~/.agents/skills`) |
726 | `skills.local.path` / `.present` / `.count` | — | `skills/` dir |
727 | `skills.opencode.path` / `.present` / `.count` | — | `.opencode/skills/` dir |
728 | `skills.claude.path` / `.present` / `.count` | — | `.claude/skills/` dir |
729 | `tools.path` / `.present` / `.count` | — | Global tools directory |
730 | `plugins.path` / `.present` / `.count` | — | Global plugins directory |
731 | `sandbox.available` | bool | Whether sandbox is supported on this OS |
732 | `sandbox.kind` | string or null | Sandbox kind (e.g. `"macos_seatbelt"`) |
733 | `storage.spillover.path` / `.present` / `.count` | — | Tool output spillover dir |
734 | `storage.stash.path` / `.present` / `.count` | — | Composer stash |
735
736 ### Example
737
738 ```json
739 {
740 "version": "0.8.9",
741 "config_path": "/Users/you/.codewhale/config.toml",
742 "config_present": true,
743 "workspace": "/Users/you/projects/codewhale-tui",
744 "api_key": {
745 "source": "secret_store_unprobed",
746 "availability": "not_probed"
747 },
748 "base_url": "https://api.deepseek.com",
749 "default_text_model": "deepseek-v4-pro",
750 "memory": {
751 "enabled": false,
752 "path": "/Users/you/.codewhale/memory.md",
753 "file_present": true
754 },
755 "mcp": {
756 "config_path": "/Users/you/.codewhale/mcp.json",
757 "present": true,
758 "servers": [
759 {"name": "filesystem", "enabled": true, "transport": "stdio", "args_count": 2, "env_count": 0, "status": "ok"}
760 ]
761 },
762 "sandbox": {
763 "available": true,
764 "kind": "macos_seatbelt"
765 }
766 }
767 ```
768
769 ## HTTP/SSE runtime API: `codewhale app-server --http`
770
771 ```bash
772 codewhale app-server --http [--host 127.0.0.1] [--port 7878] [--workers 2] [--auth-token TOKEN] [--insecure-no-auth]
773 codewhale app-server --mobile [--host 127.0.0.1] [--port 7878] [--auth-token TOKEN]
774 codewhale app-server --mobile --host ::1 [--port 7878] [--insecure-no-auth]
775 codewhale web [--port 7878]
776
777 # Compatibility aliases — identical server, serve flag names:
778 codewhale serve --http [...] [--insecure]
779 codewhale serve --mobile [...] [--insecure]
780 ```
781
782 Defaults: host `127.0.0.1`, port `7878`, 2 workers (clamped 1–8).
783
784 The server binds to `localhost` by default. Configuration is via CLI flags —
785 there is no `[app_server]` config section.
786
787 `/v1/*` routes require a bearer token unless `codewhale app-server` is started
788 with `--insecure-no-auth` on a loopback bind such as `127.0.0.1`. Mobile mode
789 is loopback-only: non-loopback hosts are rejected until Runtime has a TLS or
790 verified-overlay transport boundary. The `codewhale serve` compatibility aliases
791 use `--insecure` for the same loopback escape hatch.
792 Pass `--auth-token TOKEN` or set `CODEWHALE_RUNTIME_TOKEN=TOKEN` before starting
793 the server; `DEEPSEEK_RUNTIME_TOKEN` remains a compatibility alias. If neither
794 is set, the process generates a Runtime token for that process and does **not**
795 print it. `/health`, `/v1/runtime/info`, and an enabled static client shell
796 remain public; Runtime mutations and thread data stay behind `/v1/*`
797 authentication. `/mobile` returns 404 when mobile mode is disabled and serves
798 the unchanged static shell when it is enabled.
799
800 Authenticated clients can provide the token as `Authorization: Bearer TOKEN`,
801 `X-Codewhale-Runtime-Token: TOKEN`, the legacy
802 `X-DeepSeek-Runtime-Token: TOKEN`. Query-string and raw Runtime-token cookie
803 authentication are not supported.
804
805 ### Local browser client
806
807 `codewhale web` starts the canonical Runtime API on `127.0.0.1`, serves
808 dependency-free assets embedded in the binary, prints a single-use launch URL,
809 and asks the operating system to open that URL in the default browser. If the
810 browser does not open, the printed URL remains usable for ten minutes. The
811 command cannot bind to a non-loopback host and cannot run with Runtime auth
812 disabled.
813
814 The browser-launch URL contains a random, short-lived, one-time bootstrap
815 capability, never the Runtime token. A loopback request exchanges that
816 capability for a
817 `codewhale_web_session=…; HttpOnly; SameSite=Strict; Path=/` cookie backed by a
818 single process-local server session that expires 12 hours after the server
819 process starts, consumes the capability immediately, and redirects to `/`.
820 Reused, expired, malformed, or
821 non-loopback bootstrap attempts fail closed. The Runtime bearer token is not
822 placed in rendered HTML, browser storage, logs, URL queries/fragments, or
823 browser-launch arguments. The one-time bootstrap capability is printed in the
824 local terminal and transits the OS browser launcher's argument list. A same-user
825 process could race the browser to the exchange, which is why the capability is
826 single-use, loopback-only, and expires after ten minutes — and why a same-user
827 attacker has strictly easier local avenues than this race.
828 Web fetches require the session cookie plus an origin-scoped request proof;
829 streams use a fresh single-use ticket. The initial redirect carries the proof
830 in a fragment, which the client removes and saves in origin-scoped
831 `sessionStorage`. On reload or in a second tab, an authenticated `GET /` also
832 embeds the proof in a meta tag when `Sec-Fetch-Site` is `same-origin` or `none`
833 (direct navigation). The page uses `no-store`, disallows framing, and grants no
834 cross-origin read access. This lets a new tab recover without reusing the
835 bootstrap URL, including when storage is unavailable. Clients without Fetch
836 Metadata can only reuse the fragment or their existing stored proof; recovery
837 does not extend the server session or replace an expired cookie.
838 Cross-origin Fetch Metadata or a mismatched Origin is rejected on web API
839 requests. Explicit bearer and Runtime-token header clients keep their existing
840 behavior. Transient stream-ticket failures retry with capped backoff; HTTP
841 401/403 stops ticket retries until a fresh session is opened.
842
843 The embedded client provides a responsive thread/search rail, Runtime-owned
844 session facts, transcript and tool receipts, and a bottom composer. It can
845 create, select, rename, and archive threads; choose a provider and model for a
846 new thread without changing Runtime defaults; start or steer turns; interrupt
847 work; resolve approvals; and answer Runtime user-input requests. Selection
848 loads `GET /v1/threads/{id}` first, then opens the replayable event stream with
849 `since_seq=latest_seq`; reconnection advances from the newest accepted sequence
850 and drops duplicates or events from a stale selection. The thread detail
851 snapshot includes `pending_approvals`, `pending_user_inputs`, and
852 `pending_dynamic_tool_calls`; clients must hydrate those fields before
853 subscribing so a reload cannot strand work whose request event is at or before
854 `latest_seq`. Resolution is also published as `approval.decided`,
855 `user_input.answered`, `user_input.canceled`, `tool_call.resolved`,
856 `tool_call.canceled`, or `tool_call.timeout` for already-connected clients.
857
858 An existing thread's model, mode, permission posture, workspace, and branch are
859 display-only in this client. Files/Changes, PTY/terminal, preview, artifacts,
860 provider login or global-default switching, Fleet creation, and
861 undo/retry/restore controls are not built into this page. The native desktop
862 client covers them through the workspace-file, turn-artifact, terminal and
863 workspace-restore routes documented here.
864
865 ### Mobile control page
866
867 `codewhale serve --mobile` starts the same HTTP/SSE runtime API and serves a
868 phone-friendly control page at `/mobile`. It binds only to loopback
869 (`127.0.0.1` or `::1`); a non-loopback host is rejected because this Runtime
870 surface does not yet provide TLS or a verified overlay transport. The static
871 HTML page contains no Runtime bearer and is not itself token-gated. When
872 Runtime auth is enabled, the CLI prints a short-lived, single-use loopback
873 bootstrap URL. That capability creates a 30-minute, process-local
874 `Max-Age=1800; HttpOnly; SameSite=Strict` mobile session cookie plus origin-scoped browser
875 proofs. A sibling port that receives the host-scoped cookie cannot use it by
876 itself. The page can also exchange an explicitly entered bearer once, then
877 clears it rather than storing it in browser storage or a cookie. EventSource
878 connections use separate short-lived, single-use stream tickets.
879
880 The mobile page can list/create threads, send prompts, follow live SSE events,
881 steer or interrupt an active turn, and resolve normal tool approvals through
882 `POST /v1/approvals/{approval_id}`. It is a local-only convenience surface;
883 do not expose it directly to another device or public network until Runtime has
884 a TLS or verified transport boundary.
885
886 ### Endpoints
887
888 **Health**
889 - `GET /health`
890
891 **Sessions** (durable session manager)
892 - `GET /v1/sessions?limit=50&search=<fuzzy>&include_archived=false&archived_only=false&workspace=<path>&sort=recent|name|size`
893 - `GET /v1/sessions/summary?…` (same query params; projected row shape)
894 - `GET /v1/sessions/{id}` (add `?peek=true&entries=12` for a bounded, redacted
895 read-only peek instead of the full transcript). The full response carries
896 `turn_outcomes` when a turn ended `Failed`: `{ status, error, ended_at,
897 after_message_count }` per failure, oldest first, bounded to 64, with the
898 error text the transcript showed and secrets redacted
899 - `PATCH /v1/sessions/{id}` (`{ "title"?: string, "archived"?: bool }`)
900 - `DELETE /v1/sessions/{id}`
901 - `POST /v1/sessions/{id}/resume-thread` returns the open thread that already
902 holds the whole saved session (`200`), or seeds a new thread from it (`201`)
903 when none does, including when the session grew after that thread opened it.
904 - `GET /v1/sessions/{id}/artifacts` and `GET /v1/sessions/{id}/artifacts/{artifact_id}?offset=&limit=`
905 (see workspace files and session artifacts above; runtime-turn spills are
906 read through `GET /v1/threads/{id}/turns/{turn_id}/artifacts/{artifact_id}`)
907 - `POST /v1/sessions` (`{ "thread_id": string, "title"?: string }`) exports a
908 thread as a saved session. It is idempotent: it writes only the document
909 whose id is derived from the thread, creating it the first time (`201`) and
910 updating it after (`200`), so a retry never makes a duplicate. A session the
911 thread was resumed from is left unchanged.
912 - `PUT /v1/sessions` (`{ "thread_id"?: string, "session_id"?: string }`) saves
913 a thread's live conversation. Naming a `session_id` that another thread is
914 bound to returns `409 Conflict`.
915 - `GET /v1/sessions/repair` returns the last session-store repair summary, or
916 `null` if none has run
917
918 Sessions and threads answer the same `include_archived` / `archived_only` pair
919 with the same meaning, and `search` is the same fuzzy match (title, id,
920 workspace — substring, then subsequence) the TUI session picker and the workbar
921 Sessions list use. All three surfaces run one projection
922 (`crates/tui/src/session_projection.rs`), so a listing cannot differ between
923 the terminal and the dashboard.
924
925 `GET /v1/sessions/summary` returns rows that are field-compatible with
926 `GET /v1/threads/summary` — `id`, `title`, `preview`, `model`, `mode`,
927 `workspace`, `archived`, `updated_at` — plus `message_count`, `total_tokens`,
928 `created_at`, `parent_session_id`, and `is_current`. One caveat stated plainly:
929 `preview` is the session's recorded **title**, not its last message. Session
930 metadata does not store a last message, and reading every transcript to
931 synthesise one would make a list view an unbounded read. Full transcript
932 preview lives in the TUI session picker, which reads one selected session.
933
934 `PATCH /v1/sessions/{id}` renames and/or archives a saved session and returns a
935 lifecycle receipt shaped like the thread patch receipt:
936
937 ```json
938 {
939 "session": { "id": "…", "title": "Renamed", "archived": true, "…": "…" },
940 "changes": { "title": "Renamed", "archived": true }
941 }
942 ```
943
944 `changes` lists only what actually moved, so a no-op patch is distinguishable
945 from an applied one. Archiving is durable and reversible: an archived session
946 stays on disk and stays loadable, disappears from default listings, and is
947 never chosen by `--continue` or by auto-resume. The route is the same writer
948 the TUI picker (`e`) and `/sessions archive <id>` use — there is no second
949 archive notion.
950
951 While a session is open in an interactive Codewhale process, that process holds
952 the authoritative copy in memory and rewrites the whole document on its next
953 autosave. `PATCH`, `PUT` and `DELETE` therefore fail closed on it with
954 `409 Conflict` rather than writing something that would be silently reverted.
955 Change it in the terminal instead. The open process holds a lock on the
956 session (`sessions/.late-usage/<id>.live`), so this holds whether the request
957 reaches the API inside that process or a separate `codewhale serve`.
958
959 The session store is repaired in the background at each launch and each
960 `codewhale serve` start. The repair gives a "Recovered:" session to each thread
961 in a Runtime store that no session is bound to. It unbinds threads whose session
962 document is gone; they then load from their own turns. It moves unreadable
963 documents, empty unbound stores, and old artifact directories that no session
964 names to `sessions/.set-aside/<run>/`, and writes a `MANIFEST.jsonl` there.
965 Nothing is deleted. `GET /v1/sessions/repair` and `codewhale doctor` report the
966 last run; `codewhale doctor --repair-sessions [--dry-run]` runs one on demand.
967
968 `GET /v1/sessions/{id}?peek=true` returns a bounded, redacted, read-only view
969 instead of the transcript: at most 12 entries of at most 400 characters each
970 (`&entries=N` lowers the budget, never raises it past the cap), tool calls and
971 results summarised to a name and a size rather than inlined, and
972 credential-shaped substrings masked. `omitted_before` reports how many earlier
973 messages were dropped. The payload carries `"live": false` and deliberately has
974 no turn status, `running`, or `active` field — a saved session is a recording,
975 and live state comes only from a resumed thread's SSE stream.
976
977 **Threads** (durable runtime data model)
978 - `GET /v1/threads?limit=50&include_archived=false&archived_only=false`
979 - `GET /v1/threads/summary?limit=50&search=<optional>&include_archived=false&archived_only=false`
980 - `GET /v1/threads/running`
981 - `GET /v1/threads/{id}/notices`
982 - `DELETE /v1/threads/{id}/notices/{notice_id}`
983 - `POST /v1/threads`
984 - `GET /v1/threads/{id}`
985 - `PATCH /v1/threads/{id}` (see body shape below)
986 - `POST /v1/threads/{id}/resume`
987 - `POST /v1/threads/{id}/fork`
988 - `GET /v1/threads/{id}/receipt` — what the thread did, one entry per action
989 (read-only; shape in [RECEIPTS.md](RECEIPTS.md))
990 - `GET /v1/threads/{id}/turns/{turn_id}/receipt` — the same, for one turn;
991 `404` for an unknown thread or a turn that is not this thread's
992
993 `POST /v1/threads` accepts optional execution defaults in addition to the
994 provider, model, workspace, and permission fields:
995
996 ```json
997 {
998 "model_provider": "openai-codex",
999 "model": "gpt-5.6",
1000 "reasoning_effort": "high",
1001 "allowed_tools": ["read_file", "search"]
1002 }
1003 ```
1004
1005 `reasoning_effort` uses the canonical Runtime vocabulary (`auto`, `off`,
1006 `low`, `medium`, `high`, `xhigh`, `ultra`, or `max`; documented compatibility
1007 aliases are accepted and persisted canonically). `allowed_tools` is a
1008 model-visible allowlist. Omitting it keeps the normal configured catalog; an
1009 explicit empty array (`"allowed_tools": []`) exposes no tools to the model.
1010 Both fields are additive: older thread records and clients that omit them
1011 retain their previous behavior.
1012
1013 `GET /v1/threads/summary` is the read-only summary surface used by the VS Code
1014 Agent View. `search` matches thread `id`, `title`, and `model` (and, when the
1015 title is unset, the latest turn's input summary — the displayed title). It
1016 does not scan turn or item bodies: `preview` is filled only after a match, so
1017 a dashboard keystroke is not a whole-store read per thread. Each item includes
1018 `id`, `title`, `preview`, `model`, `mode`, `archived`, `updated_at`,
1019 `latest_turn_id`, `latest_turn_status`, plus workspace metadata:
1020
1021 ```json
1022 {
1023 "id": "thread_...",
1024 "title": "Implement MCP status count",
1025 "preview": "The TUI footer should count project MCP servers...",
1026 "model": "deepseek-v4-pro",
1027 "mode": "agent",
1028 "branch": "feature/runtime-api",
1029 "head": "abc1234",
1030 "dirty": false,
1031 "workspace": "/Users/you/projects/codewhale",
1032 "archived": false,
1033 "updated_at": "2026-06-06T05:43:00Z",
1034 "latest_turn_id": "turn_...",
1035 "latest_turn_status": "completed"
1036 }
1037 ```
1038
1039 `branch` is resolved from the thread workspace at request time and may be
1040 `null` when the workspace is not a Git repository or the branch cannot be read.
1041 `head` is the current short Git commit for that workspace when available.
1042 `dirty` is true when the workspace has staged, unstaged, or untracked changes.
1043 `workspace` is included so editor clients can show when an agent lane is working
1044 outside the current VS Code folder.
1045
1046 Thread forks are sibling runtime threads, not an in-place tree projection.
1047 `thread.forked` events include `source_thread_id`; internal backtrack-aware
1048 forks may also include `backtrack_depth_from_tail` and `dropped_turn_id`, and
1049 a fork anchored to a named turn (`/fork-at-turn`) reports them with the depth
1050 resolved from that turn, and names the first user turn it dropped (not the
1051 anchor, which a named-turn fork keeps, and not a prompt-less turn such as a
1052 manual compaction that sits between them).
1053 Thread list and summary responses remain flat in v0.8.40, so clients that need
1054 a graph should reconstruct it from events instead of assuming list order is a
1055 complete tree.
1056
1057 `GET /v1/threads/running` is the running-work accounting surface
1058 (#6180): threads with at least one queued or in-progress turn, each with
1059 `thread_id`, `model`, `title`, and `active_turns` (`turn_id` + `status`).
1060 Background-capable clients use it for quit/background decisions — one call,
1061 no inference from latest-turn status. Archive state is ignored (archiving
1062 has no quiescence gate); an empty array means no owned work is live.
1063
1064 `GET /v1/threads/{id}/notices` is the per-thread active-notice surface
1065 (#6180): the TUI-visible conditions a watch-only client must surface —
1066 `subagent-terminal` (a child settled), `elevation-needed` (a tool call is
1067 blocked on elevation), `model-notify` (the model asked the user to come
1068 back) — each with `turn_id` and a `subject` id for targeting. Notices are
1069 in-memory session state, bounded to 32 per thread (oldest evicted), and
1070 never persisted. Clearing: elevation auto-clears when its tool call
1071 completes; terminal/notify clear on `DELETE .../notices/{notice_id}`
1072 (204, unknown ids 404). Unknown threads 404 on both endpoints.
1073
1074 `archived_only=true` returns archived threads only (mutually overrides
1075 `include_archived`). Default behavior is unchanged: `include_archived=false`
1076 and `archived_only=false` returns active threads. Added in v0.8.10 (#563).
1077
1078 `PATCH /v1/threads/{id}` body — every field is optional, missing means
1079 "no change". At least one field must be present. `title` and `system_prompt`
1080 accept an empty string to clear a previously-set value. Added in v0.8.10 (#562):
1081
1082 ```json
1083 {
1084 "archived": true,
1085 "allow_shell": false,
1086 "trust_mode": false,
1087 "auto_approve": false,
1088 "model": "deepseek-v4-pro",
1089 "mode": "agent",
1090 "title": "User-set thread title",
1091 "system_prompt": "You are a useful assistant.",
1092 "model_provider": "custom",
1093 "model_provider_id": "lm-studio"
1094 }
1095 ```
1096
1097 `model_provider` switches the provider the thread's future turns use. It
1098 takes a built-in kind (`deepseek`, `xai`, ...) or a configured route name, as
1099 `/provider` does. `model_provider_id` names one exact `[providers.<id>]` table
1100 and wins over a route name. The target route is resolved and its client
1101 preflighted before anything is saved, so an unknown or credential-less
1102 provider is refused and nothing changes. Without `model`, the thread takes the
1103 new provider's default model; an `auto` thread stays `auto`. The loaded
1104 engine and conversation history are kept, and the next turn installs the new
1105 route.
1106
1107 **Turns** (within a thread)
1108 - `POST /v1/threads/{id}/turns`
1109 - `POST /v1/threads/{id}/turns/{turn_id}/steer` - inject guidance into the running turn. The response is a receipt for what actually happened, not for what was attempted; see [Steer delivery](#steer-delivery).
1110 - `POST /v1/threads/{id}/turns/{turn_id}/interrupt`
1111 - `GET /v1/threads/{id}/turns/{turn_id}/artifacts` - what the turn produced: typed references plus the workspace-delta state. See [Turn artifacts](#turn-artifacts).
1112 - `GET /v1/threads/{id}/turns/{turn_id}/artifacts/{artifact_id}?offset=&limit=&revision=` - read one reference from the workspace, the post-turn snapshot, or the session artifact directory.
1113 - `POST /v1/threads/{id}/compact` (manual compaction)
1114 - `POST /v1/threads/{id}/undo` - fork the thread with the last N turns removed (`{"depth": N}`, default 0 = last turn only); returns the forked thread plus `original_user_text` so a GUI can pre-populate the input box
1115 - `POST /v1/threads/{id}/fork-at-turn` - fork at one named user turn (`{"turn_id": "turn_…"}`, as `GET /v1/threads/{id}` reports it). The fork *keeps* that turn and every turn before it, and drops the turns after it; naming the last turn therefore keeps the whole conversation. The receipt is `/undo`'s (`thread`, `original_user_text`, `original_user_images`), carrying the *first dropped* user turn's prompt — what was asked next, even when a prompt-less turn such as a manual `/compact` sits between — so a client can put it back in the composer for editing. The source thread, its session document and the workspace are untouched, and there is no file rollback: a fork is a sibling conversation, and rewinding the workspace would rewind the branch left behind with it. Clients should name the turn instead of computing a `depth` — the transcript they render and the turn list this cuts are not the same list (steers, image-only prompts and injected handoffs each sit on one side only), and a client-side count that is off by one forks the wrong prefix while answering `201`. `400` when the turn is not a user turn of that thread.
1116 - `POST /v1/threads/{id}/patch-undo` - rolls back the files the dropped turns changed, then the same fork (`{"depth": N}`); returns `patch_result` (`files_restored`, `summary`, `snapshot_label`) alongside the forked thread. See [Workspace restore endpoints](#workspace-restore-endpoints) for ownership, the trust, admission and abort rules, and the `error.code` values of a refusal.
1117 - `POST /v1/threads/{id}/file-revert` - restore exactly one file from one restore point the thread owns (`{"path", "snapshot_id", "expected_hash"}`); never forks the conversation. See [Workspace restore endpoints](#workspace-restore-endpoints).
1118 - `POST /v1/threads/{id}/retry` - fork with the last N turns removed and immediately start a new turn (`{"depth": N, "prompt": "..."}`; `prompt` overrides the original user text, which is re-used when omitted)
1119
1120 `POST /v1/threads/{id}/turns` accepts the same optional
1121 `reasoning_effort` and `allowed_tools` fields as per-turn overrides:
1122
1123 ```json
1124 {
1125 "prompt": "Review this change without running tools.",
1126 "operation_key": "cwc-request-01J7Y6Q9W4",
1127 "reasoning_effort": "max",
1128 "allowed_tools": []
1129 }
1130 ```
1131
1132 The same `model_provider` / `model_provider_id` fields on a turn route that
1133 one turn through another provider. The saved thread keeps its provider.
1134 Without `model`, the turn uses that provider's default model (an `auto` thread
1135 stays `auto`). The override is always preflighted and is part of the
1136 `operation_key` fingerprint.
1137
1138 Resolution is deterministic: a turn override wins over the thread default,
1139 which wins over the Runtime's normal configuration. For tools, reaching normal
1140 configuration means the ordinary configured catalog; `[]` is never treated as
1141 missing. Reasoning is normalized only after the exact provider/model route is
1142 resolved, and `auto` remains a per-prompt reasoning decision even when the
1143 thread uses a fixed model. The request still enters the existing
1144 `Op::SendMessage` path and the single `Engine::run_turn` loop.
1145
1146 Image input uses the same turn path: `"images": [{"mime": "image/png",
1147 "dataBase64": "..."}]`. Clients must first observe
1148 `capabilities.turn_image_inputs: true` in `/v1/runtime/info` (or the isolated
1149 Runtime Chat relay catalog). Older HTTP runtimes ignore unknown fields, so a
1150 successful text response is not evidence that an attachment was accepted.
1151 The field is omitted when empty. It is also accepted by app-server
1152 `thread/message`, `thread/request` messages, and prompt requests; that bridge
1153 checks the underlying Runtime capability before forwarding image bytes.
1154 Legacy remote Work commands do not support images and explicitly refuse them.
1155
1156 New inline images require a named model whose exact resolved route reports
1157 `image_input: "supported"`; Auto and unknown/unsupported image routes are
1158 refused before classifier or provider dispatch. This does not change the
1159 existing trusted-local attachment behavior for routes with unknown capability.
1160 A nonempty prompt is required. Inputs are limited to 10 images, 4 MiB decoded
1161 bytes per image, 5 MiB total, and an 8 MiB JSON body. PNG, JPEG, GIF and WebP
1162 must have matching MIME, canonical padded base64 and valid bounded image
1163 content: at most 8192 pixels per dimension, 33,554,432 pixels total and 64 MiB
1164 decoder allocation. The Runtime does not fetch paths or URLs from this field.
1165 Malformed images refuse the whole turn; callers can retain the draft for
1166 correction. Relay command polling uses an 8 MiB response budget; the sender
1167 must paginate by serialized bytes without advancing past unserved commands.
1168
1169 Accepted image bytes and order are retained in the existing turn records and
1170 reconstructed after restart, import and fork. Retry retains those images even
1171 when its optional `prompt` changes the text; undo responses include
1172 `original_user_images` when present. Image-bearing records require schema v3,
1173 which older readers refuse. Text-only records and operation fingerprints retain
1174 their prior representation. Validated stored local images retain the existing
1175 5 MiB per-image ceiling and prior aggregate/count semantics on import/retry;
1176 this internal storage authority does not
1177 relax exact model or permission checks. Image bytes, MIME and order participate in request
1178 identity, so changing an image under the same operation key conflicts.
1179 Compaction can summarize older context; retaining the original attachment does
1180 not promise that every later model request includes it. Image pixels are not
1181 subject to text-secret redaction.
1182
1183 `operation_key` is an optional idempotency key for clients that may lose an
1184 HTTP response after the Runtime accepted a turn. It is scoped to the current
1185 Runtime store and thread, may contain at most 128 UTF-8 bytes, and may not be
1186 empty or contain surrounding whitespace or control characters. Omitting it
1187 preserves the legacy create-a-new-turn behavior.
1188
1189 The first accepted request durably binds a SHA-256 fingerprint of the key to
1190 the Runtime turn id and a canonical request fingerprint before sending the
1191 existing `Op::SendMessage`. An exact retry returns that original turn in the
1192 normal `{ "thread": ..., "turn": ... }` response and emits no second engine
1193 operation, item, or lifecycle sequence. Reusing the key on the same thread
1194 with a different provider/model, prompt, reasoning policy, tool allowlist or
1195 dynamic-tool schema, environment, or permission policy fails closed with
1196 `409 Conflict`. The same caller key may be used independently on another
1197 thread.
1198
1199 Only the scoped key fingerprint, request fingerprint, thread id, and turn id
1200 are stored in the Runtime's private turn-operation index. The raw key is never
1201 persisted or logged, and request bodies, credentials, and attachments are not
1202 copied into that index. Existing thread/turn persistence remains the source of
1203 the returned turn after a process restart.
1204
1205 **Exact accepted-turn lookup**
1206
1207 `GET /v1/threads/{id}/turn-operations/{operation_key}` uses the same Runtime
1208 authentication as turn submission. URL-encode each path segment. It returns
1209 `200 OK` with the existing bare `TurnRecord` (the `turn` object in the POST
1210 response), identified by that exact thread and operation key. It does not use
1211 the thread's latest turn or require the original request body or current route
1212 settings to match.
1213
1214 - `404 Not Found`: no binding exists for that thread/key, or persisted identities
1215 do not match. These cases share a generic response.
1216 - `409 Conflict`: admission holds the operation claim, or its durable binding
1217 is incomplete. Retry the lookup; this response does not authorize another turn.
1218 - `400 Bad Request`: the thread ID or operation key is malformed. The key uses
1219 the same 128-byte and whitespace/control-character rules as POST.
1220 - `500 Internal Server Error`: storage or the existing claim lock cannot be
1221 checked safely. This is not evidence that the operation is absent.
1222
1223 The lookup holds a shared read lock on the existing operation claim while
1224 reading the binding and turn. It creates no files, starts no engine, emits no
1225 events, and performs no replay or recovery. Normal Runtime startup may recover
1226 an incomplete admission before a later lookup, but GET itself never does so.
1227
1228 **Approvals**
1229 - `POST /v1/approvals/{approval_id}` with body
1230 `{ "decision": "allow" | "deny", "remember": false }`
1231
1232 `approval_id` is minted by the Runtime, not by the model or the provider. It is
1233 an opaque `approval_<32 hex>` capability, unique per prompt, bound to the thread
1234 that raised it, and single-use: the Runtime removes it when the decision is
1235 delivered, when the prompt times out, or when the turn abandons it. Clients echo
1236 the value they were given and must not construct, derive, or guess one.
1237
1238 It is deliberately **not** the provider's tool-call ID. Providers restart their
1239 call-ID counters per response, so two threads can gate calls whose raw IDs are
1240 byte-equal; keying approvals by that value let one thread's decision settle
1241 another thread's call. The endpoint therefore performs one exact match on the
1242 minted ID and has no fallback: a raw tool-call ID, an expired ID, or a replayed
1243 ID that has already been settled all return `404` and reach no engine. A `404`
1244 means the capability is not pending — it is not evidence about how the approval
1245 was resolved; read `approval.decided` for that.
1246
1247 The raw provider call ID travels separately as `tool_call_id` on
1248 `pending_approvals[]` and on the approval events. It is a correlator for
1249 attaching a prompt to the tool row it gates, and never accepted as a decision.
1250 Each thread-detail `pending_approvals[]` entry is
1251 `{ "id", "turn_id", "tool_name", "description", "intent_summary"?, "tool_call_id"?, "summary"? }`,
1252 where `id` is the capability above. `summary` (also on `approval.required`) is
1253 a one-line description of the gated call built from the tool name and its
1254 arguments only, never from model text ("Search the web for 'espresso'",
1255 "Write notes/espresso.md"); paths inside the workspace are workspace-relative.
1256 Clients show it first and keep the raw arguments behind it. For task and
1257 automation create/update, `summary` also names the requested trust mode,
1258 shell, auto-approve, mode and workspace.
1259
1260 `"remember": true` on an `allow` records a **session grant** for that tool and
1261 argument class (the approval grouping key: a shell command family for a
1262 simple, known command such as `git status` whose only options are value-free
1263 ones like `-s` or `--porcelain` — a compound, wrapper, interpreter or
1264 unrecognised command, one with any other option, or one whose arguments are
1265 what runs or is installed (`go run`, `make`, `git bisect`, package installs)
1266 is granted as its full normalized command, and a
1267 shell interact or wait call as the exact call — a patch's file set, a `fetch_url` host, an MCP tool, a `web.run` action kind — for
1268 `open`, the hosts it opened). Computer Use consent and `app_script` calls, and
1269 any tool without a class, are granted for the exact call only. A grant never
1270 changes the thread's permission posture. Later matching calls on the thread are
1271 approved without a prompt: they still emit `approval.required`, then
1272 `approval.decided` with `"auto": true` and the `grant_id`. Creating a grant
1273 emits `approval.grant_added` with `{ "grant": { "grant_id", "tool_name",
1274 "scope", "summary", "granted_at" } }`; thread detail lists live grants in
1275 `approval_grants[]`. `DELETE /v1/threads/{id}/approval-grants/{grant_id}`
1276 revokes one (emitting `approval.grant_revoked`); the next matching call
1277 prompts again. Archiving or deleting the thread ends all of its grants
1278 (archiving emits `approval.grant_revoked` for each; unarchiving does not
1279 restore them). Grants live in memory for the Runtime process: a restart
1280 forgets them, and a forced (non-bypassable) prompt is never answered by one.
1281
1282 **User input**
1283 - `POST /v1/user-input/{thread_id}/{input_id}` with body
1284 `{ "answers": [{ "id": "question-id", "label": "Choice", "value": "Choice" }] }`
1285
1286 Submitted values are delivered to the active model turn but are deliberately
1287 excluded from durable Runtime items and events. The settled tool item contains
1288 only a neutral receipt and a machine-readable `response_redacted` marker. The
1289 Runtime accepts only an exact pending `(thread_id, input_id)` request; an
1290 unknown, concurrently settling, or already settled id returns 404 and is never
1291 placed in the engine mailbox. It commits the secret-free
1292 `user_input.answered` receipt before removing the snapshot-authoritative prompt
1293 or delivering the answer to the engine. That settlement runs independently of
1294 the HTTP connection, so disconnecting after submission cannot leave a prompt
1295 half accepted. Terminal-turn cancellation follows the same receipt-before-
1296 removal ordering through `user_input.canceled`.
1297
1298 **Client-executed dynamic tools**
1299 - `POST /v1/threads/{thread_id}/turns/{turn_id}/tool-calls/{call_id}/result`
1300
1301 The thread and turn in the result route must match the pending call. A call is
1302 settled at most once; wrong-route and duplicate results return 404. Terminal
1303 lifecycle events carry identifiers and status only, never tool result content.
1304 The Runtime commits the terminal lifecycle event before making a submitted
1305 result available to the model. Result delivery, timeout, and terminal-turn
1306 cancellation race through one settlement owner, so exactly one of these events
1307 is durable for a call:
1308
1309 - `tool_call.requested` — the typed client-executed call became pending;
1310 - `tool_call.resolved` — a result was durably accepted by the Runtime
1311 (`result_accepted: true`; `success` is result metadata, but result content is
1312 excluded);
1313 - `tool_call.timeout` — no result won before the bounded wait expired;
1314 - `tool_call.canceled` — the turn terminated before a submitted result won.
1315
1316 HTTP `202 Accepted` and `tool_call.resolved` share that durable-acceptance
1317 meaning. Neither claims that the model consumed the result: a concurrent turn
1318 shutdown may close the model receiver after acceptance. Once the Runtime has
1319 accepted the result, that call is terminal and a duplicate result returns 404.
1320
1321 **Events** (SSE replay + live stream)
1322 - `GET /v1/threads/{id}/events?since_seq=<u64>&replay_limit=<n>&progress=true`
1323
1324 Cursors:
1325
1326 - `since_seq` is the per-thread cursor: events with `seq > since_seq` are sent.
1327 Omitted (and no `Last-Event-ID`), the stream starts from the beginning of the
1328 thread's history.
1329 - Every journal frame carries `id: <seq>`, so a browser `EventSource` can
1330 resume through the `Last-Event-ID` header it sends on reconnect. An explicit
1331 `since_seq` wins over the header, so a deliberate replay from `0` is never
1332 overridden by a stale id. A header value that is not a decimal integer is
1333 ignored.
1334 - `replay_limit` (at most 4096) returns only the newest tail of the requested
1335 history; `previous_seq` on the first returned event advances past exactly the
1336 omitted history.
1337
1338 Durable history parsing runs off the async server workers and reaches SSE in
1339 bounded batches of at most 256 events through a backpressured channel. Broadcast
1340 delivery is only a wake-up optimization: a lagged receiver opens the same
1341 bounded durable replay from its last accepted cursor.
1342
1343 `progress=true` adds `stream.progress` transport frames
1344 (`{schema_version, event, kind, thread_id, seq, state}`, `state` is
1345 `replaying` or `live`) at the current cursor and advertises them with
1346 `x-codewhale-event-progress: 1`. The stream reports `live` only after durable
1347 history and the already-queued live tail are both drained; broadcast-lag
1348 recovery returns it to `replaying`. Progress frames never carry a new sequence
1349 number.
1350
1351 Failures before the stream opens are ordinary HTTP errors with the JSON error
1352 body, never SSE:
1353
1354 | Status | When |
1355 | --- | --- |
1356 | `401` / `403` | Missing or wrong Runtime credential |
1357 | `404` | Unknown thread |
1358 | `400` | `replay_limit` above 4096 |
1359 | `500` | The durable history could not be opened (including a replay worker crash before the first cursor) |
1360
1361 Once the response is `200`, every end the server chooses is a final
1362 `stream.end` frame; see [Ending and resuming a thread
1363 stream](#ending-and-resuming-a-thread-stream).
1364
1365 **Snapshots** (side-git restore point listing + restore)
1366 - `GET /v1/snapshots?limit=20`
1367 - `POST /v1/snapshots/{id}/restore`
1368
1369 `/v1/snapshots` lists recent side-git restore points for the runtime workspace.
1370 `limit` defaults to `20` and must be between `1` and `100`. `POST
1371 /v1/snapshots/{id}/restore` restores workspace files from the snapshot and
1372 returns `{"restored": "<snapshot-id>"}`. It is the direct operator surface for
1373 the server's own workspace (the same action as the TUI's `/restore <N>`): it is
1374 gated by the Runtime API bearer token, not by any thread's trust flag, and it
1375 is refused with `409` while a turn is active in an overlapping workspace (see
1376 below). A `pre-restore:` safety snapshot is taken first.
1377
1378 ```json
1379 [
1380 {
1381 "id": "snap_...",
1382 "label": "post-turn:1",
1383 "timestamp": 1780730580
1384 }
1385 ]
1386 ```
1387
1388 ### Workspace restore endpoints
1389
1390 Three routes mutate workspace files from side-git snapshots. They share one
1391 admission rule and one safety net, and they differ in scope and trust.
1392
1393 | Route | Scope | Trust | Forks the thread |
1394 | --- | --- | --- | --- |
1395 | `POST /v1/snapshots/{id}/restore` | whole server workspace | bearer token only (operator action) | no |
1396 | `POST /v1/threads/{id}/patch-undo` | the files the dropped turns changed | thread `trust_mode` or `auto_approve` when files would change | yes |
1397 | `POST /v1/threads/{id}/file-revert` | exactly one regular file | thread `trust_mode` or `auto_approve`, always | no |
1398
1399 **Admission.** A restore reserves the same admission the Runtime uses for
1400 config reloads and session checkpoints, so no new turn starts and no saved
1401 history changes while files are being rewritten. If any thread already has an
1402 active turn in the same workspace, a nested checkout of it, or a parent of it,
1403 the request is refused with `409` and the message `already has an active turn`.
1404 The reservation is owned by the worker performing the Git mutation, so a client
1405 that disconnects mid-request cannot release it early; the operation completes
1406 or fails as a whole. Concurrent restores serialize. The reservation is
1407 runtime-wide: while a restore's safety snapshot and checkout run, new turns,
1408 steering, compaction and user-input delivery on every thread wait for it to
1409 finish, so a large workspace can add seconds of latency elsewhere during a
1410 restore. A thread whose workspace directory is not available (unmounted
1411 volume, disconnected share, missing directory) is refused with `409` rather
1412 than treated as having nothing to restore.
1413
1414 **Ownership.** A thread owns exactly the workspace restore points recorded on
1415 its own turns. While a turn runs, the engine reports each snapshot it takes —
1416 `pre_turn` before the turn, `tool` before and `post_tool` after each tool call
1417 that may write (every call that is not read-only: file tools, shell commands,
1418 programs, write-capable MCP tools), `post_turn` when it ends (always before
1419 the turn settles) — and the Runtime appends it to the turn record's
1420 `workspace_snapshots`, in order:
1421
1422 ```json
1423 "workspace_snapshots": [
1424 { "kind": "pre_turn", "snapshot_id": "<commit>", "tree_id": "<tree>", "session_id": "thr_1a2b3c4d" },
1425 { "kind": "tool", "snapshot_id": "<commit>", "tree_id": "<tree>", "session_id": "thr_1a2b3c4d", "tool_call_id": "call_…", "write_paths": ["src/lib.rs"], "changed_paths": [] },
1426 { "kind": "post_tool", "snapshot_id": "<commit>", "tree_id": "<tree>", "session_id": "thr_1a2b3c4d", "tool_call_id": "call_…", "changed_paths": ["src/lib.rs"] },
1427 { "kind": "post_turn", "snapshot_id": "<commit>", "tree_id": "<tree>", "session_id": "thr_1a2b3c4d", "changed_paths": [] }
1428 ]
1429 ```
1430
1431 `changed_paths` lists the workspace-relative paths whose content changed since
1432 the turn's previous receipt — what happened in the span the receipt closes; it
1433 is absent on `pre_turn` and when it could not be computed (a snapshot in
1434 between failed). `write_paths` is set on the `tool` receipt of a file tool
1435 (`write_file`, `edit_file`, `apply_patch`) to the paths the call declared, as
1436 it named them; a tool without it (a shell command) may write any path. On a
1437 user shell turn the `pre_turn` receipt carries the command's `tool_call_id`,
1438 since the command runs from it to `post_turn`.
1439
1440 Each receipt is also published as a `turn.workspace_snapshot` event (payload:
1441 the receipt). The engine runs every Runtime thread under the thread's own id,
1442 across restarts and engine eviction, so `session_id` is the thread id for turns
1443 the thread ran itself; it does not follow the thread's saved-session binding
1444 (`PUT`/`POST /v1/sessions`, resume), which only names a document. A fork clones
1445 its source's turn records and so owns the restore points of the turns it
1446 inherited. Another thread's, or a TUI session's, snapshots in the same
1447 workspace are never candidates. `tree_id` is the durable identity: a prune
1448 rebuilds the side repo and rewrites every commit id but keeps each tree, and a
1449 restore point resolves only to a stored snapshot with the same tree, session
1450 tag and kind. The count prune after each snapshot keeps the newest 50
1451 snapshots plus the newest 50 turn boundaries (`pre-turn:`/`post-turn:`), so a
1452 turn with more tool calls than that, or a burst from another thread, never
1453 pushes out a recent turn's own restore points. Turns recorded before receipts existed, turns imported by
1454 `resume-thread`, and turns run with snapshots off or unavailable have none.
1455
1456 **Safety net.** Every restore first records a `pre-restore:<target>` snapshot
1457 of the current workspace. That label is never a `/undo`, `patch-undo` or
1458 `file-revert` candidate, so the net does not change what later undos select.
1459 For `file-revert` the backup is mandatory: if it cannot be written, or the
1460 requested file is excluded from it (for example by `.gitignore`), the request
1461 fails and nothing is changed.
1462
1463 **`patch-undo`.** Undoes whole turns. For each dropped turn's `pre_turn` →
1464 `post_turn` window, the paths that differ between the two snapshots must all be
1465 the turn's own: changed only in the span of one of the turn's tool calls (a
1466 `tool` → `post_tool` span, or a shell turn's whole window), and, inside a file
1467 tool's span, a path that call declared. A path that changed while none of the
1468 turn's tools could have written it — another thread, an editor, a background
1469 process — is someone else's change, and the undo is refused rather than revert
1470 it. Each of the turn's paths goes back to its content before the first dropped
1471 turn that changed it, and nothing else is touched, so later work by the user or
1472 another thread in the same workspace survives. Then the conversation forks
1473 exactly as `/undo` does.
1474
1475 Snapshots never hold paths excluded by the workspace's `.gitignore` files or
1476 the built-in snapshot exclusions (`node_modules/`, `target/`, `dist/`, build
1477 caches, binary artifacts), nor paths outside the workspace. A dropped file-tool
1478 call that declared such a path is refused with `path_not_snapshotted`, because
1479 no snapshot can put it back. A shell command declares no paths: what it writes
1480 under an excluded path (build output, dependency installs) is outside what
1481 `patch-undo` restores and is not reported. `201` means either files were restored (`files_restored: true`, one
1482 `<action> <path>` line per file in `summary`, `snapshot_label` naming the
1483 pre-turn snapshot) or there was provably nothing to restore
1484 (`files_restored: false`): every dropped turn ran here without a tool call,
1485 changed no file, or its files are already back at their pre-turn content.
1486 Anything that cannot be restored aborts the whole undo with `409`, nothing is
1487 changed and no fork is published; `error.code` says why:
1488
1489 | `error.code` | Meaning |
1490 | --- | --- |
1491 | `restore_point_unavailable` | a dropped turn that may have changed files has no complete recorded restore point (older record, imported by `resume-thread`, snapshots off, or the snapshot failed) |
1492 | `restore_point_pruned` | the restore point is no longer in the snapshot store |
1493 | `path_not_snapshotted` | a dropped file-tool call wrote a path snapshots do not hold (ignored, built-in exclusion, or outside the workspace) |
1494 | `workspace_changed_since_turn` | a path the turns changed was changed afterwards (or between two dropped turns), a path changed while a dropped turn ran but outside its own tool calls, or a path that is not a regular file |
1495 | `restore_requires_trust` | there is something to restore and the thread is not in trusted mode or Full Access |
1496 | `workspace_unavailable` | the workspace directory is not available |
1497
1498 For the first four a client can offer the conversation-only
1499 `POST /v1/threads/{id}/undo` instead, and `file-revert` for individual files.
1500 Snapshot repository, listing or comparison failures abort with `500` and also
1501 preserve the conversation, so a turn is never dropped while its file changes
1502 stay on disk. Depth and history are validated before any file changes. If the
1503 fork cannot be persisted after files were restored, the response is a `500`
1504 that names the restored snapshot; the original thread still holds the turn and
1505 the `pre-restore:` snapshot holds the previous files.
1506
1507 **`file-revert`.** Request body:
1508
1509 ```json
1510 {
1511 "path": "src/lib.rs",
1512 "snapshot_id": "3f2a…40-or-64 hex…",
1513 "expected_hash": "sha256:<64 lowercase hex digits>"
1514 }
1515 ```
1516
1517 - `path`: workspace-relative, or absolute inside the thread workspace. The
1518 name is literal (brackets, spaces and glob characters are filename bytes;
1519 Git runs with `--literal-pathspecs`). It must name a regular file: directories,
1520 symlinks anywhere in the path, and `.git` components are `400`.
1521 - `snapshot_id`: the exact `tool` or `pre_turn` restore point of the change
1522 the user selected, from the thread's own turn records: a receipt's
1523 `snapshot_id` or `tree_id` (for a tool call, the receipt whose
1524 `tool_call_id` matches), or the current commit id `GET /v1/snapshots` lists
1525 for it. It must be a restore point recorded on one of this thread's turns;
1526 the server never picks "the newest snapshot that differs", because an
1527 unrelated newer snapshot can erase later user edits while leaving the tool's
1528 change.
1529 - `expected_hash`: `sha256:` of the current file bytes the client displayed,
1530 or `absent` when the client saw the file as deleted. It is checked before the
1531 safety backup and again immediately before the mutation.
1532
1533 Responses:
1534
1535 - `200 {"path", "action", "snapshot_id", "snapshot_label"}` — `action` is
1536 `modified`, `recreated` (file was missing) or `removed` (the snapshot does
1537 not contain the file, so the file the tool created is deleted; its parent
1538 directories are left in place).
1539 - `400`: malformed `snapshot_id`/`expected_hash`, path outside the workspace,
1540 or a path that is not a regular file on either side.
1541 - `404`: unknown thread.
1542 - `409`: thread not in trusted mode or Full Access; active turn in an
1543 overlapping workspace; workspace directory not available; snapshot unknown,
1544 pruned, not recorded on this thread's turns (another thread's or a TUI
1545 session's), or not a restore point (refresh the change record); file already
1546 matches the
1547 snapshot (nothing to revert); or the file changed after the reviewed
1548 `expected_hash` (refresh and review again). Nothing is changed in any of
1549 these cases.
1550 - `422`: missing or mistyped body fields.
1551 - `500`: Git or filesystem failure; a failure after the safety snapshot names
1552 that snapshot so the previous bytes can be recovered with
1553 `POST /v1/snapshots/{id}/restore` or `/restore`.
1554
1555 Capability probe: `GET` on the route returns `405` where the endpoint exists
1556 and `404` on an older engine; clients treat any non-`404` as available and
1557 degrade with an explanation otherwise.
1558
1559 **Compatibility stream** (one-shot, backwards-compatible)
1560 - `POST /v1/stream`
1561
1562 **Tasks** (durable background work)
1563 - `GET /v1/tasks`
1564 - `POST /v1/tasks`
1565 - `GET /v1/tasks/{id}`
1566 - `POST /v1/tasks/{id}/cancel`
1567
1568 **Automations** (scheduled recurring work)
1569 - `GET /v1/automations`
1570 - `POST /v1/automations`
1571 - `GET /v1/automations/{id}`
1572 - `PATCH /v1/automations/{id}`
1573 - `DELETE /v1/automations/{id}`
1574 - `POST /v1/automations/{id}/run`
1575 - `POST /v1/automations/{id}/pause`
1576 - `POST /v1/automations/{id}/resume`
1577 - `GET /v1/automations/{id}/runs?limit=20`
1578
1579 Create and update requests accept an optional `model`. When present, each
1580 scheduled or manually triggered run uses that model; omitting it keeps the
1581 runtime's default task model.
1582
1583 **Operate** (always-on named operation; same `OperateRecord` as CWC
1584 `20de981` / PR #284)
1585
1586 - `GET /v1/operate` — current operation + plan board
1587 - `POST /v1/operate` — create (`direction`, optional `burnRate`)
1588 - `PATCH /v1/operate` — steer direction, `burnRate`, or `leadPlan`
1589 - `PUT /v1/operate/plan` — set `leadPlan` (`{ slices: [...] }`)
1590 - `POST /v1/operate/keepalive` — observe spend / burn; never stops
1591 - `POST /v1/operate/cancel` — explicit cancel (`/v1/operate/stop` aliases);
1592 also pauses the `cw-operate` keepalive so nothing keeps spending after cancel
1593 - `POST /v1/operate/auto-merge/check` — call landed
1594 `scripts/check-auto-merge.py --repo --pr --agent` (does not merge)
1595
1596 The operation record (`current.json`) is persisted under a cross-process
1597 file lock with atomic temp+rename writes; every PATCH / keepalive / plan
1598 save reloads the latest state inside the lock, so concurrent saves merge
1599 instead of losing writes. A `PATCH` that changes `direction` invalidates
1600 the recorded `leadPlan` (workers stop executing superseded slices) and
1601 pulls the keepalive lead run forward to re-plan. `POST /v1/operate`
1602 installs the hourly `cw-operate` keepalive and kicks its first lead-plan
1603 run immediately instead of waiting out the first recurrence; credentials
1604 resolve through the normal Z.ai provider resolution (config, `api_key_env`,
1605 secret store, or provider env vars — blank values count as missing).
1606
1607 `burnRate` is `{ "kind": "usd_per_hour", "amountUsdPerHour": number }`,
1608 a positive number, or `null` (unbounded). Status is
1609 `planning | running | idle_blocked | cancelled`. Pace
1610 (`unbounded | hold | throttle | widen`) is not a status: over-target
1611 throttles, under-target widens, no wallet-cap stop. Idle-blocked only
1612 for empty direction, awaiting lead plan, missing credentials, or a
1613 human gate. Auto-merge is `scripts/check-auto-merge.py --repo … --pr …
1614 --agent …` from codewhale-ops `origin/main` (exit 0), then
1615 `scripts/auto-merge-pr.py`. Do not invent a second checker.
1616
1617 **Introspection**
1618 - `GET /v1/workspace/status`
1619 - `GET /v1/workspace/files/search?query=<partial>&limit=<1-100>` (see workspace file suggestions above)
1620 - `GET /v1/workspace/files?path=<dir>&limit=<1-2000>`, `GET /v1/workspace/files/read?path=<file>&offset=&limit=`
1621 and `PUT /v1/workspace/files` (see workspace files and session artifacts above)
1622 - `GET /v1/skills`
1623 - `GET /v1/apps/mcp/servers`
1624 - `GET /v1/apps/mcp/tools?server=<optional>`
1625
1626 Skill activation toggles are persisted under a cross-process transaction lock.
1627 Each mutation reloads and merges the latest exact-name state before an atomic
1628 write, and `GET /v1/skills` refreshes that shared state so another Codewhale
1629 process's successful toggle is visible without restarting the Runtime API.
1630
1631 **Usage** (token/cost aggregation across threads)
1632 - `GET /v1/usage?since=<rfc3339>&until=<rfc3339>&group_by=<day|model|provider|thread>`
1633
1634 `since` / `until` are inclusive RFC 3339 timestamps and may be omitted (no
1635 bound). `group_by` defaults to `day`. Buckets are sorted by ascending key.
1636 Empty time ranges produce empty `buckets` (never a 404). Cost is computed via
1637 the model→pricing map; turns whose model has no pricing entry contribute
1638 tokens but `0.0` cost. Added in v0.8.10 (#564).
1639
1640 ```json
1641 {
1642 "since": "2026-04-01T00:00:00Z",
1643 "until": "2026-04-30T23:59:59Z",
1644 "group_by": "day",
1645 "totals": {
1646 "input_tokens": 12345,
1647 "output_tokens": 6789,
1648 "cached_tokens": 0,
1649 "reasoning_tokens": 0,
1650 "cost_usd": 0.012,
1651 "turns": 42
1652 },
1653 "buckets": [
1654 {
1655 "key": "2026-04-30",
1656 "input_tokens": 1234,
1657 "output_tokens": 678,
1658 "cached_tokens": 0,
1659 "reasoning_tokens": 0,
1660 "cost_usd": 0.001,
1661 "turns": 3
1662 }
1663 ]
1664 }
1665 ```
1666
1667 ### Native client routes (GPUI desktop)
1668
1669 These families serve the GPUI desktop client over the same bearer-token
1670 transport. They reuse the runtime's existing authorities — the engine's shell
1671 manager, the durable thread store, the workspace confinement layer, the
1672 config's credential plumbing — and add no second runtime, session store,
1673 scheduler, or credential store.
1674
1675 **Terminal sessions** (the persistent Engine-owned shell)
1676
1677 The jobs family above runs one command per job. A terminal pane needs the
1678 *other* authority: the stateful PTY-backed shell the agent's own terminal
1679 tools drive, which keeps cwd and environment across inputs. These routes
1680 attach to that session and never create one — a name with no live session is
1681 `404`, because conjuring a shell from an HTTP request would give the client a
1682 terminal the Engine does not know about. Input is attributable by route:
1683 `input` is the client's writer, `terminal_send` is the agent's.
1684
1685 - `GET /v1/terminal/{name}/output?cursor=<bytes>&max_bytes=<1-64KiB>&format=
1686 <base64|text>` — the resumable byte stream. `{name, offset, next_cursor,
1687 total, dropped, encoding, data, running, exit_code}`: pass `next_cursor`
1688 back to continue; reads never consume, so several clients may hold
1689 independent cursors; `dropped` reports bytes the 512 KiB ring discarded,
1690 and a cursor past `total` is answered from `total` rather than echoed back
1691 - `POST /v1/terminal/{name}/input` — `{ "data", "encoding"? }`, `base64` by
1692 default (exact bytes) or `text` for UTF-8 → `{ "name", "written" }`
1693 - `POST /v1/terminal/{name}/resize` — `{ "rows", "cols" }` → the kernel
1694 window the child draws for
1695 - `POST /v1/terminal/{name}/kill` — end the shell; observe the exit through
1696 `output` (`running` / `exit_code`) rather than the acknowledgement
1697
1698 `GET /v1/runtime/info` advertises `terminal_stream`, `terminal_input`,
1699 `terminal_resize` and `terminal_kill`. All four are `false` on Windows and OpenHarmony builds
1700 today: the owner is Unix-only, those routes answer `501`, and a client
1701 should gate its terminal controls on these flags rather than discovering it
1702 from a failed request. Known limitations, stated because a reader would
1703 otherwise assume them: there is no `wait_ms` long poll (poll the cursor),
1704 scrollback dropped by the ring is gone with the process, a restarted Engine
1705 reports no session rather than pretending to reattach, and the
1706 `@codewhale/runtime-sdk` package has no terminal client wrapper yet — the raw
1707 routes are the contract for now.
1708
1709 **Jobs** (operator-scoped shell jobs; the terminal surface)
1710 - `GET /v1/jobs` — every live and known-stale job across all threads
1711 - `GET /v1/threads/{id}/jobs` — jobs owned by one thread's manager:
1712 model-launched, subagent-launched, and client-launched together
1713 - `POST /v1/threads/{id}/jobs` — `{ "command", "cwd"?, "timeout_ms"?,
1714 "tty"?, "env"? }` → `201 { "job" }`; runs as a background shell under the
1715 thread's projected sandbox policy. `tty: true` merges stderr into stdout
1716 and gives the command a terminal (required for interactive programs);
1717 background jobs are never killed at `timeout_ms`. A relative `cwd`
1718 resolves against the thread workspace. Outside trust mode `cwd` must stay
1719 inside it after symlinks resolve (`403` otherwise): unlike the shell tool,
1720 this route does not follow `workspace_follow_symlinks` or `/trust add`
1721 roots, so a symlink leading out of the workspace is refused. The job runs
1722 in the resolved directory that was checked (a later symlink retarget does
1723 not move it); a `cwd` that resolves to a non-UTF-8 path is a `400`
1724 - `GET /v1/threads/{id}/jobs/{job_id}` — one job's status + metadata
1725 - `GET /v1/threads/{id}/jobs/{job_id}/output?stream=<stdout|stderr>&cursor=
1726 <bytes>&max_bytes=<1-512KiB>&wait_ms=<0-30s>&format=<base64|text>` — the
1727 resumable byte stream. `{job_id, stream, offset, next_cursor, total,
1728 dropped, encoding, data, status, exit_code, done}`: pass `next_cursor`
1729 back to continue; `wait_ms` long-polls for new bytes on a running job;
1730 `done` means a terminal status and nothing left past the cursor
1731 - `POST /v1/threads/{id}/jobs/{job_id}/stdin` — `{ "data", "encoding"?,
1732 "close"? }`: `data` is UTF-8 text by default or `base64`, `close: true`
1733 sends EOF; works for PTY and piped jobs → `204`
1734 - `POST /v1/threads/{id}/jobs/{job_id}/kill` — bounded SIGTERM → SIGKILL
1735 escalation on the process group → `{ "job", "result" }` with the final
1736 snapshot
1737
1738 Reads are non-consuming: several clients may hold independent cursors, and
1739 polling never steals output from the engine's own delta consumer. The
1740 buffer is bounded with exact drop accounting — a reader whose `cursor`
1741 falls behind the retained window gets `offset` past it and `dropped > 0`,
1742 and must re-anchor. Evicted jobs keep a tail snapshot which the output
1743 route serves as the final retained window. Jobs are scoped to the thread
1744 that created them and are killed when that thread is removed; the engine's
1745 background commands use the same per-thread manager, so `GET /v1/jobs` is
1746 also how a client sees model-spawned work.
1747
1748 **Commands** (typed command catalog, APPS-28)
1749 - `GET /v1/commands` — `{commands: [...]}`: every registered slash command,
1750 builtin and user, as the TUI's own registry holds it. Per entry: `name`,
1751 `aliases`, `summary` and `usage` (English source text — localizing is the
1752 client's surface), `subcommands` (the literal verbs the usage line
1753 declares), `takes_arguments`, `kind` (`builtin` registered code, or `user`
1754 expanding a stored template), `binding` (`host` runs locally and never
1755 reaches the model; `prompt` expands into the request the model sees),
1756 `discovery` (`primary` / `advanced` / `compatibility`, builtins only),
1757 `hidden` for rows the product does not advertise, and `shadowed_by` /
1758 `shadowed_aliases` where a user command has taken a builtin's spelling.
1759
1760 Each entry also carries the composer argument shape, computed the way the
1761 TUI composer computes it, so a client does not re-derive it from `usage`
1762 (#6230):
1763 - `requires_argument` — the usage line mentions any argument, required or
1764 optional.
1765 - `requires_required_argument` — the usage line has a `<required>` argument
1766 outside every `[optional]` group.
1767 - `composer_wants_trailing_space` — accepting the command leaves a trailing
1768 space for its arguments.
1769 - `palette_runs_directly` — the palette runs the command on selection
1770 instead of pasting it into the composer.
1771 - `show_in_empty_discovery` — the command is listed when the slash menu
1772 opens with no filter text.
1773
1774 User commands derive these from `takes_arguments`: their arguments are
1775 never required, a template that takes arguments waits in the composer and
1776 one that does not runs directly, and a `hidden` template stays out of empty
1777 discovery.
1778
1779 The same registry the TUI palette reads, so a desktop palette can be
1780 checked against it instead of drifting from it. Two rules a client must
1781 respect: a `binding: "host"` row is never submitted as a model prompt, and
1782 a user command shadowing a builtin name wins that spelling.
1783
1784 **Hooks**
1785 - `GET /v1/hooks[?thread_id=...]` — `{workspace, enabled, hooks: [...],
1786 problems: [...]}`: the hook set Runtime API threads run for that workspace
1787 (the server workspace, or the named thread's). Per entry: `name`, `event`,
1788 `command` (credential-shaped values masked; every URL keeps only its scheme and host), `background`, `timeout_secs`,
1789 and `source` (`global` user config, `plugin` reviewed plugin, `project`
1790 trusted and approved `.codewhale/hooks.toml`). `problems` lists hooks
1791 rejected or warned about at load, one line each.
1792
1793 Every Runtime thread builds its engine with this set: `tool_call_before`
1794 can deny a call, `shell_env` applies to shell tools, and `tool_call_after`
1795 and `on_error` (for a failed tool) fire as observers, as they do in the TUI.
1796 Clients read this route instead of keeping a hook table of their own.
1797
1798 **Context** (per-thread context pressure, APPS-90)
1799 - `GET /v1/threads/{id}/context` — `input_tokens` (the conservative live
1800 estimate the visible meter uses), `billed_input_tokens` (last
1801 provider-counted prompt size when one exists), `window_tokens`,
1802 `output_cap_tokens`, `input_budget_ceiling`, `available_input_tokens`,
1803 `compaction_trigger_tokens`, `usage_percent` and `pressure`. Served by the
1804 live engine via `Op::GetContextBudget`. Every numeric field is nullable —
1805 a route that cannot express a bounded window reports `null` rather than an
1806 invented number — and `live: false` marks responses where the engine
1807 could not be loaded and only the store-recorded route's static window
1808 resolved.
1809
1810 **Git** (workspace repository operations, APPS-106)
1811 - `GET /v1/git` — status detail: `git_repo`, `branch`, `head`
1812 (abbreviated, for display), `head_oid`, `index_token`, `revision`,
1813 `ahead`/`behind`, counts, per-file porcelain `files[]`
1814 (`{path, index, worktree, staged, status, old_path?, rev}`), `branches`,
1815 `remotes`. `files[].path` and `old_path` are workspace-relative, the same
1816 frame the write routes take; in a workspace that is a subdirectory of its
1817 repository, rows outside the workspace are not listed (the counts stay
1818 repository-wide). A rename out of the workspace appears as its source
1819 deletion. Normal Git filters and untracked settings apply. Untracked
1820 directories stay collapsed; only a row naming the workspace itself is
1821 expanded to individually addressable files. The precondition tokens are
1822 opaque:
1823 - `head_oid` — the full HEAD commit id; `null` on an unborn branch or when
1824 HEAD cannot be read
1825 - `index_token` — the whole index (mode, blob, stage and path of every
1826 entry, repository-wide). A stat-only refresh by `git status` does not
1827 change it. Status remains best-effort on repository read failures;
1828 unavailable tokens are `null`, and broken HEADs cannot satisfy guards
1829 - `files[].rev` — one row: its index entries plus the working-tree state
1830 of every file it covers (including a rename's in-workspace source).
1831 Content tokens (`c-…`) include file bytes, the executable bit on Unix,
1832 symlink targets, and submodule HEAD/status. Submodule dirty contents
1833 are summarized by porcelain status, not recursively hashed; ordinary
1834 stage/discard do not write those contents. A read hashes at most
1835 64 MiB / 4,096 files; rows past that budget or containing a file over
1836 16 MiB carry a display-only size-and-mtime token (`s-…`). Those tokens
1837 cannot guard writes. Unreadable paths, paths beneath symlinked
1838 directories, special files and broken nested repositories have `rev: null`;
1839 other rows retain their tokens. If a broken tracked submodule aborts
1840 porcelain, ordinary rows are recovered without submodule recursion and
1841 the unreadable submodule is shown with `status: "unknown"`, `worktree: "?"`
1842 - `revision` — the whole tree: `head_oid`, `index_token`, every row's
1843 `rev`, including the working-tree state of rows outside a subdirectory
1844 workspace. `null` when any row is stat-only or unreadable, a repository
1845 read is incomplete, or untracked paths are hidden; whole-tree guarded
1846 writes are unavailable in that case. Clients must not silently omit a guard
1847 - `GET /v1/changes` — the same porcelain `files[]` projection plus
1848 `head_oid`, `index_token` and `revision`, minus repo chrome
1849 (branches/remotes): one authority, so the change list can never disagree
1850 with the status read
1851 - `GET /v1/diff?path=` — one file's unified `diff` against `base` (`HEAD`,
1852 or the empty tree on an unborn branch — which reads staged adds as new
1853 files). Covers staged+unstaged in one patch; `truncated` reports the
1854 512 KiB cap. An untracked file answers `untracked: true` with an empty
1855 diff — the client reads the file itself rather than mistaking it for
1856 unchanged
1857 - `GET /v1/workspace/diff?limit=` — whole-tree patch (default 256 KiB,
1858 max 4 MiB) plus a complete `--numstat` `files[]` inventory
1859 (`{path, added, deleted}`) so every changed row renders even when the
1860 patch is truncated
1861 - `GET /v1/git/graph?limit=` — bounded commit rows (`id`, `short`,
1862 `parents`, `author`, `timestamp`, `refs`, `subject`); an unborn branch is
1863 an empty graph, not an error
1864 - `POST /v1/git/stage` `{ "paths": [...] }` or `{ "all": true }`;
1865 `POST /v1/git/unstage` same; `POST /v1/git/discard` `{ "paths": [...] }`
1866 (tracked paths only — no `all`, an untracked path fails closed);
1867 `POST /v1/git/commit` `{ "message", "all"? }`; stage, unstage, discard
1868 and commit also take an optional `expect` (below); `POST /v1/git/push`
1869 `{ "remote"?, "set_upstream"? }` (`remote`, or `origin` when only
1870 `set_upstream` is given, must name a configured remote);
1871 `POST /v1/git/branch`
1872 `{ "name", "create"? }`
1873
1874 Diffs and precondition token reads run through the hardened review command
1875 (filters, fsmonitor, hooks, lazy fetches and replace-objects neutralized).
1876 Porcelain status uses normal Git, matching the workspace counts and filters;
1877 writes run through the non-interactive command path
1878 (`GIT_TERMINAL_PROMPT=0`, BatchMode ssh) so a credential or host-key prompt
1879 can never hang a request. Path lists are workspace-relative under the same
1880 confinement as the file routes (traversal → 400, `.git` → 403), passed after
1881 `--` under `--literal-pathspecs`, so `src/*` names a file called `*` and is
1882 never a glob. (Unstage of the whole tree uses the `:/` root pathspec.)
1883 Mutations answer `{ok, output, status, current}`: the refreshed status and
1884 the full `GET /v1/git` detail, so a client re-reads nothing after an
1885 operation and can chain the next write from fresh tokens. A workspace that
1886 is not a repository answers `404`.
1887
1888 *Preconditions.* Stage, unstage, discard and commit accept
1889 `expect: { head?, index?, revision?, files? }`, built from the last
1890 `GET /v1/git`. Each present field is checked and an absent one is not;
1891 `head: null` means "HEAD must still be unborn". Without `expect` (or with
1892 `expect: {}`) a write behaves exactly as before. Recommended use:
1893
1894 | Operation | `expect` |
1895 | --- | --- |
1896 | stage / unstage `paths` | `{head, files: {path: rev}}` |
1897 | stage / unstage `all` | `{head, revision}` |
1898 | discard (always — it destroys edits) | `{head, files: {path: rev}}` |
1899 | commit | `{head, index}` |
1900 | commit `all` | `{head, revision}` |
1901
1902 Malformed preconditions answer `400` before anything runs: `head` must be a
1903 40- or 64-hex id or `null`; `index` and `revision` 64 hex (an explicit
1904 `revision: null` is refused); `files` values a content-safe `c-` rev (`s-`
1905 tokens answer `400`). `files` keys
1906 (workspace-relative; `dir/` and `dir` are the same key)
1907 must name exactly the requested paths, so no path is left unguarded by
1908 accident. `files` is refused with `all: true` (use `revision`) and on
1909 commit. An unknown key inside `expect` is rejected like any other unknown
1910 field. When the repository no longer matches, the route writes nothing and
1911 answers `409`:
1912
1913 ```json
1914 { "error": { "message": "The repository changed since it was read (HEAD moved; src/a.rs changed). Nothing was written; refresh and review again.",
1915 "status": 409, "code": "git_state_changed" },
1916 "stale": ["head", "files"], "stale_paths": ["src/a.rs"],
1917 "current": { "...": "the GET /v1/git detail" } }
1918 ```
1919
1920 `stale` lists the components that moved (`head`, `index`, `files`,
1921 `revision`); `stale_paths` the `files` keys whose `rev` changed. A client
1922 re-renders from `current`, keeps the user's selection, and asks again.
1923
1924 Stage, unstage, discard, commit and branch from one runtime are serialized,
1925 so a check and its write are atomic with respect to that runtime's other
1926 windows; a second concurrent write answers `409` with
1927 `error.code: "git_busy"` rather than queueing behind a long commit hook.
1928 Push is not serialized: it only moves a remote ref and can wait on the
1929 network for up to 120 s. The lock does not cover processes outside this
1930 runtime — a terminal, an editor, or Codewhale's own agent tools — which can
1931 still change the repository in the moment between the check and git taking
1932 `index.lock`; a truly concurrent git write then fails on git's own
1933 `index.lock` (a `400` carrying git's message). A compare-and-swap commit via
1934 `commit-tree` and `update-ref` would close that window but skip the
1935 repository's hooks, which a Review-sheet commit must run, so it is not used.
1936
1937 **Diagnostics** (read-only logs, crashes, process — APPS-103)
1938 - `GET /v1/logs` → `{sources: [{dir, files: [{name, size, modified}]}]}` —
1939 the runtime's log directory plus `audit.log[.1]` from the codewhale home,
1940 newest first, capped
1941 - `GET /v1/logs/{name}?offset=<bytes>&limit=<bytes>&tail=<bytes>` →
1942 `{name, size, modified, offset, bytes, truncated, encoding, content}` —
1943 one bounded window; `tail` reads from the end and is mutually exclusive
1944 with `offset`; `truncated` means bytes remain after the returned window
1945 (a tail read at EOF is `false`), `encoding` is `utf-8` or `base64`
1946 - `GET /v1/crashes`, `GET /v1/crashes/{name}` — the same list/read contract
1947 over the crash-dump directories (`~/.codewhale/crashes`, legacy
1948 `~/.deepseek/crashes` merged)
1949 - `GET /v1/process` → `{pid, version, commit, started_at, uptime_seconds,
1950 executable, rss_bytes}` — `rss_bytes` only where the platform reports it
1951 (Linux `/proc`); absent rather than fabricated elsewhere
1952
1953 These routes package what already exists on disk for a client-side export;
1954 there is no telemetry upload route and no second log store. Names are
1955 basename-validated (no separators, no `..`), listings are capped, reads are
1956 bounded windows, and symlinks are never followed — a client bundles the
1957 files itself.
1958
1959 **Targets and remote posture** (APPS-50)
1960 - `GET /v1/targets` → `{targets: [self], remote: {supported: true,
1961 attach: "client", probe: "POST /v1/remote/connect"}, ssh: {…},
1962 cloud: {…}}` — this runtime's own record as the attachable target plus
1963 per-surface ownership; the runtime keeps no persistent target registry,
1964 so `POST /v1/targets` and `POST /v1/targets/switch` answer
1965 `501 Not Implemented` — target selection is client-owned and a switch
1966 must never move a running task server-side
1967 - `GET /v1/remote` → `{bind_host, port, loopback_only, reachable_from_lan,
1968 auth_required, mobile, tls}` — this listener's reachability posture.
1969 `tls` is always `false`: the API has no TLS terminator, so non-loopback
1970 reachability assumes a verified overlay (VPN/mesh), never plain LAN trust
1971 - `POST /v1/remote/connect` `{ "endpoint": "http://host:port" }` — probes a
1972 candidate remote's unauthenticated `GET /v1/runtime/info` (origin only;
1973 any pasted path is discarded). Answers `{ok, remote: {endpoint,
1974 runtime_api_version, codewhale_version, auth_required, …}, attach:
1975 "client"}` on success, and `{ok: false, reason: "unreachable" |
1976 "not a Codewhale runtime" | …}` as data on failure. URLs carrying
1977 credentials are refused with 400 — the remote's token is configured
1978 client-side, and a connect route that forwarded one would be an
1979 exfiltration primitive
1980 - `GET /v1/ssh`, `GET /v1/cloud` → `{supported: false, owner:
1981 "codewhale-control-plane", reason}`; `POST /v1/ssh/connect` and
1982 `POST /v1/cloud/attach` → `501`: SSH workspace provisioning and hosted
1983 cloud computers belong to the Apps control plane (ASCII Box for Managed
1984 Computer), not to a second authority inside Core
1985
1986 A remote Codewhale is a `serve --http` runtime with a token — that is the
1987 whole attach model. These routes describe and probe it; they never execute
1988 a remote request on the local machine.
1989
1990 **LSP** (workspace language intelligence, APPS-93)
1991 - `GET /v1/lsp` — capability: `enabled`, supported `languages` with their
1992 server commands, `custom_languages`, operations, poll and diagnostic caps
1993 - `GET /v1/diagnostics?path=` — file diagnostics
1994 - `GET /v1/definition?path=&line=&character=` (1-based)
1995 - `GET /v1/references?path=&line=&character=` (1-based)
1996 - `GET /v1/symbols?path=&query=` — empty query returns document symbols
1997
1998 One lazily-built workspace-level `LspManager` serves these; engine threads
1999 keep their own per-thread managers for the post-edit hook, and a server that
2000 never serves an LSP route never spawns a language server. `path` is
2001 workspace-relative under the same confinement as the file routes. Normal
2002 absence is data: no language server, a disabled `[lsp]` config, or a timeout
2003 answers `200` with `ok: false` and a machine-readable `reason`
2004 (`no_server`, `lsp_disabled`, `lsp_error`); malformed input is a 400 and a
2005 missing file a 404.
2006
2007 **Voice** (host dictation, APPS-98)
2008 - `GET /v1/voice` — capability: `available`, detected `recorder` command,
2009 resolved `asr` `{kind, model}`, `modes`, `send_phrases`,
2010 `max_record_seconds`
2011 - `POST /v1/voice/dictate` — record then transcribe → `{ ok, text }`
2012 - `POST /v1/voice/send` — same capture with the "send it" / 发送/發送
2013 suffix contract: `send: true` tells the client to submit (empty `text`
2014 with `send: true` means submit the client's current draft)
2015 - `POST /v1/voice/control` `{ "composer": "draft text" }` — assisted
2016 dictation that shows the model the composer text; `assisted: false` in the
2017 response means a free ASR backend (local whisper/Groq) handled the audio
2018 and the composer context was never seen
2019
2020 The runtime owns the host microphone and the ASR dispatch — the same
2021 implementation the TUI's `/voice` commands run, headless. Recording is one
2022 blocking capture per host (requests serialize; the loser gets
2023 `ok:false`/`no_speech`, not a fought-over device). Provider ASR resolves its
2024 key lazily so local-whisper and Groq paths work without provider auth.
2025 Interim and final transcription use the selected ASR backend. If local whisper
2026 or Groq fails, the error stays on that backend: the runtime never retries the
2027 recording or composer text with the active model provider. Select provider ASR
2028 explicitly to use that route.
2029 Failure is data: `no_recorder`, `no_speech`, `no_provider_auth`,
2030 `transcription_failed`. `CODEWHALE_DISABLE_VOICE=1` is an operator
2031 kill-switch — a headless `serve --http` host reports `available: false` and
2032 every dictate call fails closed.
2033
2034 ## Provider and model selection
2035
2036 These three routes are how a GUI renders a model picker whose contents are true
2037 for *this* runtime instead of guessed from a version snapshot. They were
2038 undocumented until 2026-08-04, which cost a desktop integration a day: the
2039 client probed `/v1/models`, `/v1/runtime/models`, and `/v1/runtime/providers`
2040 (all correctly 404) and concluded the capability did not exist.
2041
2042 ### `GET /v1/providers`
2043
2044 ```json
2045 {
2046 "current": "modelstudio-token-plan",
2047 "providers": [
2048 {
2049 "id": "modelstudio-token-plan",
2050 "model_provider_id": "modelstudio-token-plan",
2051 "display_name": "Alibaba Cloud Model Studio",
2052 "default_model": "qwen3.8-max",
2053 "has_model_catalog": true,
2054 "credentialState": "configured"
2055 }
2056 ]
2057 }
2058 ```
2059
2060 `current` is the active generic provider id. Only the active entry carries an
2061 exact identity: an active built-in normally repeats its canonical id in
2062 `model_provider_id`, while an active named custom route has `current` set to
2063 `custom` and the exact configured key there (for example `lm-studio`). Other
2064 entries have a null exact id; a null id on the active `custom` entry identifies
2065 the released legacy root-level custom route. Preserve both fields from the
2066 selected entry and send a non-null exact id back as `POST /v1/threads`'s
2067 `model_provider_id`; dropping a named custom id would collapse the selection to
2068 the legacy root custom route. `credentialState` is a stable, non-secret
2069 projection of the Runtime's existing structural credential classification:
2070
2071 - `configured`: credential material is structurally available;
2072 - `login_required`: the route needs login or a usable login capability;
2073 - `missing`: an API-style credential is unavailable;
2074 - `no_auth`: the route explicitly disables credential use;
2075 - `local`: the exact route is local and keyless;
2076 - `legacy`: the compatibility route cannot be classified more precisely.
2077
2078 For an active named custom provider, this state is calculated from the exact
2079 route named by `model_provider_id`, not from a generic custom-provider default.
2080 It deliberately collapses saved-key and imported-token details into
2081 `configured`, and login/consent-source details into `login_required`.
2082
2083 The response never includes endpoint URLs, credential environment-variable
2084 names, filesystem paths, credential values, consent-source details, or token
2085 metadata. `credentialState` is not a provider canary: `configured` does not
2086 prove endpoint reachability, credential validity, model entitlement, or a
2087 successful request. The models route below is also only a selection catalog; a
2088 non-empty list does not prove that the route can currently serve a request.
2089
2090 ### `GET /v1/providers/{id}/models`
2091
2092 ```json
2093 {
2094 "provider": "deepseek",
2095 "models": [
2096 {
2097 "id": "deepseek-v4-flash-vision-exp",
2098 "image_input": "supported",
2099 "reasoning_effort": "unknown",
2100 "reasoning_effort_levels": [],
2101 "reasoning_effort_source": null
2102 }
2103 ]
2104 }
2105 ```
2106
2107 For an exact configured route, supply `?model_provider_id=vision-work` and
2108 require the response to echo that same `model_provider_id`. The Runtime resolves
2109 that identity under the requested provider kind before reading model support.
2110 Unknown or mismatched identities return `400`. Named pagination cursors bind the
2111 configuration identity, endpoint and catalog snapshot; changing any of those
2112 requires restarting pagination. Omitting the query preserves the legacy catalog
2113 projection and omits the identity echo.
2114
2115 The catalog for one provider. Returns `400` for an unknown id, and for the
2116 legacy `deepseek-cn` alias, which has no provider metadata — use `deepseek`.
2117 An empty `models` array means the Runtime has no discoverable or configured
2118 model ids for that provider; it does not report credential presence.
2119
2120 The ids returned here are exactly the values accepted by `POST /v1/threads`'s
2121 `model` field and by the switch route below. `image_input` is the exact resolved
2122 provider/model route's capability state: `supported`, `unsupported`, or
2123 `unknown`. Keep `unknown` unknown rather than inferring from the model name or
2124 wire protocol. `supported` describes the model route; it does not mean a given
2125 client implements an image-upload control.
2126
2127 `reasoning_effort` uses the same three capability states and describes whether
2128 the exact model's metadata publishes a selectable effort ladder.
2129 `reasoning_effort_levels` contains only canonical, recognized active effort levels
2130 from that metadata. Off and provider synonyms such as none are excluded: the
2131 Apps/Chat protocol treats off as omission, which does not prove support for an
2132 explicit provider disable command. A model capable of reasoning may still have
2133 an unknown active effort ladder.
2134 Codex levels are also excluded when native compatibility would change their
2135 wire value (currently minimal and auto). This projection does not change native
2136 compatibility behavior or advertise a tier the Runtime cannot send unchanged.
2137 No levels are inferred from a provider-wide default or a familiar model name
2138 on a custom endpoint. `reasoning_effort_source` identifies `catalog`,
2139 `codex_cli_cache`, or `codex_app_server`; missing, stale, and unrecognized model
2140 metadata stays unknown. Codex roster metadata describes the external CLI's
2141 roster, not proof that a separately configured Runtime credential belongs to
2142 the same account or that an authentication boundary is approved.
2143
2144 Pass `?model_provider_id=<exact configured id>` when selecting a named route.
2145 The Runtime validates the provider kind and exact identity together, returns
2146 `model_provider_id` alongside that route's model list, and leaves the active
2147 configuration unchanged. An empty or unknown requested identity, or a mismatched kind, returns
2148 `400`; it never falls back to another named route.
2149
2150 The Runtime Chat relay publishes the same effort fields in camelCase
2151 (`reasoningEffort`, `reasoningEffortLevels`, `reasoningEffortSource`). These
2152 model facts do not enable tool execution or establish account entitlement.
2153
2154 For a thread-scoped choice, send the provider fields from the selected entry
2155 alongside the selected model. Omit `model_provider_id` when it is null:
2156
2157 ```json
2158 {
2159 "model_provider": "custom",
2160 "model_provider_id": "lm-studio",
2161 "model": "local-vision-model"
2162 }
2163 ```
2164
2165 This creates one thread on the exact named custom route without changing the
2166 Runtime's provider or model defaults.
2167
2168 ### `PUT /v1/providers/{id}/key` — write-only credential
2169
2170 ```json
2171 // request
2172 { "key": "sk-…" }
2173
2174 // response
2175 { "provider": "openai-codex", "stored": true, "backend": "keychain",
2176 "credentialState": "configured", "configPath": "/…/config.toml" }
2177 ```
2178
2179 Stores a provider API key through the same transactional write as
2180 `codewhale auth set --provider <id> --api-key-stdin`: the secret store under
2181 the provider write lock, plus the `[providers.<id>] auth_mode` metadata
2182 marker persisted to the config document and mirrored into the live runtime
2183 config so `GET /v1/providers` reports the new state immediately. `backend`
2184 names which secret backend holds the key and `configPath` which config
2185 document carries the marker (the user-global file when the ambient config
2186 is workspace-scoped).
2187
2188 The key is never returned — there is no read route for credential material,
2189 and neither the key nor its length appears in the response, errors, or
2190 logs; the response carries only the readiness projection
2191 (`credentialState`). An unknown provider id, the `deepseek-cn` legacy
2192 alias, an empty key, a key over 4 KiB, or one containing control characters
2193 is `400`. `credentialState: "local"` after a successful write is honest
2194 output for a keyless local route: the key is stored, but the route
2195 classifies as not needing one.
2196
2197 ### `DELETE /v1/providers/{id}/key` — clear a Codewhale-owned credential
2198
2199 ```json
2200 // response
2201 { "provider": "openai", "cleared": true, "credentialState": "missing" }
2202 ```
2203
2204 Clears the credential through the same shared owner as
2205 `codewhale auth clear`: the config document is snapshotted and restored if
2206 its save fails, the secret store is only touched once that save has landed,
2207 and the cleared markers are mirrored into the live runtime config so
2208 `GET /v1/providers` reports `missing` on the next read rather than after a
2209 restart.
2210
2211 Clearing an already-clear route returns `cleared: true` — a client retrying
2212 a revoke must not be told something went wrong. If the config entry is
2213 cleared but the secret backend refuses the delete, the route answers `500`
2214 and names the slot: reporting success while the key is still in the keyring
2215 would be a lie about a security action.
2216
2217 ### Credential ownership: `credentialSource` and `credentialWritable`
2218
2219 Both credential verbs refuse a route whose credential Codewhale does not
2220 own, and `GET /v1/providers` carries the same classification so a client can
2221 disable its control *before* submitting instead of failing late:
2222
2223 | `credentialSource` | `credentialWritable` | Meaning |
2224 | --- | --- | --- |
2225 | `secret_store` | `true` | Codewhale's own durable backend. The only writable source. |
2226 | `config` | `false` | A literal key in a config file, which still wins at request time. |
2227 | `external_auth` | `false` | An active external consent (OAuth) owns the credential. |
2228 | `none` | `false` | The route sends no credential, or has no credential slot. |
2229
2230 When `credentialWritable` is `false`, `credentialWritableReason` carries
2231 user-facing copy naming the owner, and both `PUT` and `DELETE` answer `409`
2232 with that same reason. The classification is structural: it reads declared
2233 auth mode, consent state and the *kind* of any configured `api_key` value,
2234 and never resolves a secret, an environment value, or an auth command. It is
2235 a class and never a value, a path, or an environment variable name.
2236
2237 ### `POST /v1/providers/{id}/switch`
2238
2239 ```json
2240 // request (model is optional; omit to take the provider default)
2241 { "model": "qwen3.8-max" }
2242
2243 // response
2244 { "provider": "modelstudio-token-plan", "model": "qwen3.8-max",
2245 "message": "…", "persisted": true }
2246 ```
2247
2248 **Use this rather than simulating a switch with repeated `POST /v1/config`
2249 writes plus a reload.** Provider and model move together here, the change is
2250 validated against the provider's catalog before it is applied, and `persisted`
2251 reports whether it was written to config or applied to the live session only.
2252 Rejects an unknown provider id and the `deepseek-cn` alias with `400`.
2253
2254 ## Runtime data model
2255
2256 The runtime uses a durable Thread/Turn/Item lifecycle.
2257
2258 - **ThreadRecord** — `id`, `created_at`, `updated_at`, `model`,
2259 `model_provider` (generic kind), `model_provider_id` (optional exact configured
2260 route), `workspace`, `mode`, `task_id`, `system_prompt`, `latest_turn_id`,
2261 `latest_response_bookmark`, `archived`
2262 - **TurnRecord** — `id`, `thread_id`, `status` (`queued|in_progress|completed|
2263 failed|interrupted|canceled`), `effective_provider`, `effective_model`,
2264 `effective_billing_surface`, timestamps, duration, usage, error summary,
2265 `artifacts` and `workspace` (see [Turn artifacts](#turn-artifacts))
2266 - **TurnItemRecord** — `id`, `turn_id`, `kind` (`user_message|agent_message|
2267 tool_call|file_change|command_execution|context_compaction|status|error`),
2268 lifecycle `status`, `metadata`, `artifacts` and the legacy
2269 `artifact_refs` projection
2270
2271 Events are append-only with a global monotonic `seq` for replay/resume.
2272
2273 `effective_billing_surface` is a non-secret classification derived from the
2274 endpoint that served the turn. Recognized StepFun routes use `stepfun-payg` or
2275 `stepfun-plan`; unknown and custom endpoints leave it unset. The raw base URL is
2276 not persisted in `TurnRecord`.
2277
2278 ### Restart semantics
2279
2280 - If the process restarts while a turn or item is `queued` or `in_progress`,
2281 the recovered record is marked `interrupted` with an `"Interrupted by
2282 process restart"` error.
2283 - The trailing newline is an event append's commit marker. On startup, a final
2284 JSONL fragment without that delimiter is truncated and fsynced even when its
2285 bytes form valid JSON; it is an uncommitted append, and its already-reserved
2286 sequence number is not reused. Newline-terminated malformed records are not
2287 identifiable crash debris and continue to fail closed during replay.
2288 - If a terminal turn record reached disk but its terminal event sequence did
2289 not, the first async read reconciles any unresolved dynamic calls as
2290 `tool_call.canceled` and then emits one `turn.completed`. Existing terminal
2291 call and turn receipts are detected and never duplicated.
2292 - Turn-operation bindings survive restart. Retrying the same `operation_key`
2293 and request returns the original recovered turn (including an `interrupted`
2294 turn recovered from an in-progress process exit); mismatched reuse remains a
2295 conflict. A crash-created binding that never acquired a turn is discarded at
2296 startup because engine submission happens only after both records are
2297 durable.
2298 - Task execution performs its own recovery on top of the same persisted
2299 thread/turn store.
2300
2301 ### Approval model
2302
2303 - The `auto_approve` flag applies to the runtime approval bridge and engine
2304 tool context. When enabled for a thread/turn/task, approval-required tools
2305 are auto-approved in the non-interactive runtime path, shell safety checks
2306 run in auto-approved mode, and spawned sub-agents inherit that setting.
2307 - When omitted, `auto_approve` defaults to `false`.
2308 - [Authorization order](AUTHORIZATION_ORDER.md) describes where typed rules,
2309 registered tool requirements, safety floors, repository law, approval
2310 transport, and sandbox enforcement sit relative to one another.
2311
2312 ### SSE event stream
2313
2314 The SSE event payload shape for `/v1/threads/{id}/events`:
2315
2316 ```json
2317 {
2318 "schema_version": 1,
2319 "seq": 42,
2320 "previous_seq": 38,
2321 "event": "item.delta",
2322 "kind": "item.delta",
2323 "thread_id": "thr_1234abcd",
2324 "turn_id": "turn_5678efgh",
2325 "item_id": "item_90ab12cd",
2326 "timestamp": "2026-02-11T20:18:49.123Z",
2327 "created_at": "2026-02-11T20:18:49.123Z",
2328 "payload": {
2329 "delta": "partial output",
2330 "kind": "agent_message"
2331 }
2332 }
2333 ```
2334
2335 Compatibility notes:
2336
2337 - `schema_version` is the HTTP/SSE envelope schema version. It is independent of
2338 the runtime store schema used for persisted thread/turn/event records.
2339 - `event` remains the SSE event name in existing clients; it is preserved as-is.
2340 - `kind` mirrors `event` in the stable envelope for typed clients.
2341 - `seq` is allocated globally across all Runtime threads. Consequently, gaps
2342 between a thread's events are normal when other threads interleave. On this
2343 per-thread SSE stream, `previous_seq` is the sequence of the last event
2344 delivered for this thread (or the requested replay cursor for the first
2345 event); clients detect loss by comparing it with their accepted per-thread
2346 cursor, not by requiring `seq == previous_seq + 1`. Sequence allocation is
2347 also not rewound after an append is transactionally rolled back, so a retry
2348 can intentionally skip an unused value without implying a missing event.
2349 - `thread.started`, `turn.started`, and `turn.completed` are emitted as SSE event
2350 names exactly as before.
2351 - `timestamp` remains the canonical event time for schema version 1. `created_at`
2352 is an equivalent alias for clients that use `created_at` naming elsewhere; do
2353 not require both fields to be present.
2354
2355 ### Ending and resuming a thread stream
2356
2357 Whenever the server ends a `/v1/threads/{id}/events` stream that already
2358 returned `200`, the last frame is `stream.end`, sent exactly once and always
2359 (it does not need `progress=true`):
2360
2361 ```
2362 event: stream.end
2363 data: {"schema_version":1,"event":"stream.end","kind":"stream.end","thread_id":"thr_1234abcd","reason":"replay_failed","last_seq":42,"retryable":true}
2364 ```
2365
2366 - It is a transport frame, not a journal event: it has **no `seq`** and **no
2367 SSE `id:`**. Clients that acknowledge on `seq` skip it, and a browser's
2368 `Last-Event-ID` stays on the last real event.
2369 - `last_seq` is the stream's cursor when it ended: the last journal `seq`
2370 delivered on this connection, or the effective start cursor if none was
2371 (after a `replay_limit` tail, that is already past the omitted history). It
2372 is exactly the `since_seq` that resumes with no loss and no repeats.
2373 - `retryable` says whether resuming from `last_seq` can succeed. Key the client
2374 on it, not on the reason list. Treat an unknown `reason` by its `retryable`.
2375 - There is no free-text message. The underlying error is in the Runtime log,
2376 and it can contain store paths.
2377
2378 | `reason` | Meaning | `retryable` |
2379 | --- | --- | --- |
2380 | `replay_failed` | The durable history read that feeds the opening replay failed, including a replay worker crash after the first cursor | `true` |
2381 | `catch_up_failed` | After broadcast lag, the durable re-read from the stream's cursor could not be opened or failed | `true` |
2382 | `runtime_shutdown` | The Runtime API server is stopping (SIGINT, SIGTERM, or SIGHUP; Ctrl+C or Ctrl+Break on Windows). Open streams get this frame within a bounded drain window before the process exits | `true` |
2383
2384 Every `200` response carries `x-codewhale-stream-end: 1`. With that header, an
2385 EOF **without** `stream.end` means the connection or the Runtime process died
2386 without the server choosing to end the stream: a network or proxy drop, a
2387 crash, or a kill that allows no drain (for example `SIGKILL`). A Runtime that
2388 predates this frame sends no header; there EOF stays ambiguous, so treat it as
2389 connection loss.
2390
2391 Client resume rule:
2392
2393 1. Keep `cursor` = the `seq` of the last journal frame you accepted. Ignore
2394 frames with `seq <= cursor`. If a frame's `previous_seq` is not your
2395 `cursor`, you missed events: reload the thread snapshot rather than trust
2396 local state.
2397 2. On `stream.end` with `retryable: true`, reconnect with
2398 `since_seq = last_seq` after a bounded backoff, and tell the user what the
2399 Runtime said (for example, "Runtime is shutting down — reconnecting").
2400 With `retryable: false`, stop, show the reason, and fall back to the
2401 snapshot.
2402 3. On EOF or a transport error without `stream.end`, reconnect with
2403 `since_seq = cursor` after a bounded backoff and show a connection problem,
2404 not a Runtime error.
2405 4. A `401`/`403` or `404` before the stream opens is terminal (a credential or
2406 thread problem). Retry `5xx` with backoff.
2407 5. Never reconnect from `0` to "start over": replay is idempotent only by
2408 cursor.
2409
2410 Fleet streams (`/v1/fleet/runs/{run_id}/events`) keep their own end frames,
2411 `fleet.stream.error {retryable}` and `fleet.replay.cursor_unavailable`. They
2412 differ from `stream.end`: they carry no cursor, because a Fleet client resumes
2413 from the opaque `cursor` of the last Fleet event it accepted.
2414
2415 ### Steer delivery
2416
2417 Putting a steer into the engine's mailbox is not the same as the model reading
2418 it. The engine discards a steer whose turn has already moved on, and an
2419 interrupted or failed turn drops whatever it had queued. The API reports the
2420 engine's real verdict rather than the attempt:
2421
2422 - The item is persisted `queued` when the steer is accepted into the mailbox.
2423 - **Delivered.** The engine committed the text into the turn's record: the item
2424 becomes `completed`, `steer_count` rises, and `turn.steered` + `item.completed`
2425 are emitted. `POST .../steer` returns `200` with that turn.
2426 - **Not delivered.** The turn moved on, was interrupted, or failed first: the
2427 item becomes `canceled`, `steer_count` does not rise, and `turn.steer_dropped`
2428 is emitted carrying `input`, `reason`, and the settled `item`. `POST .../steer`
2429 returns `409`, so a client can keep the user's text and resend it rather than
2430 clearing a composer over guidance that was never seen.
2431 - **Still pending.** A steer sent while the engine is inside a long tool call
2432 cannot settle until that call returns, and the request does not hang for it.
2433 After a short wait `POST .../steer` returns `200` with the item still `queued`;
2434 the eventual `turn.steered` or `turn.steer_dropped` event carries the verdict.
2435
2436 A client that treats `200` as "the model saw it" is therefore wrong in the third
2437 case: read the item's status, or wait for the event.
2438
2439 Common event names: `thread.started`, `thread.forked`, `turn.started`,
2440 `turn.lifecycle`, `turn.steered`, `turn.steer_dropped`, `turn.interrupt_requested`,
2441 `turn.completed`, `turn.artifacts`, `item.started`, `item.delta`, `item.completed`,
2442 `item.failed`, `item.interrupted`, `approval.required`, `approval.decided`,
2443 `approval.timeout`, `user_input.required`, `user_input.answered`,
2444 `user_input.canceled`, `tool_call.requested`, `tool_call.resolved`,
2445 `tool_call.timeout`, `tool_call.canceled`, `sandbox.denied`,
2446 `turn.workspace_snapshot`, `runtime.store_failure`.
2447
2448 `runtime.store_failure` is the runtime reporting a fault in the operator's own
2449 on-disk state: a thread, turn, or item record under the session's runtime
2450 store could not be read, parsed, or written. The payload carries `operation`
2451 (`read` | `parse` | `write`), `record_kind` (`thread` | `turn` | `item`),
2452 `record_id`, `path`, the full `error` chain, the root-cause `reason`, a
2453 `next_action` (which file to move aside, or where to check free space and
2454 permissions), and a one-line `message`. When `terminal` is `true`, the turn's
2455 own record is unreadable or unwritable and no `turn.completed` will follow;
2456 clients waiting on that turn should treat it as failed.
2457
2458 Agent-message and reasoning deltas are materialized into the item projection
2459 before their corresponding `item.delta` event is sequenced. To avoid an fsync
2460 for every provider fragment, adjacent deltas are coalesced to configured bounds
2461 of at most 32 ms or approximately 16 KiB before publication (an indivisible
2462 upstream chunk can itself exceed the byte target). A process crash inside that
2463 unpublished window can lose the recent suffix; no durable event claims that
2464 suffix existed. Once an `item.delta` is durable, snapshots at or beyond its
2465 cursor include the same materialized prefix.
2466
2467 `approval.required` events may include a `matched_rule` string when an
2468 execution-policy rule caused the prompt. This field is explanatory metadata for
2469 clients and does not grant or persist permissions.
2470
2471 `approval.required`, `approval.decided`, and `approval.timeout` carry two
2472 distinct identifiers. `approval_id` is the Runtime-minted, single-use capability
2473 described under **Approvals** — the only value `POST /v1/approvals/{id}` accepts
2474 — and `approval.required` also repeats it in the legacy `id` field for older
2475 clients. `tool_call_id` is the provider's raw tool-call ID, present for
2476 correlation only. Automatically resolved prompts (thread `auto_approve`, and the
2477 Auto-Review posture, which never opens a modal) mint an `approval_id` as well, so
2478 the field has one meaning on every path; those IDs register no waiter and are
2479 inert against the endpoint. Clients must never treat `tool_call_id` as an
2480 approval capability or assume it is unique across threads.
2481
2482 The thread event stream forwards these payloads intact. The compatibility turn
2483 stream carries `approval_id`, its `id` alias and `tool_call_id`; the pending
2484 snapshot carries the same capability and correlator so reconnecting clients can
2485 attach an approval prompt to its tool row.
2486
2487 ## Security boundary
2488
2489 - **Localhost by default**. The server binds to `127.0.0.1` by default.
2490 `--mobile` is also loopback-only and rejects a non-loopback host until a TLS
2491 or verified-overlay transport boundary exists. The runtime does not provide
2492 user isolation or TLS.
2493 - **Optional token guard**. `--auth-token` or `DEEPSEEK_RUNTIME_TOKEN`
2494 requires a matching bearer token for `/v1/*` routes. This is a local
2495 convenience guard, not a replacement for TLS, VPN, or a trusted reverse
2496 proxy on public networks.
2497 - **No provider-token custody**. The server never returns the API key. The
2498 `api_key.source` capability field reports `env`, `config`, or `missing` —
2499 never the key itself.
2500 - **No hosted relay**. The app-server is a local process under the user's
2501 control. There is no cloud component.
2502 - **Capability responses** never leak secrets, file contents, or session
2503 message bodies. They report *metadata*: presence, counts, status flags.
2504
2505 ### CORS allow-list
2506
2507 The runtime API ships with a built-in dev-origin allow-list:
2508 `http://localhost:3000`, `http://127.0.0.1:3000`, `http://localhost:1420`,
2509 `http://127.0.0.1:1420`, `tauri://localhost`. To add additional origins (e.g.
2510 when developing a UI on Vite's default `:5173`), use any of:
2511
2512 - CLI flag (repeatable): `codewhale serve --http --cors-origin http://localhost:5173`
2513 - Env var (comma-separated): `DEEPSEEK_CORS_ORIGINS="http://localhost:5173,http://localhost:8080"`
2514 - Config (`~/.codewhale/config.toml`):
2515 ```toml
2516 [runtime_api]
2517 cors_origins = ["http://localhost:5173"]
2518 ```
2519
2520 User-supplied origins **stack on top of** the built-in defaults; they do not
2521 replace them. Wildcard origins are not supported — the explicit allow-list
2522 model is preserved. Cross-origin preflights advertise only `Authorization`,
2523 `Content-Type`, `Accept`, `X-Codewhale-Runtime-Token`, and the compatibility
2524 `X-DeepSeek-Runtime-Token` request header; custom request headers are not
2525 allowed. Added in v0.8.10 (#561), tightened in v0.9.1 (#4454).
2526
2527 ## Managed Fleet Runtime and SDK helpers
2528
2529 The Runtime SDK lives in `npm/runtime-sdk` and is exposed as
2530 the `@codewhale/runtime-sdk` workspace package. It is deliberately thin: every
2531 helper calls the local Rust Runtime API and therefore cannot bypass Codewhale's
2532 sandbox, approval prompts, provider configuration, or fleet ledger authority.
2533
2534 ```js
2535 import { createRuntimeClient } from "@codewhale/runtime-sdk";
2536
2537 const client = createRuntimeClient({
2538 baseUrl: "http://127.0.0.1:7878",
2539 token: process.env.CODEWHALE_RUNTIME_TOKEN,
2540 });
2541
2542 const created = await client.createFleetRun({
2543 target: "this_computer",
2544 roles: [{ name: "reviewer" }, { name: "verifier" }],
2545 workflow: {
2546 id: "release-check",
2547 kind: "parallel",
2548 tasks: [
2549 { id: "review", name: "Review", instructions: "Review locally.", worker: { role: "reviewer" } },
2550 { id: "verify", name: "Verify", instructions: "Verify locally.", worker: { role: "verifier" } },
2551 ],
2552 },
2553 });
2554
2555 // POST /runs only prepares durable work. This call crosses the launch gate.
2556 await client.startFleetRun(created.run.id);
2557
2558 let cursor;
2559 for await (const event of client.fleetEvents(created.run.id, { after: cursor })) {
2560 if (event.cursor) cursor = event.cursor;
2561 if (event.event === "fleet.replay.cursor_unavailable") {
2562 // Reload getFleetRun(created.run.id), then reconnect without the old cursor.
2563 }
2564 }
2565 ```
2566
2567 The managed path is deliberately two-step. `POST /v1/fleet/runs` validates and
2568 persists the run and queue without starting a worker. A separate authenticated
2569 `POST /start` activates it and schedules the executor driver; its `202` response
2570 reports `leased: 0` because the driver performs all leasing after it owns the
2571 run. Creation requires named roles, one task owner per role, a `parallel`
2572 Workflow, and an explicit Runtime target. v0.9.4 executes
2573 only `this_computer`; `another_computer` and `cloud` return `501` rather than
2574 silently executing locally. Worker IDs are generated per run; caller-assigned
2575 `worker_specs` return `501` until custom workers can be given collision-free
2576 managed identities. Parallel tasks with overlapping effective write roots are
2577 rejected before the run is journaled. Managed `security_policy` overrides also
2578 fail closed until that document can be enforced end to end; executable
2579 authority comes from each named role's tool posture and bounded task workspace
2580 scope.
2581
2582 Fleet helpers cover this HTTP surface:
2583
2584 | Helper | Runtime API route |
2585 |---|---|
2586 | `createFleetRun(spec)` | `POST /v1/fleet/runs` |
2587 | `startFleetRun(runId)` | `POST /v1/fleet/runs/{run_id}/start` |
2588 | `listFleetRuns()` | `GET /v1/fleet/runs` |
2589 | `getFleetRun(runId)` | `GET /v1/fleet/runs/{run_id}` |
2590 | `listFleetWorkers(runId)` | `GET /v1/fleet/runs/{run_id}/workers` |
2591 | `getFleetWorker(workerId)` | `GET /v1/fleet/workers/{worker_id}` |
2592 | `interruptWorker(workerId)` | `POST /v1/fleet/workers/{worker_id}/interrupt` |
2593 | `stopWorker(workerId)` | `POST /v1/fleet/workers/{worker_id}/stop` |
2594 | `restartWorker(workerId)` | `POST /v1/fleet/workers/{worker_id}/restart` |
2595 | `stopFleetRun(runId)` | `POST /v1/fleet/runs/{run_id}/stop` |
2596 | `replayFleetEvents(runId, options)` | `GET /v1/fleet/runs/{run_id}/events/replay` |
2597 | `fleetEvents(runId, options)` | `GET /v1/fleet/runs/{run_id}/events` (SSE) |
2598
2599 `stopWorker` durably cancels that worker's active task and leaves the rest of
2600 the Fleet running. `interruptWorker` is the compatibility name for the same
2601 attempt-fenced cancellation transition. `stopFleetRun` cancels every queued or
2602 active task and marks the whole run cancelled.
2603
2604 Replay covers aggregate run/task transitions and privacy-bounded individual
2605 worker transitions. Event bodies omit prompts, tool call IDs, completion text,
2606 artifact paths/checksums, and cancellation identities; bounded failure reasons
2607 pass through secret redaction. `cursor` is opaque and stable across ordinary
2608 appends and Runtime restarts. Clients reconnect with `after=<cursor>`. A fresh
2609 request returns a bounded newest tail and marks `history_truncated` when older
2610 history exists. Ledger compaction can remove an old cursor; the JSON endpoint
2611 then returns `409`, while the SSE endpoint emits
2612 `fleet.replay.cursor_unavailable`, so the client reloads the current run
2613 projection instead of accepting a silent gap.
2614
2615 `GET /v1/runtime/info` advertises `fleet_run_create`, `fleet_run_start`,
2616 `fleet_event_replay`, `fleet_event_stream`, and `fleet_local_target`. Older
2617 runtimes without a requested route still produce a typed SDK
2618 `RuntimeCapabilityError`.
2619
2620 Verification:
2621
2622 ```bash
2623 npm test --workspace @codewhale/runtime-sdk
2624 ```
2625
2626 ## Agent Run Receipts
2627
2628 Sub-agent lanes persist compact run receipts in
2629 `.codewhale/state/subagents.v1.json`. The Runtime API exposes those receipts as
2630 a read-only inspection surface:
2631
2632 | Operation | Endpoint |
2633 |---|---|
2634 | List persisted agent runs | `GET /v1/agent-runs` |
2635 | Inspect one run | `GET /v1/agent-runs/{run_id}` |
2636 | Stop one run | `POST /v1/agent-runs/{run_id}/cancel` |
2637
2638 The response is the same worker-record shape surfaced by `agent` receipts:
2639 `spec.run_id`, `actor_kind`, lifecycle `status`, bounded `events`,
2640 `follow_up`, `takeover`, `artifacts`, `usage`, and `verification`. `run_id`
2641 falls back to the worker id for older records, and `{run_id}` may be either the
2642 run id or the worker id.
2643
2644 These endpoints do not start or steer sub-agents. The API surface exists so
2645 app/editor/headless clients can inspect the same handoff receipts that the TUI
2646 and parent model see, and stop a run they are showing.
2647
2648 `POST /v1/agent-runs/{run_id}/cancel` takes no body. It stops the run through
2649 the same session-scoped path as the TUI's stop and the `agent/cancel` tool:
2650 descendants stop with it, and a write-scoped child's changed files are named in
2651 its result rather than dropped. It answers with the worker record:
2652
2653 - `200` when the record is terminal (stopping an already-finished run is a
2654 no-op that returns its receipt);
2655 - `202` when the owning engine accepted the stop but has not recorded the
2656 terminal receipt within a few seconds; poll `GET /v1/agent-runs/{run_id}`;
2657 - `404` for an unknown run;
2658 - `409` when the run belongs to a session this runtime is not hosting (for
2659 example a separate terminal session); stop it from that session.
2660
2661 ## Session lifecycle (native UI supervision)
2662
2663 | Operation | Endpoint |
2664 |---|---|
2665 | List sessions | `GET /v1/sessions` |
2666 | List session summaries | `GET /v1/sessions/summary` |
2667 | Get session | `GET /v1/sessions/{id}` |
2668 | Rename / archive session | `PATCH /v1/sessions/{id}` |
2669 | Delete session | `DELETE /v1/sessions/{id}` |
2670 | Session store repair summary | `GET /v1/sessions/repair` |
2671 | Resume into thread | `POST /v1/sessions/{id}/resume-thread` |
2672 | Create thread | `POST /v1/threads` |
2673 | List threads | `GET /v1/threads` |
2674 | Attach to events | `GET /v1/threads/{id}/events?since_seq=0` |
2675 | Send message | `POST /v1/threads/{id}/turns` |
2676 | Steer | `POST /v1/threads/{id}/turns/{turn_id}/steer` |
2677 | Interrupt | `POST /v1/threads/{id}/turns/{turn_id}/interrupt` |
2678 | Compact | `POST /v1/threads/{id}/compact` |
2679
2680 ## Compatibility tests
2681
2682 Contract snapshots live in `crates/protocol/tests/`. Run:
2683
2684 ```bash
2685 cargo test -p codewhale-protocol --test parity_protocol --locked
2686 ```
2687
2688 This validates that the app-server's event schema hasn't drifted from the
2689 documented contract. CI runs this on every push to `main` and on release tags.
2690
2691 The app-server stdio control surface has its own drift guard — the advertised
2692 `capabilities` method set is pinned in `crates/app-server/src/lib.rs`:
2693
2694 ```bash
2695 cargo test -p codewhale-app-server capabilities
2696 ```
2697
2698 Before a release, run the headless smoke (stdio probe + optional provider
2699 matrix, no secrets leaked):
2700
2701 ```bash
2702 scripts/release/app-server-smoke.sh --matrix # dry-run plan
2703 bash scripts/release/app-server-smoke.test.sh # parser self-test (fake binary)
2704 ```
2705
2705 lines MARKDOWN