返回 CodeWhale
TUI_DECONSTRUCTION.md
根目录 / docs / design / TUI_DECONSTRUCTION.md
1 # TUI deconstruction
2
3 The deliverable is an independently buildable headless runtime and a terminal
4 client of that runtime. Preserve working behavior while moving ownership out
5 of `codewhale-tui`. A lower line count alone does not establish the split.
6
7 ## Audited baseline, 2026-09-09
8
9 Source inspection and offline Cargo metadata at `ce737266683b` found:
10
11 | Source | Physical Rust lines |
12 | --- | ---: |
13 | `crates/tui/src` | 971,321 in 817 files |
14 | `crates/tui/src/tui` | 266,378 |
15 | `crates/tui/src/tools` | 153,713 |
16 | `crates/tui/src/core` | 58,378 |
17 | `crates/tui/src/commands` | 60,951 |
18 | `crates/tui/src/lib.rs` | 20,327 |
19 | All of `crates/core/src` | 5,475 |
20
21 Counts include comments, blank lines, tests, and source files that may not be
22 compiled. Dedicated test-source files account for 191,058 lines within the
23 TUI total; additional inline tests remain in other files. The directory named
24 `tui` also contains domain logic. These are ownership clues, not production
25 LOC, a complete compiler dependency graph, or a language-port estimate.
26
27 The current dependencies explain the blockage:
28
29 - CLI imports TUI for runtime dispatch and route preferences.
30 - `core/engine.rs` imports approval policy, context thresholds, attachment
31 parsing, and roster construction from `tui/`.
32 - `core/events.rs` carries the roster row and `session_manager.rs` persists
33 the durable context reference; both now name their owning crate
34 (`crate::agent_roster::AgentRosterRow`, `codewhale_core::ContextReference`)
35 rather than a `tui::` re-export.
36 - `tools/subagent` imports engine policy/catalog functions, and implements
37 its own repeated model-request/tool-result cycle in `run_subagent`.
38 - `crates/core` owns request construction and some runtime/session services;
39 the main `Engine::run_turn` remains inside the TUI crate. The comment in
40 `tui/src/core/mod.rs` claiming the engine has moved is incorrect.
41 - `core/protocol_parity.rs` exhaustively projects internal operations/events,
42 but explicitly has no production consumers. Reuse or retire it during the
43 migration; its existence does not establish client convergence.
44
45 `single_turn_loop.rs` currently counts functions named `run_turn`. It does
46 not detect `run_subagent`'s execution cycle. A passing name scan is therefore
47 insufficient evidence of one execution implementation.
48
49 ## Runtime split sequence (2026-09-25, supersedes the order below)
50
51 The build-graph split runs first; the protocol-client contract follows it.
52 Nothing here is a claim that a step has landed: `git log` and the ratchet
53 (`scripts/runtime-boundary-baseline.json`) are the record.
54
55 1. **Rails.** `scripts/split/module_graph.py` computes the runtime closure
56 (everything reachable from `core`, `tools`, `runtime_api`,
57 `runtime_threads`, `client`, `llm_client`, `config` and
58 `session_manager` without passing through `tui`, `commands`,
59 `remote_control`, `context_report`, `composer_*` or `lib.rs`) and counts
60 every reference from it into UI code, UI libraries, modules that move
61 later, and UI intra-doc links. `scripts/check-command-crate-boundaries.py`
62 runs it in CI: counts may only go down (RS-0).
63 2. **Create `crates/runtime` (`codewhale-runtime`) with live code.** The
64 first batch is the leaf modules that reference nothing outside the batch
65 in production or test code and touch no UI library. The crate is
66 publishable, because the published `codewhale-tui` depends on it (RS-2).
67 3. **Cut the upward edges, one blocker per slice**, each lowering the
68 ratchet: palette's `ratatui` dependency becomes an optional default
69 feature (RS-3); the one `host_terminal` port carries every terminal side
70 effect the runtime needs, starting with raw-mode suspension around
71 interactive children (RS-4); voice capture splits from its slash command
72 (RS-5); the context-window formatter moves to `utils` (RS-6);
73 notification payload and sound policy move down while delivery (OSC
74 writes, taskbar, title) stays in the TUI behind the port (RS-7); then the
75 engine's terminal-chrome calls, auto-review and risk policy into
76 `core::authority`, the remaining engine leaks, the command catalog, the
77 `lib.rs` helpers, and the test-only references.
78 4. **Move the strongly connected core in one rename-only change** (about 64
79 modules; a crate cannot hold half a cycle), after landing its visibility
80 and rustdoc edits in place. Then `runtime_api` and the modules above the
81 core, then the headless surfaces, then delete the path alias.
82 5. **Then the client contract** (behavioral, tracked separately): one engine
83 owner in `runtime_threads`, runtime-owned queue and steer, approvals as
84 server requests, a `runtime-client` facade, one protocol method table,
85 and loop convergence.
86
87 `crates/runtime` must never depend on `codewhale-tui`, `codewhale-cli`,
88 ratatui, crossterm, `ansi-to-tui` or a terminal component kit. The TUI is the
89 one terminal-output owner: the runtime reaches the terminal only through
90 `host_terminal`, which the composition root installs for every host it
91 launches today. Withholding it from stdio hosts (ACP, MCP server,
92 app-server) is a later one-line change, not a code move.
93
94 ### Where this sequence departs from the rest of this document
95
96 These three departures are deliberate; each is safe for the stated reason.
97
98 1. **The config hub does not have to leave first.** Once the upward edges are
99 cut, `config` sits inside the runtime's strongly connected component, so
100 it moves with it. Splitting config into `codewhale-config` (#6034, #6143)
101 becomes internal runtime work afterwards instead of a precondition.
102 2. **Converging the two loops is not a precondition for moving the engine.**
103 That rule assumed the engine would move while `run_subagent` stayed in the
104 TUI, leaving two loops in two crates. Here both move into the same crate,
105 so convergence stays a client-contract step and the one-loop guard keeps
106 scanning every crate.
107 3. **The move uses one path alias**, a single root
108 `use codewhale_runtime::{...};` block in the TUI `lib.rs`, instead of
109 rewriting every caller in the moving change. A rename-only move lets
110 in-flight branches rebase across it; a content rewrite of every caller
111 conflicts with all of them. The alias is the only shim: no wrapper types,
112 no per-item re-exports except the ones `crates/cli` needs until the
113 headless surfaces move, and it is deleted by a published rewrite script
114 once the moves are done.
115
116 The destination below still describes the intended layering, with one
117 change: the engine, turn loop and tools move into `codewhale-runtime`
118 together (there is no separate `codewhale-engine` crate), because the
119 engine, tools, config and client form one dependency cycle today.
120 `crates/core` stays the lower request-construction layer:
121 `codewhale-command-contract` depends on it, and the runtime needs the
122 contract, so putting the runtime into `core` would be a Cargo cycle.
123
124 ## Intended ownership
125
126 This is the proposed destination, not a claim that the boundaries exist now.
127 Reuse existing crates; introduce only the three cohesive runtime libraries
128 below, with real consumers and all replaced paths migrated in each slice.
129
130 | Owner | Responsibility |
131 | --- | --- |
132 | `codewhale-tui` | Terminal lifecycle, rendering, input, pickers, terminal command presentation. No provider I/O, policy decisions, durable store, or agent loop. |
133 | `codewhale-cli` | Argument parsing, launch/composition, headless command presentation. Existing binary names remain compatible. |
134 | `codewhale-app-server` | HTTP/SSE and stdio transport adapters over the same runtime. Reconcile the embedded Runtime API and existing app-server; preserve external routes and auth. |
135 | New `codewhale-runtime` | Session/thread lifecycle, scheduling, recovery, child supervision, and composition of engine and stores. No model/tool execution loop. Used in process by terminal and server hosts. |
136 | New `codewhale-engine` | The shared parent/child execution implementation, context/compaction, tool dispatch, cancellation, approvals, and typed events. |
137 | New `codewhale-models` | Provider clients, live catalog/pricing resolution, and model routing against canonical config facts. Consolidate existing `agent` catalog consumers instead of retaining a second seeded registry. |
138 | Existing `config`, `secrets`, `execpolicy` | Canonical schema/route identity, credential storage/access, and policy decisions. UI labels stay outside these owners. |
139 | Existing `tools`, `mcp`, `hooks` | Tool contracts and implementations, extension transports, and hook execution. Agent/task tools call runtime capabilities; they do not own another agent loop. |
140 | Existing `state`, `protocol`, `core` | Persistence, shared wire/domain records, and request construction. Move the existing core runtime service owner into the runtime library as its callers migrate. |
141
142 Dependency direction: terminal/server -> runtime -> engine -> provider/tool
143 implementations and shared lower-level crates. Tools must not import the
144 concrete engine or runtime host. Use narrow service capabilities at the
145 composition boundary where a tool needs scheduling or agent control; do not
146 introduce a generic service-locator framework or a trait for every helper.
147
148 The TUI can retain in-process channels through the existing `EngineHandle`,
149 `Op`, and `Event` seams. HTTP remains a transport for external clients, not a
150 mandatory hop for local terminal use. Keep wire DTOs separate from internal
151 operations containing reply channels or resolved capabilities.
152
153 ## Model judgment and test-time compute
154
155 Founder clarification, September 9: the harness should let the model decide
156 when a goal, plan, delegation, further investigation, or verification is useful.
157 Provide enough reasoning and tool-feedback opportunities for that judgment.
158 Do not interpret unwanted automatic goals as a request to forbid inferred goals.
159
160 The current goal path has conflicting authorities: `operate_goal_from_prompt`
161 classifies an instruction with verb/question heuristics before inference, while
162 `CreateGoalTool::description` tells the model to require explicit goal requests.
163 `runtime_handoff` additionally tells the model the host already created a goal.
164 Those three policies must become one model-facing contract with runtime-owned
165 state transitions. Explicit `/goal` commands remain a direct user control.
166
167 Reference inspection was local, not a claim about every upstream version:
168
169 | Snapshot | Useful evidence |
170 | --- | --- |
171 | Codex `45eec73b11` (2026-09-09) | `ext/goal/src/spec.rs` leaves goal tool selection to the model but instructs explicit user/system intent. Base instructions let the model choose when planning helps. Goal state and continuation live outside the terminal renderer. |
172 | Kimi Code `1414d4602` (2026-08-13; older snapshot) | `agent-core` exposes `CreateGoal` with completion criteria and a separate reusable turn loop. Creation guidance accepts explicit autonomous-outcome requests or host goal intake. Thinking effort is mapped against model capabilities. |
173 | DSH `c389f96bf3` (2026-09-08) | `goal/tool-goal` explicitly permits inferring a long-running objective from a direct human request. Execution validates top-level human-turn provenance and exact state revisions. Goal state, goal tools, and continuation scheduling are separate consumers. |
174
175 DSH most directly matches the requested goal discretion. Its runtime validates
176 who may mutate state; the model judges whether persistence benefits the task.
177 Codewhale should preserve that distinction without copying DSH's package count.
178
179 Implementation packet, to coordinate separately from mechanical extraction:
180
181 1. Remove host-side semantic goal classification. Present the request, session
182 state, tools, and existing goal to the model before deciding on persistence.
183 2. Revise the existing goal-tool and Operate guidance together: infer a goal
184 when the requested outcome warrants durable continuation and has a useful
185 completion criterion; answer, investigate, or perform ordinary multi-step
186 work without a goal when that suffices. Honor corrections and opt-outs.
187 Explicit user controls and model actions use the same goal state owner.
188 3. Treat test-time compute as reasoning effort, useful tool-feedback rounds,
189 and evidence-driven revision. `auto_reasoning::select` currently chooses
190 effort using message keywords; this is another semantic heuristic to
191 replace. Preserve explicit route/effort choices. Let the lead allocate
192 supported effort and execution budgets to work, with additional effort
193 requested for later steps when new evidence makes that useful. Reuse
194 `RequestTuning` and existing runtime/tool contracts; do not add an
195 always-on classifier or a second agent loop ahead of every prompt.
196 4. Keep accounting, supported provider limits, permission checks, input
197 provenance, durable state, and cancellation in Rust. Emit current state,
198 remaining authorized resources, and tool/test results as compact feedback.
199 A goal does not grant new spend or execution authority. Preserve the pinned
200 prefix and append changing feedback to history.
201 5. Qualify judgment with model-driven sessions, not only deterministic mocks.
202 Compare matched tasks at explicit effort/resource settings: a greeting,
203 architecture discussion, one-file repair, large migration, unrelated followup,
204 mid-run correction, false success evidence, cancellation, and repeated
205 failure. Judge objective quality, useful continuation, task completion,
206 verification quality, latency, tokens/cost, and correct stopping. Do not
207 score a run better merely for creating a goal or taking more steps.
208
209 Start with the existing model's ordinary reasoning/tool loop. Add independent
210 review or multiple candidate attempts only where measured failures and task
211 stakes justify the extra compute. No model-evaluation runs, provider spend, or
212 performance gains were established by this source audit.
213
214 ## Implementation order (before 2026-09-25)
215
216 The runtime split sequence above replaces this order where they differ.
217 Every packet names the predecessor, all consumers, changed dependency edges,
218 and its verification. One owner handles shared manifests and integration.
219 Keep unrelated active work intact; follow the current workspace authority.
220
221 1. **Remove upward domain dependencies.** Finish the existing `AppMode` and
222 `ApprovalMode` migration by pointing runtime consumers at their actual
223 `config`/`execpolicy` owners. Move durable context-reference records out of
224 file-mention UI; retain composer completion there. Separate worker receipt
225 data from roster glyphs/layout. Move reasoning preference and approval
226 policy out of UI modules, preserving exact route/credential identity.
227 `ApiProvider` and `ProviderKind` currently differ for legacy table identity;
228 do not replace one with the other through a lossy cast.
229 2. **Extract provider and tool foundations by cohesive subsystem.** Consolidate
230 config schema and catalog facts as each affected consumer migrates. Move
231 provider adapters with their tests into `codewhale-models`; reuse the
232 existing model-client seam. Grow `tools`, `mcp`, `hooks`, and `state` in
233 place. Move ordinary file/shell/MCP capabilities first. Leave agent
234 orchestration with the execution owner until step 3; moving all of
235 `tools/subagent` into a leaf tool crate would preserve a dependency cycle.
236 3. **Converge parent and child execution.** Inventory and preserve child
237 budgets, route pins, permissions, tool activation, steering, parking,
238 checkpoints, nested work, and terminal fan-in. Adapt children to the
239 existing engine, then remove the old child model/tool cycle. Use actual
240 parent/child call-path and behavior evidence; the function-name guard
241 alone is not acceptance. Do not couple this semantic migration with the
242 mechanical engine file move.
243 4. **Move the engine and shared host.** Once the runtime-to-UI dependencies
244 are gone, move the existing execution implementation into
245 `codewhale-engine`, with its owning unit tests. Establish one runtime host
246 for terminal, exec, server, scheduling, and recovery. Migrate the existing
247 core runtime and thread-manager consumers fully, preserving on-disk
248 formats, replay cursors, locks, authority, and exact-once terminal events.
249 No second runtime store or speculative replacement turn loop.
250 5. **Finish the clients.** Fold the embedded HTTP API into the existing
251 app-server transport surface over `codewhale-runtime`. Move argument parsing
252 and headless command presentation out of TUI `lib.rs` into CLI. Slash
253 commands retain presentation in TUI and call the same runtime operations.
254 Prune dependencies, temporary re-exports, and obsolete implementations.
255 Only then resize remaining UI files according to actual responsibilities.
256
257 The first bounded source packet is step 1's existing mode imports and durable
258 context-reference records, including every caller. Subsequent packets are
259 chosen from the remaining dependency graph, not from a target crate count.
260
261 Moving a definition and repointing its internal callers belong to one complete
262 packet. Mechanical movement and behavioral changes should remain separately
263 reviewable, but do not land temporary re-export shims with their last consumers
264 left for an unspecified future migration. Keep shims only for genuine external
265 compatibility contracts and identify that contract.
266
267 ## Verification and completion
268
269 - Preserve the model-facing runtime receipt, prompt/cache-prefix semantics,
270 tool names/order, serialized records, and public commands for mechanical
271 moves. A deliberate behavior fix names and tests the intended difference.
272 - Move private unit tests with their implementation. A `#[path]` module split
273 remains in the same compilation unit; it does not reduce the test binary or
274 establish faster builds. Do not make internals public just to relocate tests.
275 - Exercise local mock-provider parent and child turns, tool approval and
276 denial, streaming, cancel/steer, explicit goals, pause/resume, reconnect,
277 restart recovery, and terminal fan-in at the affected boundaries.
278 - Goal persistence is independent of Plan/Act/Operate. Let the model decide
279 when persistent tracking benefits the requested work, using context and
280 adequate reasoning time. Remove the host verb heuristic; do not replace it
281 with a blanket explicit-command-only restriction. Explicit user opt-outs,
282 cancellation, and authorized resource limits remain binding.
283 - The headless runtime and its tests must build with **no transitive dependency
284 on `codewhale-tui`, ratatui, or crossterm**. Desktop and terminal consume the
285 same lifecycle and execution authority.
286 - Measure warmed edit/build/test cycles and peak memory for a provider edit,
287 tool edit, terminal-renderer edit, and locale edit before and after. Record
288 compiler/profile/features/cache state. A leaf change must not recompile the
289 unrelated TUI library test unit to run that leaf's own tests. Relinking a
290 final application is a separate cost; no unmeasured speedup promises.
291 - Use existing `scripts/dev-test.sh` / `scripts/dev-cargo.sh` and focused checks
292 during packets. Integration uses the required repository gates with actual
293 counts. Local tests, full gates, hosted CI, installed artifacts, and PTY
294 behavior remain separate evidence. Follow the workspace's human gates for
295 publication, deploys, and spend.
296
297 `docs/BUILD_PERFORMANCE.md` retains historical measurements. Its older B3 and
298 micro-crate candidate lists are superseded by this dependency-led sequence.
299 The old all-at-once preconditions, test-file-move speed claim, and mandatory
300 uncompleted two-PR shim sequence are retired. A paused migration reports the
301 remaining monolith and unresolved consumers explicitly.
302
302 lines MARKDOWN