返回 CodeWhale
TOOL_SURFACE.md
根目录 / docs / TOOL_SURFACE.md
1 # Tool surface
2
3 > 阅读简体中文版:[zh_hans/TOOL_SURFACE.md](zh_hans/TOOL_SURFACE.md)。
4
5 This document describes the current model-facing tool contract. The v0.9.1
6 cutover that produced it is recorded in `docs/RUNTIME_SIMPLIFICATION_DESIGN.md`;
7 read the workspace version from `Cargo.toml`, not from this line. The registry
8 remains larger than the first-turn catalog so
9 saved transcripts can replay and uncommon capabilities can be loaded on demand.
10 The model should learn one canonical name for each common operation.
11
12 Implementation sources:
13
14 - `crates/tui/src/core/engine/tool_catalog.rs` owns the eager/deferred catalog.
15 - `crates/tui/src/tools/registry.rs` registers canonical tools and hidden aliases.
16 - `crates/tui/src/tools/{file,file_tool,shell}.rs` own the small foreground
17 primitive behavior and schemas; the other native tools remain searchable.
18 - `docs/RUNTIME_SIMPLIFICATION_DESIGN.md` records the v0.9.1 cutover and receipt.
19
20 ## Default-active contract
21
22 New turns start with eleven eager native names plus synthetic `tool_search`:
23
24 1. `read`
25 2. `write`
26 3. `edit`
27 4. `bash`
28 5. `agent`
29 6. `workflow`
30 7. `todo_write`
31 8. `create_goal`
32 9. `get_goal`
33 10. `update_goal`
34 11. `load_skill`
35 12. `tool_search` (synthetic, always active)
36
37 The eleven native names are `DEFAULT_ACTIVE_NATIVE_TOOLS` in
38 `crates/tui/src/core/engine/tool_catalog.rs`, pinned by
39 `default_active_contract_keeps_discovery_and_core_tools_eager`. An authority
40 boundary may remove `agent` at the maximum child depth, but route size alone
41 must not change this core vocabulary.
42
43 The direct schemas deliberately stay small:
44
45 | Tool | Input | Purpose |
46 |---|---|---|
47 | `read` | `path`, optional `offset`, optional `limit` | Read a bounded file window with explicit continuation or truncation notices. |
48 | `write` | `path`, `content` | Create or replace a file. |
49 | `edit` | `path`, `edits` | Apply one or more unambiguous text replacements against one original snapshot. |
50 | `bash` | `command`, optional `timeout` | Run one cancellable foreground shell command and return a bounded tail. |
51 | `agent` | delegated task and optional scope/context controls | Start or inspect focused child work. |
52 | `workflow` | plan/script/source_path plus run controls | Coordinate multi-agent phases with dependencies and completion checks. |
53 | `todo_write` | complete replacement list of `{content, status}` items | Keep optional, agent-owned progress notes for genuinely multi-step work. |
54 | `create_goal` | objective plus optional budget | Start the session goal the turn works toward. |
55 | `get_goal` | none | Read the active goal and its progress. |
56 | `update_goal` | terminal status | Mark the goal complete or blocked. |
57 | `tool_search` | `query`, optional matching controls | Discover policy-allowed deferred tools and add selected schemas to this conversation's toolbox. |
58
59 Mode is an authority decision, not a synonym system. Plan, Work, and Operate
60 use the same primitive identities. Plan centrally refuses `write`, `edit`, and
61 `bash`; Work and Operate still pass those calls through approval, sandbox,
62 trusted-path, repository-law, and managed-policy gates. Full Access changes
63 ordinary approval behavior but does not bypass hard safety or repository law.
64
65 `update_plan` remains registered only for saved-artifact compatibility and is
66 not model-visible. `tasks`, `Git`, `Run`, `Web`, `remember`, and other
67 specialized capabilities are searchable rather than first-turn ceremony.
68
69 ## Local code execution
70
71 `code_execution` (Python) and `js_execution` (Node.js) use the same
72 permission-aware local launcher as workspace task gates and test runs. Ordinary
73 approval retains the effective session policy. To retry a denied exact call
74 with wider permissions, supply `sandbox_permissions` (`workspace-write` or
75 `danger-full-access`) and a nonempty `justification`; Ask mode requires explicit
76 one-shot user approval. The approved policy affects only that call.
77
78 Both tools read source from stdin, return stdout/stderr/return_code, and keep
79 the existing timeout and process-tree cleanup. They are not persistent REPLs;
80 tracebacks refer to stdin rather than a temporary script path. The separate
81 RLM/REPL kernel is not changed by this launcher integration.
82
83 Platform/backend limits of the shared launcher still apply: no available local
84 wrapper means workspace-write remains unenforced, while read-only is refused.
85 An external sandbox session cannot use these local interpreter tools and must
86 use its shell execution path. See [sandbox limits](SANDBOX.md).
87
88 ## Deferred and dynamic tools
89
90 `Web` is conditional and deferred. It is discoverable through `tool_search`
91 only when the active policy and runtime backend permit it. Read-only
92 children retain its read-only search/fetch evidence path; read-only authority
93 does not mean "unable to research."
94
95 The durable `github`, `automation`, and `rlm` action families are also deferred
96 by default. `rlm` owns `open`, `eval`, `configure`, and `close` actions for a
97 persistent local Python session (a subprocess with a scrubbed environment, not
98 an OS sandbox). Inline ```` ```repl ```` fences in a reply run in the same kind of
99 kernel only when `code_execution` is on the turn's surface (never in Plan mode),
100 only when the fence opens its own line, and only after `code_execution`'s
101 approval under the session posture. Feature-gated native tools may be added to
102 the active or deferred catalog only when their implementation and host
103 dependencies are available.
104
105 MCP tools are dynamic. Successfully connected servers register names such as
106 `mcp_<server>_<tool>` from `~/.codewhale/mcp.json`; a failed or disabled server
107 must not be presented as available. MCP and plugin tools are deferred unless a
108 user explicitly names them in `[tools].always_load`.
109
110 For an unstarted configured server, an MCP-focused `tool_search` first
111 performs bounded discovery under the current turn's server/tool ceiling.
112 It searches real server-provided schemas after connection, rather than
113 inventing `mcp_*` definitions. A search inside `execute_tools` describes
114 those schemas without activating them; a direct search uses the existing
115 bounded activation cache. Ordinary unrelated searches leave optional servers
116 unstarted. The CLI's standalone MCP inspection commands own a separate pool.
117
118 ### Code mode (`execute_tools`)
119
120 `execute_tools` is engine-injected, alongside the synthetic interpreter tools.
121 It runs a JavaScript program whose only host surface is
122 `await tools.call(name, args)`, and it is the default way to compose several
123 tool calls — MCP and plugin tools included — without round-tripping every
124 intermediate result through the conversation. It is hidden from Plan mode and
125 refused under a worker authority envelope.
126
127 - **One gate.** In a session turn every nested call is sent back to the turn
128 loop and planned exactly like a direct call: deny/allow lists, preparation
129 (MCP `readOnlyHint`/`destructiveHint`), `tool_call_before` hooks, ask-rules,
130 Auto-Review, repository law, and the Computer Use consent refusal. MCP calls
131 run through the session MCP pool. Approving the program grants nothing, so
132 `execute_tools` itself is auto-approved. If the permission posture changes
133 while a program runs, its remaining nested calls are refused (an approved
134 call survives only an equal or broader posture, as for a direct call) and
135 the model retries them under the new posture.
136 - **Approvals suspend the program.** A nested call that needs approval raises
137 the normal approval card (named `execute_tools program call: ...`) and the
138 program waits; allow resumes it, deny fails only that nested call as an
139 exception the program can catch. Time spent waiting on a decision does not
140 count against the program's run deadline, which is the turn's remaining
141 wall clock.
142 - **Receipts.** The result lists every nested call with its decision (`auto`,
143 `approved`, `denied`, `refused`) and status (`ok`, `failed`, `refused`,
144 `in_flight`). A program that hits its deadline still returns the receipt;
145 `in_flight` calls were cancelled and may have partially run. Each nested
146 result is `{content, metadata, truncated}`; an oversized result keeps that
147 shape, and `truncated` names the original size and the spillover file with
148 the full output.
149 - **Discovery without re-pinning.** Inside a program,
150 `tools.call('tool_search', {query})` returns matching deferred tools with
151 their input schemas and does not activate them, so the request's tool array
152 and the session-pinned prefix do not change.
153 - **Stays direct:** `agent`, `workflow`, `request_user_input`, nested
154 `execute_tools`, interactive shells, sandbox escalation, Computer Use
155 consent and scripts, and MCP sign-in (`mcp_<server>_authenticate`).
156
157 Code mode is on by default (`[features] code_mode = true`), which makes
158 `execute_tools` eager from the first request; direct tools and `tool_search`
159 stay available either way. Set `code_mode = false` (or run with
160 `--disable code_mode`) to go back to deferring `execute_tools` behind
161 `tool_search`. The flag is session configuration, so the prompt prefix stays
162 stable within a session. Without an engine turn (sub-agents), a program keeps
163 the conservative profile: read-only, auto-approved native calls only, no MCP.
164
165 ### Conversation toolbox cache
166
167 A successful search activation is remembered by name for the current
168 conversation. The cache holds at most eight deferred names and 16 KiB of
169 serialized schemas, evicts least-recently-used entries, and revalidates every
170 entry against the current catalog and policy before advertising it again. A
171 session sync clears it. The cache cannot resurrect a removed, denied, or
172 newly-eager tool.
173
174 Each subagent gets its own policy-filtered deferred catalog, always-present
175 `tool_search`, and bounded activation cache. Forked messages and instructions
176 remain in context, but the child cache starts empty and discovers tools locally;
177 neither forked context nor a cache can become a discovery allowlist. A child can
178 still search every tool its own authority permits, including Web search/fetch
179 for read-only research roles.
180
181 ## Inspect the model-client request tool payload
182
183 Run `/tools` after a model turn to inspect a bounded projection of the exact
184 tool field in the latest prepared model-client request. `/tools json` emits the
185 same evidence as bounded machine-readable JSON. Both formats open in a pager;
186 they are not copied into transcript history. `/tool-studio` remains a human-
187 command compatibility alias; it is not a model tool.
188
189 The snapshot distinguishes an absent tool field from a present empty array. It
190 reports the exact model-client tool JSON byte count and SHA-256 digest only when
191 measurement fits the one-MiB inspection bound; larger payloads stay unavailable.
192 Provider adapters may transform, sanitize, or omit those fields while building
193 a provider-specific wire body, so `/tools` marks provider delivery and the wire
194 payload unavailable. Capture and rendering are bounded: retained schemas,
195 descriptions, caller lists, catalog rows, turn IDs, and payload measurement all
196 carry explicit truncation, omission, or unavailable receipts. The snapshot stays
197 in memory only for the current session and is replaced on each prepared request.
198
199 Provider, model, approval, registry provenance, and runtime capability metadata
200 are not fields in the request tool schema. `/tools` therefore reports them as
201 unavailable instead of joining against mutable state or inferring values. Use
202 the separate route and permission receipts for those facts.
203
204 ## Modes and permission postures
205
206 Modes and permission postures are separate controls:
207
208 - **Plan** keeps the stable primitive vocabulary but centrally refuses shell
209 execution and file mutation.
210 - **Work** is ordinary interactive execution.
211 - **Operate** uses the same direct-tool authority as Work. Small work stays
212 direct; multi-step delegation uses a compact Workflow plan with dependencies,
213 bounded scopes, and completion evidence. Fleet manages the same sub-agents
214 and roles. One bounded, independent task can use a direct agent; `followup`
215 reuses that agent for continued work.
216 - **Ask**, **Auto-Review**, and **Full Access** control approval behavior within
217 an action-capable mode. They never widen Plan into write or shell access.
218
219 See `docs/MODES.md` for the full mode and posture contract.
220
221 ## Compatibility names
222
223 The model-facing contract is the lowercase core above. Saved v0.9.x
224 transcripts and protocol clients may still call exact hidden compatibility
225 names such as `File`, `Bash`, and the older single-operation file names. Those
226 names never enter a new model catalog or `tool_search` result.
227
228 Compatibility is execution compatibility, not fuzzy aliasing: an exact legacy
229 call must reach the handler for its legacy schema. It must not be rewritten
230 into a small lowercase primitive whose input shape is different. Unknown or
231 retired names still fail closed instead of guessing a destination.
232
233 Specialized native families such as `Git`, `Run`, and `Web` are not aliases for
234 the lowercase core. They remain real, policy-filtered deferred tools and are
235 loaded through `tool_search` when needed.
236
237 ## Long-running work
238
239 `bash` runs one cancellable foreground command. It does not carry background,
240 TTY, wait, interact, or cancel action fields. Stateful process and terminal
241 control is specialized functionality that must be discovered explicitly; it
242 does not enlarge the first-turn shell schema.
243
244 Use `tasks` when the work itself needs a durable lifecycle, structured gates,
245 artifacts, replayable timelines, or a stable task id. Large tool results should
246 remain behind bounded handles or artifacts instead of being copied wholesale
247 into the parent transcript.
248
249 ## Parallel fan-out
250
251 The sub-agent capacity source of truth is
252 `crates/tui/src/config/subagent_limits.rs`:
253
254 - default configured concurrency: **64**;
255 - maximum configured concurrency: **128**;
256 - maximum admitted running-plus-queued work: **1024**.
257
258 These are capacity ceilings, not advice to dispatch every available slot. A
259 manager should use the smallest useful fan-out, preserve a single owner for
260 fan-in, and verify worker receipts before reporting combined completion.
261
262 RLM child-query batching is a different, cheaper cost class. Its
263 `sub_query_batch` helper accepts 1–16 one-shot children inside a live `rlm`
264 session; it is not a substitute for tool-carrying `agent` workers.
265
266 ## Human inspection: `/tools` (`/tool-studio`)
267
268 `/tools` renders a **read-only, bounded human projection** of the tool field of
269 the request that was prepared for one `(turn, step)`. It is not a second
270 registry and not an execution surface.
271
272 **The seam.** The snapshot is built in `crates/tui/src/core/engine/turn_loop.rs`
273 immediately after `MessageRequest` is constructed, from `request.tools` — the
274 same value the model client is handed. The engine resolves the surrounding
275 per-turn data once in `engine.rs` (`ToolSurfaceContext`: flattened registry
276 facts, the MCP pool's own server attribution, the engine-injected catalog names,
277 and the resolved model client's receipt) and passes it as plain data, so the
278 per-step seam never re-locks the MCP pool or holds a tool object.
279
280 **Turn and step identity.** The tool set can differ between steps of a turn, so
281 each snapshot is stamped with turn id and step and each seam emits its own. The
282 TUI keeps only the latest (`SessionState.last_tool_request_snapshot`). Before
283 the first seam there is no snapshot and `/tools` says so rather than rebuilding
284 a registry in the UI.
285
286 Two kinds of fact are kept apart:
287
288 - **Wire facts** come from the prepared request: name, description, schema,
289 `defer_loading` / `strict` / `allowed_callers` / `cache_control`, byte
290 accounting, and the catalog digest.
291 - **Surface facts** come from the `ToolSurfaceContext`: provenance
292 (`builtin` / `plugin` / `mcp` / `synthetic` / `unknown`), MCP server identity,
293 declared capabilities, declared approval requirement, and model visibility.
294
295 Contract:
296
297 - **One digest.** `active_tool_catalog_sha256`
298 (`crates/tui/src/core/engine/preview.rs`) is the single definition of the
299 active-tool-catalog hash. The request manifest publishes it as
300 `ToolSurfaceFacts::active_tool_catalog_sha256` and `/tools` reports the same
301 value for the same prepared request; neither surface keeps a hash of its own.
302 - **Nothing is guessed.** MCP server identity is shown only when the real pool
303 attributed that exact model tool name. `McpPool::mcp_model_tool_name` is the
304 single definition shared by the model catalog and the human attribution, and
305 an ambiguous name (two servers colliding on one model name) resolves to no
306 server. Synthetic provenance comes from
307 `default_synthetic_catalog_tool_names`, which is asserted against the engine's
308 own `is_synthetic_catalog_tool` predicate. A transmitted tool with no registry
309 entry reports `capabilities: unknown`, never "none".
310 - **Provider availability follows the resolved client.** It comes from
311 `Engine::tool_surface_provider_receipt`, never from "a tool registry exists".
312 With no client the receipt is `unavailable` even when the registry is full.
313 - **Unknown shrinks, it does not vanish.** `unavailable_for_this_request` always
314 contains `provider_wire_payload`: nothing on this path observes what the
315 provider adapter finally transmits. It additionally contains `provider` and
316 `model` without a resolved client, and `provenance` / `capabilities` /
317 `approval` when no surface context was captured.
318 - **Absent stays distinct from empty.** A request with no tools field is not a
319 request with an empty tools array; an unresolved field is `unknown` with a
320 reason, not a default.
321 - **Bounded.** Rendering is capped by tool count (32), name, description, schema
322 bytes, allowed-caller count, and a payload measurement bound, each with an
323 explicit truncation or omission receipt. Registered tools that this request
324 does *not* carry are reported as a bounded name list plus an exact count
325 rather than expanding the projection.
326 - **Inert.** The snapshot lives beside the transcript, never in
327 `session.messages`, so it cannot enter a model request or perturb the
328 provider's prefix cache. It never executes a tool, never reads credentials,
329 never reorders the catalog, and is never registered as a model-callable tool.
330 - **Delivery is never claimed.** The capture happens before connection setup, so
331 `delivery_status` stays `unknown`.
332
333 ## Release verification
334
335 Do not infer the public surface from handler function names. Verify the model
336 catalog and alias visibility at the exact candidate SHA:
337
338 ```bash
339 python3 scripts/measure-runtime-contract.py
340 cargo test -p codewhale-tui --lib --locked core::engine::tests::default_active_contract_keeps_discovery_and_core_tools_eager -- --exact
341 cargo test -p codewhale-tui --lib --locked tools::file_tool::tests::primitive_schemas_are_separate_and_small_contract_shaped -- --exact
342 cargo test -p codewhale-tui --lib --locked tools::shell::tests::lowercase_bash_schema_is_small_contract -- --exact
343 cargo test --locked -p codewhale-tui --lib core::engine::tests::print_mode_tool_catalog_metrics -- --ignored --exact --nocapture
344 ```
345
346 Check the test names against the source before trusting a green run: `cargo test`
347 exits 0 with "0 passed; N filtered out" when a filter matches nothing, so a
348 misspelled filter is indistinguishable from a pass. Each `--exact` command
349 above must report `1 passed` (the ignored metrics test reports `1 passed`
350 only because `--ignored` selects it); `0 passed` means the filter matched
351 nothing and the check did not run.
352
353 The provider-free receipt must report the eleven default-active names listed
354 above. A separate repository-wide tool count may include deferred, dynamic,
355 feature-gated, and compatibility-only registrations; it is not the number of
356 tools placed in the first-turn model catalog.
357
357 lines MARKDOWN