| 1 | # Tool surface |
| 2 | |
| 3 | > 阅读简体中文版:[zh_hans/TOOL_SURFACE.md](zh_hans/TOOL_SURFACE.md)。 |
| 4 | |
| 5 | This document describes the current model-facing tool contract. The v0.9.1 |
| 6 | cutover that produced it is recorded in `docs/RUNTIME_SIMPLIFICATION_DESIGN.md`; |
| 7 | read the workspace version from `Cargo.toml`, not from this line. The registry |
| 8 | remains larger than the first-turn catalog so |
| 9 | saved transcripts can replay and uncommon capabilities can be loaded on demand. |
| 10 | The model should learn one canonical name for each common operation. |
| 11 | |
| 12 | Implementation sources: |
| 13 | |
| 14 | - `crates/tui/src/core/engine/tool_catalog.rs` owns the eager/deferred catalog. |
| 15 | - `crates/tui/src/tools/registry.rs` registers canonical tools and hidden aliases. |
| 16 | - `crates/tui/src/tools/{file,file_tool,shell}.rs` own the small foreground |
| 17 | primitive behavior and schemas; the other native tools remain searchable. |
| 18 | - `docs/RUNTIME_SIMPLIFICATION_DESIGN.md` records the v0.9.1 cutover and receipt. |
| 19 | |
| 20 | ## Default-active contract |
| 21 | |
| 22 | New turns start with eleven eager native names plus synthetic `tool_search`: |
| 23 | |
| 24 | 1. `read` |
| 25 | 2. `write` |
| 26 | 3. `edit` |
| 27 | 4. `bash` |
| 28 | 5. `agent` |
| 29 | 6. `workflow` |
| 30 | 7. `todo_write` |
| 31 | 8. `create_goal` |
| 32 | 9. `get_goal` |
| 33 | 10. `update_goal` |
| 34 | 11. `load_skill` |
| 35 | 12. `tool_search` (synthetic, always active) |
| 36 | |
| 37 | The eleven native names are `DEFAULT_ACTIVE_NATIVE_TOOLS` in |
| 38 | `crates/tui/src/core/engine/tool_catalog.rs`, pinned by |
| 39 | `default_active_contract_keeps_discovery_and_core_tools_eager`. An authority |
| 40 | boundary may remove `agent` at the maximum child depth, but route size alone |
| 41 | must not change this core vocabulary. |
| 42 | |
| 43 | The direct schemas deliberately stay small: |
| 44 | |
| 45 | | Tool | Input | Purpose | |
| 46 | |---|---|---| |
| 47 | | `read` | `path`, optional `offset`, optional `limit` | Read a bounded file window with explicit continuation or truncation notices. | |
| 48 | | `write` | `path`, `content` | Create or replace a file. | |
| 49 | | `edit` | `path`, `edits` | Apply one or more unambiguous text replacements against one original snapshot. | |
| 50 | | `bash` | `command`, optional `timeout` | Run one cancellable foreground shell command and return a bounded tail. | |
| 51 | | `agent` | delegated task and optional scope/context controls | Start or inspect focused child work. | |
| 52 | | `workflow` | plan/script/source_path plus run controls | Coordinate multi-agent phases with dependencies and completion checks. | |
| 53 | | `todo_write` | complete replacement list of `{content, status}` items | Keep optional, agent-owned progress notes for genuinely multi-step work. | |
| 54 | | `create_goal` | objective plus optional budget | Start the session goal the turn works toward. | |
| 55 | | `get_goal` | none | Read the active goal and its progress. | |
| 56 | | `update_goal` | terminal status | Mark the goal complete or blocked. | |
| 57 | | `tool_search` | `query`, optional matching controls | Discover policy-allowed deferred tools and add selected schemas to this conversation's toolbox. | |
| 58 | |
| 59 | Mode is an authority decision, not a synonym system. Plan, Work, and Operate |
| 60 | use the same primitive identities. Plan centrally refuses `write`, `edit`, and |
| 61 | `bash`; Work and Operate still pass those calls through approval, sandbox, |
| 62 | trusted-path, repository-law, and managed-policy gates. Full Access changes |
| 63 | ordinary approval behavior but does not bypass hard safety or repository law. |
| 64 | |
| 65 | `update_plan` remains registered only for saved-artifact compatibility and is |
| 66 | not model-visible. `tasks`, `Git`, `Run`, `Web`, `remember`, and other |
| 67 | specialized capabilities are searchable rather than first-turn ceremony. |
| 68 | |
| 69 | ## Local code execution |
| 70 | |
| 71 | `code_execution` (Python) and `js_execution` (Node.js) use the same |
| 72 | permission-aware local launcher as workspace task gates and test runs. Ordinary |
| 73 | approval retains the effective session policy. To retry a denied exact call |
| 74 | with wider permissions, supply `sandbox_permissions` (`workspace-write` or |
| 75 | `danger-full-access`) and a nonempty `justification`; Ask mode requires explicit |
| 76 | one-shot user approval. The approved policy affects only that call. |
| 77 | |
| 78 | Both tools read source from stdin, return stdout/stderr/return_code, and keep |
| 79 | the existing timeout and process-tree cleanup. They are not persistent REPLs; |
| 80 | tracebacks refer to stdin rather than a temporary script path. The separate |
| 81 | RLM/REPL kernel is not changed by this launcher integration. |
| 82 | |
| 83 | Platform/backend limits of the shared launcher still apply: no available local |
| 84 | wrapper means workspace-write remains unenforced, while read-only is refused. |
| 85 | An external sandbox session cannot use these local interpreter tools and must |
| 86 | use its shell execution path. See [sandbox limits](SANDBOX.md). |
| 87 | |
| 88 | ## Deferred and dynamic tools |
| 89 | |
| 90 | `Web` is conditional and deferred. It is discoverable through `tool_search` |
| 91 | only when the active policy and runtime backend permit it. Read-only |
| 92 | children retain its read-only search/fetch evidence path; read-only authority |
| 93 | does not mean "unable to research." |
| 94 | |
| 95 | The durable `github`, `automation`, and `rlm` action families are also deferred |
| 96 | by default. `rlm` owns `open`, `eval`, `configure`, and `close` actions for a |
| 97 | persistent local Python session (a subprocess with a scrubbed environment, not |
| 98 | an OS sandbox). Inline ```` ```repl ```` fences in a reply run in the same kind of |
| 99 | kernel only when `code_execution` is on the turn's surface (never in Plan mode), |
| 100 | only when the fence opens its own line, and only after `code_execution`'s |
| 101 | approval under the session posture. Feature-gated native tools may be added to |
| 102 | the active or deferred catalog only when their implementation and host |
| 103 | dependencies are available. |
| 104 | |
| 105 | MCP tools are dynamic. Successfully connected servers register names such as |
| 106 | `mcp_<server>_<tool>` from `~/.codewhale/mcp.json`; a failed or disabled server |
| 107 | must not be presented as available. MCP and plugin tools are deferred unless a |
| 108 | user explicitly names them in `[tools].always_load`. |
| 109 | |
| 110 | For an unstarted configured server, an MCP-focused `tool_search` first |
| 111 | performs bounded discovery under the current turn's server/tool ceiling. |
| 112 | It searches real server-provided schemas after connection, rather than |
| 113 | inventing `mcp_*` definitions. A search inside `execute_tools` describes |
| 114 | those schemas without activating them; a direct search uses the existing |
| 115 | bounded activation cache. Ordinary unrelated searches leave optional servers |
| 116 | unstarted. The CLI's standalone MCP inspection commands own a separate pool. |
| 117 | |
| 118 | ### Code mode (`execute_tools`) |
| 119 | |
| 120 | `execute_tools` is engine-injected, alongside the synthetic interpreter tools. |
| 121 | It runs a JavaScript program whose only host surface is |
| 122 | `await tools.call(name, args)`, and it is the default way to compose several |
| 123 | tool calls — MCP and plugin tools included — without round-tripping every |
| 124 | intermediate result through the conversation. It is hidden from Plan mode and |
| 125 | refused under a worker authority envelope. |
| 126 | |
| 127 | - **One gate.** In a session turn every nested call is sent back to the turn |
| 128 | loop and planned exactly like a direct call: deny/allow lists, preparation |
| 129 | (MCP `readOnlyHint`/`destructiveHint`), `tool_call_before` hooks, ask-rules, |
| 130 | Auto-Review, repository law, and the Computer Use consent refusal. MCP calls |
| 131 | run through the session MCP pool. Approving the program grants nothing, so |
| 132 | `execute_tools` itself is auto-approved. If the permission posture changes |
| 133 | while a program runs, its remaining nested calls are refused (an approved |
| 134 | call survives only an equal or broader posture, as for a direct call) and |
| 135 | the model retries them under the new posture. |
| 136 | - **Approvals suspend the program.** A nested call that needs approval raises |
| 137 | the normal approval card (named `execute_tools program call: ...`) and the |
| 138 | program waits; allow resumes it, deny fails only that nested call as an |
| 139 | exception the program can catch. Time spent waiting on a decision does not |
| 140 | count against the program's run deadline, which is the turn's remaining |
| 141 | wall clock. |
| 142 | - **Receipts.** The result lists every nested call with its decision (`auto`, |
| 143 | `approved`, `denied`, `refused`) and status (`ok`, `failed`, `refused`, |
| 144 | `in_flight`). A program that hits its deadline still returns the receipt; |
| 145 | `in_flight` calls were cancelled and may have partially run. Each nested |
| 146 | result is `{content, metadata, truncated}`; an oversized result keeps that |
| 147 | shape, and `truncated` names the original size and the spillover file with |
| 148 | the full output. |
| 149 | - **Discovery without re-pinning.** Inside a program, |
| 150 | `tools.call('tool_search', {query})` returns matching deferred tools with |
| 151 | their input schemas and does not activate them, so the request's tool array |
| 152 | and the session-pinned prefix do not change. |
| 153 | - **Stays direct:** `agent`, `workflow`, `request_user_input`, nested |
| 154 | `execute_tools`, interactive shells, sandbox escalation, Computer Use |
| 155 | consent and scripts, and MCP sign-in (`mcp_<server>_authenticate`). |
| 156 | |
| 157 | Code mode is on by default (`[features] code_mode = true`), which makes |
| 158 | `execute_tools` eager from the first request; direct tools and `tool_search` |
| 159 | stay available either way. Set `code_mode = false` (or run with |
| 160 | `--disable code_mode`) to go back to deferring `execute_tools` behind |
| 161 | `tool_search`. The flag is session configuration, so the prompt prefix stays |
| 162 | stable within a session. Without an engine turn (sub-agents), a program keeps |
| 163 | the conservative profile: read-only, auto-approved native calls only, no MCP. |
| 164 | |
| 165 | ### Conversation toolbox cache |
| 166 | |
| 167 | A successful search activation is remembered by name for the current |
| 168 | conversation. The cache holds at most eight deferred names and 16 KiB of |
| 169 | serialized schemas, evicts least-recently-used entries, and revalidates every |
| 170 | entry against the current catalog and policy before advertising it again. A |
| 171 | session sync clears it. The cache cannot resurrect a removed, denied, or |
| 172 | newly-eager tool. |
| 173 | |
| 174 | Each subagent gets its own policy-filtered deferred catalog, always-present |
| 175 | `tool_search`, and bounded activation cache. Forked messages and instructions |
| 176 | remain in context, but the child cache starts empty and discovers tools locally; |
| 177 | neither forked context nor a cache can become a discovery allowlist. A child can |
| 178 | still search every tool its own authority permits, including Web search/fetch |
| 179 | for read-only research roles. |
| 180 | |
| 181 | ## Inspect the model-client request tool payload |
| 182 | |
| 183 | Run `/tools` after a model turn to inspect a bounded projection of the exact |
| 184 | tool field in the latest prepared model-client request. `/tools json` emits the |
| 185 | same evidence as bounded machine-readable JSON. Both formats open in a pager; |
| 186 | they are not copied into transcript history. `/tool-studio` remains a human- |
| 187 | command compatibility alias; it is not a model tool. |
| 188 | |
| 189 | The snapshot distinguishes an absent tool field from a present empty array. It |
| 190 | reports the exact model-client tool JSON byte count and SHA-256 digest only when |
| 191 | measurement fits the one-MiB inspection bound; larger payloads stay unavailable. |
| 192 | Provider adapters may transform, sanitize, or omit those fields while building |
| 193 | a provider-specific wire body, so `/tools` marks provider delivery and the wire |
| 194 | payload unavailable. Capture and rendering are bounded: retained schemas, |
| 195 | descriptions, caller lists, catalog rows, turn IDs, and payload measurement all |
| 196 | carry explicit truncation, omission, or unavailable receipts. The snapshot stays |
| 197 | in memory only for the current session and is replaced on each prepared request. |
| 198 | |
| 199 | Provider, model, approval, registry provenance, and runtime capability metadata |
| 200 | are not fields in the request tool schema. `/tools` therefore reports them as |
| 201 | unavailable instead of joining against mutable state or inferring values. Use |
| 202 | the separate route and permission receipts for those facts. |
| 203 | |
| 204 | ## Modes and permission postures |
| 205 | |
| 206 | Modes and permission postures are separate controls: |
| 207 | |
| 208 | - **Plan** keeps the stable primitive vocabulary but centrally refuses shell |
| 209 | execution and file mutation. |
| 210 | - **Work** is ordinary interactive execution. |
| 211 | - **Operate** uses the same direct-tool authority as Work. Small work stays |
| 212 | direct; multi-step delegation uses a compact Workflow plan with dependencies, |
| 213 | bounded scopes, and completion evidence. Fleet manages the same sub-agents |
| 214 | and roles. One bounded, independent task can use a direct agent; `followup` |
| 215 | reuses that agent for continued work. |
| 216 | - **Ask**, **Auto-Review**, and **Full Access** control approval behavior within |
| 217 | an action-capable mode. They never widen Plan into write or shell access. |
| 218 | |
| 219 | See `docs/MODES.md` for the full mode and posture contract. |
| 220 | |
| 221 | ## Compatibility names |
| 222 | |
| 223 | The model-facing contract is the lowercase core above. Saved v0.9.x |
| 224 | transcripts and protocol clients may still call exact hidden compatibility |
| 225 | names such as `File`, `Bash`, and the older single-operation file names. Those |
| 226 | names never enter a new model catalog or `tool_search` result. |
| 227 | |
| 228 | Compatibility is execution compatibility, not fuzzy aliasing: an exact legacy |
| 229 | call must reach the handler for its legacy schema. It must not be rewritten |
| 230 | into a small lowercase primitive whose input shape is different. Unknown or |
| 231 | retired names still fail closed instead of guessing a destination. |
| 232 | |
| 233 | Specialized native families such as `Git`, `Run`, and `Web` are not aliases for |
| 234 | the lowercase core. They remain real, policy-filtered deferred tools and are |
| 235 | loaded through `tool_search` when needed. |
| 236 | |
| 237 | ## Long-running work |
| 238 | |
| 239 | `bash` runs one cancellable foreground command. It does not carry background, |
| 240 | TTY, wait, interact, or cancel action fields. Stateful process and terminal |
| 241 | control is specialized functionality that must be discovered explicitly; it |
| 242 | does not enlarge the first-turn shell schema. |
| 243 | |
| 244 | Use `tasks` when the work itself needs a durable lifecycle, structured gates, |
| 245 | artifacts, replayable timelines, or a stable task id. Large tool results should |
| 246 | remain behind bounded handles or artifacts instead of being copied wholesale |
| 247 | into the parent transcript. |
| 248 | |
| 249 | ## Parallel fan-out |
| 250 | |
| 251 | The sub-agent capacity source of truth is |
| 252 | `crates/tui/src/config/subagent_limits.rs`: |
| 253 | |
| 254 | - default configured concurrency: **64**; |
| 255 | - maximum configured concurrency: **128**; |
| 256 | - maximum admitted running-plus-queued work: **1024**. |
| 257 | |
| 258 | These are capacity ceilings, not advice to dispatch every available slot. A |
| 259 | manager should use the smallest useful fan-out, preserve a single owner for |
| 260 | fan-in, and verify worker receipts before reporting combined completion. |
| 261 | |
| 262 | RLM child-query batching is a different, cheaper cost class. Its |
| 263 | `sub_query_batch` helper accepts 1–16 one-shot children inside a live `rlm` |
| 264 | session; it is not a substitute for tool-carrying `agent` workers. |
| 265 | |
| 266 | ## Human inspection: `/tools` (`/tool-studio`) |
| 267 | |
| 268 | `/tools` renders a **read-only, bounded human projection** of the tool field of |
| 269 | the request that was prepared for one `(turn, step)`. It is not a second |
| 270 | registry and not an execution surface. |
| 271 | |
| 272 | **The seam.** The snapshot is built in `crates/tui/src/core/engine/turn_loop.rs` |
| 273 | immediately after `MessageRequest` is constructed, from `request.tools` — the |
| 274 | same value the model client is handed. The engine resolves the surrounding |
| 275 | per-turn data once in `engine.rs` (`ToolSurfaceContext`: flattened registry |
| 276 | facts, the MCP pool's own server attribution, the engine-injected catalog names, |
| 277 | and the resolved model client's receipt) and passes it as plain data, so the |
| 278 | per-step seam never re-locks the MCP pool or holds a tool object. |
| 279 | |
| 280 | **Turn and step identity.** The tool set can differ between steps of a turn, so |
| 281 | each snapshot is stamped with turn id and step and each seam emits its own. The |
| 282 | TUI keeps only the latest (`SessionState.last_tool_request_snapshot`). Before |
| 283 | the first seam there is no snapshot and `/tools` says so rather than rebuilding |
| 284 | a registry in the UI. |
| 285 | |
| 286 | Two kinds of fact are kept apart: |
| 287 | |
| 288 | - **Wire facts** come from the prepared request: name, description, schema, |
| 289 | `defer_loading` / `strict` / `allowed_callers` / `cache_control`, byte |
| 290 | accounting, and the catalog digest. |
| 291 | - **Surface facts** come from the `ToolSurfaceContext`: provenance |
| 292 | (`builtin` / `plugin` / `mcp` / `synthetic` / `unknown`), MCP server identity, |
| 293 | declared capabilities, declared approval requirement, and model visibility. |
| 294 | |
| 295 | Contract: |
| 296 | |
| 297 | - **One digest.** `active_tool_catalog_sha256` |
| 298 | (`crates/tui/src/core/engine/preview.rs`) is the single definition of the |
| 299 | active-tool-catalog hash. The request manifest publishes it as |
| 300 | `ToolSurfaceFacts::active_tool_catalog_sha256` and `/tools` reports the same |
| 301 | value for the same prepared request; neither surface keeps a hash of its own. |
| 302 | - **Nothing is guessed.** MCP server identity is shown only when the real pool |
| 303 | attributed that exact model tool name. `McpPool::mcp_model_tool_name` is the |
| 304 | single definition shared by the model catalog and the human attribution, and |
| 305 | an ambiguous name (two servers colliding on one model name) resolves to no |
| 306 | server. Synthetic provenance comes from |
| 307 | `default_synthetic_catalog_tool_names`, which is asserted against the engine's |
| 308 | own `is_synthetic_catalog_tool` predicate. A transmitted tool with no registry |
| 309 | entry reports `capabilities: unknown`, never "none". |
| 310 | - **Provider availability follows the resolved client.** It comes from |
| 311 | `Engine::tool_surface_provider_receipt`, never from "a tool registry exists". |
| 312 | With no client the receipt is `unavailable` even when the registry is full. |
| 313 | - **Unknown shrinks, it does not vanish.** `unavailable_for_this_request` always |
| 314 | contains `provider_wire_payload`: nothing on this path observes what the |
| 315 | provider adapter finally transmits. It additionally contains `provider` and |
| 316 | `model` without a resolved client, and `provenance` / `capabilities` / |
| 317 | `approval` when no surface context was captured. |
| 318 | - **Absent stays distinct from empty.** A request with no tools field is not a |
| 319 | request with an empty tools array; an unresolved field is `unknown` with a |
| 320 | reason, not a default. |
| 321 | - **Bounded.** Rendering is capped by tool count (32), name, description, schema |
| 322 | bytes, allowed-caller count, and a payload measurement bound, each with an |
| 323 | explicit truncation or omission receipt. Registered tools that this request |
| 324 | does *not* carry are reported as a bounded name list plus an exact count |
| 325 | rather than expanding the projection. |
| 326 | - **Inert.** The snapshot lives beside the transcript, never in |
| 327 | `session.messages`, so it cannot enter a model request or perturb the |
| 328 | provider's prefix cache. It never executes a tool, never reads credentials, |
| 329 | never reorders the catalog, and is never registered as a model-callable tool. |
| 330 | - **Delivery is never claimed.** The capture happens before connection setup, so |
| 331 | `delivery_status` stays `unknown`. |
| 332 | |
| 333 | ## Release verification |
| 334 | |
| 335 | Do not infer the public surface from handler function names. Verify the model |
| 336 | catalog and alias visibility at the exact candidate SHA: |
| 337 | |
| 338 | ```bash |
| 339 | python3 scripts/measure-runtime-contract.py |
| 340 | cargo test -p codewhale-tui --lib --locked core::engine::tests::default_active_contract_keeps_discovery_and_core_tools_eager -- --exact |
| 341 | cargo test -p codewhale-tui --lib --locked tools::file_tool::tests::primitive_schemas_are_separate_and_small_contract_shaped -- --exact |
| 342 | cargo test -p codewhale-tui --lib --locked tools::shell::tests::lowercase_bash_schema_is_small_contract -- --exact |
| 343 | cargo test --locked -p codewhale-tui --lib core::engine::tests::print_mode_tool_catalog_metrics -- --ignored --exact --nocapture |
| 344 | ``` |
| 345 | |
| 346 | Check the test names against the source before trusting a green run: `cargo test` |
| 347 | exits 0 with "0 passed; N filtered out" when a filter matches nothing, so a |
| 348 | misspelled filter is indistinguishable from a pass. Each `--exact` command |
| 349 | above must report `1 passed` (the ignored metrics test reports `1 passed` |
| 350 | only because `--ignored` selects it); `0 passed` means the filter matched |
| 351 | nothing and the check did not run. |
| 352 | |
| 353 | The provider-free receipt must report the eleven default-active names listed |
| 354 | above. A separate repository-wide tool count may include deferred, dynamic, |
| 355 | feature-gated, and compatibility-only registrations; it is not the number of |
| 356 | tools placed in the first-turn model catalog. |
| 357 |