返回 CodeWhale
runtime-contract-budget.json
根目录 / scripts / runtime-contract-budget.json
1 {
2 "_comment": "One-way numeric ceilings and exact structural identities for the provider-free runtime contract. Decreases pass; increases or identity changes fail. Lock in decreases with: python3 scripts/check-runtime-contract-budget.py --update The v0.9.8 child-receipt restore grew every production tool surface by 1496 schema bytes / 374 estimated tokens (agent tool). The v0.9.8 workshop read/tool-result byte fields then grew every production tool surface by 371 schema bytes / 93 estimated tokens. Both raises are explicit maintainer decisions; identities stay on the pre-raise digests only if the name set is unchanged \u2014 re-measure on Linux CI if Lint reports identity drift. The v0.9.8 pinned session prefix added the <context_update> sentence to the base prompt (5848 -> 6084 bytes, every representative stage re-hashed), and the host-side Workflow/Goal verbs plus honest child posture grew the tool catalog (active 16531 -> 16602 bytes, full 71473 -> 72371); both are explicit v0.9.8 maintainer decisions measured from the release train. The v0.9.9 configured-skills change hides only custom configured-root paths, preserves discoverable default-root paths, normalizes Windows prompt separators, and trims 50 redundant skills-prompt bytes. The skill/memory/goal/handoff identities were re-measured without raising any ceiling. Explicit maintainer decision for #5473/#5492. The v0.9.10 full surfaces intentionally add the safe read_media tool; their measured schemas remain below the prior byte/token ceilings. Representative prompt byte metrics now use the same host-independent normalized text as their identities; the normalized base is 6089 bytes. The v0.9.11 model-visible sub-agent surface intentionally retires six legacy agents/* tools in favor of the canonical agent tool; all affected schema and prompt metrics decrease. The v0.9.12 plugin prompt-match slice intentionally adds the request_plugin_install tool to the full tool surfaces (plan full: +518 schema bytes / +130 estimated tokens / 29 -> 30 tools) so a strong prompt match can surface the human review CTA; explicit maintainer decision for #5663/#5579. The v0.9.13 profile pins a non-executed bare bash shell so interpreter guidance is reproducible across hosts. The duplicate tts catalog entry is intentionally hidden; speech remains canonical and the alias remains available for saved-transcript dispatch. Explicit v0.9.13 maintainer decision (2026-09-08): after removing 1426 repeated guidance bytes and pinning the bash-v2 fixture, accept only the measured tool byte/token ceilings from all-features macOS source e27735bb63c897f88061c71567701fd971f5d396, verified libtest SHA-256 5e8cbe213f32c4ecdec63494c4de5e31857b4a40134edf7b21a55bca926b1b38: active 13274/3319 in every mode, Plan full 39885/9972, Act/Operate full 67603/16901, with no margin. Against the prior budget, active +390 bytes is agent -41 plus retained bash command syntax +431. Plan full also retains Git commit_plan +253, update_goal progress +583, github bounded local-report guidance +127, review complete-input refusal +35, and send_later dispatching status +14. Act/Operate full instead has github +2151 and additionally speech +230, hidden tts -2120, and tasks/automation exact model-route fields +274 each. The older budget predates v0.9.12: that tag had already removed 361 agent bytes and added the two 274-byte route fields; the retained initial increase versus the tag is 751 source-attributed bytes (agent +320, bash +431), not the +390 budget delta. Only the seven active definitions form the initial request; full catalogs include deferred tools. Estimated tokens use the existing bytes/4 heuristic, not provider usage or billing. Prompt, representative-context, skill-discovery and tool-name identities/ceilings are unchanged. Explicit v0.9.13 maintainer decision (2026-09-09): source ccc5dadfa2279545bf084d37cff3617e41ceaae2 intentionally exposes create_goal, get_goal and update_goal before continuation, so all three initial surfaces now contain ten tools. Measure exact source 4648d148eea64782be857eda6952af2c539cbfcc with the hosted macOS all-features libtest SHA-256 4485c88c7a8b66b8bb9a135807321e417a1e266457ecb74cffc7dfb92f850fc4: four exact provider-free metric tests pass. Active schemas are exactly 17847 bytes / 4462 estimated tokens (+4573 / +1143 for the three eager goal definitions); Plan full is 40597 / 10150 and Act/Operate full is 68315 / 17079, with no margin. The +712 full-catalog bytes are request_user_input guidance +358, update_goal state-change guidance +98, list_dir home-relative path guidance +56, explicit review max_passes schema +197, and three defer_loading true-to-false values +3. The three active name sets/digests, their counts, and measured active/full byte/token ceilings change; full name identities and all prompt, representative-context and skill-discovery measurements remain unchanged. This updates the earlier seven-tool initial-request receipt; bytes/4 remains an estimate, not measured provider usage or billing. The one-way numeric and exact identity gates remain enforced. Explicit v0.9.13 maintainer decision (2026-09-10): the +765 active/full tool-schema bytes since source 4648d148ee are exactly source-attributed \u2014 the agent tool's followup/parked-child continuation guidance (b6fad79373: +144 action description, +43 message parameter, +169 resume_from parameter) and the read tool's real output budget (e7f7c71e2c: +164 description, +245 for the new max_bytes parameter). No tool enters or leaves any surface: every name-set identity, count, and all prompt, representative-context and skill-discovery measurements are unchanged; only the measured byte/token ceilings move, to active 18612/4653 in every mode, Plan full 41362/10341, and Act/Operate full 69080/17270, with no margin. bytes/4 remains an estimate, not measured provider usage or billing. The one-way numeric and exact identity gates remain enforced. Explicit v0.9.13 maintainer decision (2026-09-13): source 5bfe88c8e8f45266ea1f9abce87c76b2eed894af intentionally keeps the native workflow tool eager on Plan/Act/Operate first-turn surfaces (DEFAULT_ACTIVE_NATIVE_TOOLS; commit 9e49d0918). Active name sets gain `workflow` (10\u219211 tools); active schema ceilings move to 29402/7351 with margin pending exact Linux --update lock-in. Full catalogs already advertised workflow; only a small defer_loading true\u2192false spelling bump is reserved (+32 bytes / +8 tokens). Prompt, representative-context, and skill-discovery measurements are unchanged. Do not remove workflow from Plan. bytes/4 remains an estimate, not measured provider usage or billing. The one-way numeric and exact identity gates remain enforced. Explicit v0.9.13 maintainer decision (2026-09-13, CI lock-in after b4d48e9a4): Linux Lint run 34766373771 measured the intentional eager `workflow` surface at active 31438/7860 in every mode, Plan full 50801/12701, and Act/Operate full 78513/19629. Prior ceilings (29402/7351 active, Plan full 41394/10349, Act/Operate full 69112/17278) under-counted the workflow schema body plus defer_loading true\u2192false on full catalogs; name-set identities are unchanged and `workflow` stays on Plan. Lock ceilings to those measured values with no margin. bytes/4 remains an estimate, not measured provider usage or billing. The one-way numeric and exact identity gates remain enforced. Explicit v0.9.14 maintainer decision (2026-09-16): #5715 intentionally adds the two bounded read-only recall tools `session_get` and `session_search` to the Act/Operate full catalogs (50 -> 52 tools), so both full name-set identities and their digests move to 1203d192385fd2b02227ef9e4212e5379bc5cbb813a373f12539406b5958aef1. Plan full is unchanged: the session tools are not offered there. Measured on macOS aarch64 all-features from source 55a9e1b778fa; per the 2026-09-13 precedent the exact byte/token ceilings must be re-locked from a Linux Lint run if CI reports drift. bytes/4 remains an estimate, not measured provider usage or billing. The one-way numeric and exact identity gates remain enforced. Explicit v0.9.14 maintainer decision (2026-09-17, #6319): re-measure from source b0866943b4 on macOS aarch64 all-features; per the 2026-09-13 precedent the exact byte/token ceilings must be re-locked from a Linux Lint run if CI reports drift. No tool enters or leaves any surface (52 tools; every name-set identity unchanged). Representative stages and system prompt grow exactly +742 bytes per stage/mode (base 6089 -> 6831, prompt 6084 -> 6826) for the deliberate dc32272f15 Bearing article plus mandate-first scope law; every stage re-hashed. Tool growth is exactly source-attributed per tool (measured per-tool at 55a9e1b778fa vs HEAD): agent +670 (2ca54ea8da #6282 output-token-cap parameter, 3c9c62571a spawn-requirement docs, 85d8dc7501 #6278 exact_files sentence, less 3dcb41f5d2 #6272 release-clause trim and the a7a8bdb338 token-allowance description removal), workflow +424 (a035914336 #6232 plan-child cwd property, serialized twice via phases/items and top-level children/items), Git +867 (b89349286f #6298 merge_tree verify surface), Run +158 (233da9fb76 #6296 bounded cwd), read +91 (f6fb5f42d1 #6283 size/truncated/line_count response fields), create_goal +76 (4de9e9e281 model-decides-goals description rewrite). Active 31453 -> 32714 (+1261 = agent +670, workflow +424, read +91, create_goal +76); Plan full 50816 -> 52944 (+2128 = active +1261 plus Git +867); Act/Operate full 79531 -> 81817 (+2286 = Plan full +2128 plus Run +158). bytes/4 remains an estimate, not measured provider usage or billing. The one-way numeric and exact identity gates remain enforced. Explicit v0.9.14 maintainer decision (2026-09-19 overnight): Code Mode Phase-1 advertises `execute_tools` on Act/Operate full catalogs (52 -> 53 tools; Plan full unchanged). Default Direct mode keeps it deferred (`defer_loading: true`), so active surfaces are unchanged. Name-set identity digests move with the sorted name list; measured tool JSON is 884 bytes (+885 including the catalog comma) so Act/Operate full ceilings lock to 82702/20676. bytes/4 remains an estimate, not measured provider usage or billing. The one-way numeric and exact identity gates remain enforced. Explicit v0.10.0 source reconciliation (2026-09-19): retain the existing agent cwd parameter (+256 serialized bytes in every surface) and tasks create name parameter (+121 bytes in Act/Operate full only), both already present before the preceding execute_tools lock-in. Linux CI run 35490467762 measured active 32970/8243, Plan full 53200/13300, and Act/Operate full 83079/20770. Each increase is attributed to exactly one schema property; no margin, tool identity, prompt or representative-context change. These are schema byte/4 estimates, not billed tokens. Explicit v0.10.0 maintainer decision (2026-09-22): 91b5898a2 intentionally makes `load_skill` eager so the `## Skills` instruction in the base prompt is actually callable \u2014 it was deferred behind tool_search, so the prompt told the model to call a tool that was not in the array it was printed beside. load_skill is read-only, so it joins the active sets of Plan, Act and Operate (11 -> 12 tools; +816 schema bytes / +204 estimated tokens in every mode, active ceilings lock to 33786/8447). Full catalogs move only by the defer_loading true->false spelling plus that schema body (+245 bytes / +61-62 tokens). Representative stages re-hash and each move +149 bytes with the total +37 tokens: dc32272f15 landed the Bearing article and mandate-first scope law, and 91b5898a2 rewrote the Skills paragraph every stage renders. No other name-set identity, count, prompt or skill-discovery measurement changes. Measured from 6bafe6e72 on macOS aarch64 all-features; per the 2026-09-13 precedent the exact byte/token ceilings must be re-locked from a Linux Lint run if CI reports drift. bytes/4 remains an estimate, not measured provider usage or billing. The one-way numeric and exact identity gates remain enforced. Explicit maintainer decision (#6562, PR #6583, 2026-09-25): code mode for MCP is on by default ([features] code_mode = true), so `execute_tools` is eager on the Act/Operate first-turn surfaces (12 -> 13 tools; active 33876/8469 -> 35406/8852, +1530 bytes / +383 estimated tokens) and its definition text now documents nested MCP/plugin calls through the direct-call gate (Act/Operate full 83384/20846 -> 84029/21008, +645 bytes). Plan is unchanged: execute_tools is not offered there. `code_mode = false` restores the prior deferred surface. Measured on macOS aarch64 all-features; re-lock from a Linux Lint run if CI reports drift. Re-measured after merging main 2026-09-25: active 35392/8848 (-14, main's 4b0f54edb fork_context trim), Act/Operate full 83737/20935 (-292: the same -14 plus -278 not attributed per tool in this merge); decreases locked with --update, identities unchanged. bytes/4 remains an estimate, not measured provider usage or billing. Explicit maintainer decision (PR #6589, 2026-09-25): the agent tool declares `runtime` (local|cloud; cloud only proposes a /dispatch job the person confirms) and `remote` (github|cnb|gitee) as schema properties so the model can request a cloud proposal without an undeclared field. Both descriptions were trimmed (-303 bytes) before locking, which also keeps the parent prompt+catalog surface under PARENT_SURFACE_BYTE_CEILING; the remaining +284 bytes / +71 estimated tokens land on every surface that carries agent (Plan/Act/Operate active and full: active 35676/8919 Act/Operate, 34146/8537 Plan; full 84021/21006 Act/Operate, 53785/13447 Plan). No tool enters or leaves any surface; identities unchanged. Measured on macOS aarch64 all-features; re-lock from a Linux Lint run if CI reports drift. bytes/4 remains an estimate, not measured provider usage or billing.",
3 "document_kind": "codewhale.runtime_contract_budget",
4 "representative_context": {
5 "fixture_id": "representative-v1",
6 "stages": {
7 "base": {
8 "bytes": 7231,
9 "identity_sha256": "5130806e324482a4b5ec28ae6fc408846b894844db9b7c38f3cfcf48e2ac63d8"
10 },
11 "goal": {
12 "bytes": 9415,
13 "delta_bytes": 81,
14 "identity_sha256": "7f004aa61453193c3ad328efb86bc70ad57d86ee5f28af5e4cb05d228a18d817"
15 },
16 "handoff": {
17 "bytes": 9803,
18 "delta_bytes": 388,
19 "identity_sha256": "350a47b56a210bf3cd027204d9db5e791e2e2c26af5ab37ca514415ba1527c65"
20 },
21 "instructions": {
22 "bytes": 7605,
23 "delta_bytes": 131,
24 "identity_sha256": "85bd54a1ca30028f852c6a531c6b1aa04b6b971f1ed3498c66e942f055b276c1"
25 },
26 "memory": {
27 "bytes": 9334,
28 "delta_bytes": 963,
29 "identity_sha256": "d21379b0fa308b121649e5d756da4d8b6e31d5a0b10e2d87b5511197a7959bcd"
30 },
31 "project": {
32 "bytes": 7474,
33 "delta_bytes": 243,
34 "identity_sha256": "918a55fd0c6da2a1334cdad93b2da61bcee660723ebc0443cd0a7b045979a0f8"
35 },
36 "skill": {
37 "bytes": 8371,
38 "delta_bytes": 766,
39 "identity_sha256": "1cb100b5a923dead5fe2486539a37e9ab11ab03c82cb00bfa8eb3de2cb5c4a38"
40 }
41 },
42 "system_prompt_blocks": 6,
43 "total_bytes": 9803,
44 "total_tokens_est": 2451
45 },
46 "schema_version": 1,
47 "skill_discovery": {
48 "first_delta": {
49 "directories_visited": 1,
50 "root_discovery_calls": 1,
51 "skill_md_read_attempts": 1
52 },
53 "second_delta": {
54 "directories_visited": 0,
55 "root_discovery_calls": 0,
56 "skill_md_read_attempts": 0
57 }
58 },
59 "system_prompt": {
60 "modes": {
61 "act": {
62 "mode_instructions_bytes": 0,
63 "mode_instructions_tokens_est": 0,
64 "system_prompt_blocks": 4,
65 "system_prompt_bytes": 7226,
66 "system_prompt_tokens_est": 1807
67 },
68 "operate": {
69 "mode_instructions_bytes": 0,
70 "mode_instructions_tokens_est": 0,
71 "system_prompt_blocks": 4,
72 "system_prompt_bytes": 7226,
73 "system_prompt_tokens_est": 1807
74 },
75 "plan": {
76 "mode_instructions_bytes": 0,
77 "mode_instructions_tokens_est": 0,
78 "system_prompt_blocks": 4,
79 "system_prompt_bytes": 7226,
80 "system_prompt_tokens_est": 1807
81 }
82 }
83 },
84 "tool_catalog": {
85 "execution_shell": "bash",
86 "modes": {
87 "act": {
88 "active": {
89 "bytes": 35676,
90 "identity_sha256": "cc8f1f208bcf83261451f1ff67f7bf65534616cfd50be0102f50aba009a9e59d",
91 "tokens_est": 8919,
92 "tool_names": [
93 "agent",
94 "bash",
95 "create_goal",
96 "edit",
97 "execute_tools",
98 "get_goal",
99 "load_skill",
100 "read",
101 "todo_write",
102 "tool_search",
103 "update_goal",
104 "workflow",
105 "write"
106 ],
107 "tools": 13
108 },
109 "full": {
110 "bytes": 84151,
111 "identity_sha256": "45e989bbe5ac0bb1f2d9084361c021009f30a06539f132ebd4fd2331a1bb1954",
112 "tokens_est": 21038,
113 "tool_names": [
114 "Git",
115 "Run",
116 "Web",
117 "agent",
118 "apply_patch",
119 "automation",
120 "bash",
121 "create_goal",
122 "diagnostics",
123 "edit",
124 "execute_tools",
125 "file_search",
126 "fim_edit",
127 "finance",
128 "get_goal",
129 "github",
130 "grep_files",
131 "handle_read",
132 "harness",
133 "list_dir",
134 "load_skill",
135 "lsp",
136 "note",
137 "notify",
138 "project_map",
139 "read",
140 "read_media",
141 "request_plugin_install",
142 "request_user_input",
143 "retrieve_tool_result",
144 "revert_turn",
145 "review",
146 "send_later",
147 "session_get",
148 "session_search",
149 "speech",
150 "task_shell_start",
151 "task_shell_wait",
152 "tasks",
153 "terminal/cancel",
154 "terminal/reset",
155 "terminal/run",
156 "terminal/send",
157 "terminal/wait",
158 "todo_write",
159 "tool_search",
160 "tui_help",
161 "update_goal",
162 "validate_data",
163 "verify",
164 "web.run",
165 "workflow",
166 "write"
167 ],
168 "tools": 53
169 }
170 },
171 "operate": {
172 "active": {
173 "bytes": 35676,
174 "identity_sha256": "cc8f1f208bcf83261451f1ff67f7bf65534616cfd50be0102f50aba009a9e59d",
175 "tokens_est": 8919,
176 "tool_names": [
177 "agent",
178 "bash",
179 "create_goal",
180 "edit",
181 "execute_tools",
182 "get_goal",
183 "load_skill",
184 "read",
185 "todo_write",
186 "tool_search",
187 "update_goal",
188 "workflow",
189 "write"
190 ],
191 "tools": 13
192 },
193 "full": {
194 "bytes": 84151,
195 "identity_sha256": "45e989bbe5ac0bb1f2d9084361c021009f30a06539f132ebd4fd2331a1bb1954",
196 "tokens_est": 21038,
197 "tool_names": [
198 "Git",
199 "Run",
200 "Web",
201 "agent",
202 "apply_patch",
203 "automation",
204 "bash",
205 "create_goal",
206 "diagnostics",
207 "edit",
208 "execute_tools",
209 "file_search",
210 "fim_edit",
211 "finance",
212 "get_goal",
213 "github",
214 "grep_files",
215 "handle_read",
216 "harness",
217 "list_dir",
218 "load_skill",
219 "lsp",
220 "note",
221 "notify",
222 "project_map",
223 "read",
224 "read_media",
225 "request_plugin_install",
226 "request_user_input",
227 "retrieve_tool_result",
228 "revert_turn",
229 "review",
230 "send_later",
231 "session_get",
232 "session_search",
233 "speech",
234 "task_shell_start",
235 "task_shell_wait",
236 "tasks",
237 "terminal/cancel",
238 "terminal/reset",
239 "terminal/run",
240 "terminal/send",
241 "terminal/wait",
242 "todo_write",
243 "tool_search",
244 "tui_help",
245 "update_goal",
246 "validate_data",
247 "verify",
248 "web.run",
249 "workflow",
250 "write"
251 ],
252 "tools": 53
253 }
254 },
255 "plan": {
256 "active": {
257 "bytes": 34146,
258 "identity_sha256": "df6676989a677fb08fecc8bf7ae12caf4fa88e0cae3d143cea6a4da4382b746c",
259 "tokens_est": 8537,
260 "tool_names": [
261 "agent",
262 "bash",
263 "create_goal",
264 "edit",
265 "get_goal",
266 "load_skill",
267 "read",
268 "todo_write",
269 "tool_search",
270 "update_goal",
271 "workflow",
272 "write"
273 ],
274 "tools": 12
275 },
276 "full": {
277 "bytes": 53772,
278 "identity_sha256": "ac8af1f4988199825be7b00b054c258724a44074b1d4e6de6c92ade7c1cffe63",
279 "tokens_est": 13443,
280 "tool_names": [
281 "Git",
282 "Web",
283 "agent",
284 "automation",
285 "bash",
286 "create_goal",
287 "diagnostics",
288 "edit",
289 "file_search",
290 "get_goal",
291 "github",
292 "grep_files",
293 "handle_read",
294 "list_dir",
295 "load_skill",
296 "notify",
297 "read",
298 "read_media",
299 "request_plugin_install",
300 "request_user_input",
301 "review",
302 "send_later",
303 "tasks",
304 "todo_write",
305 "tool_search",
306 "update_goal",
307 "validate_data",
308 "web.run",
309 "workflow",
310 "write"
311 ],
312 "tools": 30
313 }
314 }
315 },
316 "surface_profile": "production-default-builtins-no-mcp-no-host-interpreters-bash-v2"
317 },
318 "_codex_0101_remeasure": "2026-10-01: lock the composed candidate measurement after the audited composition-required schema repair (46835a2fc, D04-11) and shared finance-timeout documentation (7c36620d4, D03-m3). Their parent-surface causal controls are recorded in 6655b191e. Act/Operate full catalog measured 84151 bytes / 21038 estimated tokens versus stale 83988 / 20997; no tools or surface identities changed. Plan/active surfaces and other budgets use their measured values. Estimate only, not provider usage. Receipt: CW/artifacts/codex-0101-takeover-20260930/root-runtime-contract-receipt.json."
319 }
320
320 lines JSON