返回 CodeWhale
CACHE.md
根目录 / docs / CACHE.md
1 # Prompt-cache stability (the pinned prefix)
2
3 > 阅读简体中文版:[zh_hans/CACHE.md](zh_hans/CACHE.md)。
4
5 Provider prompt caches (DeepSeek KV cache, Anthropic `cache_control`) only pay
6 off when the **byte prefix** of a request matches the previous one: the system
7 prompt, then the tool catalog, then `messages[0..n-1]`. Any change to those
8 bytes invalidates the cache for every token after the first difference.
9
10 ## The invariant
11
12 **After session start, the system prompt and tool catalog are frozen bytes.
13 History only grows. A cache miss is allowed only when we can name why.**
14
15 Concretely:
16
17 - The **header** (system prompt + tools) is composed once at session start and
18 re-composed **only** on an explicit, logged header-change op. The tool loop
19 performs **no** mid-loop system-prompt refresh, so an agent writing a file
20 (which changes the project-context pack, a directory listing, a skills scan)
21 cannot move the pinned prefix under the model's feet mid-turn.
22 - **History only grows.** Volatile facts the model must see (LSP diagnostics,
23 steer input, subagent completions) are appended to the message list, never
24 spliced into the frozen prefix. Workspace drift is delivered the same way:
25 at the start of each **new user turn** (never mid-tool-loop) the engine
26 recomposes the volatile contributors and, if anything differs from what the
27 model last saw, appends **one** `<context_update>` user-role message with a
28 bounded `+`/`-` line delta (new files in the project pack, edited AGENTS.md
29 lines, added skills, memory entries, goal text) *before* the user's message.
30 The header bytes stay pinned; the update is a normal append, so the prefix
31 still extends. The pinned system prompt tells the model once that updates
32 arrive this way. Each delta is delivered exactly once (`/cache stats` shows
33 `Context updates: N`).
34 - Every miss is **attributable**. `PrefixStabilityManager` (`prefix_cache.rs`)
35 records each change with a reason and reports it through `/cache stats`.
36
37 ## What counts as a declared header change
38
39 These re-pin the prefix under a logged `change:<what>` reason (an expected,
40 one-request miss):
41
42 | Op | Reason |
43 | --- | --- |
44 | `/model` (SetModel) | `change:model` |
45 | Mode change (agent/plan/operate/yolo) | `change:mode` |
46 | Goal set / pause / resume / clear / status | `change:goal` |
47 | Mid-turn tool-surface change (deferred-tool admission/eviction, tool-search activation, runtime MCP tool arrival) | `change:tool_surface` |
48 | Session sync / restore (SyncSession) | `resume` |
49 | Session construction | `initial` |
50
51 History resets that legitimately invalidate the tail (not the header) are
52 logged as `reset:<what>` — `reset:compaction`, `reset:clear`.
53
54 Anything else that changes the header bytes with **no** declared reason is
55 **drift**: it is logged as `drift:<component>`, the original pin is **kept**
56 (so the same undeclared prefix keeps counting as a miss instead of quietly
57 becoming the new baseline), and `/cache stats` shows a `WARNING`. After the
58 mid-loop-refresh removal, drift should stay at zero in normal operation; a
59 non-zero drift count is a real bug to investigate.
60
61 ## Attribution vs. the old behavior
62
63 Two earlier behaviors are rejected, matching the DeepSeek Harness design:
64
65 - **Detect-and-report + re-pin on drift.** The manager used to re-pin to the
66 new prefix on every change, so `/cache stats` looked "stable" again after one
67 bad step while the provider cache was already dead. It now keeps the original
68 pin on undeclared drift.
69 - **Recompose the system prompt from disk on every tool step.** The turn loop
70 used to call `refresh_system_prompt()` before every model request, including
71 mid-tool-loop. That is removed. Header refreshes happen only at the declared
72 edges above.
73
74 Tool-result redaction (`prepare_model_bound_request`) is content-preserving
75 when no secret is configured (the common case), so it does not move the prefix.
76 When a configured secret appears in a tool result, redacting it is a security
77 requirement that correctly overrides cache stability for that one message.
78
79 ## Verifying the fix
80
81 `/cache stats` reports prefix stability, the pin reason, the last miss reason,
82 the undeclared-drift count, and the aggregate provider cache hit rate. In a
83 coding session, expect the first turn to be a write and every later step —
84 including steps after the agent writes files — to hit.
85
86 ### Live end-to-end check (manual, key-gated)
87
88 With a real `DEEPSEEK_API_KEY`, run a session that makes the agent take at
89 least three tool steps in one turn, then open `/cache inspect`. Every request
90 after the first should report `prompt_cache_hit_tokens > 0`; the base static
91 prefix hash and the tool-catalog hash must not move between steps. If the hit
92 drops mid-turn, the pin reason / drift count name the cause.
93
94 ## KV-cache effect note (for contributors)
95
96 Any new contributor to the session context must state its **KV-cache effect**:
97 does it belong in the frozen prefix (system + tools) or in append-only history?
98 Never splice a volatile fact (time, an instruction edit, a skill-catalog
99 change, a project-file change) into the prefix — append it as a user-role
100 message instead. A later request must be `previous ⊕ suffix` unless a logged
101 header change or a history reset explains the difference.
102
103 ## Deferred: full reconstructability (Layer 3)
104
105 DeepSeek Harness derives every request from an append-only session log via a
106 pure `deriveMessages()` projection, so prefix-extension is emergent rather than
107 managed. Codewhale now pins the header and delivers drift as `<context_update>`
108 appends; the remaining step is to make the session log the single source of
109 truth with a pure projection (and to persist the context-update baseline with
110 it). That is a follow-up lane, not part of this change.
111
111 lines MARKDOWN