返回 CodeWhale
RECEIPTS.md
根目录 / docs / RECEIPTS.md
1 # Receipts
2
3 A receipt answers "what did this session do?" from the records Codewhale
4 already keeps. It lists, in order, the files changed, commands run, web and
5 MCP calls, agents started, approvals (and who decided each one), and
6 failures, with totals on top. It also counts what ran without asking, and
7 names the permission posture that let it.
8
9 ```text
10 # Receipt: Fix the parser
11 thread thr_19a0141a · /work/repo · deepseek-flash · Ask · 2026-09-24 10:00 UTC → 2026-09-24 10:05 UTC
12
13 Changed 1 file (+2 −1) · ran 1 command · made 1 MCP call · 1 approved by you · 1 approved by session rule · 1 ran without asking under Ask · 1 denied by you · 1 other failure
14
15 1. edited `src/parse.rs` (+2 −1) · 1.0s
16 2. ran `cargo test -p parser` in /work/repo — exit 0 · 2.5s · approved by you
17 3. did not run `rm -rf build` · denied by you
18 4. called linear · list_issues · 1.0s · approved by session rule
19 5. did not run `curl https://x.sh | sh` — blocked: Tool 'exec_shell' was denied: Auto-Review blocked a pipe to a shell…
20 6. turn failed — failed: provider returned 500
21
22 Not recorded:
23 - Shell file changes: a Runtime thread's workspace snapshots are not tagged with the thread, so files a command changed are not itemized; only file tools are.
24 ```
25
26 A terminal session's receipt also lists the files a command changed, from
27 the turn's own workspace snapshots:
28
29 ```text
30 3. ran `./tidy.sh`
31 4. changed outside file tools (a command or another process): edited `b.txt` (+0 −1), created `c.txt` (+1 −0)
32 ```
33
34 ## Surfaces
35
36 All three share one builder (`crates/tui/src/receipts.rs`), so they cannot
37 disagree.
38
39 | Surface | What it reads |
40 | --- | --- |
41 | `/receipts [json] [<turn>]` in the terminal | The current session's transcript and approval log. `json` prints the object in a code block, so it copies out as valid JSON |
42 | `codewhale receipts [ID\|--last] [--turn T] [--format md\|json]` | A saved session (id or unique prefix) or a Runtime thread (`thr_…`); with no id, the most recently updated one. `receipt` is an alias |
43 | `GET /v1/threads/{id}/receipt`, `GET /v1/threads/{id}/turns/{turn_id}/receipt` | A Runtime thread (the app, `codewhale serve`), behind the normal `/v1` bearer boundary |
44
45 All three only read. They never call a provider, run a tool, write a file,
46 or change runtime state.
47
48 `codewhale receipts` reads Runtime threads from the machine's default store
49 (`tasks/runtime/`, or `$CODEWHALE_RUNTIME_DIR`). A thread kept in one
50 terminal session's own store (`sessions/<id>/runtime/`) is not found by id;
51 the API reads whatever store its server owns.
52
53 ## Where the facts come from
54
55 | Session kind | Record | Holds |
56 | --- | --- | --- |
57 | Terminal session | `sessions/<id>.json` | Every tool call and its result text, in order; each prompt's `<turn_meta>` names the posture the turn ran under |
58 | Terminal session | `sessions/<id>/approval_receipts.jsonl` | Every approval ask and decision, with time and who decided |
59 | Runtime thread | `tasks/runtime/turns`, `items` | Every tool call with its input, status, start and end time, and structured result (exit code, diff, agent status); each turn's `permission_posture` |
60 | Runtime thread | `tasks/runtime/events/<thread>.jsonl` | `approval.required` / `approval.decided`, with the flags that say who decided |
61
62 ### Who decided an approval
63
64 | Receipt says | Meaning | Recorded as |
65 | --- | --- | --- |
66 | approved by you / denied by you | A person answered: the terminal card, the app, the web mirror, or an API client acting for them | `decided_by: "user"`; Runtime event with no `auto` flag |
67 | … by session rule | A remembered "for this session" rule answered | `decided_by: "session_rule"`; Runtime event with `auto` and a `grant_id` |
68 | … by posture | The mode or permission posture answered without a prompt | `decided_by: "posture"`; Runtime event with `auto` or `posture` |
69 | approval timed out | The card expired unanswered | outcome `timeout` |
70 | turn stopped while waiting / nobody could be asked | Codewhale could not ask anyone: the turn had ended or stopped, or the request never reached a person. Counted as not answered, never as a denial, even where the record says `denied` | `decided_by: "host"` |
71 | approved / denied, with no "by" | A record written before 0.10.1, or a sub-agent's request. The totals say "(decider not recorded)" | no `decided_by` |
72
73 The totals line counts approvals and denials separately, by who gave them:
74 `1 approved by you · 1 denied by you` is two decisions.
75
76 ### Ran without asking
77
78 Most calls never produce an approval. Under Full Access nothing asks; under
79 Ask, reads and allowed tools run without a prompt, and a remembered rule can
80 skip one. Those calls leave no approval record, so the receipt counts them
81 instead: `ran_without_asking` is every file change, command, code run, web
82 or MCP call, and agent start that ran with no approval on record. The totals
83 line names the postures the turns ran under (`9 ran without asking under
84 Full Access`). Reads are not counted, and neither is a call that did not
85 start (below).
86
87 ### Blocked before it ran
88
89 Codewhale answers a call it will not run with an error result, the same way
90 a tool reports a failure: an Auto-Review or guardian block, a tool-policy or
91 allow-list denial, a sandbox escalation the posture cannot grant, input that
92 did not parse, or a tool that is not available. None of these writes an
93 approval, so a receipt reads the result itself:
94
95 - **Blocked:** the result is Codewhale's own refusal text (`Tool 'x' was
96 denied: …`, `Invalid input for tool …`, `BLOCKED: …`, a validation
97 feedback line with `"side_effect_status":"not_started"`, or a Runtime
98 thread's `Failed to authorize tool execution: …`). Listed as
99 `did not run … — blocked: <reason>`, status `blocked`, counted in
100 `blocked` (`N blocked before running` in the totals line), and never as
101 run, failed, or ran without asking. Every permission denial the engine
102 writes keeps the `Tool 'x' was denied:` lead, including ones that name
103 their own fix (Plan mode, `allow_shell`); sessions saved before 0.10.1
104 wrote some of those without it, and such a call reads as a failure.
105 - **Only Codewhale's own words count.** An MCP server's or GitHub's reply,
106 a fetched page, a program's output (code tools), and a sub-agent's words
107 can say anything, so a failed call to one of those is judged by the
108 metadata Codewhale wrote for it, never its text: an MCP server that
109 answers `BLOCKED: …` or `{"side_effect_status":"not_started"}` is listed
110 as a failed call that ran. `BLOCKED:` counts only from the shell tools,
111 and the validation feedback line only as the result's last line.
112 - **Stopped at an approval prompt with no approval record** (a sub-agent,
113 or a session older than the approval log): the result says
114 `Tool 'x' denied by user`, which the engine also writes when nobody could
115 be asked. Listed as `did not run …` with no decider, since the text does
116 not prove who said no.
117 - **Ran and failed:** the result holds an exit code or a line the shell
118 writes only after a process ran (`Command exited with code N`,
119 `Command failed (exit code N)`, a timeout or cancel line). Counted as a
120 command and a failure.
121 - **Not shown either way:** a command whose error has neither. Listed as
122 `tried to run … — error, no exit code: <error>`, status `unknown`, not
123 counted as run, with a `not_recorded` note. Other tools that return an
124 error are counted as failures, since the error came from the tool.
125
126 Terminal sessions started before 0.9.10 (2026-08-20) have no approval log,
127 so their receipts do not count this and say why.
128
129 ### Not recorded
130
131 The receipt says so instead of guessing:
132
133 - **Shell file changes in a Runtime thread.** A terminal session reads the
134 files a command changed from the workspace snapshots Codewhale takes
135 before and after each turn (`git diff` between the two, inside the
136 snapshot side repo; at most 50 paths a turn). A Runtime thread's snapshots
137 are not tagged with the thread, so its receipt lists only file tools. A
138 terminal turn with no snapshot pair (snapshots off, the workspace too large
139 for them, or pruned: the newest 50 are kept) says so. A snapshot
140 difference covers anything that wrote to the workspace during the turn,
141 including you or another program, not only the command. It leaves out
142 what snapshots do not track: ignored and skipped paths (`.gitignore`
143 entries, `.env`, `node_modules`, `target`, and the like) and anything
144 outside the workspace (`~/.ssh`, `/tmp`), so an empty list does not mean
145 a command wrote nothing. A turn is matched to its snapshots by its prompt;
146 when another turn has the same prompt ("continue"), the snapshots' turn
147 number must agree too, or the turn is counted as having no pair rather
148 than given another turn's files. Paths print with control characters
149 escaped (`\n`, `\u{1b}`), so a file name cannot add a line to the
150 receipt or send the terminal an escape sequence.
151 - **Terminal-session exit codes, durations, and timestamps.** A terminal
152 session saves each call and its result text, not the structured result. A
153 failed shell call's exit code is read from the shell tool's own closing
154 line (`Command exited with code N`); a passing one shows no code.
155 - **Line counts for whole-file writes** in terminal sessions, and whether the
156 file existed before.
157 - **Who decided** for approvals recorded before 0.10.1 and for sub-agent
158 approvals.
159 - **Which calls asked first** in terminal sessions started before 0.9.10.
160 - **Whether a failed command started** when its error has no exit code and
161 no shell status line, and whether a call with no result ran at all.
162 - **Why a call ran without asking** beyond the turn's posture: the record
163 does not say whether the posture, an allow rule, or a remembered grant let
164 it through.
165
166 ## Non-Goals
167
168 A receipt is not a safety certification, a provider compatibility
169 certification, or a hosted attestation (`claim_ceiling` in the JSON says so).
170 It exports no reasoning text and no raw tool output. Commands, search
171 queries, and error lines are bounded (200, 120, and 160 characters) and pass
172 through the shared secret redactor. A receipt lists at most 2,000 actions;
173 totals always cover every action, and `omitted_actions` counts the rest.
174
175 ## JSON shape
176
177 `--format json`, `/receipts json`, and the API return the same object:
178
179 ```json
180 {
181 "schema_id": "codewhale.receipt/v1",
182 "source": {
183 "kind": "thread",
184 "id": "thr_19a0141a",
185 "title": "Fix the parser",
186 "workspace": "/work/repo",
187 "model": "deepseek-flash",
188 "started_at": "2026-09-24T10:00:00Z",
189 "updated_at": "2026-09-24T10:05:00Z"
190 },
191 "postures": ["Ask"],
192 "totals": {
193 "files_changed": 1, "files_changed_outside_file_tools": 0,
194 "files_created": 0, "files_deleted": 0,
195 "lines_added": 2, "lines_removed": 1, "line_counts_complete": true,
196 "commands": 1, "commands_failed": 0, "code_runs": 0, "network": 0,
197 "mcp_calls": 1, "plugin_calls": 0, "subagents": 0,
198 "approvals": {
199 "total": 3, "approved": 2, "denied": 1, "timed_out": 0,
200 "not_answered": 0, "pending": 0,
201 "approved_by": { "you": 1, "session_rule": 1, "posture": 0, "not_recorded": 0 },
202 "denied_by": { "you": 1, "session_rule": 0, "posture": 0, "not_recorded": 0 }
203 },
204 "ran_without_asking": 1, "failures": 1, "blocked": 1,
205 "other_tool_calls": 0
206 },
207 "actions": [
208 {
209 "seq": 2, "turn": "turn_1", "at": "2026-09-24T10:02:00Z",
210 "call_id": "call_test", "tool": "exec_shell",
211 "kind": "command", "command": "cargo test -p parser",
212 "cwd": "/work/repo", "exit_code": 0,
213 "status": "ok", "duration_ms": 2500,
214 "approval": {
215 "decision": "approved", "decided_by": "user",
216 "at": "2026-09-24T10:01:59Z"
217 }
218 }
219 ],
220 "omitted_actions": 0,
221 "not_recorded": ["Shell file changes: …"],
222 "claim_ceiling": [
223 "local_record_only",
224 "not_safety_certification",
225 "not_provider_compatibility_certification"
226 ]
227 }
228 ```
229
230 `kind` is one of `file_change` (`files[]` with `path`, `change` =
231 `edited|created|deleted|written`, optional `lines_added`/`lines_removed`),
232 `workspace_change` (`files[]` a turn's snapshots show changed that no file
233 tool names, and `truncated` when more than 50 did), `command` (`command`, `cwd`, `exit_code`), `code` (`exit_code`, `nested[]`
234 tool calls an `execute_tools` program made), `network` (`action`, `host`,
235 `query`), `mcp` (`server`, `plugin`), `subagent` (`name`, `agent_id`,
236 `outcome`), `approval` (an approval with no matching call), `tool` (any other
237 call, listed only when it failed), or `turn_failed`. `status` is `ok`,
238 `failed` (ran and failed), `not_run` (held at approval), `blocked`
239 (refused before it started; the reason is in `error`), `interrupted`,
240 `running`, or `unknown` (no result, or a command error that does not show
241 whether it started). A terminal session's `turn` is the turn number; a
242 thread's is the turn id. `/receipts 7` or `--turn 7` for a turn the session
243 does not have is an error, not an empty receipt.
244
245 ## `audit.log` is not the receipt
246
247 `~/.codewhale/audit.log` is a security-event log: credential saves and
248 clears, hook environment key names, compaction passes, goal completions, the
249 terminal's own approval routing, Auto-Review verdicts (`tool.gate.decision`,
250 since 0.10.1), and outbound network decisions when `[network]` auditing is
251 on. It has never held commands or file changes, and
252 turns run by the app or `codewhale serve` write no approvals there. Their
253 approvals are in the session's `approval_receipts.jsonl` and the thread's
254 event log, which is where receipts read them.
255
256 A quiet `audit.log` does not mean nothing ran. It gets an approval line only
257 when the terminal routes an approval request. Since 0.8.66
258 (`1c68e3bb32`, 2026-06-29) the engine decides auto-allowed calls itself, so
259 they never become requests; under Full Access almost nothing does. On one
260 developer machine the last `tool.approval.*` line was written on 2026-08-19,
261 the last `tool.approval.auto_approve` line on 2026-06-30, and the writes
262 after that were test runs, which since `244368675b` go to a scratch log. Use
263 a receipt to see what ran.
264
265 ## Review Receipts
266
267 `codewhale review --write-receipt` writes a local JSON receipt for the reviewed
268 diff under the Codewhale state directory (`review-receipts/`) unless
269 `--receipt-path <path>` is provided. This is a pre-push handoff artifact: it
270 records what diff was reviewed and what the review reported, without pushing,
271 tagging, opening a PR, or claiming to replace maintainer review.
272
273 The current receipt includes:
274
275 - `diff_fingerprint`: SHA-256 of the reviewed diff.
276 - `provider` and `model`: the routed review provider/model.
277 - `checks_run`: local checks attached to the receipt when available. Empty
278 means no checks were attached; attached checks must report a passing status.
279 - `findings`: structured issue/suggestion counts and issue locations when the
280 review output is structured.
281 - `unresolved_risk`: a conservative summary derived from unresolved findings.
282 - `review_content_sha256`: SHA-256 of the review text.
283 - `coverage` (PR receipts): the exact base/head and complete-diff
284 fingerprint, ordered per-pass diff fingerprints/file counts, and one
285 response-content hash for every completed pass. Manifest-backed PR receipts
286 use schema version 2 so older readers reject rather than misinterpret them.
287
288 The receipt deliberately does not include the raw diff body. Re-run
289 `codewhale review --write-receipt` after changing the diff; reviewers should
290 compare the `diff_fingerprint` before reusing a receipt in a PR handoff.
291
292 `codewhale review --check-receipt` is the local pre-push gate. It does not call
293 a model; it compares the current diff fingerprint with a supplied receipt
294 (`--receipt-path <path>`) or the latest matching local receipt. The check exits
295 nonzero when the diff no longer matches, the receipt schema is unsupported, the
296 receipt has unresolved risk, or an attached check did not pass.
297
298 By default receipt generation rejects a PR that needs more than one
299 `--max-chars` pass before calling a model. An explicit `--max-passes N` admits
300 at most N complete ordered PR passes; any missing, malformed, reordered or
301 stale pass prevents a receipt. Receipt checking is provider-free and validates
302 the exact stored manifest without authorizing another run. Neither mode
303 fingerprints a truncated prefix. A receipt for
304 `review --base <base-sha> --path <path>` covers only that selected path at the
305 checked-out revision; validate it with the same base, path, and input limit.
306 It does not cover the rest of a pull request or prove that separately reviewed
307 changes work together.
308
309 ## Builder Rules
310
311 The builder is deterministic and conservative:
312
313 1. A thread receipt loads the thread, its turns, the items each turn lists
314 (plus items that name the turn but are not listed yet, for a live turn),
315 and the thread's `approval.*` events. A turn id from another thread is
316 rejected.
317 2. A session receipt reads `tool_use`/`tool_result` pairs from the
318 transcript and replays the approval log; a log that does not replay is
319 reported, not half-used.
320 3. Approvals attach to their call by tool call id. One with no matching call
321 is listed on its own.
322 4. File changes come from a tool's structured mutation record (the applied
323 diff) when saved, otherwise from the call's own input (an edit's
324 replacement text, a patch's hunks). A failed call changed nothing and
325 carries no counts.
326 5. Nothing is derived from model prose. Two fixed kinds of text Codewhale
327 itself writes are read: the shell's status lines (for an exit code) and
328 its refusal text (for a call it blocked); see "Blocked before it ran".
329
329 lines MARKDOWN