返回 CodeWhale
LIVE_SMOKE.md
根目录 / docs / LIVE_SMOKE.md
1 # Opt-in live smoke runs
2
3 > 阅读简体中文版:[zh_hans/LIVE_SMOKE.md](zh_hans/LIVE_SMOKE.md)。
4
5 This page is **manual, opt-in, and never automated.** Nothing in CI, no test,
6 no build script, and no skill runs these commands. The repository's automated
7 suite is provider-free by design; see
8 [`crates/tui/assets/skills-catalog-matrix.json`](../crates/tui/assets/skills-catalog-matrix.json)
9 and the catalog-matrix tests for what is actually asserted without a provider.
10
11 Run this only when you want to answer one narrow question: *does a real route
12 to a real model return a well-formed receipt on this machine?*
13
14 ## What a live smoke run does and does not prove
15
16 | Question | Answered here? |
17 | --- | --- |
18 | Does the receipt record the provider/model I asked for? | Yes. This does not by itself prove which network endpoint handled the request. |
19 | Does the run emit an inspectable route/usage receipt? | Yes, when the harness reaches that stage. |
20 | Does one response prove my account's entitlement state? | **No.** Provider configuration, authentication/entitlement, and harness behavior remain candidate causes until corroborated. |
21 | Does the model semantically pick the right skill? | **No.** Not measured. |
22 | Is skill registry/catalog/alias behavior correct? | **No** — that is the provider-free suite's job. |
23
24 Treat these as investigation starting points, not proven failure classes:
25
26 - **Provider error response** — HTTP 401/403, unknown-model, quota, or region
27 errors can reflect the configured provider/endpoint, credential
28 authentication or entitlement, provider availability, or a harness
29 routing/request defect. The response alone does not distinguish them.
30 - **Receipt or process anomaly** — wrong `provider`/`model` in the receipt,
31 missing receipt fields, a crash, or failure to use the isolated state
32 directory is evidence to investigate the harness, but still needs a minimal
33 reproduction or other corroboration before assigning the cause.
34
35 ## Isolation rules these snippets follow
36
37 1. `env -i` clears the inherited environment, so your ambient `HOME`,
38 `CODEWHALE_HOME`, and `*_API_KEY` values are not forwarded. Only variables
39 listed explicitly on the `env` line survive.
40 2. Only `CODEWHALE_HOME` points at the task-specific throwaway directory, so
41 Codewhale config, sessions, and the bundled skill install land in scratch
42 state. `HOME` is intentionally left unset; the smoke run never repurposes it.
43 3. You name the credential variable yourself (`CW_SMOKE_CRED_VAR`). Nothing is
44 guessed from the provider.
45 4. The isolated child reads the secret with echo disabled, restores the prior
46 terminal state on `EXIT`, `INT`, `HUP`, or `TERM`, and exports it only in
47 that child. The value is not persisted to disk or placed in a command
48 argument or shell history.
49 5. `PATH` is forwarded explicitly, and is the only host variable carried over.
50
51 Portable `sh` is used throughout; `stty` and `mktemp -d` are the only non-POSIX
52 niceties and both exist on macOS and mainstream Linux.
53
54 ## Step 1 — create the throwaway state (both runs)
55
56 ```sh
57 CW_SMOKE_CODEWHALE_HOME="$(mktemp -d)" || exit 1
58 mkdir -p "$CW_SMOKE_CODEWHALE_HOME/tmp"
59 echo "scratch Codewhale state: $CW_SMOKE_CODEWHALE_HOME"
60 ```
61
62 ## Step 2 — name the credential variable
63
64 `CW_SMOKE_CRED_VAR` must be the variable name the provider expects. Codewhale
65 reads `MOONSHOT_API_KEY` (or `KIMI_API_KEY`) for the Moonshot/Kimi route and
66 `DEEPSEEK_API_KEY` for the DeepSeek route.
67
68 ```sh
69 CW_SMOKE_CRED_VAR="MOONSHOT_API_KEY" # you choose this; nothing is inferred
70 ```
71
72 The run command prompts for the value inside its isolated child process. It
73 does not create a credential file.
74
75 ## Step 3a — run A: Kimi K3
76
77 `kimi-k3` is a model id this build knows about. The configured provider and its
78 resolved endpoint determine the route: `--provider moonshot` selects the
79 configured Moonshot route; selecting `opencode_go` would select that separately
80 configured route. The account does not choose between them, and the harness
81 does not switch between them based on a response. Set
82 `CW_SMOKE_PROVIDER` / `CW_SMOKE_MODEL` for the route you intend to exercise. A
83 model-not-found response is an unclassified result until the provider/endpoint
84 configuration, credential access, and harness request are corroborated.
85
86 ```sh
87 CW_SMOKE_PROVIDER="moonshot"
88 CW_SMOKE_MODEL="kimi-k3"
89 CW_SMOKE_EFFORT="medium"
90 CW_SMOKE_PROMPT="Reply with exactly: SMOKE OK"
91
92 env -i \
93 PATH="$PATH" \
94 TMPDIR="$CW_SMOKE_CODEWHALE_HOME/tmp" \
95 CODEWHALE_HOME="$CW_SMOKE_CODEWHALE_HOME" \
96 CW_SMOKE_CRED_VAR="$CW_SMOKE_CRED_VAR" \
97 sh -c '
98 CW_SMOKE_STTY_STATE="$(stty -g)" || exit 1
99 restore_terminal() {
100 stty "$CW_SMOKE_STTY_STATE" 2>/dev/null || :
101 }
102 trap "restore_terminal" EXIT
103 trap "restore_terminal; exit 129" HUP
104 trap "restore_terminal; exit 130" INT
105 trap "restore_terminal; exit 143" TERM
106
107 printf "Paste value for %s (input hidden): " "$CW_SMOKE_CRED_VAR" >&2
108 stty -echo || exit 1
109 if ! IFS= read -r CW_SMOKE_CRED; then
110 printf "\nCredential input failed.\n" >&2
111 exit 1
112 fi
113 restore_terminal
114 trap - EXIT HUP INT TERM
115 unset CW_SMOKE_STTY_STATE
116 printf "\n" >&2
117
118 export "$CW_SMOKE_CRED_VAR=$CW_SMOKE_CRED"
119 unset CW_SMOKE_CRED
120 exec codewhale exec \
121 --provider "$1" --model "$2" --reasoning-effort "$3" --json "$4"
122 ' sh "$CW_SMOKE_PROVIDER" "$CW_SMOKE_MODEL" "$CW_SMOKE_EFFORT" "$CW_SMOKE_PROMPT"
123 ```
124
125 ## Step 3b — run B: a second provider/model (DeepSeek)
126
127 Set `CW_SMOKE_CRED_VAR="DEEPSEEK_API_KEY"`, then:
128
129 ```sh
130 CW_SMOKE_PROVIDER="deepseek"
131 CW_SMOKE_MODEL="deepseek-v4-pro"
132 ```
133
134 …and re-run the identical `env -i …` block from step 3a; it prompts for a fresh
135 credential value. Running the *same* command shape against two providers is the
136 point: a difference in outcome is an observation to investigate, not proof of
137 route, entitlement, or harness correctness. Provider/endpoint configuration,
138 credentials, provider health, and the generated request all remain possible
139 explanations.
140
141 ## Step 4 — optional: tool-and-reasoning receipt
142
143 The `--json` one-shot above records the resolved route claimed by the harness;
144 it does not independently prove which endpoint handled the request. To also see
145 tool-catalog and reasoning receipts, use the streaming form (still inside the
146 same `env -i` wrapper, substituting the `exec` line):
147
148 ```sh
149 exec codewhale exec --auto --max-turns 3 \
150 --output-format stream-json \
151 --provider "$1" --model "$2" --reasoning-effort "$3" "$4"
152 ```
153
154 ## Step 5 — what to record
155
156 From the `--json` one-shot receipt:
157
158 | Field | Expectation |
159 | --- | --- |
160 | `mode` | `one-shot` |
161 | `provider` | exactly the `--provider` you passed |
162 | `model` | exactly the `--model` you passed |
163 | `success` | `true` |
164 | `output` | the model's text; content is *not* a pass/fail criterion |
165
166 From the `stream-json` metadata receipt:
167
168 | Field | Expectation |
169 | --- | --- |
170 | `provider`, `model` | match the flags you passed |
171 | `route_source` | records *why* that route was chosen |
172 | `reasoning_tokens` | present when the receipt reports reasoning; absence can reflect model/provider behavior, configuration, or a harness omission and needs corroboration |
173 | `tool_catalog_sha256` | present when a tool surface was offered |
174 | `approval_posture`, `sandbox_posture` | match the flags you passed |
175 | `duration_ms`, `input_tokens`, `output_tokens` | present for a completed run |
176
177 Report the receipt fields. Do **not** paste the credential, the key file, or
178 raw provider error bodies (they can echo request headers).
179
180 ## Step 6 — clean up
181
182 ```sh
183 rm -rf "$CW_SMOKE_CODEWHALE_HOME"
184 unset CW_SMOKE_CODEWHALE_HOME CW_SMOKE_CRED_VAR \
185 CW_SMOKE_PROVIDER CW_SMOKE_MODEL CW_SMOKE_EFFORT CW_SMOKE_PROMPT
186 ```
187
188 ## Scope note
189
190 A green live smoke run is evidence that the configured live attempt completed
191 today. By itself it does not prove endpoint identity, durable account
192 entitlement, or the absence of a harness defect; corroborate those claims
193 separately. It says nothing about skill selection, alias resolution, locale
194 routing, or prompt budget — all of which are covered deterministically and
195 provider-free in `crates/tui/src/skills/catalog_matrix.rs`.
196
196 lines MARKDOWN