返回 CodeWhale
OPERATIONS_RUNBOOK.md
根目录 / docs / OPERATIONS_RUNBOOK.md
1 # codewhale Operations Runbook
2
3 > 阅读简体中文版:[zh_hans/OPERATIONS_RUNBOOK.md](zh_hans/OPERATIONS_RUNBOOK.md)。
4
5 This runbook covers practical debugging and incident response for the local CLI/TUI runtime.
6
7 ## Quick Triage
8
9 1. Confirm binary + config:
10 - `cargo run -- --version`
11 - `cat ~/.codewhale/config.toml` (or inspect configured profile)
12 2. Enable verbose logs:
13 - `RUST_LOG=codewhale_tui=debug cargo run`
14 - For HTTP retries/reconnects: `RUST_LOG=codewhale_tui::client=debug cargo run`
15 3. Capture current state:
16 - `ls ~/.codewhale/sessions`
17 - `ls ~/.codewhale/sessions/checkpoints`
18 - `ls ~/.codewhale/tasks`
19
20 ## Incident: Turn Hangs or Stream Stops
21
22 Symptoms:
23 - TUI remains in loading state
24 - partial assistant output with no completion
25
26 Checks:
27 1. Inspect retry/health logs (`codewhale_tui::client`)
28 2. Verify endpoint connectivity:
29 - `curl -sS https://api.deepseek.com/beta/models -H "Authorization: Bearer $DEEPSEEK_API_KEY"`
30 3. Confirm no local sandbox/permission deadlock in tool output
31
32 Actions:
33 1. If a foreground shell command is running, press `Ctrl+B` to move it to the background (the turn keeps running and the command becomes a background job under `/jobs`); use `Ctrl+C` instead if you want to cancel the turn.
34 2. If the command was started in the background, ask the assistant to use `Bash` with `action: "cancel"` and the returned process id.
35 3. Use `Esc` or `Ctrl+C` to interrupt the current turn when you want to stop the request itself.
36 4. Retry prompt; if still failing, restart TUI.
37 5. On restart, verify the previous queued/in-flight runtime turn is shown as interrupted rather than left in a running state.
38
39 ## Incident: Network Outage / Offline Behavior
40
41 Expected behavior:
42 - New prompts are queued while offline mode is active
43 - Queue state persists per session to
44 `~/.codewhale/sessions/checkpoints/<session-id>.offline_queue.json`; a legacy
45 global `offline_queue.json` is adopted once on upgrade
46
47 Checks:
48 1. Open queue in TUI: `/queue list`
49 2. Confirm persisted queue file exists and updates timestamp
50
51 Actions:
52 1. Restore connectivity
53 2. Re-send queued entries (from `/queue edit <n>` + Enter, or normal input flow)
54 3. Ensure queue file clears when queue is empty
55
56 ## Incident: Crash Recovery Needed
57
58 Expected behavior:
59 - Each session checkpoints to `~/.codewhale/sessions/checkpoints/<session-id>.json`;
60 a legacy `latest.json` is still read for recovery but is no longer written
61 - Startup begins a fresh session unless `--resume`/`--continue` is supplied
62
63 Actions:
64 1. Resume prior work explicitly via `codewhale --resume <id>` (alias
65 `codewhale resume <id>`; `codewhale --continue` recovers the newest
66 interrupted checkpoint for the workspace) or `Ctrl+R` in TUI
67 2. If checkpoint inspection is needed, inspect `checkpoints/<session-id>.json` (or a leftover legacy
68 `latest.json`) for schema mismatch/details
69 3. If schema is newer than binary supports, upgrade binary or remove stale checkpoint
70
71 ## Incident: Persistent State Schema Errors
72
73 Symptoms:
74 - Errors like `schema vX is newer than supported vY`
75
76 Affected stores:
77 - sessions (`~/.codewhale/sessions/*.json`)
78 - runtime thread/turn/item records
79 - tasks (`~/.codewhale/tasks/tasks/*.json`)
80
81 Actions:
82 1. Confirm binary version and migration expectations
83 2. Back up the state directory before editing
84 3. Either:
85 - run with a newer compatible binary, or
86 - archive incompatible records and regenerate state
87
88 ## Incident: MCP/Tool Execution Failures
89
90 Checks:
91 1. Validate `~/.codewhale/mcp.json` schema and server command paths
92 2. Confirm server process can start manually
93 3. Check sandbox denials in TUI history / logs
94
95 Actions:
96 1. Retry with required approvals (or YOLO only when appropriate)
97 2. Temporarily disable failing MCP server and isolate issue
98 3. Re-enable after verification with `/mcp` diagnostics
99
100 ## Post-Incident Checklist
101
102 1. Preserve logs and relevant state files
103 2. Record trigger, impact, and mitigation
104 3. Add or update regression tests (retry/recovery/schema)
105 4. Update this runbook and architecture docs if behavior changed
106
106 lines MARKDOWN