返回 CodeWhale
AUTOMATIC_WORKFLOWS.md
根目录 / docs / AUTOMATIC_WORKFLOWS.md
1 # Automatic Workflows
2
3 > 阅读简体中文版:[zh_hans/AUTOMATIC_WORKFLOWS.md](zh_hans/AUTOMATIC_WORKFLOWS.md)。
4
5 You do **not** need to write a `.workflow.js` file to coordinate agents. Operate
6 handles small or tightly coupled work directly. Multi-step delegation starts
7 with a compact Workflow plan: named steps, dependencies, bounded scopes, and
8 completion checks. Workflow runs the same sub-agents that Fleet configures and
9 manages, passing results and evidence between dependent steps. One bounded,
10 independent task can use a direct background agent; follow-up work should reuse
11 that agent with `followup`. Act/Agent can still use the optional soft-auto
12 policy described below.
13
14 Related docs:
15
16 - [Workflow Authoring](WORKFLOW_AUTHORING.md) — checked-in scripts and IR
17 - [Fleet + Workflow Tutorial](FLEET_WORKFLOW_TUTORIAL.md) — manual fleet paths
18 - [Configuration](CONFIGURATION.md) — `[workflow]` knobs
19 - [Sandbox](SANDBOX.md) — what the Workflow VM cannot do
20
21 ## Soft-auto in Act/Agent
22
23 1. **You ask naturally** — “audit every crate for unsafe,” “scout then implement,”
24 “compare these two providers in parallel.”
25 2. **Codewhale decides in Act/Agent** — broad, independent, or staged work can
26 trigger Workflow; one-file edits, simple commands, and pure Q&A do not.
27 3. **It tells you first** — e.g. “This looks set up for a Workflow — three scouts
28 then one verifier.”
29 4. **Optional setup** — if one or two facts would change the plan (read-only vs
30 writes, scope, child count), it opens the **`request_user_input`** modal
31 (structured multiple choice, not a long free-form interview).
32 5. **Launch** — structured `plan` JSON (goal / phases / children) or a short
33 inline script. Parallel branches use `parallel()` partial-success semantics.
34
35 In Operate, those same asks use a compact Workflow plan when they need multiple
36 delegated steps. The plan makes parallel work, dependency handoffs, and the
37 evidence needed to finish visible together. Small or tightly coupled work can
38 stay in the parent under the active tool and approval policy; one bounded,
39 independent task can use a direct agent. Continue an existing agent with
40 `followup` when the task remains the same. You can always type `/workflow` to
41 request orchestration explicitly.
42
43 ## Read-only auto-start vs write approval
44
45 `[workflow]` config (see `config.example.toml`):
46
47 | Knob | Default | Meaning |
48 |------|---------|---------|
49 | `automatic` | `true` | Soft-auto orchestration is enabled |
50 | `auto_start_read_only` | `true` | Read-only plans may start without a write-approval card |
51 | `require_approval_for_writes` | `true` | Gates the plan-approval card for writes / elevated starts |
52 | `max_children` / `max_concurrent` / `max_depth` | `1000` / `16` / `5` | Task count, concurrent children, and plan-structure (IR shape, not spawn depth) ceilings. The Runtime child-delegation budget (default 3, hard ceiling 8) is a separate quantity with an overlapping name; see `docs/SUBAGENTS.md`. |
53 | `default_token_budget` | `0` | Shared admission cap for a run and its children; `0` = none — set it or pass `token_budget` on the call to bound spend |
54
55 Elevated work (writes, shell beyond read-only, network, secrets, worktrees, high
56 budget) surfaces an approval card with goal, child summary, capability flags,
57 and budget before launch (#4126) when `require_approval_for_writes` is on.
58 That flag only gates the card. Session-level auto-approve (YOLO / Full Access /
59 `bypass`) still skips it, the same as other ordinary `Required` tools.
60 Writes inside a running VM `task()` step are the VM runtime contract
61 (sandbox, `writeAuthority`, parent tool policy) — this flag does not re-ask
62 for each child write.
63
64 Worktree isolation and write ownership are separate. A write-capable `task()`
65 (`type: "implementer"`, or `writeAuthority: "workspace_write"` /
66 `"worktree_write"`) may declare repo-relative `writeRoots`, `exactFiles`, or
67 `coordinationContracts`; with none, the spawn boundary claims its
68 `deliverables`, or else the workspace root (`.`), exactly as a plain Agent
69 spawn does. The coordination ledger refuses a second live writer whose claim
70 overlaps, so when a script fans out parallel writers in one checkout it should
71 give each one disjoint `writeRoots` or `exactFiles`. Read-only roles cannot
72 declare write authority. `worktree: true` selects isolation but does not
73 silently grant mutation authority. A prompt-only general task is read-only.
74 `dependencies` and `acceptance` carry bounded child-specific prerequisites and
75 observable completion checks; they are not a parent-transcript copy.
76
77 When a workflow runs from a workspace containing multiple repositories, a
78 child that needs shell or file access must set `cwd` to the repository-relative
79 directory it should use. The host validates that the directory exists inside
80 the parent workspace before dispatch. Use `worktree: true` for isolated writes;
81 `cwd` selects an existing checkout and does not grant write authority or
82 isolation by itself.
83
84 ## Controlling a run
85
86 `/workflow status [run_id]`, `/workflow cancel [run_id]`, and `/workflow
87 settings` are answered by Codewhale itself from the run journal and the live
88 run state — they never spend a model turn, so a status check is free and a
89 cancel lands even while the model is busy. `/workflow cancel` with no id stops
90 the only running workflow.
91
92 Starting work is review-first. `/workflow <objective>` and bare `/workflow`
93 ask the model for a bounded, tool-less proposal; `/workflow run
94 <path/to/x.workflow.js>` prepares a review of that exact checked-in source.
95 Neither form executes anything. After reviewing the proposal, run `/workflow
96 confirm` to launch the latest reviewed draft. The `[workflow]` table above is
97 read from your `config.toml` for every launch decision (auto-start,
98 write-approval card, child limits); `/workflow settings` prints the effective
99 values with what each one does. Reloading `config.toml` refreshes that table
100 for both settings and the workflow tool.
101
102 `/workflows` opens the run dashboard: every run this workspace's journal
103 keeps for the session — running and finished — newest first. Each row shows
104 the status token, the run's label, elapsed time, child count, and latest
105 progress; `Enter` opens the detail pane (run id, phases, the child roster
106 with per-child state, recent progress, and the error/result summary). `x`
107 cancels the selected running run through the same host path as `/workflow
108 cancel`, `r` re-reads the journal, and `Esc` closes. The dashboard never
109 launches anything — orchestration authority stays with `/workflow`.
110
111 ## What you see while it runs
112
113 - **Workflow panel** — phases, children, status, budget
114 - **Compact history card** — one calm row that expands for detail
115 - **One artifact per delegated unit** — no duplicate “delegate + tool card”
116 - **Typed child identity** — labels/roles; no “unknown child” in the default UI
117
118 Cancel stops the run and child agents. Completed activity can persist across the
119 session (and across restarts when configured).
120
121 ## Sandbox guarantees
122
123 The Workflow JS VM has **no** filesystem, shell, network, env, imports, clock, or
124 randomness. Allowed host calls: `task`, `parallel`, `pipeline`, `phase`, `log`,
125 `budget`, `args`. Real work happens in sub-agents / fleet under normal tool and
126 approval policy. See [Sandbox](SANDBOX.md).
127
128 ## Synthesis and compatibility
129
130 - Prefer `responseSchema` on children that must return structured fields.
131 - Ordinary failed parallel slots become `null` (partial success); filter them
132 before synthesizing one operator-facing summary. A `responseSchema` mismatch
133 is a contract failure and intentionally fails the run instead of being
134 silently converted to `null`.
135 - A `null` slot is no longer anonymous. `parallel()` and `pipeline()` attach a
136 non-enumerable `errors` array to the result — `[{ index, kind, message }]`,
137 ordered by index — so a synthesizer can say *why* a slot is missing. The
138 array's own contents and JSON encoding are unchanged.
139 - `kind` is one of `admission`, `budget`, `cancelled`, `agent`, `schema`,
140 `driver` (assigned by the host where the failure happened) or `script` (the
141 script threw it). Read it from the thrown `Error`'s `.kind`; it is never
142 inferred from message text, so a child's own prose cannot forge a kind.
143 - `opts.mode` selects the contract: `settled` (default — today's behavior),
144 `fail-fast` (reject the whole fan-out with the first non-fatal slot error),
145 or `partial` (resolve every non-cancellation failure to
146 `{ __taskError: { index, kind, message } }`). An unrecognized mode throws
147 rather than quietly reading as `settled`.
148 - A run whose every task failed is recorded as **failed**, not as a partial
149 success, even when the script itself returned a value.
150 - Workflow token budgets govern admission and aggregate accounting. Once
151 exhausted they reject later or descendant spawns, but children already
152 running in parallel can reconcile aggregate usage above the hint because
153 providers report usage only at response boundaries.
154 - Compatibility paths remain: `script`, `source_path` (checked-in
155 `.workflow.js` / `.workflow.ts`), and structured `plan`.
156
157 ## When automatic stays off
158
159 Automatic Workflow is suppressed for:
160
161 - One-file edits and tiny one-step asks
162 - Simple commands / factual questions
163 - Highly interactive design conversations
164 - Risky writes without a clear decomposition
165 - Plans that would exceed `max_children` / `max_depth` (refused before launch)
166
167 In those cases Codewhale uses direct tools or a single `agent` instead.
168
169 ## Example scenarios (#4131)
170
171 Checked-in example workflows cover four automatic-Workflow scenarios:
172
173 1. Read-only repo audit
174 2. Staged bug fix with worktree implementer + verifier
175 3. Partial failure and synthesis
176 4. Cancellation mid-run
177
178 Fixtures: [`docs/examples/dogfood-automatic/`](examples/dogfood-automatic/).
179 Panel regression tests use the `dogfood_` prefix in
180 `crates/tui/src/tui/widgets/workflow_panel.rs`.
181
181 lines MARKDOWN