返回 DeepSeek-Reasonix
RECOVERY_VALIDATION.md
根目录 / docs / RECOVERY_VALIDATION.md
1 # Protocol and recovery validation / 协议与恢复验证
2
3 Current policy (2026-09-24): provider transport/service failures stop after the
4 first failed request. Automatic transport retries and recovery waiting have
5 been removed. The experiments and retry counts below describe the historical
6 implementation; see [the current contract](AGENT_CORE_SIMPLIFICATION.md).
7
8 当前策略(2026-09-24):供应商连接或服务失败后直接结束请求,已取消自动传输
9 重试和恢复等待。下文的实验结果及重试次数属于当时的实现;当前行为以
10 [核心契约](AGENT_CORE_SIMPLIFICATION.zh-CN.md)为准。
11
12 Implementation date / 实施日期: 2026-09-05.
13
14 ## Reference scope / 参考范围
15
16 This document covers provider-protocol recovery: retries and stream
17 resumption. Transcript persistence, session versions, and concurrent writers
18 are described in [`SESSION_RECOVERY_AND_PARALLELISM.md`](SESSION_RECOVERY_AND_PARALLELISM.md)
19 and [`SESSION_OWNERSHIP.md`](SESSION_OWNERSHIP.md).
20
21 本文只覆盖 provider 协议层的恢复(重试与断流续传)。会话持久化、会话版本与
22 并发写者见上述两份文档。
23
24 Pi is pinned to `9841914c71a74d81abe07f751aefd271fd924e63`. The executable
25 comparison uses its `packages/ai/src/utils/retry.ts` `retryAssistantCall`
26 helper, with an injected zero delay. It is not an end-to-end comparison of the
27 Pi runtime, and does not measure either product's production recovery rate.
28
29 Pi 固定为上述提交;可执行对照使用其重试辅助函数,测试中等待设为零。
30 这不是完整 Pi Agent 的端到端对照,也不是线上恢复率或费用基准。
31 主会话持续等待是 Reasonix 的扩展,不是 Pi 的默认策略。
32
33 | Fault / 故障 | Pi helper requests / 请求数 | Reasonix requests / 请求数 | Outcome / 结果 |
34 | --- | ---: | ---: | --- |
35 | Two temporary service failures, then success / 两次临时服务故障后成功 | 3 | 3 | Automatically completes / 自动完成 |
36 | Exhausted quota / 配额耗尽 | 1 | 1 | Stops immediately / 立即停止 |
37 | Persistent interrupted stream / 持续断流 | 4 | 4 | Finite failure; no endless regeneration / 有限失败,不无限重生成 |
38
39 For the recoverable fixture, additional requests are 2 and manual continuation
40 is 0. Reasonix's scheduled quick backoff is 2 + 4 seconds; the persistent failure
41 fixture schedules 2 + 4 + 8 seconds. Test clocks avoid actually waiting that long.
42 Missing provider usage remains unknown: these fixtures do not establish token
43 cost, real recovery latency, or a statistically meaningful recovery percentage.
44
45 可恢复用例增加两次请求,人工继续次数为零;名义退避总时长为 6 秒。
46 持续失败用例的名义退避总时长为 14 秒。测试时钟跳过实际等待。
47 缺失的供应商用量保持未知,不把它记为零费用;这些用例不构成 token
48 成本、真实恢复耗时或有统计意义的恢复率测量。
49
50 ## Deterministic coverage / 确定性覆盖
51
52 - Compatible missing reasoning: no regeneration before the tool, and one tool
53 execution. / 兼容协议缺少 reasoning 不额外生成,工具执行一次。
54 - Mixed network and replay failures share four total attempts; cancellation
55 prevents a late completion from starting tools. / 混合失败共用四次请求上限,
56 取消后的迟到响应不能启动工具。
57 - Main conversations wait after quick retry exhaustion; subagents, planners,
58 partial streams and unknown tool outcomes cannot enter that wait.
59 / 主会话可持续等待;子任务、规划、部分断流和未知工具结果不能进入该状态。
60 - Auxiliary calls use 2/4/8-second finite backoff, aggregate usage, and suppress
61 failed partial text. / 辅助调用有限退避、汇总用量、排除失败的部分文本。
62 - A failed durable intent prevents mutation, including directory creation.
63 Verification checks every recorded target again before skipping a write.
64 Conflicts, changed symlink destinations, unavailable/replaced transports and
65 unknown evidence versions cannot prove success. / 意图持久化失败不开始写入或
66 创建目录;跳过写入前重新核验所有目标,冲突、符号链接换目标、原通道不可用或
67 被替换、未知证据版本均不能证明成功。
68 - Raw future-version write evidence survives serialization and is excluded
69 from model messages. / 未知版本原始证据往返保留,不进入正常模型消息。
70
71 Tests: `internal/agent/pi_recovery_test.go`,
72 `internal/agent/write_recovery_test.go`,
73 `internal/provider/recovery_test.go`,
74 `internal/provider/auxiliary_recovery_test.go`,
75 `internal/tool/builtin/write_recovery_test.go`, plus existing protocol,
76 checkpoint, session-generation, and frontend stream suites.
77
78 ## Validation boundaries / 验证边界
79
80 At the initial local-validation stage no live credential was available. The
81 official endpoint follow-up below supersedes that limitation, but does not
82 reproduce the reporter's exact Windows session or every custom gateway.
83 The real Serve-page check was attempted through the available in-app and Edge
84 browser channels; both blocked the loopback test URL before loading the page.
85 Frontend type and stream-state tests are separate from a real UI smoke test.
86 Native Desktop visual behavior still requires that smoke test.
87
88 初始本地验证阶段没有可用密钥;下方补充了官方端点实测,但仍未复现反馈者的
89 完整 Windows 会话,本地与官方端点通过不等同于所有中转场景均已解决。
90 尝试通过内置浏览器和 Edge 检查实际 Serve 页面时,两条通道均在加载前拦截
91 本地测试地址。前端类型与流状态测试不能替代真实界面检查,原生 Desktop
92 的视觉表现仍需完成界面冒烟验证。
93
94 ## Local checks / 本地检查结果
95
96 Passed / 已通过:
97
98 - Root module: `go test -p 1 ./... -timeout 180s`; the final affected Agent,
99 Provider, built-in tool and event-wire packages were also rerun successfully.
100 - Desktop module: `go test ./... -timeout 240s`.
101 - `make lint` and `git diff --check`.
102 - Frontend `tsc --noEmit` and `test:stream`.
103 - Targeted `-race` checks across Agent, Provider, built-in tools and Controller,
104 including cancellation, unknown writes, auxiliary retry budgets, concurrent
105 snapshots and session-generation changes.
106
107 根模块与 Desktop 全量测试、最终受影响包复测、静态检查、前端类型与流状态
108 测试均通过。恢复与取消、辅助重试、文件核验、并发快照、会话切换的定向
109 race 检查通过。真实界面与服务端验证仍受上节边界约束。
110
111 ## Anthropic compatibility follow-up / Anthropic 兼容性补充
112
113 The adapter distinguishes native Claude signature requirements from unknown
114 Anthropic-compatible gateways. Complete unsigned native non-tool history can be
115 converted to assistant text in the request view; tool activity, mixed proofs,
116 redacted data and incomplete reasoning remain protected. Gateways preserve
117 received unsigned thinking when replay is enabled, and explicit DeepSeek
118 contracts remain strict. No empty signatures are fabricated.
119
120 Additional deterministic HTTP/SSE fixtures in
121 `internal/agent/anthropic_compatibility_e2e_test.go` verify:
122
123 - Missing or unsigned thinking on a custom adaptive gateway: two HTTP requests
124 for one tool round and the final answer; one tool execution, no regeneration.
125 - A server rejection of unsigned history: three HTTP requests, one tool
126 execution, with completed-tool facts in the repaired request and original
127 reasoning retained locally.
128 - A compatible native text conversion does not claim the missing-reasoning
129 recovery incident. Adapter tests cover immutable history, idempotent
130 conversion, signed/redacted block preservation, and rejection of unsafe input.
131
132 新增测试区分原生 Claude、未知 Anthropic 网关和显式 DeepSeek 契约。模拟网关
133 缺失或返回 unsigned thinking 时,一个工具轮加最终回答共两次 HTTP 请求,工具
134 执行一次;模拟服务端拒绝旧 thinking 时,共三次请求,工具仍只执行一次,修复
135 请求携带已完成事实,本地保留原始 reasoning。兼容文本转换不消耗严格恢复预算。
136
137 These are simulated servers, not live endpoint acceptance tests. No DeepSeek,
138 Anthropic or OpenCode Go API credentials were available for this follow-up.
139 本轮为模拟服务端验证;环境未提供上述供应商密钥,不能据此宣称 #9808 的反馈者
140 端点或所有兼容网关已通过实测。
141
142 ## Official endpoint investigation (2026-09-05) / 官方端点实测
143
144 This follow-up uses a user-authorized credential only against
145 `api.deepseek.com`, on `deepseek-v4-flash` and `deepseek-v4-pro`. It covers
146 `/chat/completions`, `/responses`, and `/anthropic/v1/messages`. Credentials
147 are supplied in memory to isolated test processes; no credential file is
148 created. Only synthetic marker tools and confined temporary-file writes are
149 available. No shell, MCP, credential reader, or other provider is exposed.
150
151 本次使用用户授权的官方密钥,覆盖 Flash、Pro 及三种协议。密钥仅在测试进程中
152 传递;模型只能调用固定标记工具或临时目录内的文件写入,不提供 shell、MCP 或
153 凭据读取能力。以下区分原始服务端契约、真实模型加本地故障注入、确定性回归。
154
155 ### Raw replay contract / 原始回放契约
156
157 The initial 44 direct, non-streaming HTTP probes all returned 200. Those probes
158 reused provider-issued call IDs. A second set of 30 probes replaced call IDs,
159 with full-reasoning controls to establish that the replacement IDs themselves
160 were valid. Twenty-two returned 200; eight deliberately invalid requests
161 returned the expected 400. Both models produced the same distinctions:
162
163 | Historical input / 历史输入 | Chat | Anthropic Messages | Responses |
164 | --- | --- | --- | --- |
165 | Original call IDs; omit reasoning / 原始调用 ID,省略 reasoning | 200 | 200 | 200 |
166 | Replacement call IDs; full reasoning / 替换调用 ID,完整 reasoning | 200 | 200 | 200 |
167 | Replacement call IDs; omit reasoning / 替换调用 ID,省略 reasoning | 400 `reasoning_content` | 400 `content[].thinking` | 400 `reasoning_text` |
168 | Replacement call IDs; explicit empty field/block/lists / 替换调用 ID,显式空字段、块或列表 | 200, empty string | 200, empty thinking block | 400, empty content/summary lists |
169
170 All rejection messages require the named content to be passed back. The
171 ID-dependent difference is evidence of server behavior, not proof of its
172 internal storage/cache implementation or a durability guarantee. A successful
173 request with an original call ID is **not** sufficient evidence that missing
174 reasoning will always be accepted after restart, ID normalization, or gateway
175 translation. Do not fabricate opaque Responses items or extend Chat's empty
176 field rule to other protocols from these results.
177
178 三种协议在替换调用 ID 后均能复现真实的回放 400。原始 ID 下成功,不能证明
179 历史重载、ID 转换或网关转发后仍可省略 reasoning;服务端内部如何找回这些
180 信息未验证。Responses 的“空列表”与合法的 reasoning item 并不等价。
181
182 The official [thinking-mode guide](https://api-docs.deepseek.com/guides/thinking_mode/)
183 continues to require full historical reasoning with tools. The
184 [Anthropic compatibility guide](https://api-docs.deepseek.com/guides/anthropic_api/)
185 describes Messages compatibility. Healthy history therefore retains all
186 received proof. This investigation does not relax native Claude signatures or
187 explicit strict DeepSeek Anthropic contracts.
188
189 官方文档仍要求带工具请求完整回传历史 reasoning。正常历史继续保留真实内容;
190 本次没有放宽原生 Claude 签名要求,也没有修改显式严格 Anthropic 契约。
191
192 ### Defects found and fixed / 实测发现及修复
193
194 1. An EOF inside a JSON data line was a fatal decode error in Chat and Messages.
195 The shared stream scanner now distinguishes an unterminated JSON prefix from
196 a malformed complete event. The former uses bounded stream recovery; the
197 latter still fails. Partial tool calls never execute.
198 2. The actual Responses rejection names `reasoning_text`, which the replay-error
199 parser did not recognize. It now enters the existing bounded history repair,
200 preserving completed-tool facts and excluding invalid protocol history.
201 3. Request-only and byte-estimated usage could lose the unknown-usage flag.
202 Missing provider usage now remains unknown through estimation and aggregation;
203 request counting and known token telemetry remain available.
204 4. Strong history repair retained completed-tool names but discarded their
205 outputs, causing real models to repeat the tool or be unable to answer.
206 Recovery now includes bounded original model-visible results as escaped,
207 explicitly untrusted JSON; raw/local-only output stays excluded. Only a
208 repaired fault prefix changes; healthy requests are untouched.
209 5. The write-intent hook was attached to the outer execution context while
210 dispatch used the already prepared tool context. It is now installed on the
211 actual dispatch context after permission and preparation. A failed intent
212 checkpoint prevents the write from starting.
213
214 发现并修复五处遗漏:断流 JSON 误分类、Responses 回放错误漏识别、未知 usage
215 标记丢失、历史修复丢掉实际工具结果、写入持久化钩子未传入真正执行上下文。
216 只有故障修复视图新增有界工具结果;正常提示词、工具 schema、字段顺序与健康
217 历史不变。回归测试还验证了结果转义、RawContent 排除、持久化失败禁止写入。
218
219 Deterministic regressions: `internal/provider/stream_scanner_test.go`,
220 `internal/provider/reasoning_replay_error_test.go`,
221 `internal/agent/stream_fragment_recovery_test.go`, and
222 `internal/agent/cancel_test.go`. Live entrypoints are build-tagged `live` and
223 credential-gated. The old live missing-reasoning expectations were updated to
224 assert zero extra generation for compatible Chat/Responses turns. Independent
225 search tests now pin the supplied process credential explicitly instead of
226 silently skipping because an isolated home has no global credential file.
227
228 ### Measurements and qualification / 指标与验收结果
229
230 Worktree base: `1b4f9ae8324413d04ae272ceadb86ad49ffded2e`, plus local changes;
231 these results do not describe a published release. All paid calls were
232 sequential. Root package tests replace retry sleeps with a controllable test
233 sleeper, so the following live latency numbers exclude the production 2/4/8s
234 backoff. The proxy buffers a real upstream response before injecting a fault;
235 this is not a measurement of a naturally occurring server outage.
236
237 | Suite / 用例组 | Observed result / 结果 |
238 | --- | --- |
239 | Raw wire contract / 原始 HTTP 契约 | 74 requests: 66 HTTP 200 and 8 expected HTTP 400; 25,469 input and 1,860 output tokens reported across metered responses. Anthropic input includes cache-read/create tokens. The eight 400s have unknown usage. |
240 | Recovery matrix / 恢复矩阵 | 54 cases matched their expected outcomes: 122 client HTTP attempts, 116 official upstream requests, 46 actual marker-tool executions. Six 503s were local injection without upstream calls. |
241 | Recoverable faults / 可恢复故障 | 20/20 continued automatically, zero manual continuation, one extra attempt per fault. This includes six stream cuts, six temporary 503s, six actual server replay rejections and two single missing-thinking strict turns. |
242 | Recovery latency / 恢复耗时 | Whole recoverable scenarios: median 2.949s, range 0.994–6.591s; this includes normal tool/final requests, not just the recovery request, and excludes production retry sleeps. |
243 | Protective stops / 保护性停止 | Six cancellations executed no tool; two persistent strict Anthropic missing-thinking cases stopped after two requests and executed no tool. These are expected stops, not successful automatic recoveries. |
244 | Compatible missing reasoning / 兼容缺失 reasoning | Eight Chat/Responses missing-once/persistent cases completed with two requests and one tool execution; no reasoning regeneration. |
245 | Matrix usage / 矩阵用量 | 41,161 input and 4,347 output tokens recorded, including local estimates for incomplete requests. Twenty-four cases retained unknown-usage metadata. Extra retry tokens were not independently metered; no exact monetary total is claimed. |
246 | Conversation continuity / 连续会话 | Six Flash/Pro × protocol combinations, six user turns each, one save/load each: 72 requests, 36 tool executions, zero retries, 97,336 input / 2,053 output tokens. Healthy history, tools and settings stayed byte-stable across requests and reload. |
247 | Cache / 缓存 | Continuity aggregate: 88,320 cached / 97,336 input tokens (90.74%). A separate large-tool-output Chat test preserved its bounded stable prefix; its last two turns hit 9,472/9,611 and 9,856/9,882 tokens. Local RawContent sentinel was absent from requests. |
248 | Write effect before result checkpoint / 已写入但结果未保存 | Three protocols passed: one durable intent and one actual disk write each. Reload produced an unknown-result placeholder, not a fabricated completed result; verification prevented a second disk write. |
249 | Independent search / 独立搜索 | Search returned eight structured sources; default Chat → independent Messages search → Chat final completed. Standalone search reported one request, 17,238 input / 939 output tokens. |
250
251 Additional earlier checks passed: 20 Responses Flash/Pro tool loops, Chat and
252 Responses reasoning-removal probes, official Messages tool/history/search
253 round-trips, and cancel-after-tool save/load continuation.
254
255 真实恢复矩阵的 54 个场景均达到各自预期;其中 20 个可恢复故障自动继续,8 个
256 取消或严格协议持续缺失场景按设计停止,不能将后者算成“自动恢复成功”。主矩阵
257 没有重复执行工具。统计使用固定合成任务,不能据此估算所有真实任务的恢复率。
258 重试等待在测试中被替换,因此这些时长不是生产环境故障的真实等待时间。
259
260 A subsequent repair-boundary check found that anchoring the overlay to the
261 already stripped history could omit a trailing removed tool pair. The boundary
262 now anchors to the original source history. Completed-result facts identify
263 originating user turns, avoiding treating old work as fulfillment of a new task.
264 The deterministic follow-up asserts that the next request retains the repaired
265 view and the actual completed output.
266
267 后续还修正了历史修复边界:定位到原始历史,而不是已经删除工具轮的结果视图。
268 这样下一轮不会立即重新带回刚删除的错误协议历史;结果标记所属用户轮次,避免
269 把历史工作误当成新任务已经完成。
270
271 **Observed model variability:** the six real post-repair continuation checks
272 completed, but one Flash/Chat sample requested the read-only marker a second
273 time (five requests / two executions instead of four / one), failing the strict
274 no-repeat assertion. Three diagnostic repeats of that exact case then passed
275 without a duplicate; the added counters showed one execution before and after
276 continuation in those repeats. The first observation remains a limitation,
277 not erased by the passing repeats. No generic same-arguments deduplication was
278 added: a new request can legitimately require a fresh read. The verified
279 no-repeat disk-write result must not be generalized to every tool or model call.
280
281 补充的六组“修复后再继续”均完成任务,但其中一个 Flash/Chat 样本重复调用了
282 一次只读标记工具,未达到严格的零重复断言;随后三次定向复测没有复现。
283 该观察仍保留为边界,不因复测通过而抹去。没有按相同参数永久去重,因为用户
284 新请求可能需要重新读取。文件未重复落盘不等于所有模型都不会重复调用工具。
285
286 Final deterministic checks passed: root and Desktop module suites, targeted
287 race checks for stream/usage/replay/write-intent recovery and controller write
288 checkpoint reload, `make lint` (0 Go issues; existing repolint baseline unchanged),
289 and `git diff --check`. Live tests remain opt-in under the `live` build tag;
290 model-dependent no-repeat assertions can fail as described above.
291
292 最终根模块、Desktop 全量测试及恢复/取消/写入检查点的定向 race 检查通过,静态
293 检查与差异检查通过。真实测试受 `live` 标签保护;上述依赖模型选择的严格零重复
294 断言仍可能失败。
295
296 This is official API evidence on macOS plus controlled local faults. It does
297 not validate the reporter's complete Windows session, native Claude, arbitrary
298 custom gateways, unknown shell/MCP effects, native Desktop visual behavior, or
299 Pi/OpenCode against the same live account. Issue #9808 cannot be declared solved
300 for every deployment from these samples; no issue closure, push or release was
301 performed.
302
303 本次覆盖 macOS 下官方 API 及本地可控故障,不等同于反馈者完整 Windows 会话、
304 任意中转、原生 Claude、未知 Shell/MCP 副作用或原生界面的全面实测。没有关闭
305 issue、提交、推送或发布,不能据此宣称 #9808 的所有部署场景均已解决。
306
306 lines MARKDOWN