返回 DeepSeek-Reasonix
AGENT_CORE_SIMPLIFICATION.md
根目录 / docs / AGENT_CORE_SIMPLIFICATION.md
1 # Agent Core Simplification
2
3 This document tracks the agent-core simplification effort: the behavioral
4 contract of the simplified loop, the metrics used to compare before/after, and
5 where each contract item is tested. The Chinese version lives in
6 `AGENT_CORE_SIMPLIFICATION.zh-CN.md`.
7
8 ## Target loop
9
10 ```text
11 build request
12 -> provider stream
13 -> clean final: done
14 -> tool call: execute, next step
15 -> transport/service/empty-response failure: return error, user may retry
16 -> unhandled error: explicit failure
17 ```
18
19 ## Product decisions
20
21 - Connection failures, HTTP errors (including 429/5xx), interrupted streams and
22 empty responses stop the current attempt. There is no automatic backoff,
23 minute-long recovery wait, or unchanged-request replay, including auxiliary
24 search/summary requests. Keep the original error, partial display output and
25 actual request usage; another user submission can try again. Targeted
26 protocol/context repairs remain bounded and separate from transport retries.
27
28 - Normal requests are executor-only; the planner is opt-in (`planner_model`).
29 - Normal agents ship with synthetic continuation disabled; Goal, review,
30 guardian, and typed-report flows keep their own constraints.
31 - Compaction defaults to a single summary; chunked/tree-reduce recovery is
32 explicit only (manual `/compact` and marked recovery workflows).
33 - Final readiness, tool safety, cancellation, budgets, and explicit whole-file
34 read pauses stay as hard boundaries. Ordinary partial reads do not freeze
35 independent work; see [Read evidence lifecycle](READ_EVIDENCE_LIFECYCLE.md).
36 - Old configs and old session state stay readable for one release; the new
37 runtime never executes the old fallbacks.
38
39 ## Metrics baseline
40
41 Reuse the existing usage and e2ebench instrumentation; no new fallback
42 telemetry is added. Capture before/after per phase with:
43
44 - `reasonix run --metrics <path>` — per-run `RunMetrics`: token/cost totals,
45 `usage_by_source` (executor/planner/subagent/compaction/... request calls),
46 `retries`, `compactions`, `steps`.
47 - `go run ./cmd/e2ebench -task <task> -json` — per-task request counts,
48 `usage_by_source`, trajectory digest (stream retries, reasoning replays,
49 empty-final retries, TTFT, requests by source), wall time, cache hit/miss.
50
51 | Metric | Source |
52 |---|---|
53 | model requests per normal turn | `usage_by_source["executor"].Calls` / trajectory `ExecutorRequests` |
54 | planner requests | `usage_by_source["planner"].Calls` / trajectory `PlannerRequests` |
55 | reviewer/evaluator/guardian requests | `usage_by_source["recovery_reviewer"|"goal_evaluator"]`, guardian assessment usage |
56 | synthetic continuations | executor requests per clean turn above 1; trajectory `EmptyFinalRetries` |
57 | stream retries | trajectory `StreamRetries` / `RunMetrics.Retries` |
58 | compaction requests and summary spans | `RunMetrics.Compactions`, compaction telemetry notice (`spans=`, `reqs=`) |
59 | first text latency | trajectory `TTFTMs` |
60 | turn latency | `RunMetrics.DurationMs` / bench `WallMs` |
61 | tool execution success | bench task solved rate / `SolvedThenBroken` |
62 | protocol-error terminations | trajectory retry-exhausted outcomes |
63
64 Compare at least: normal Q&A, normal code edit, long-context/tool-heavy
65 (`context-pressure` tasks). Goal per phase: a normal request is one executor
66 request chain, no implicit planner/reviewer/evaluator requests, no extra
67 continuation after a clean final, default compaction never enters multi-span
68 summaries, and hard-safety failure rates do not increase.
69
70 ## Contract tests
71
72 New consolidated suite: `internal/agent/agent_contract_test.go`.
73
74 | Contract item | Test |
75 |---|---|
76 | clean final makes exactly one model request | `TestContractCleanFinalMakesOneModelRequest` |
77 | tool call executes, loop advances | `TestContractToolCallAdvancesToNextStep` |
78 | explicit retry retains thinking without a degraded mode | `TestContractExplicitRetryPreservesThinking` |
79 | provider failure stops after one request in every agent role | `TestProviderFailureReturnsWithoutWaitingOrRetrying` |
80 | interrupted output survives until the user retries | `TestInterruptedStreamStopsUntilUserRetries` |
81 | clean final adds no synthetic continuation | `TestContractCleanFinalAddsNoSyntheticContinuation` |
82 | reasoning-only clean stop completes | `TestRunAcceptsReasoningOnlyFinalAnswer` |
83 | empty response leaves retry to the user | `TestEmptyResponseLeavesRetryToUser` |
84 | zero content returns an error without committing an empty message | `TestRunStopsOnZeroContentWithoutCommittingEmptyMessages` |
85 | strict-provider missing reasoning: one frozen-request retry | `TestRunSilentlyRecoversMissingToolCallReasoning` and the replay suites in `loop_e2e_test.go`/`retry_e2e_test.go` (#9776 repair) |
86 | incomplete-read gate | `incomplete_read_test.go` |
87 | final readiness | `final_readiness_test.go` |
88 | cancellation | `cancel_test.go` |
89 | task/token/cost budgets | `run_budget_test.go` |
90 | tool permission/malformed args | `argument_validation_test.go`, gate tests |
91 | ordinary request skips planner | `TestCoordinatorOrdinaryRequestDoesNotCallPlanner`, `TestDecidePlannerRouteExplicitOnly` |
92 | planner prose without `submit_plan` fails | `TestCoordinatorPlanAndExecuteRequiresSubmittedPlan` |
93 | planner failure does not run executor | `TestCoordinatorFailsClosedWhenPlannerFails` |
94 | unusable `planner_model` is a config error | `TestBuildFailsWhenPlannerModelIsUnresolvable` |
95 | ordinary Agent adds no todo continuation | `TestStandardTodoContinuationDisabledByDefault` |
96 | pressure compaction stays single-summary | `TestPressureCompactionDoesNotCallChunkedFold` |
97 | missing guardian/recovery models fail closed | `TestBuildFailsWhenGuardianModelIsUnresolvable`, `TestBuildFailsWhenRecoveryModelIsUnresolvable` |
98
99 ## Gates
100
101 Every phase runs:
102
103 ```bash
104 go test -count=1 ./internal/agent/... ./internal/control/... ./internal/config/... ./internal/boot/...
105 go test -race ./internal/agent/... ./internal/provider/...
106 go vet ./...
107 go run ./tools/repolint
108 ```
109
109 lines MARKDOWN