返回 ppt-master
07_classifier.md
1 [Transition] Here's one of the most interesting design decisions: what the classifier actually sees.
2
3 We strip the assistant's own text — so the agent can't talk the classifier into a bad call with persuasive rationalizations. We also strip tool results, which is the primary defense against prompt injection.
4
5 What's left? Just user messages and the bare tool call commands. The classifier judges what the agent did, not what the agent said.
6
7 [Pause]
8
9 There's a useful side effect: being reasoning-blind makes this complementary to chain-of-thought monitoring. One catches bad actions, the other catches bad reasoning. Together they're stronger than either alone.
10
11 Key points: ① Strip assistant text and tool results ② Judge actions, not words ③ Complementary to CoT monitoring
12 Duration: 1.5 minutes
12 lines MARKDOWN