feat(reward-memory): recall guidance before outbound messages - #3968
Conversation
Signed-off-by: huangruiteng <huangrt01@163.com>
huangruiteng
left a comment
There was a problem hiding this comment.
Request changes conclusion (author-owned PR; GitHub blocks formal self-review)
English verdict: REQUEST_CHANGES
动机
这个 PR 解决的是一个真实的发送前决策缺口:LoopX 目前可以积累 Reward Memory,但 Lark send / reply 在真正写出消息前不会自动取回与当前收件人、内容和目的相关的历史指导。结果是 Agent 可能在已有“先尝试安全替代方案、再升级求助”等经验时仍直接发出消息。新行为把召回放在已存在的身份校验、群成员校验和 provider dry-run 之后、实际 provider 写入之前,并把 review digest 绑定到 profile、chat、placement、文本和 purpose;这比在更上层做一个泛化 wrapper 更贴近真正的副作用边界。
改动思路
入口仍是 lark-inbox send/reply。CLI 根据 goal_id + agent_id 解析默认关闭的 Reward Memory 实验;只有 automatic_recall 和 outbound_message.before_send 都启用、且 corpus 的 peer_ref 精确匹配当前 Agent 时才安装 hook。发送器完成原有 destination/profile、身份、成员和 provider preview 校验后生成 intent digest,hook 对明确语料执行一次有界 soft-preference 召回,并返回 outbound_guidance_review_v0。存在 guidance 且不是 urgent 时,首次调用在 provider write 之前停在 agent_review_required;Agent 用精确 digest 重试后才继续。禁用、无配置、空召回、provider 不可用或 urgent 都保留原发送路径,召回不授予发送权限,也不投影为用户 gate。
具体改动
- 生产代码新增约 231 行:
reward_memory/outbound.py负责 scope、checkpoint、query、digest 和 review 决策;lark_inbox.py增加两个 CLI 参数、结果渲染和 send/reply 接线;inbox_reply.py在 provider write 前调用 hook 并把结果附回 receipt。 - 文档新增约 143 行,说明中英文配置、两步调用和边界;测试新增 225 行,覆盖 digest 绑定、urgent/provider failure、默认关闭、send/reply 写前停止、真实 CLI 接线与 opaque turn id。
agent_turn_recall的小修正不再把 opaqueturn_instance_id当 ISO 时间,而以当前 UTC 时间作为observed_at,并透传reason_code。
关键代码讲解
outbound_guidance_hook(loopx/capabilities/reward_memory/outbound.py)是能力所有者:它验证精确 Agent scope,只接受soft_preference,计算与 intent/identity/purpose/guidance 绑定的 digest,并明确保持grants_new_action_authority=false。_deliver_lark_inbox_outbound(loopx/extensions/lark/inbox_reply.py)仍是副作用所有者:旧的身份、成员、mention 和 provider-preview 检查不变;hook 只插在 dry-run 成功与实际 send 之间,因此未确认 guidance 时不会发生外部写入。handle_lark_inbox_command(loopx/cli_commands/lark_inbox.py)把 capability 接到两个已发布入口,但这里也是当前阻塞点:它直接读取新 Namespace 字段,没有兼容旧的直接调用者。
对主干的风险
有 1 个阻塞问题。精确 head 6a09e96fef10fc0aa1ffd2a31bd194bf707b758b 的必需 pytest check 为红;失败可在本地稳定复现:
uv run --with pytest python -m pytest -q tests/capabilities/test_outbound_guidance.py tests/extensions/test_lark_event_collector_routing.py::test_cli_send_resolves_route_and_forwards_safe_delivery_flags
结果为 1 failed, 10 passed。已有 caller test_cli_send_resolves_route_and_forwards_safe_delivery_flags 直接构造 argparse.Namespace,其中没有新字段;handle_lark_inbox_command 在 loopx/cli_commands/lark_inbox.py:665 读取 args.message_purpose 时抛出 AttributeError。这意味着默认关闭并没有对所有既有 handler 调用路径保持可观察兼容,也使仓库要求的全量测试失败。最小修复是在 handler 边界用 getattr(args, "message_purpose", "unspecified") 与 getattr(args, "reviewed_guidance_digest", None)(send/reply 两处统一处理),或同步迁移所有受支持的直接 caller;回归测试应保留一个没有新属性的旧 Namespace,并断言 hook 为 None、原参数继续转发。其余聚焦能力测试、diff check、build、sign-off、dependency review 和 Windows check 均通过;SonarCloud 的非阻塞 analysis job 也红,但独立 SonarCloud Code Analysis 为绿。
我的整体评价
能力边界和机制规模总体相称:它复用了现有发送器与 Reward Memory runtime,没有新建发送权限或持久状态;默认关闭、精确 scope、intent-bound digest、urgent/failure fail-open,以及“Agent review 不是用户审批”的语义都表达清楚。当前不能给出无阻塞结论,因为 exact head 破坏了已有 handler caller 并使必需 pytest 红。修复 Namespace 向后兼容、补上对应回归测试并让 required checks 转绿后,我预期可以快速复审;没有要求扩大能力范围或重做架构。
Signed-off-by: huangruiteng <huangrt01@163.com>
Review follow-up — blocker fixedThe reported legacy No remaining blocker found in the reviewed diff. The original failure was reproduced before the fix ( Product and architecture judgment
Validation
Merge decision: self-merged with explicit owner authorization as |
- 6fc4723 lark goal topic reconnect (loopx-project#3983) - 9d70e1e reward-memory 外发召回指引 (loopx-project#3968) - 8ed49d7 chat idempotency keys (loopx-project#3981) - 62a799c chat attached completion closeout - 92b03ac status contract refresh / 7bb1eb1 workspace index bound - 中继 merge commits 文档融合(中文唯一): - periodic-report-v0: 上游英文化+新增 limit/offset 有界窗口说明 -> 中文并入 - agent_turn_recall README: 以上游英文全文为源重译(新增 Turn 契约/使用/Freshness) - reward_memory README/OUTBOUND: 以上游官方全中文版(README.zh-CN.md/ OUTBOUND.zh-CN.md)升为主文件,删除 zh-CN 文件与英文原版 - repair-patterns: 上游仅整文件英化(内容逐行对应无新增)-> 保留中文版 - lark-goal-topic-connection-smoke: 采用上游实现(连接嵌套/connection_id) Co-Authored-By: Claude Code <noreply@anthropic.com>
|
Design question: guidance recall for Goal Topic auto-replies Thanks for #3968 — the digest-bound review loop is a clean boundary. OUTBOUND.md explicitly leaves one case open ("other tools and Goal Topic auto-replies are not intercepted by this integration", On the CLI path the loop closes because a CLI operator can read the guidance and rerun with
One smaller note either way: the |
|
Thanks @now-ing — the coverage gap is real, and the three trade-offs are well identified. I would welcome a focused follow-up prototype, but recommend a fourth direction: let the answering agent consume the guidance and review the exact outbound candidate inside the existing session workflow. One clarification: #3968's acknowledgement is by the executing agent, not a human CLI operator. Recommended contract
Pre-draft context delivery and final-candidate acknowledgement prove different things. The former should be labeled as context delivery, not represented by the existing exact-message review digest. Neither proves reasoning quality or factual correctness. The proposed alternatives
On Scope and acceptancePlease keep this separate from the reconnect/root-reuse fix. Start with one real direct-session caller and a narrow preparation/review seam; do not add a universal outbound interceptor or make the low-level reply helper launch models. Reward Memory owns recall; the session owns reasoning; the Lark sender owns transport checks. A typed decision contract is useful, but a broad TS migration is not required for this prototype. Acceptance should cover default-off parity, exact identity isolation, guidance actually reaching the model before its decision, changed-message/destination/guidance invalidation, bounded failure/retry, and zero duplicate send/early ACK on replay. Use an isolated session harness and synthetic transport, not live group messages. I checked the merged contract and current call paths for this reply; this is a design recommendation, not a claim that auto-reply recall is already implemented or runtime-qualified. |
Add the v1.0.0 timeline entry at its authoritative tag timestamp, covering the signed desktop in-app update path (#3994), the claim-neutral Todo correction through the TypeScript transaction (#4005) with promoted claim retry identity (#3987), multi-agent Goal Channels (#3969), and digest-bound reward-memory outbound recall (#3968). Signed-off-by: now-ing <now-ing@users.noreply.github.com>
Summary
reward_memory.outboundcapability that recalls relevant guidance at the actual Lark send/reply boundaryValidation
78 passedacross outbound guidance, agent-turn recall, and Lark send/reply suitesreward-memory-recall-application-smoke: okgit diff --check origin/main...HEADpassedBoundaries
Scope review
The capability stays with the existing Reward Memory owner and integrates only at the shipped Lark send/reply call sites. A broader automatic wrapper for every possible outbound provider is intentionally not introduced here.
This PR is submitted for review and is not being merged as part of this request.