R2 first slice: add worker lifecycle state projection for agent_management - #4678
Conversation
huangruiteng
left a comment
There was a problem hiding this comment.
English verdict: REQUEST_CHANGES at 560c299e5c3546faaeaf4bdf564eb647ec8c751b. The new lifecycle_state is derived from real existing facts and its 23 unit tests pass, but it is a second, undocumented state vocabulary on a row that already has a documented one: it reports blocked for a worker whose current_todo is runnable, and it reports executing for a watch-only monitor lane. Extend the documented vocabulary (and its protocol doc) instead of adding a parallel field, and scope blocked to the current todo.
动机
R2(小团队持续执行)确实缺一个能区分"已注册但空闲"和"可以真正启动"的 worker 读模型:agent_management_projection_v0 目前的 state 只有 running / waiting / blocked / monitoring / scope_wait / stale / unknown,无法回答"这个 worker 现在能不能被拉起"。这一点我认可,问题不在目标,而在落点。
本 PR 的净效果是:每个 agent 行多了一个 lifecycle_state(六态)和一个可选的 session_binding,输入全部来自既有事实(registry、todo、run_history.coordination.thread_agent_bindings、activity 时间戳),没有新增写路径,truth_contract.projection_is_writable 仍为 false。作为 R2 的第一片只读切片是合理的增量,但它改动的是一份已经被文档化的跨消费者契约,而文档没有跟着动。
改动思路
- 入口:
loopx status→build_agent_management_projection(status_payload)(loopx/control_plane/agents/management_projection.py),只读投影,无副作用。 - 新增一份并行状态推导:
_agent_lifecycle_state()与既有_agent_state()并排存在,两者都在同一个循环里对每一行计算。既有state的三条规则(current 被 blocked → blocked、任一 open todo 被 blocked → blocked、monitor → monitoring)中,前两条被逐字复制进新函数。 - 新增
session_bindings索引(agent_id → thread_id/host_surface),只为 addressable/bound 两态服务,取自既有 coordination 段落,没有第二份绑定真相。 - 新增的六态里有三个(addressable / bound / launchable)确实是既有词汇没有的,这一点是真实需求;可避免的是"再开一个字段"而不是"扩充已被文档化的那个字段"。
唯一被改动的产品文件仍是同一份投影,架构上没有越界;apps/presentation/dashboard 与 loopx/status.py 的消费者没有被同步(对加法字段可以接受,但文档契约必须同步)。
具体改动
578 insertions / 1 deletion,三个文件:产品逻辑 104 行,单测 325 行,示例 smoke 149 行。
关键代码讲解
1. _agent_lifecycle_state()(loopx/control_plane/agents/management_projection.py:494)
六态优先级为 blocked > executing > bound > launchable > addressable > registered。阻断点在这里:blocked 分支是
if any(
_todo_status(todo) == "blocked" or todo.get("task_class") == "blocker"
for todo in open_todos
):
return WORKER_LIFECYCLE_STATE_BLOCKED它不判断这条 blocked todo 是不是 current。我在 reviewed head 上直接调用 build_agent_management_projection(一条可推进的 current_todo + 一条无关的 blocked maintenance todo)得到:
| 字段 | 观测值 |
|---|---|
current_todo |
todo_runnable |
state |
running |
lifecycle_state |
blocked |
blocked_on |
todo_blocked_other |
而 docs/reference/protocols/agent-management-projection-v0.md 对这一点的原文承诺是:"The blocker remains visible without changing todo ownership or making the whole peer appear blocked." 也就是说,R2 想用这个字段做"blocked worker 不可执行"的判断,但结果是任何持有一条无关 blocked todo 的 worker 都会被判成 blocked,这个切片本要解锁的持续执行循环反而会因此停摆。现有 23 个测试全部通过,是因为每条 blocked 测试用的都是"同一条 todo 既是 current 又被 blocked",恰好绕开了这个决定性用例。
2. session_bindings 索引(同文件,新增于 agent 循环之前)
从 run_history.goals[].coordination.thread_agent_bindings 里收集 agent_id → {thread_id, host_surface},这是 addressable/bound 的唯一来源,取值没有另起来源,缺绑定就自然回落到 registered/launchable,这一处是干净的。
3. _agent_state()(同文件:491,本 PR 未改)
它才是文档化的那一个:文档列出的七个取值都在这里产生,包括 monitoring。本 PR 没有扩展它,于是同一行出现两个都能回答"这个 agent 现在处于什么状态"的字段。实测第二个分歧:一条 watch-only 的 continuous_monitor todo 下,state=monitoring 而 lifecycle_state=executing(因为 _last_activity 是新的)。
对主干的风险
- [P1,阻断] 可推进的 worker 被判成 blocked。 触发态 → 观测 → 最小修复如上:把 blocked 分支限定为"被 blocked 的是 current todo",无关阻断继续只放在
blocked_on,并把这个混合用例补进测试。 - [P1,阻断] 同一行上出现第二份状态词汇且未修订契约。 协议文档只声明
state,本 PR 未改文档(只动了三个文件),新词表还丢掉了monitoring/scope_wait/stale/unknown。最小修复:把真正新的取值扩充进文档化的state词汇(或明确宣告新字段取代旧字段并退休旧的),保持一份 blocked 规则,并更新docs/reference/protocols/agent-management-projection-v0.md。 - [P2]
executing复用了"陈旧认领告警"阈值。STALE_CLAIM_THRESHOLD_HOURS = 36在本模块原本用于stale_claim_hint的告警线,现在被直接当作"最近活跃"窗口,于是闲置一天半的 worker 也会被标成 executing。建议为本字段命名一个活跃度阈值,或复用一个语义正确的既有规则。 - [P2] 新增 smoke 可能校验到别的 checkout。
examples/worker-lifecycle-state-smoke.py里sys.path.insert(0, str(Path(__file__).parent))插入的是examples/目录而不是仓库根,因此import loopx会落到环境里已安装的那一份。我在 reviewed head 上按文档的方式直接运行时它报了ImportError: cannot import name 'WORKER_LIFECYCLE_STATE_ADDRESSABLE';只有显式把 PYTHONPATH 指向该 checkout 才通过——也就是说这个 smoke 无法证明"跑的就是本 PR 的代码"。建议改为根目录(parents[1])。 - [P3] 死分支。
if current and not _is_done(current): return LAUNCHABLE与其后的if open_todos: return LAUNCHABLE返回值相同,前者不可能改变结果。
另外该 head 相对 main 是 BEHIND,评审时 12 个 check 仍在 pending(已通过的 9 个包含 Sign-off、build、真实 PostgreSQL authority 与 stage2c)。这不是本次判断的依据,但合并前需要 rebase 后重跑 exact-head 校验;上面的阻断项都来自我在该 head 上直接调用投影的复现,不依赖 CI。
我的整体评价
方向对、取证扎实(23 个单测 + smoke,输入全部来自既有事实,没有引入第二份真相的输入),我认可这是 R2 的合理第一片。但"没有第二份输入来源"不等于"没有第二份状态权威":现在同一行上有两个状态字段,它们已经在真实数据上互相矛盾,其中一个还会把可推进的 worker 说成 blocked。这种情况越早收敛越便宜——把新取值并入已文档化的 state 词表(或明确替换它)、让 blocked 只描述 current todo、把文档一起改掉,这个切片就可以用更小的机制抵达同样的 R2 目标。因此按 formal REQUEST_CHANGES 处理。
复现方式(reviewed head 560c299e):
python -m pytest tests/control_plane/test_agent_lifecycle_state.py -q→ 23 passedPYTHONPATH=<checkout> python examples/worker-lifecycle-state-smoke.py→ 六态全过(未设 PYTHONPATH 时 ImportError)- 直接调用
build_agent_management_projection,输入为"一条 claimed/open 的 current todo + 一条无关 blocked todo",观测见上表
huangruiteng
left a comment
There was a problem hiding this comment.
动机
这是 R2(小团队持续执行)的第一个切片:给 agent_management_projection_v0 增加 lifecycle_state,用来区分"已注册但空闲"的 worker 和"真的可以被调度起来"的 worker。上一轮我在 560c299e 上提了 5 条(2×P1 + 2×P2 + 1×P3);这一轮作者推了新 head 0f9964e7(fix(control-plane): address reviewer feedback on worker lifecycle state),逐条对照处理。
已完成的部分确实按契约修对了:mixed-blocked 那条 P1、executing 复用 36h stale 阈值那条 P2、smoke 的 sys.path 与硬编码时间戳那条 P2、以及 launchable 死分支那条 P3。剩下两条仍阻塞。
改动思路
- blocked 语义收窄:删掉"任一 open todo 是 blocked/blocker 就判 blocked"的分支,只保留"当前 todo 是 blocked/blocker"。这与协议文档
blocker remains visible without making the whole peer appear blocked一致,也让blocked_on继续负责"可见但不拖垮整行"。 - 阈值解耦:新增
EXECUTING_ACTIVITY_THRESHOLD_HOURS = 8,与STALE_CLAIM_THRESHOLD_HOURS = 36分开,避免把"过期认领告警阈值"当成"活动新鲜度"。 - smoke 自证:
sys.path改成仓库根(parent.parent),并用相对当前时间的_recent_activity()取代写死时间戳,跑法不再依赖外部PYTHONPATH。 - 删除死分支:
current分支与后续open_todos分支返回值相同,去掉冗余。
具体改动
关键代码讲解
loopx/control_plane/agents/management_projection.py::_agent_lifecycle_state:blocked 判定从any(open_todos)收到current,并在注释里直接引用协议原句;这是本次最有价值的改动,它把"整行看起来 blocked"的误报关掉了。loopx/control_plane/agents/management_projection.py::EXECUTING_ACTIVITY_THRESHOLD_HOURS:新增独立常量(8h),executing分支改读它,STALE_CLAIM_THRESHOLD_HOURS回到它原本的告警用途。examples/worker-lifecycle-state-smoke.py::_recent_activity:把写死的2026-09-17T15:00:00+00:00换成相对时间,smoke 不再随日历失效。tests/control_plane/test_agent_lifecycle_state.py::test_unrelated_blocked_todo_does_not_block_worker:新增回归,用todo_runnable+todo_blocked_maintenance两个 id 锁定"当前可跑就不该整行 blocked"。
对主干的风险
- P1(阻塞):新 head 没有 DCO trailer,
Sign-off检查是红的。gh pr checks 4678显示Sign-off=fail,job 明确报Commit 0f9964e719fe7f622ccaf555b1c9cb2961c6da7a is missing a valid Signed-off-by trailer;git log --format='%(trailers:key=Signed-off-by)'在0f9964e7上为空(560c299e有)。仓库规定每个提交都必须带 trailer。修法:git commit --amend -s(或git rebase --signoff)后 force-push。 - P1(阻塞):同一行上仍然有两个状态词汇,且文档没跟上。
docs/reference/protocols/agent-management-projection-v0.md:93把state定义为running|waiting|blocked|monitoring|scope_wait|stale|unknown,全文没有lifecycle_state;本 PR 的 diff 里也没有任何文档文件。也就是说这一行现在有两个都能回答"这个 worker 现在处于什么状态"的字段,且互相矛盾。我在这个 head 上直接调投影复现:- watch-only monitor todo →
state=monitoring,lifecycle_state=executing; - 没有近期活动的 monitor todo →
state=monitoring,lifecycle_state=launchable; - 无 todo →
state=unknown,lifecycle_state=registered。
最小修法:在协议文档里声明lifecycle_state的词表与它与state的关系(包括 monitor/等待类行怎么映射),或者直接把新词并入state;两条 blocked 规则只能留一条。
- watch-only monitor todo →
- head 状态:
0f9964e7仍BEHIND main,test-shard (2)(3)(4)在我审的时候还 pending,因此这个 head 不能声明"CI 全绿"。这两点不构成阻塞结论的依据(DCO 直接读提交、词表冲突是直接调函数复现的),但合并前需要重跑。
已修项我逐条验证过:pytest tests/control_plane/test_agent_lifecycle_state.py -q → 24 passed(含新增回归);python examples/worker-lifecycle-state-smoke.py → All 6 lifecycle states verified,且这次不需要外部 PYTHONPATH;mixed-blocked 探针在新 head 上返回 launchable 而不再是 blocked。
我的整体评价
方向对了、四条反馈修得干净,唯一还欠的是"把新词表变成受文档约束的契约"——这正是 R2 这类 readiness 阶梯最容易在后续被两个消费者读出两种答案的地方,加上 DCO 这条硬门槛,本轮仍需改动后再审。改完(补 trailer、补文档并把两套词表关系写清)我会在新 head 上重新审。
English verdict: REQUEST_CHANGES - head 0f9964e fixes four of the five earlier findings (current-todo-scoped blocked rule with a new regression test, a dedicated 8h executing threshold, a self-contained smoke, and the dead launchable branch), but two blockers remain: the commit carries no Signed-off-by trailer so the required Sign-off check is red, and lifecycle_state still adds a second, undocumented state vocabulary that contradicts the documented state field on the same row (monitoring vs executing, monitoring vs launchable, unknown vs registered) with no protocol-doc change in the diff; validated by 24 passing tests, the smoke printing all six states, and direct projection probes at this head.
|
已根据 reviewer 反馈完成修复: P1 阻断性问题:
P2 重要问题:
P3 次要问题:
新增测试:
24 个生命周期测试全部通过。 |
|
已根据最新评审结果完成修复: P1 阻断性问题解决:
变更总结:
|
huangruiteng
left a comment
There was a problem hiding this comment.
动机
R2(小团队持续执行)的目标是真的:注册不等于有可执行能力,投影只给 running/waiting/unknown/blocked 这种粗粒度状态,协调者无法判断一个 peer 到底是"没活干""有活但没会话""有会话且有活"还是"正在执行"。把生命周期从现有事实(注册表成员、todo 认领、会话绑定、活动时间戳)派生出来,方向正确,而且没有引入第二份真相。
问题在于交付方式:这一版不是"扩展词表",而是替换了 state 已有的取值。改动后 running、waiting、monitoring、unknown 全部不可达(scope_wait、stale 本来也不可达),而 state 已经被另一个控制面消费者按字面量集合匹配。于是这次改动在没有报错的情况下关掉了一条真实路径。
改动思路
- 状态链
blocked > executing > bound > launchable > addressable > registered是合理的:blocked 优先,executing 需要活动时间在 8 小时内,bound/launchable 由thread_agent_bindings区分,最后落到 registered。 - 但新词表写进了原有
state字段(PR 描述里说的"新增lifecycle_state字段"在 diff 里并不存在,只有未使用的_agent_lifecycle_state函数),同时没有任何 opt-in。 - 更关键的是没有跟随消费者:
loopx/control_plane/quota/task_orchestration.py:375的 peer 准入规则仍然是peer_state.get("state") not in {"running", "monitoring"}。
具体改动
关键代码讲解
management_projection.py::_agent_state:新的六值派生(原running/monitoring/waiting/unknown分支被整体替换)。management_projection.py::_agent_lifecycle_state(第 553 行起):与_agent_state逐字相同的副本,生产代码没有任何调用点,只有新单测调用它。session_bindings采集:从run_history.coordination.thread_agent_bindings读取,供has_session_binding和新增的session_binding行字段使用。task_orchestration.py:363,375:peer 准入把投影state与{running, monitoring}做集合匹配——这一版之后该匹配永远为假。
复现(同一 payload,两个 revision)
main d8e7af141 : projected state=running execution_state=ready eligible_peer_lanes=1
head 4f6be775 : projected state=executing execution_state=blocked blocked_peer_lanes[0].reason_codes=[peer_runtime_not_active]
即:一个"注册 + 有近期活动的 open advancement todo + 有会话绑定"的 peer,在 main 上是 running 并让 lane 就绪,在本 head 上是 executing,直接落进 blocked_peer_lanes。原因是同一个 dict 由 _peer_runtime_state_by_agent(408-423 行)喂给准入规则,而规则仍是两值集合。
顺带一个语义反转:只挂 continuous_monitor 的 worker,main 上是 monitoring(无论活动新旧),本 head 上近期活动报 executing、活动过旧报 launchable——"有认领、且久未活动"被读成"可启动"。
对主干的风险
两条阻塞发现(REQUEST_CHANGES):
- [P1] 新词表静默关闭 registered-peer 编排。 触发:任何启用
peer_task_coordination的协调者目标出现候选 peer lane。结果:execution_state恒为blocked,reason_codes=[peer_runtime_not_active]。这条路径没有异常、没有重试、也没有告警——peer_runtime_not_active读起来像一个正常的"peer 尚未活跃"结论,而不是契约不匹配。本 PR 的单测(30 passed)与既有准入测试(21 passed)全都通过,因为tests/control_plane/test_task_orchestration_admission.py:19的 fixture 手写"state": "running",一个真实构建器已经产不出来的输入。 - [P1] 必需检查 Sign-off 红。 job
105458403233:##[error]Commit 0f9964e719fe7f622ccaf555b1c9cb2961c6da7a is missing a valid Signed-off-by trailer.第三个 commit 签了,第二个没签;再加一个 commit 不会让这项检查变绿,需要git rebase --signoff(或对该 commit--amend -s)后强推。
非阻塞:
- [P2] 单测覆盖的是死代码。
_agent_lifecycle_state与_agent_state逐字相同且无生产调用点,342 行单测跑的是副本;真正走到的路径只有 155 行 smoke 覆盖。建议删副本,并把单测指向build_agent_management_projection。 - [P3] 文档与实现不一致。
agent-management-projection-v0.md:93-95只是把五个新值追加进同一个"one of"列表,读者无法区分退役值与在线值;session_binding行字段、EXECUTING_ACTIVITY_THRESHOLD_HOURS = 8的阈值都没有写;PR 描述声称新增lifecycle_state字段,diff 里并没有。
验证记录(exact head 4f6be77):examples/worker-lifecycle-state-smoke.py ok(六状态全过);pytest tests/control_plane/test_agent_lifecycle_state.py tests/control_plane/test_agent_management_material_capability_gate.py -q → 30 passed;pytest tests/control_plane/test_task_orchestration_admission.py -q → 21 passed;agent-management-projection-contract-smoke 与 agent-management-observability-mvp-smoke ok;main/head 对照复现见上。未验证:没有在启用 peer_task_coordination 的真实目标上端到端观察。
我的整体评价
派生本身写得干净,六个状态的优先级和 smoke 都站得住;但它改的是一个默认打开的投影契约,而唯一按字面量消费该字段的控制面路径没有被一起搬过去,于是"更细的状态"换来了"peer 编排永远不激活"。这类改动不能只靠新增单测来背书——单测跑的是副本,准入 fixture 手写了一个构建器再也产不出的值,所以仓库全绿而真实路径已断。建议:要么让 state 保持向后可达(新词表另开字段),要么在同一改动里更新 task_orchestration.py:375、协议文档,并补一个用 build_agent_management_projection 构建投影的准入回归测试;同时修掉 DCO trailer。
English verdict: REQUEST_CHANGES - head 4f6be77 derives the R2 lifecycle states cleanly, but it replaces the default state vocabulary that loopx/control_plane/quota/task_orchestration.py:375 matches against {"running","monitoring"}: the same payload returns execution_state=ready at main and execution_state=blocked with peer_runtime_not_active at this head, so every registered-peer lane is silently dropped while 30 lifecycle/projection tests, 21 admission tests and the projection smokes all stay green (the admission fixture hand-writes state=running, and the new unit suite covers a dead duplicate _agent_lifecycle_state no production code calls); a stale monitor-only worker now also reads launchable instead of monitoring, the protocol doc keeps advertising six unreachable tokens and never documents the new session_binding field or the 8-hour executing threshold, and the required Sign-off check fails because commit 0f9964e has no Signed-off-by trailer.
Add typed lifecycle_state field to agent_management_projection_v0 with six mutually-exclusive states derived from existing facts only: - registered: agent in registry, no binding or todo - addressable: has session binding but no active todo - bound: has session binding and active todo - launchable: has active todo, no session binding - executing: has active todo with recent activity (within stale threshold) - blocked: current todo is blocked or a blocker (highest priority) The projection reads registry membership, todo claims, session bindings (run_history.coordination.thread_agent_bindings), and activity timestamps. It does not introduce a second source of truth. Includes 23 unit tests covering each state, priority ordering, and negative cases, plus a smoke script verifying all six states. Signed-off-by: Xiao Deshi <xiaods@gmail.com>
- [P1] Fix blocked state to only check current todo, not all open todos. This ensures workers with runnable current_todo but unrelated blocked maintenance todos are not incorrectly marked as blocked, per the protocol contract: "blocker remains visible without making the whole peer appear blocked." - [P2] Add EXECUTING_ACTIVITY_THRESHOLD_HOURS (8 hours) separate from STALE_CLAIM_THRESHOLD_HOURS (36 hours). The stale threshold is for warnings, while the activity threshold defines "recent" activity. - [P2] Fix smoke script PYTHONPATH to use repo root (parents[1]) instead of examples/ directory, ensuring it tests the PR code not installed version. - [P3] Remove dead branch in _agent_lifecycle_state where two LAUNCHABLE returns were unreachable together. - Add test case for unrelated blocked todo scenario to prevent regression. All 24 lifecycle state tests pass. Smoke test verifies all six states. Signed-off-by: Xiao Deshi <xiaods@gmail.com>
…ecycle states Extends the documented agent_management_projection_v0 state field to include six R2 worker lifecycle states, eliminating the parallel undocumented lifecycle_state field. This ensures the projection uses a single, documented state vocabulary as per the protocol contract. New state values: - registered: agent in registry, no binding or todo - addressable: has session binding but no active todo - bound: has session binding and active todo - launchable: has active todo, no session binding - executing: has active todo with recent activity (within 8h activity threshold) - blocked: current todo is blocked or a blocker (only checks current, not all todos) State priority (highest first): blocked > executing > bound > launchable > addressable > registered Improvements: - P1: Block check scopes to current_todo only, not all open todos - P2: New EXECUTING_ACTIVITY_THRESHOLD_HOURS = 8h separates from STALE_CLAIM_THRESHOLD_HOURS - P2: Smoke script PYTHONPATH fixed to use repo root - P3: Removed dead branch (unreachable LAUNCHABLE return) - Added test for unrelated blocked todo scenario All 24 tests pass. Smoke test verifies all six states. Signed-off-by: Xiao Deshi <xiaods@gmail.com>
Signed-off-by: Xiao Deshi <xiaods@gmail.com>
Signed-off-by: Xiao Deshi <xiaods@gmail.com>
4f6be77 to
3ebadc0
Compare
|
已按最新 review 修订,当前 head:
验证:54 tests passed;worker lifecycle、projection contract、observability MVP、live-status synthetic/bundled 四项 smoke 通过;semantic vocabulary drift、Ruff、diff whitespace 检查通过。修复前新回归出现 11 failures,修复后全部通过。 前端 schema 接受 state 字符串并透传;旧 workerStateLabel 无调用点,本次无需前端源码或打包资产改动。未验证真实 worker 启动、多轮执行、packaged browser 或 Lark transport;此 PR 不宣称完整 R2 已验收。相邻收敛已落实为删除重复推导、建立真实生产者/消费者测试。请在此新 head 复审,保留维护者合并边界。 |
huangruiteng
left a comment
There was a problem hiding this comment.
动机
这轮是复审:我此前在 4f6be775 上给过 REQUEST_CHANGES,两条阻塞分别是——(1) 新状态词表替换了 state 取值,而 peer 准入仍按字面量集合匹配 {running, monitoring},导致所有候选 peer lane 静默变成 peer_runtime_not_active;(2) DCO 未过(中间 commit 缺 trailer)。此外还有两条非阻塞:单测跑的是没人调用的死副本 _agent_lifecycle_state,以及文档仍列举 6 个不可达状态、session_binding 与 8 小时阈值未记录。作者随后推了两个 commit 专门处理。
改动思路
- 保留旧准入策略而不是换策略:消费者集合扩成
{running, monitoring, executing, bound, launchable}。这正是旧词表下"当前有非 monitor 的未完成工作即running"所覆盖的三种情形在细分后的并集;blocked/registered/addressable/waiting仍被排除,且 liveness、staleness、激活能力与依赖就绪仍是彼此独立的门。 - 把折叠掉的两个状态还回来:
monitoring(当前是 monitor,不看活动新旧)与waiting(其它非 open 的当前工作,如 deferred)在活动细分之前判定;活动窗口同时加了0 <= age下界,未来时间戳不再算"最近活动"。 - 删掉死副本:
_agent_lifecycle_state及其 438 行"测试"整体移除,单测改为直接覆盖真正发货的路径。 - 文档补齐我上一轮点名的缺口:新增 "Worker state refinement and compatibility" 一节——列出实际发出的 8 个状态、保留为 legacy 的 4 个、明确"没有并行的
lifecycle_state字段"、给出各状态的观测事实表,并显式声明默认读投影发生了变化。
具体改动
关键代码讲解
task_orchestration.py:375:注释写清"新工作状态细化保留旧的 running 准入策略",并把集合扩为 5 个 token——这是修掉 P1 的那一行。management_projection.py::_agent_state:判定顺序改为 blocked > monitoring > waiting > executing > bound > launchable > addressable > registered。management_projection.py:删除_agent_lifecycle_state(-79 行净减),单测随之收缩 438 行。agent-management-projection-v0.md:新增兼容性小节,并补上session_binding字段说明(含"绑定不证明可执行能力、缺失也不证明宿主停止")。
我复现的关键对照(同一 payload)
4f6be775(上一版): projected state=executing execution_state=blocked blocked_peer_lanes[0].reason_codes=[peer_runtime_not_active]
3ebadc009(本 head): projected state=executing execution_state=ready eligible_peer_lanes=1
即"注册 + 近期 open advancement todo + 会话绑定"的 peer 在上一版被系统性排除、在本 head 恢复可用。另跑:pytest tests/control_plane/test_agent_lifecycle_state.py test_agent_management_material_capability_gate.py test_task_orchestration_admission.py → 53 passed;examples/worker-lifecycle-state-smoke.py → 六个状态全过;rg _agent_lifecycle_state 在本 head 已无任何引用。
对主干的风险
无阻塞发现。剩余的是那处耦合本身:准入集合仍是消费者里的字面量列表,将来再加状态 token 需要同步扩展——文档把"实际发出的词表"集中写在一处,算是缓解手段,若要彻底消除可考虑生产者与消费者共享同一份声明。此外该 head 仍需 update branch,我读取时还有 12 项检查在跑(当时无失败项)。
我的整体评价
这是对我上一轮两条阻塞的完整回应,而且是按方向最小地改:没有为了过检查而把消费者换成"只判 liveness",而是忠实保留旧 running 的准入语义;没有把死副本留着同步,而是删掉并把测试指向真实路径;还把文档缺口一次补齐。我用真实构建器加真实准入函数做了前后对照,而不是只看单测——上一版所有套件与两个 smoke 全绿却真实断链,这次必须用对照说话。可以接受。
English verdict: APPROVE - head 3ebadc0 fixes both blocking findings from my previous review at their cause: the peer-admission consumer now admits executing, bound and launchable beside running/monitoring (the union the old coarse running observation covered, with liveness, staleness, activation capability and dependency readiness still independent gates), the derivation restores monitoring and waiting ahead of activity refinement and bounds the activity window below by zero, the dead duplicate _agent_lifecycle_state and its 438-line suite are removed so the unit tests exercise the shipped function, and the protocol document defines the emitted and accepted legacy vocabularies, the observation table, the default change and that there is no parallel lifecycle_state field; I verified the fix with the real builder and the real admission function (4f6be77: executing/blocked/peer_runtime_not_active; 3ebadc0: executing/ready/1 eligible lane), 53 control-plane tests and the lifecycle smoke pass, and the only remaining exposure is that admission is still a literal token list that a future vocabulary addition must extend.
Outcome
Refine the read-only
agent_management_projection_v0.statevocabulary for roadmap R2 while preserving registered-peer orchestration. The prior revision made every active peer ineligible because the producer emittedexecutingwhile the consumer accepted onlyrunningandmonitoring.Changes and compatibility
_agent_lifecycle_stateduplicate. Tests now callbuild_agent_management_projectionand pass its actual output intoapply_task_orchestration_contract.monitoringregardless of activity age and classify deferred current work aswaiting. Only the current Todo can make the row blocked or recently executing; unrelated blocked maintenance remains visible inblocked_on.runningrows intoexecuting,bound, orlaunchable. Peer admission accepts these refinements plus legacyrunning/monitoring, retaining stale-claim, activation-capability and dependency gates.registeredoraddressable. Document emitted versus legacy states, optionalsession_binding, the inclusive eight-hour current-work activity window, and rejection of future activity timestamps.These are observations, not runtime readiness or launch authorization. In particular,
launchablealone does not prove capacity, lease ownership, or permission. There is no parallellifecycle_state, new writer, scheduler, or configuration switch.Validation
git diff --checkpassed.575c0a633; all five branch commits carry DCO sign-offs.Entry points: status projection and quota peer admission. Dashboard schema accepts state strings and passes worker observations through; the old
workerStateLabelhelper has no caller, so no UI source or packaged asset changes are needed. No Lark delivery code changes. Packaged-browser/Lark transport and actual worker launch were not exercised; no end-to-end R2 qualification is claimed.Future-facing pass: deleted the duplicate state derivation and replaced private-helper tests with production-path regressions. R2 remains partial: sustained execution, dependency transfer, restart/fence recovery and independent acceptance remain with the roadmap's existing R2 work.
This control-plane change requires maintainer review and merge.