Skip to content

fix(session-runtime): keep raw material out of compact projections - #3935

Merged
huangruiteng merged 7 commits into
loopx-project:mainfrom
songoow:codex/session-runtime-typed-raw-material
Sep 5, 2026
Merged

huangruiteng merged 7 commits into
loopx-project:mainfrom
songoow:codex/session-runtime-typed-raw-material

Conversation

@songoow

@songoow songoow commented Sep 4, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • add typed key classification for compact runtime fields, metrics, raw material, and unclassified input;
  • make exact, whole-word, and exact word-sequence raw evidence outrank generic pointer/count and token-metric shortcuts;
  • keep only explicit public-safe collisions such as trace_id, message_id, conversation_id, log_count, prompt_tokens, and prompt_token_count compact while fail-closing neighboring raw-shaped keys;
  • stop generic message values from becoming user-todo, validation, or blocker summaries, and document the boundary with focused projection coverage.

Behavior and authority

This changes the session-runtime read-only projection by omitting raw or unclassified values that were previously eligible through broad fallback keys. It does not add write authority, scheduling behavior, or a new provider contract.

Validation

  • python -m pytest tests/test_session_runtime_key_classification.py -q - 89 passed
  • session-runtime projection, adapter-doc, and public-safety read-model smokes - passed
  • loopx canary premerge --from-git-diff --git-diff-base refs/remotes/official/main - 17/17 passed, no manual holds
  • Ruff, diff hygiene, Python compile, and public/private boundary checks - passed
  • current full required CI - feature checks pass; Linux is blocked by the latest-main contract regression fixed in fix(review): align behavior disclosure contract #3940, and the Windows stale-lock timing case does not reproduce on the same main snapshot

Validated after rebasing onto official main at e26928f24.

Future-facing scope pass

Applied: classification precedence and projection fallback decisions now live in the existing session-runtime boundary, with no speculative provider or compatibility layer added.

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

动机

这个 PR 处理 session-runtime readonly projection 的真实边界误判:原实现用任意子串 denylist 扫 key name,导致 token_count、trace_id、catalog_id、login_at 这类合法计数/指针被当成 raw material,从而错误关闭 agent_can_continue;同时 messages、api_key、body、file_path、diff 等实际原始材料又可能漏检。exact head e82e3972514e4ece1c7397a489e037e78f6ba7ea 将 key 分类收敛成 compact / raw_material / unclassified 三态,希望既避免 false positive,又让 credential、transcript、log、local path 和 raw output 稳定 fail closed。

改动思路

实现先把 key 按 _、- 和 camelCase 切成词,再按显式 compact keys、显式 raw keys、suffix/metric 规则及 typed raw-word categories 分类。builder 只据 key name 判断,不读取 raw value;raw key 会让 continuation 关闭并记录 category,未知 key 仅进入 bounded unclassified_key_names 而不阻塞。message 也从通用 first-screen fallback 中移除并归为 transcript,防止原始事件消息被复制到 validation/blocker/user-todo 摘要。运行历史 compactor 新增两个 bounded boundary list,文档与 smoke/test 同步说明行为。

正向路径上,显式 compact summary/action、时间戳、*_id/*_count/*_at 指针及 token usage metrics 继续生成可执行 first screen;负向路径上,message/transcript/credential/path/raw output 的值不进入 projection,并关闭 agent continuation。builder 目前仍是 pure builder、没有 production caller,因此这次是 future-facing contract hardening,不新增写权限或 provider 行为。

具体改动

  • loopx/session_runtime.py 新增 KeyState、RawMaterialCategory、KeyClassification、classify_session_runtime_key 与 _classify_keys,并将分类结果接入 projection boundary 和 continuation verdict。
  • loopx/control_plane/runtime/session_runtime.py 在 run compaction 时保留最多 8 个 raw_material_categories / unclassified_key_names。
  • tests/test_session_runtime_key_classification.py 覆盖 67 个 compact/raw/unclassified、word precedence、message 不复制和 bounded compaction 情形。
  • readonly projection smoke 新增合法 lookalike 与漏检 raw-key 的正反例;adapter/protocol 文档补充三态协议和 token/trace/log 的精确语义。

关键代码讲解

  1. classify_session_runtime_key 是核心 decision boundary:先 exact contract,再 typed raw rule,最后处理 pointer suffix 和 tokens metric。
  2. _classify_keys 跨 sessions/events/outcomes/gates/artifacts/decision_results 汇总 key names 和 typed categories,raw value 不参与分类。
  3. build_session_runtime_readonly_projection 用 raw_keys 同时驱动 boundary evidence、agent_can_continue 和 remediation action。
  4. compact_session_runtime_boundary 用现有 public-safe list compactor进一步裁剪新增的 bounded lists,避免 run projection 无界增长。

对主干的风险

我发现一个 P1 边界优先级缺口:classify_session_runtime_key() 在检查 raw whole words 之前直接把任何以 id/ids/ref/refs/count/at 结尾的 key 判为 compact(当前代码第 308-309 行)。因此 secret_id、password_id、transcript_id、raw_id、stdout_id、api_key_id 都会被视为安全指针,尽管同一个 exact head 明确规定“raw-material word takes precedence”,而且 raw key value 是否真是 pointer 没有独立类型证据。实测 classify_session_runtime_key("secret_id") 返回 COMPACT,projection 的 raw_material_detected 为 false、continuation 仍可继续。这会让接入方只需给敏感载荷一个 pointer-like 后缀就绕过 typed boundary。

最小修复应把 raw exact/whole-word precedence 放在 generic suffix shortcut 前;如果确实需要允许 trace_id、message_id 等安全指针,应使用显式 pointer allowlist/typed input contract,而不是“最后一个词像 pointer”就覆盖 credential/transcript/raw-output 证据。请加入上述组合 key 的反事实测试,并让文档准确描述例外。

独立验证已完成:focused pytest 67 passed;readonly projection、adapter doc 和 public-safety readmodel 三个 smoke 均通过;git diff --check 通过。GitHub required pytest 在最终判定前仍为进行中。现有测试覆盖了 tokens_password 的 raw precedence,却没有覆盖 raw word + pointer suffix,所以没有捕获这个缺口。

我的整体评价

用 typed word-level classification 替代 substring denylist 的方向正确,改动也集中在既有 readonly boundary,scope 和体量与问题基本相称;移除通用 message fallback 尤其有价值。但这个 PR 的核心承诺就是 raw material 不得跨 compact projection,而通用 pointer suffix 当前优先于 raw word,正好破坏了该承诺。因此 exact head 本轮结论是 REQUEST_CHANGES。修复 precedence、补齐组合 key 反事实覆盖并让 required checks 通过后,可以重新评审。

English verdict: REQUEST_CHANGES for exact head e82e3972514e4ece1c7397a489e037e78f6ba7ea. Typed word-level classification is the right fix, but the generic pointer-suffix shortcut runs before raw-word classification, so keys such as secret_id, password_id, transcript_id, raw_id, and api_key_id are incorrectly accepted as compact. Make raw evidence take precedence or use an explicit typed pointer allowlist, add counterfactual tests, and rerun required checks. Independent validation: 67 focused tests and three public smokes passed; diff check passed.

@songoow

songoow commented Sep 4, 2026

Copy link
Copy Markdown
Collaborator Author

CI triage note: the Linux pytest failure is outside this PR diff and reproduces unchanged on the PR base 751765be6 with:

python -m pytest tests/control_plane/test_split_root_todo_writeback_fence.py::test_monitor_poll_writeback_allows_when_only_registry_root_is_fenced -q

The base commit own pytest check also fails: https://github.com/huangruiteng/loopx/actions/runs/33895057706/job/101095626072

This PR changes only the six session-runtime projection/docs/test files. Its focused 67-test suite, projection/doc/public-safety smokes, and 17-check premerge canary pass. I have not folded an unrelated Todo-authority fix into this PR.

@songoow
songoow force-pushed the codex/session-runtime-typed-raw-material branch from e82e397 to 0c84e89 Compare September 4, 2026 18:15
@songoow

songoow commented Sep 4, 2026

Copy link
Copy Markdown
Collaborator Author

Rebased this PR onto the latest official main 9761b6e0d; new exact head is 0c84e8949.

The focused branch surface remains green: 67 classification tests, all three read-model/doc/public-safety smokes, and the standard premerge canary (17/17, 0 failures, 0 manual holds).

The required pytest failure remains reproducible, unchanged, in both this branch and a fresh detached worktree at exact official main 9761b6e0d:

tests/control_plane/test_split_root_todo_writeback_fence.py::test_monitor_poll_writeback_allows_when_only_registry_root_is_fenced

Both runs fail through monitor_poll_writeback.py -> todos.py -> local_authority.py; none of those files are touched by this six-file session-runtime projection PR. I am keeping that unrelated main regression out of this branch so its reviewer boundary stays intact.

@songoow
songoow force-pushed the codex/session-runtime-typed-raw-material branch from 0c84e89 to 15121d8 Compare September 4, 2026 18:52
@songoow

songoow commented Sep 4, 2026

Copy link
Copy Markdown
Collaborator Author

Rebased again onto exact official main bf5dda59c; new head is 15121d817.

The previously reported required pytest regression is fixed by the newly merged upstream commit edc168277; the exact split-root case now passes on this branch. The session-runtime scope remains green: 67 focused tests and all three projection/doc/public-safety smokes pass.

The standard premerge gate now has a different upstream-only result: 16 checks pass and todo-readmodel-boundary-smoke.py fails identically on this branch and exact official main because the just-merged authority-shadow adapter reverse-imports top-level loopx.status. That one-file baseline repair is isolated in #3937 rather than mixed into this PR.

@songoow
songoow force-pushed the codex/session-runtime-typed-raw-material branch from c0a5bfd to 82d482d Compare September 4, 2026 19:47
@songoow

songoow commented Sep 4, 2026

Copy link
Copy Markdown
Collaborator Author

已处理 owner review 指出的 pointer-suffix precedence 缺口,当前 exact head 为 82d482d0c6295a50d48b86a597f8e043390966d5:

  • raw exact key、whole word 和 multi-word key phrase 现在都先于通用 *_id / *_ref / *_count / *_at shortcut;
  • trace_id、message_id、log_count 改为显式 public-safe collision,避免用通用 suffix 覆盖 raw 证据;
  • secret_id、password_id、transcript_id、raw_id、stdout_id、api_key_id、access_token_ref、tool_result_ref 均新增反事实覆盖;
  • projection-level 测试证明这些字段会设置 typed raw category、关闭 agent_can_continue / must_attempt_work,且 marker value 不会被复制。

验证:focused pytest 80 passed;3 个 session-runtime/public-safety smokes 通过;基于最新 official main@e26928f24 的 standard premerge canary 17/17 通过、无 manual hold;Ruff、diff/public boundary 均通过。请按当前 head 重新评审。

@songoow

songoow commented Sep 4, 2026

Copy link
Copy Markdown
Collaborator Author

新增一个 follow-up head 211db620c,在 owner 要求的 raw 优先规则之上补齐了几个真实的紧凑键,请一并重审。

为什么还要改一次。 raw 词优先于通用 *_id/*_count 后缀之后,我用同一套探针复扫了修复 head,发现三个常见的合法键落到了 raw 侧:prompt_tokens、prompt_token_count(标准 provider usage 计数)和 conversation_id(会话指针)。它们会置 raw_material_detected、关闭 agent_can_continue,并把首屏换成 remediation action,但并没有任何原始材料。

怎么改。 只把这三个键加进已有的显式 EXPLICIT_COMPACT_COLLISION_KEYS(原有 trace_id、message_id、log_count),不动任何通用规则。为证明例外足够窄,新增了紧邻键的 fail-closed 反事实:prompt_id、prompt_text_tokens、conversation_ref、conversation_log_count、messages_count 继续判为 transcript raw。两份 session-runtime 文档逐字列出例外。

刻意保持 fail-closed、未加入例外的键(供 reviewer 判断):stdout_count、diff_count、path_count、transcript_count、log_at。它们在仓库内没有出现,语义上更接近对原始材料的计数或时间戳,宁可误关 continuation 也不放行。

验证。 focused pytest 89 passed;readonly projection、adapter-doc、public-safety read-model 三个 smoke 通过;uvx ruff check 干净;基于最新 official main@5c56d3868 的 standard premerge canary 17/17、无 manual hold。

@songoow

songoow commented Sep 4, 2026

Copy link
Copy Markdown
Collaborator Author

Current-head required-CI failures are both outside this PR's session-runtime diff:

The current session-runtime head remains green locally: 89 focused tests, all three projection/doc/public-safety smokes, Ruff, and 17/17 standard canary. I am keeping both unrelated failures out of this PR; rerun the required workflow after #3940 lands.

Replace the substring denylist in build_session_runtime_readonly_projection
with a word-level, typed rule: keys split on _, -, and camelCase and match
exact keys or whole words only. Each key lands in one of three states:
compact (allowed), raw material (flagged with a RawMaterialCategory), or
unclassified (reported in unclassified_key_names, never blocking).

The substring rule flagged usage metrics and pointers such as token_count,
max_tokens, trace_id, catalog_id, and login_at, which switched
agent_can_continue and must_attempt_work off for compact input, while it
missed messages, prompt, api_key, password, body, file_path, diff, and
stdout_tail. This is contract hardening ahead of the first producer; the
builder has no production callers yet.

Additive boundary fields raw_material_categories and unclassified_key_names
survive run compaction as bounded lists. schema_version is unchanged.

Signed-off-by: song <liusongstep@gmail.com>
Smoke: assert the compact look-alike keys (token counts, *_id, *_at) keep
agent_can_continue on, unclassified keys are reported without blocking, and
each formerly missed raw-material key is flagged with its value never copied.

Pytest: parametrized classifier contract per state and category, substring
non-matching, unclassified bound, and compaction passthrough of the two new
bounded boundary lists.

Signed-off-by: song <liusongstep@gmail.com>
State the word-level classification rule and its three states (compact, raw
material, unclassified) in the adapter guide and the projection v0 protocol,
including the token, trace, and log word rules.

Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
…sions

After raw evidence started to outrank the generic pointer/count suffix
shortcut, three common compact keys fell on the raw side because their
leading word is also a transcript word: `prompt_tokens` and
`prompt_token_count` (standard provider usage metrics) and
`conversation_id` (a session pointer). Each would set
`raw_material_detected`, close `agent_can_continue`, and replace the first
screen with the remediation action although no raw material is present.

Add the three keys to `EXPLICIT_COMPACT_COLLISION_KEYS`, the reviewable
allowlist that already carries `trace_id`, `message_id`, and `log_count`.
No generic rule changes: `prompt_id`, `prompt_text_tokens`,
`conversation_ref`, `conversation_log_count`, and `messages_count` stay raw
and are now covered as fail-closed counterfactuals next to the exceptions.
Both session-runtime docs list the exceptions verbatim.

Validation:
- python3 -m pytest tests/test_session_runtime_key_classification.py -q -> 89 passed
- readonly projection, adapter-doc, and public-safety read-model smokes -> ok
- uvx ruff check on the changed Python -> clean

Signed-off-by: song <liusongstep@gmail.com>
@songoow
songoow force-pushed the codex/session-runtime-typed-raw-material branch from 211db62 to 1a04696 Compare September 5, 2026 02:00
@songoow

songoow commented Sep 5, 2026

Copy link
Copy Markdown
Collaborator Author

Rebased onto official main@b2cf4312d so the required pytest no longer inherits the contract-phrase failure fixed in #3940. New exact head 1a04696e6; git range-diff against 211db620c reports zero non-identical commits, and the focused suite remains 89 passed. The earlier windows-powershell failure on 211db620c was the test_file_lock stale-holder reclaim timing out on Windows and passed on the previous head; it re-runs with this push.

huangruiteng
huangruiteng previously approved these changes Sep 5, 2026

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

本轮结论:APPROVE,未发现阻塞性问题。此前 raw evidence 被 generic pointer suffix 覆盖的问题已经修复。有一条非阻塞 P2 文档建议如下;本轮只审查,不执行合并。

审查 head:1a04696e69233511115af21d50ffea9bc2510459。完整差异相对 base b2cf4312dbb20abc43a4feb84bb713ec8f02c3d8,覆盖全部 6 个文件(+602/-30)。

动机

原来的 session-runtime 只读投影以任意子串判断 raw material,会把 token_count、catalog_id、login_at 等合法计数或引用误判为不安全,关闭 agent continuation;同时某些 message、body 和 API key 形式并未被识别。另一个直接问题是通用 message 曾被用于 user todo、validation、blocker 的摘要回退,使原始事件消息可能进入首屏。这不是新增 agent runtime,而是修正既有公开 pure-builder 合同和下游 compact read model 的边界行为。

本次修订还解决了上一轮的具体反例:secret_id、password_id、transcript_id、raw_id、stdout_id、api_key_id 现在先命中 raw 分类,不再因为后缀像指针而被放行。合法 usage metrics 和显式安全碰撞键仍可继续。收益是减少错误阻塞,并让生产者看到分类诊断;非目标包括读取真实 host 原始材料、运行 session、授予权限或改变真实 quota ledger。

改动思路

入口仍为 build_session_runtime_readonly_projection(...),接收 sessions、events、outcomes、gates、artifacts、decision_results 六组 compact mapping。首屏摘要提取与 key-name 分类各自完成既有职责:前者仅从明确字段读取文本,后者把每个键归入 compact、raw_material 或 unclassified,并在 raw 类别下给出 credential、transcript、log、local_path、raw_output 证据。

分类顺序现在是显式 compact 合同/碰撞例外 → exact raw key → whole-word 或 exact word-sequence raw evidence → 通用 pointer/count 与 token-metric shortcut → unclassified。api_key_id 因包含完整 api_key 词序列而被拒绝;catalog_id 不会因为包含 log 子串而误伤。未知值不会自动复制,未知键诊断在 builder 截到 24 项,下游 compactor 再截到 8 项。

正向路径:生产者给 compact next_action 和 prompt_tokens,builder 返回可继续首屏,run projection 经过 compact_session_runtime_projection_from_run,由 attention routing 和 status summary 消费。负向路径:任意一组出现 secret_id,builder 标记 raw,关闭 agent_can_continue 和 must_attempt_work,返回提供 compact 摘要的修复动作;raw value 不进入返回内容。保留原有纯函数与 compactor 边界,比引入 provider、存储层或另一个权限状态机更合适。

具体改动

  • loopx/session_runtime.py:+228/-26,增加 typed classification、显式碰撞键与词序列匹配;删除三处 message 回退,生成两项新增 boundary 诊断。
  • loopx/control_plane/runtime/session_runtime.py:+8/-3,用同一个 public-safe list compactor 处理 raw keys、categories、unclassified keys;没有新增存储或写操作。
  • tests/test_session_runtime_key_classification.py:222 行,89 个测试覆盖分类、优先级、首屏 omission、continuation 和有界压缩。
  • 既有 readonly projection smoke:增加 94 行数据驱动正反例,保留实际 attention/status read-model 消费路径,而不是再增加一套独立展示入口。
  • 两份 adapter/protocol 文档:+50/-1,说明三态、raw precedence、安全碰撞与未知键行为;下面的 P2 对其中一句过宽表述提出修正。

关键代码讲解

  1. classify_session_runtime_key(loopx/session_runtime.py:328):输入键先规范化;显式安全键是有意例外,其余 raw whole-word/phrase 必须先于后缀和 metric。只返回 KeyClassification,不查看或输出输入值。调用方 _classify_keys 据 enum 汇总,不再依赖任意子串。
  2. _classify_keys(同文件第 357 行):遍历六组 mapping,去重并排序 raw 键、类别和未知键;未知诊断限制为 24 项,不把 unknown 当 raw,也不自行授予运行权限。调用方 builder 仍负责 continuation 决策。
  3. build_session_runtime_readonly_projection(同文件第 485 行):既有 human-gate 判断先生成 user/agent 首屏,raw 结果再约束可继续条件;开放 human gate 和 raw evidence 均可阻止 continuation。message 不再作为三类摘要来源,但明确的 summary、validation_summary 等合同字段保持可用。返回 projection,不执行工具、不写 host。
  4. compact_session_runtime_boundary(loopx/control_plane/runtime/session_runtime.py:110):只为已有 boundary 压缩保留字段,新增两类列表使用现有 public-safe 过滤器与 8 项上限。它由 readonly projection compactor 调用,再进入 run compaction/status;不会重新解释 raw values,也不是新调度权威。

对主干的风险

P2:收窄文档对 message* 的保证。 adapter 文档第 179 行 写着除例外外 every other prompt*, message*, or conversation* key stays raw。实际 message 只在 exact raw keys 中,whole-word 表包含的是 messages;独立探针验证 message_text 为 unclassified,message_count、message_ref 为 compact,首屏仍可继续。建议把文档改成明确的 whole-word/exact-key 规则并列出这些邻近键的语义,或在确实需要时用 typed 规则与对应测试补齐;不要为满足通配文字恢复任意子串 denylist。这些邻近键的 synthetic value 均未被复制,所以这是合同描述偏差,不是本轮发现的值泄漏,也不否定已修复的旧 blocker。

范围判断需要区分 producer 与 consumer:builder 本来就是已公开的纯函数,仓库内直接调用仍在 tests/smoke,没有证明新增 live-host producer;不能宣传为新的端到端运行时接入。下游 compactor 则有真实 run-compaction、attention-routing、status 调用。本 PR 没有新模块、CLI、provider 或 activation gate,属于既有合同硬化,不要求为了这次修复制造新的接入框架。默认行为变化已披露;read-only 不是 default-off 的同义词。本次无可选能力启用声明,也没有改 actor 生命周期或增加 authority。

独立验证:89 个仓库 pytest + 12 个额外探针通过;探针覆盖六组输入的旧反例、三处 message 回退、五组既有 compact 行为相对 base 的逐对象 parity(仅排除新增诊断字段)、metric 例外及未知值 omission。readonly projection、adapter-doc、public-safety readmodel 三个 smoke 均通过;其中 projection smoke 走 attention queue、status summary 和 Markdown renderer。Ruff 与 diff hygiene 通过。最终远端为 8 SUCCESS / 3 SKIPPED / 0 pending,三个跳过的是 deploy、upload-release、publish-pypi;未把它们计为执行成功。未重复运行完整跨平台 CI、真实 host ingest 或最终 main 集成;没有 PostgreSQL/authority-store 改动,也未触碰活动 goal 状态做测试。

我的整体评价

这轮修订兑现了旧 review 的具体要求:raw evidence precedence 已恢复,显式 token/pointer 碰撞可用,raw message 不再成为首屏摘要。约 236 行生产新增主要是可审阅的常量、typed 分类与诊断,测试/文档约 366 行;对于既有公开边界修正,整体体量相称,没有新增持久状态或迁移成本。

future-facing pass 已体现在既有 owner 内集中 precedence、复用列表压缩,避免建立第二份规则;目前不需要扩大重构。后续接入真实 host 时仍应验证 producer 的 compact 字段保证,不能把键名分类视作对任意值的内容审计。本轮 APPROVE,附非阻塞 P2 文档建议;发布前已再次确认 exact head 未变。批准评审不等于执行合并。

English verdict: APPROVE for exact head 1a04696e69233511115af21d50ffea9bc2510459. The previous raw-evidence/pointer-suffix blocker is fixed, and generic message fallbacks no longer populate first-screen summaries. One non-blocking P2: narrow the documented message* guarantee to match the actual exact-key/whole-word classifier; the tested neighboring values are omitted, not leaked. Independent validation: 89 repository tests, 12 additional probes, three smokes, Ruff and diff hygiene passed; final remote checks show 8 successes and 3 publication/deployment skips, with no pending checks. No live-host ingest or final-main integration is claimed, and no merge was performed.

…ed rule

Review P2 on the previous head: the adapter document claimed that every
other `prompt*`, `message*`, or `conversation*` key stays raw, but the
classifier only treats `message` as an exact raw key and `messages`,
`prompt`, `prompts`, and `conversation` as raw whole words. State that rule
verbatim and name the neighbouring keys: `prompt_id`, `prompt_text`, and
`conversation_ref` stay raw; `message_count` and `message_ref` are compact
pointers; `message_text` is reported as unclassified. No substring denylist
is reintroduced and no classification behavior changes; the three neighbours
are added to the existing parametrized tests so the document describes a
tested contract.

Validation:
- python3 -m pytest tests/test_session_runtime_key_classification.py -q -> 92 passed
- session-runtime adapter-doc and readonly projection smokes -> ok

Signed-off-by: song <liusongstep@gmail.com>
@songoow

songoow commented Sep 5, 2026

Copy link
Copy Markdown
Collaborator Author

已按 P2 收窄文档措辞,新 exact head 299cf5127b9b5a3c7a086e4685ac9b349d35b771,只改 adapter 文档一段和测试,无分类行为变化。

文档现在逐字描述 typed 规则:transcript 证据只匹配精确键 message 和整词 messages、prompt、prompts、conversation;并点名邻近键的语义:prompt_id、prompt_text、conversation_ref 保持 raw,message_count、message_ref 为 compact pointer,message_text 报告为 unclassified。没有恢复任何子串 denylist。为了让这段描述成为被验证的合同,把 message_count、message_ref 加进 compact 参数表,message_text 加进 unclassified 参数表。

验证:focused pytest 92 passed;adapter-doc 与 readonly projection smoke 通过;Ruff 通过。

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

动机

这个 PR 用 typed、word-level 的 key classifier 替换 session-runtime readonly projection 中易误判的子串规则:既要保留 token_count、trace_id、catalog_id 等合法 compact 指标/指针,又要让 credential、transcript、log、local path 和 raw output 稳定 fail closed,并阻止通用 message 内容进入 user-todo、validation 或 blocker 首屏摘要。此前评审发现 raw word 会被 generic pointer suffix 覆盖;后续修订已解决该 blocker,上一轮仅留下文档把 message* 保证写得比实现更宽的 P2。

改动思路

分类边界仍集中在 classify_session_runtime_key:显式 compact 例外优先,随后匹配 exact raw key、raw whole word/word sequence,再处理 pointer/count 与 token metric shortcut,最后落到 unclassified。这样 secret_id、api_key_id 等不会借后缀绕过 raw evidence,而显式安全碰撞仍可投影。当前增量进一步把文档收窄到真实 typed rule:message 是 exact raw key,不是任意 message* 前缀;相邻 token 根据规则分别成为 compact 或 unclassified。

具体改动

  • classify_session_runtime_key 及 _classify_keys 对六组 runtime mapping 只检查 key name,汇总 bounded raw categories 和 unclassified names,不读取或复制 raw value。
  • build_session_runtime_readonly_projection 用 raw evidence 关闭 continuation,并移除三处通用 message 摘要 fallback。
  • compact_session_runtime_boundary 沿用 public-safe list compactor,对新增诊断列表继续限长。
  • 当前 head 相对上次已批准的实现只改 2 个文件、+10/-2:adapter 文档明确列出 prompt_id/prompt_text/conversation_ref 为 raw,message_count/message_ref 为 compact,message_text 为 unclassified;测试同步锁定这三类邻近行为。

关键代码讲解

  • classify_session_runtime_key:typed precedence 的唯一决策点,raw whole-word/phrase 在通用 suffix shortcut 之前生效。
  • _classify_keys:汇总分类证据并做有界去重,不让未知键或原始值直接成为调度 authority。
  • build_session_runtime_readonly_projection:将 raw evidence 转换为只读 continuation gate,同时保持首屏只读取明确 compact 字段。
  • compact_session_runtime_boundary:将诊断压缩进既有 run/status read model,不新增长期无界状态。

对主干的风险

此前 blocker 与 P2 都已闭环。secret_id、password_id、transcript_id、raw_id、stdout_id、api_key_id 会先命中 raw;message 仍为 exact raw,文档现在不再暗示所有 message* 都 raw。新增反事实明确说明 message_count/message_ref 是 compact pointer/count,而 message_text 在没有 typed 证据时保持 unclassified;这些值不会被复制进 projection。

当前独立 worktree 上 focused suite 为 92 passed,GitHub 11 个 checks 全部为 success/skipped publication steps,无 pending 或 failure。范围仍是既有 pure builder 和 downstream compactor 的合同硬化,没有新增 provider、写权限、actor lifecycle 或 opt-in activation。残余风险在于未来真实 host producer 必须遵守 compact 字段 contract,不能把 key-name classifier 当作任意 value 的内容审计;这属于接入责任,不是当前 diff 的合入 blocker。

我的整体评价

这轮在不回退到任意子串 denylist 的前提下,准确修正了文档和 typed classifier 的边界关系,并增加邻近 key 的测试锁定。核心安全优先级、首屏 omission、bounded diagnostics 与公开说明目前一致,改动体量也与问题相称。当前 exact head 未发现新的阻断项,支持合入。

English verdict: APPROVE

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

动机

这个 PR 修正 session-runtime readonly projection 的双向边界误判:旧的任意子串 denylist 会把 token_count、catalog_id、login_at 等合法指标或指针误判为 raw material,同时漏掉 messages、api_key、body、file_path、diff 等真实原始材料;通用 message 还可能被首屏摘要 fallback 复制。受影响的是调用公开 pure builder 的 producer,以及消费 compact run projection 的 attention/status read model。最小可行修复不是继续扩充子串表,而是在现有 owner 内建立可审阅的 typed、word-level 分类与明确例外。

改动思路

入口仍是 build_session_runtime_readonly_projection,权威输入是六组 compact mapping 的 key contract,classify_session_runtime_key 是唯一分类决策点。它先认显式 compact 字段与安全碰撞,再识别 exact raw key、raw whole word/word sequence,之后才应用 pointer/count 与 tokens metric shortcut,最后归入 unclassified。正向路径中 compact summary、合法 pointer 和 usage metric 保持可继续;负向路径中 credential/transcript/log/path/raw-output key 只留下 bounded key/category 证据,关闭 continuation,原始值不进入 projection。失败修复责任仍在 producer:提供明确 compact summary,而不是让 read model猜测内容。

具体改动

  • 生产:loopx/session_runtime.py 为 +228/-26,增加 typed 三态分类、显式安全碰撞和 raw precedence,并删除 user todo、validation、blocker 三处 message fallback;loopx/control_plane/runtime/session_runtime.py 为 +8/-3,将新增诊断列表纳入既有有界 compactor。
  • 测试/示例:新增 226 行 focused pytest,并扩充 94 行 readonly smoke,覆盖 compact/raw/unclassified、pointer 反事实、值不复制、continuation 与 bounded compaction。
  • 文档:两份公开 adapter/protocol 文档共 +54/-1,披露三态、优先级、显式碰撞和邻近 message* 行为。没有 generated 或机械移动,也没有新增 CLI、provider、存储、迁移或写权限。

关键代码讲解

  • classify_session_runtime_key(loopx/session_runtime.py:328):以 KeyState/RawMaterialCategory 返回 typed 结论;raw evidence 在 generic suffix 前,secret_id、api_key_id 不再绕过边界。
  • _classify_keys(同文件 :357):跨六组输入汇总去重后的 raw key/category 与最多 24 个 unknown key;只读 key name,不复制 value。
  • build_session_runtime_readonly_projection(同文件 :485):把 raw evidence 转成 continuation gate 和 remediation action,并只从明确 compact 字段生成首屏。
  • compact_session_runtime_boundary(loopx/control_plane/runtime/session_runtime.py:110):复用 public-safe list compactor,把三类诊断都压到最多 8 项后交给现有 run/status 消费链。

对主干的风险

未发现阻断项。此前 raw word 被 pointer suffix 覆盖的 blocker 已修复;上一轮文档把任意 message* 都描述为 raw 的 P2 也已收窄并用 message_count、message_ref、message_text 反事实锁定。最强残余风险是未来 producer 把实际原始值放进语义上声明为 compact 的字段,或出现未知 raw synonym;key-name classifier 不是内容审计,因此接入方仍必须遵守 compact 字段合同。回滚可局限于 pure builder/compactor,无持久状态迁移。

独立 exact-head 验证:focused pytest 92 passed;readonly projection、adapter-doc、public-safety read-model smokes通过;Python compile 和 diff hygiene 通过。standard risk-based premerge 执行 4 个 direct checks、8 个 catalog canaries、8 个 risk-profile smokes及 1 个 public-boundary scan,共 17/17 选中检查通过,0 warning、0 manual hold。GitHub 当前 11 个 checks 全为 success 或发布步骤 skip,无 pending/failure,且显示 CLEAN/MERGEABLE。default-off isolation 不适用:这里没有 opt-in 声明;authority semantics、domain-neutral obligation 和 guidance/enforcement 均未改变。默认行为变化已在 PR、测试和两份文档中披露。

我的整体评价

完整 diff 约 236 行生产变更、320 行测试/示例和 54 行文档,体量主要来自显式 typed 规则与正反例,对修复公开安全边界是相称的。active consumer 是 run compaction、attention routing 和 status;builder 本身仍是公开 pure contract,未虚构 live-host producer。future-facing pass 已通过集中 precedence、复用 compactor 完成,不需要再加抽象。当前 exact head 299cf5127b9b5a3c7a086e4685ac9b349d35b771 新鲜、验证完整,支持合入。

English verdict: APPROVE at exact head 299cf51.

@huangruiteng

Copy link
Copy Markdown
Collaborator

Merge readiness

  • Exact head: 299cf5127b9b5a3c7a086e4685ac9b349d35b771
  • Changed surfaces: session-runtime readonly key classification, compact run-boundary diagnostics, focused tests/smoke, and matching public adapter/protocol documentation.
  • Focused validation: 92 passed; readonly projection, adapter-doc, and public-safety read-model smokes passed; changed Python compile and diff hygiene passed.
  • Risk-based premerge: standard gate ran 4 direct checks plus 17 selected checks (8 catalog canaries, 8 risk-profile smokes, 1 public-boundary scan); all passed with 0 warnings and 0 manual holds.
  • Remote status: 11 GitHub checks are successful or intentionally skipped publication steps; no pending/failure; exact-head review is APPROVED; GitHub reports MERGEABLE/CLEAN.
  • Public/private boundary: the automated boundary scan passed. The diff contains public contracts, synthetic fixtures, tests, and docs only; no credentials, private state, local paths, raw evidence, or generated logs.
  • Scope judgment: this hardens an existing read-only projection/compaction contract. It does not add a live producer, write authority, provider, schema migration, scheduler behavior, or activation gate.
  • Future-facing pass: classification precedence remains in one typed owner and compaction reuses the existing public-safe bounded-list helper; no additional abstraction is justified before another active caller needs it.

Merge decision: ready for the authorized maintainer self-merge.

@huangruiteng
huangruiteng merged commit 3dfe874 into loopx-project:main Sep 5, 2026
11 checks passed
@songoow
songoow deleted the codex/session-runtime-typed-raw-material branch September 16, 2026 05:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants