Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 12 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -299,6 +299,18 @@ with a nested `uv run`, rewrite historical execution receipts, or commit a
generated `uv.lock` as part of an unrelated change. See the testing and quality
guide for the validation layers and the source-checkout environment boundary.

### Evidence-Based Budget Decisions

When a size, structure, latency, or similar regression budget fails, follow the
[budget decision guide](docs/development/testing-and-quality.md#budget-failure-decisions).
Establish what the limit protects, measure the same base/head workload, and
inspect consumer value and true redundancy before choosing compaction, keeping
the limit, or an evidence-backed increase. Preserve decision semantics and
compatibility; neither a historical ceiling nor a green revised test is the
objective. Record the tradeoff in existing PR validation/review evidence, not a
new approval workflow. Hard limits and frozen experiment/promotion criteria
retain their owning authority and cannot be reclassified to erase a failure.

### Refactor Real-Path Validation

Before delivering a refactor, validate the affected production entrypoint and
Expand Down
74 changes: 69 additions & 5 deletions docs/development/testing-and-quality.md
Original file line number Diff line number Diff line change
Expand Up @@ -526,13 +526,15 @@ profile 的条目表示 owner 待澄清,并不等于该测试没有价值。
The interface budget gate measures stable command scenarios and compares the
candidate checkout with its base. It catches accidental payload growth,
duplicated diagnostics, and hot-path fields that silently return after a
refactor. Budget changes are contract changes: update the implementation and
the expectation together, explain every added or removed semantic field, and
request owner review when the default agent-facing projection changes.
refactor. Budgets protect useful, bounded decisions; historical ceilings are
not optimization targets. Explain added or removed semantic fields and use the
decision process below when changing a budget. Default agent-facing projection
changes remain subject to the existing owner review.

接口预算门会测量稳定命令场景,并比较 candidate 与 base,捕获意外膨胀、重复诊断
以及重构后悄悄回到热路径的字段。预算变化就是合同变化:实现与期望必须一起修改,
逐项解释新增或删除的语义字段;默认 agent-facing 投影变化时需要 owner review。
以及重构后悄悄回到热路径的字段。预算保护有用且有界的决策,历史上限不是优化目标。
解释新增或删除的语义字段,预算调整按下述流程判断;默认 agent-facing 投影变化
仍遵循既有的 owner review。

The full diagnostic packet remains an explicit drill-down surface. Moving a
field off the default path is acceptable only when the default still tells the
Expand All @@ -541,6 +543,68 @@ agent what to do and how to request the omitted detail.
完整诊断包保留为显式 drill-down。只有默认路径仍能告诉 agent 下一步做什么、以及
如何请求被省略细节时,字段才能移出默认热路径。

### Budget Failure Decisions

Classify the limit by its owning contract before deciding how to repair a
failure. This applies to output size/structure and latency regression budgets;
it does not grant execution quota, spending, or provider authority.

先根据所属合同判断上限的性质,再选择修复方式。这适用于输出尺寸、结构和延迟
回归预算,不授予执行配额、费用额度或 provider 权限。

| Limit / 上限 | Decision boundary / 决策边界 |
| --- | --- |
| Hard external or authorized limit / 外部或授权硬上限 | Respect the transport, storage, resource or owner constraint; a test edit cannot raise it. / 遵守传输、存储、资源或 owner 约束,改测试不能扩容。 |
| Regression budget / 回归预算 | A measured guard against unexplained growth; preserving, reducing or increasing it needs consumer and cost evidence. / 用测量防止无解释增长;保持、压缩或扩容都依据消费者价值与成本。 |
| Presentation cap / 展示上限 | Bound returned detail, not source truth; preserve selection, totals, completeness and drill-down. / 限制返回细节而非源事实;保留选择、总数、完整性与下钻路径。 |

1. **Measure the same contract.** Record base/head revisions, workload, metric
and measurement boundary. Compact JSON characters, UTF-8 bytes, nested keys,
emitted stdout and tokens are different metrics. For latency, retain the
sample window, workload and distribution; acknowledge noise. Keep the
failing scenario and original result; do not shrink fixture populations,
scan roots or sampling depth to obtain a pass.
2. **Inspect information value and redundancy.** Name the current consumer and
decision each changed field supports. Remove derivable or unused copies when
the consumer contract permits it. Similar rows in different lanes may serve
different consumers; deduplicating them needs caller migration/parity, not
just equal JSON. Prefer bounded summaries and reachable cold paths for
detail. Do not delete identity, completeness, safety or settlement semantics,
shorten names solely to pass, or build a reference framework for tiny savings.
3. **Choose and disclose the tradeoff.** Compare compaction, retaining the
ceiling, and a justified increase; a combination is valid. For an increase,
explain the remaining useful cost, old/new ceiling, measured headroom and
expected variation or scale. No universal headroom percentage is required.
Update the owning contract and test expectation together, rerun the original
scenario and affected semantic/scale checks, and report both the original
failure and new result in existing PR validation and review evidence. A
budget-only change need not force unrelated code cleanup. Frozen experiment
or promotion thresholds stay fixed for that result; revised thresholds belong
to a new qualification, never a relabeled historical pass.

1. **同口径测量。** 记录 base/head、负载、指标和测量边界。紧凑 JSON 字符、UTF-8
字节、嵌套键数、真实 stdout 和 token 不可互换。延迟要保留样本窗口、负载和
分布,并承认噪声。保留失败场景与原结果,不缩小 fixture、扫描范围或采样深度。
2. **分析信息价值与真实冗余。** 说明变化字段服务哪个消费者、哪个决策。合同允许时
删除可推导或无人使用的副本;不同 lane 中相同的数据可能服务不同消费者,去重
需要调用方迁移和语义等价验证。详情优先使用有界摘要和可达冷路径。不能删身份、
完整性、安全或结算语义,不能只为过线缩字段名,也不为微小收益制造引用框架。
3. **选择并披露取舍。** 比较压缩、保持上限、合理扩容,也可组合使用。扩容需说明
保留信息的价值与成本、新旧上限、实测余量和预期波动或规模,不规定统一余量比例。
合同和测试同步修改,重跑原场景及受影响的语义/规模检查,在既有 PR 验证和评审
证据中同时保留原失败与新结果。纯预算调整不必捆绑无关清理。已冻结的实验或
promotion 阈值不能追溯放宽;新阈值属于新一轮验证,不能改写历史结论。

The PR-review packet's `semantic_alignment` rule consumes this evidence through
the existing `validation_matrix` and `observable_semantics` rows. It does not
add a separate budget receipt or approval gate. The result checker verifies
evidence structure and verdict consistency; the reviewer still judges whether
the measurements and tradeoff are sound.

PR-review 的 `semantic_alignment` 通过既有 `validation_matrix` 和
`observable_semantics` 使用这些证据,不增加独立预算回执或审批门。结果校验器检查
证据结构和结论一致性;测量是否可信、取舍是否合理仍由评审判断。

## Decision Replay And Issue #2191 / 决策回放与 #2191

Issue #2191 is the reference pattern for a cross-layer control-plane
Expand Down
12 changes: 8 additions & 4 deletions docs/reference/contracts/interface-budget-contract.md
Original file line number Diff line number Diff line change
Expand Up @@ -74,8 +74,8 @@ status collection.

The canonical emitted-output inventory and current characterization ceilings
live in `loopx.control_plane.testing.cli_output_budget`. Those ceilings are
regression baselines, not target sizes: preserving a large current value makes
unreviewed growth fail while a later optimization lowers the ceiling. Tests
regression baselines, not target sizes: unexplained growth fails, while measured
consumer value can justify compaction or a reviewed increase. Tests
also record UTF-8 bytes, line count, JSON parseability, pretty-print overhead,
semantic anchors, collection-growth slope, and bootstrap duplication. Every
declared agent-facing surface must name an owner, consumer action, and cold-path
Expand Down Expand Up @@ -141,8 +141,12 @@ Restraint rules for new fields:
decision summary into a hot-path surface.
2. A hot-path field must answer a current consumer action. If the consumer only
says "nice to inspect", keep the field in the cold path.
3. A new nested object must either stay within the nested budget above or retire
/ compact an older field in the same surface.
3. For a new nested object or a budget failure, compare compaction, retaining the
limit, and an evidence-backed increase using the
[budget decision guide](../../development/testing-and-quality.md#budget-failure-decisions).
Update the owning contract and tests together when the budget changes;
preserve semantic checks and the original measurement scope. Similar
objects with different consumers are not automatically redundant.
4. Do not add prompt branches to compensate for an unclear payload. Clarify the
status/quota/review-packet contract instead.
5. If a short worker would need to read more than one hot-path payload before it
Expand Down
16 changes: 13 additions & 3 deletions loopx/capabilities/pr_review_queue/review_contract.py
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@
from typing import Any

# Increment when review requirements change without changing the packet shape.
REVIEW_POLICY_REVISION = 6
REVIEW_POLICY_REVISION = 7

REQUIRED_FINAL_SECTIONS = [
"动机",
Expand Down Expand Up @@ -285,8 +285,18 @@ def build_review_execution_contract(*, wait_for_ci: bool = True) -> dict[str, An
"approval and must name the contract, PR trigger, observed evidence, minimum "
"repair and rerun command. Advisory cannot override a required CI failure "
"or concrete blocking finding. Do not promote future/advisory RFC properties "
"to current obligations. Reject hiding a failure by renaming a symbol, "
"raising a budget, narrowing the scan root or registering an unrelated value. "
"to current obligations. For budget failures or changes, distinguish "
"hard limits from regression budgets and presentation caps. Reuse "
"validation_matrix and observable_semantics for base/head measurements under the same workload and metric; "
"assess consumer value, true redundancy, compatibility cost and headroom. "
"Evidence-backed budget increases are valid when the owning contract and "
"tests change together; compare them with compaction or retaining the limit. "
"Preserve the original failure and any remaining evidence gaps. Hard limits "
"still require their owner's authority. Frozen experiment or promotion thresholds "
"cannot be relaxed to relabel an existing result as passing. "
"Reject hiding a failure through unjustified budget increases, "
"deleting decision semantics, cosmetic renaming, narrowing the scan root or "
"comparison workload, or registering an unrelated value. "
"Docs-only reviews may supply this same row when contract impact is found."
),
},
Expand Down
2 changes: 1 addition & 1 deletion skills/loopx-pr-review/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -47,7 +47,7 @@ Do not pipe the only copy through `jq`. When an exhaustive request has
`result_completeness.complete=false`, rerun with its `recommended_limit` before
reviewing.

Require execution `policy_revision == 6`; a schema name alone is insufficient.
Require execution `policy_revision == 7`; a schema name alone is insufficient.
If missing or unequal, do not publish APPROVE; a conservative REQUEST_CHANGES is
allowed only when it names the incompatible-policy evidence gap. Do not retain
expired temporary worktree overrides. Honor explicit runtime pins, but report
Expand Down
1 change: 1 addition & 0 deletions skills/loopx-self-repair/references/repair-patterns.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,7 @@ teaches a reusable control-plane lesson.

| Pattern | Symptoms | Evidence To Read | Likely Root | Durable Repair |
| --- | --- | --- | --- | --- |
| `budget_metric_overfitting` | A budget failure triggers automatic expansion, or mechanical compaction that removes useful semantics or breaks consumers. | Owning limit, matched base/head measurements, consumer/caller contract, original failure and revised validation. | A regression metric became the objective; historical ceilings or green tests replaced semantic judgment. | Follow the [budget decision guide](../../../docs/development/testing-and-quality.md#budget-failure-decisions), compare true redundancy, compatibility cost and justified headroom, and repair the existing contract/tests and review evidence. Preserve hard limits and frozen qualification results. |
| `skill_import_recreation` | Duplicate LoopX skills return after successful cleanup; imported entries display a fallback brand casing. | Compare installed files and metadata with source-host skills; correlate file creation times with structured host import receipts. | A later external-host import recreates command facades in another discovered root, omits display metadata, and bypasses installer reconciliation. | Attribute the writer from import receipts without guessing the human initiator; exclude already-installed LoopX skills from later imports and rerun managed reconciliation. Ensure standalone workflow entry installation writes Codex metadata, previews missing metadata repair, and records the full installed tree. Preserve user metadata and exact-host invocation behavior. |
| `skill_discovery_split_ownership` | Duplicate skill names, conflicting PR-review routes, or canonical and legacy aliases appear together. | Enumerate discovered roots, resolve directory symlinks, compare skill hashes, managed markers, install receipts, and generated metadata. | Workflow and command installers wrote independently to overlapping host roots; dedupe was optional, omitted the bare entry name, or retired copies without proving a replacement. | Repair the shared installer reconciliation and every active installation path; preserve user changes and rich workflows, retire managed aliases from the Codex picker, test repeated installs and custom profiles, then verify a fresh host catalog. Do not treat deleting one visible duplicate or changing invocation policy as a durable fix. |
| `skill_discovery_scope_eligibility_conflation` | Project delivery rejects a reusable workflow, and adding a project marker unexpectedly removes connection, repair, or review instructions from global installation. | Canonical scope markers, default shell/CLI and packaged install output, project-copy readback, and doctor required workflows. | One marker was treated as both exclusive project eligibility and default discovery; content richness was mistaken for project authority. | Declare reusable workflows global and capability-local workflows project; accept both explicit declarations for project copies while rejecting missing/unknown markers. Preserve global command routes, existing activation gates, and packaged resource parity. Never repair project delivery by hiding bootstrap or repair instructions from unconnected projects. |
Expand Down
15 changes: 14 additions & 1 deletion tests/capabilities/test_pr_review_contract.py
Original file line number Diff line number Diff line change
Expand Up @@ -125,7 +125,7 @@ def test_execution_contract_owns_deep_review_requirements() -> None:
"compatibility_only",
"unknown",
}
assert "raising a budget" in semantic["rule"]
assert "unjustified budget increases" in semantic["rule"]
assert semantic["fields"] == ["checked_scope", "impact_reason", "verdict"]
assert semantic["fields_by_verdict"]["not_applicable"] == []
assert "analysis_limit" in semantic["fields_by_verdict"]["advisory"]
Expand Down Expand Up @@ -456,6 +456,18 @@ def test_public_cli_delivers_state_review_without_claiming_it_was_performed(caps
"introduced_or_newly_enforced_state"
)
assert requirements["observable_semantics"]["state_projection_counterfactuals"]["cases"]
# The shipped packet must support justified growth as well as compaction;
# neither smaller output nor a passing revised ceiling proves correctness.
budget_rule = requirements["semantic_alignment"]["rule"]
for obligation in (
"hard limits from regression budgets and presentation caps",
"base/head measurements under the same workload and metric",
"consumer value, true redundancy, compatibility cost and headroom",
"Evidence-backed budget increases are valid",
"deleting decision semantics",
"Frozen experiment or promotion thresholds",
):
assert obligation in budget_rule
assert packet["pull_requests"]
reviewed_code = False
for item in packet["pull_requests"]:
Expand All @@ -469,6 +481,7 @@ def test_public_cli_delivers_state_review_without_claiming_it_was_performed(caps
reviewed_code = True
assert evidence["repository_reuse"] == {"status": "unverified"}
assert evidence["observable_semantics"] == {"status": "unverified"}
assert evidence["semantic_alignment"] == {"status": "unverified"}
assert reviewed_code


Expand Down
Loading