Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,16 @@ and relevant domain acceptance; do not make every fix wait for every RFC or
invent a roadmap id. Check latest `main`, related PRs and canonical Todos so an
older task description cannot override corrected direction or duplicate work.

For recurring operational or performance problems, connect the demonstrated
failure to the owning roadmap/RFC acceptance before choosing a repair. Separate
caller overhead, shared typed semantics/transport and provider-specific costs;
prefer the existing common contract where behavior is shared. Follow the
[optimization evidence guide](docs/development/testing-and-quality.md#roadmap-aligned-optimization).
Distinguish an interim mitigation from closing the owning acceptance: faster
lookup, a larger timeout or successful promotion alone does not qualify sustained
operation. Reconcile the existing checkpoint when the evidence changes it;
do not add a parallel roadmap or require unrelated RFC work for a bounded fix.

Carry one compact delivery brief from task to PR: goal/source, current gap,
observable result, owning boundary and decisive acceptance evidence. Reuse the
existing task/PR fields; keep private Goal state out of public artifacts.
Expand Down
2 changes: 1 addition & 1 deletion docs/architecture/rfcs/loopx-overall-roadmap-v0.md
Original file line number Diff line number Diff line change
Expand Up @@ -44,7 +44,7 @@ P0 blocks correctness or continuity in the current user journey. P1 enables repe
| **S7 Budget, scheduling and fleet scale · P0 observation/P1–P2 expansion** | Quota/scheduler and partial usage aggregates exist; full provider cost, distributed reservations and hundred-Agent concurrency need evidence | Separate configured budget, admission, consumption and estimates; unknown is not zero and replay cannot double-charge. R7 pagination/bounded summaries and [complete-history transport](typescript-control-plane-migration-v0.md), including refresh/replay/single-debit evidence beyond the RPC limit; provider/host limits, fairness, backpressure, event wake and isolation; report registration/activity/throughput and cost per accepted outcome separately |
| **S8 Capabilities, extensions and domain integration · P1/P2** | Capability catalog, extension lifecycle, hooks, engineering/research/content/office capabilities and computer-use contracts exist | First exercise the shared control plane with existing issue-fix/PR-review and material/research callers. Every provider has readiness/version/permissions/default-off/uninstall/rollback/isolation and real-entry evidence. New domain effects start with one simulated operation, not a marketplace or workflow DSL |
| **S9 Identity, authority, privacy and trust · continuous P0/P1–P2 remote** | Public/private scope, capability gates, fencing and confirmation contracts belong to existing owners | R1/R3 cover sender/audience/artifact scope and stale authority; R6 authenticates tenant/Goal/actor/host, rotation/revocation and least privilege. Qualify credential custody, untrusted tool/document inputs, dependency supply chain, audit retention/deletion and vulnerability response through real paths; roles/messages/memory mint no write authority |
| **S10 Reliability, diagnostics and operations · P0/P1** | Recovery/canary, read-only diagnostics prototype and DSH event adapter exist; C0/C1, overhead and full operations qualification are open | Failure classification→observable state→recovery drill→regression prevention; process/storage/network/delivery failures and data growth. Freeze SLO/RPO/RTO/capacity/retention boundaries and measure before qualification. Runbooks include upgrade, restore, stop and human takeover; test counts do not prove recovery |
| **S10 Reliability, diagnostics and operations · P0/P1** | Recovery/canary, read-only diagnostics prototype and DSH event adapter exist; C0/C1, overhead and full operations qualification are open | Failure classification→observable state→recovery drill→regression prevention; process/storage/network/delivery failures and data growth. Use [bounded repair lookup and targeted diagnostics](../../../skills/loopx-self-repair/references/targeted-diagnostics.md) to reduce redundant reads above the provider boundary; measure backend-specific cold/warm reads, writes and lock waits separately. Freeze SLO/RPO/RTO/capacity/retention boundaries and measure before qualification. Runbooks include upgrade, restore, stop and human takeover; test counts do not prove recovery |
| **S11 Evaluation and scientific research · continuous P1/P2 research** | Benchmark toolkit, Explore, long-horizon portfolio and ten frontier-science tracks have designs/partial implementations | Pin native/passive/governed arms, model/harness/budget/task split and evaluator; report native scores, cost, failures, attention and uncertainty. Prioritize sequential evidence, continuation and stride; memory, formal kernel, curriculum/evolution, active experiments and multiscale state follow T01–T10 gates without automatic production treatment |
| **S12 Release, developer experience and community governance · P0 hygiene/P1** | Install/source validation, registration, DCO/PR, test layers, contributor routes and bilingual docs exist | Qualify first work and upgrade/rollback from clean machines/release artifacts; host/OS support follows the release contract. Reduce localization/test/review effort for useful changes; preserve exact-head evidence, fixtures, compatibility, maintainer routing and contributor credit; retire duplicate protocols/stale evidence |
| **S13 Adoption, ecosystem and sustainability · P1 discovery/P2 pilots** | Public adoption loop, showcases, licensing/governance and observer-first product contract exist; paid PMF is unproven | Gather independent first/repeat usage and exit reasons; reproducible cases and pilots with fixed budgets/acceptance/rollback. Retain reusable adapters/delivery guides. Account for model/compute/storage/support and maintenance costs; only repeated demand justifies commercial hosting/support/distribution decisions, with no invented SLA or open-source-term change |
Expand Down
58 changes: 58 additions & 0 deletions docs/development/testing-and-quality.md
Original file line number Diff line number Diff line change
Expand Up @@ -560,6 +560,64 @@ agent what to do and how to request the omitted detail.
完整诊断包保留为显式 drill-down。只有默认路径仍能告诉 agent 下一步做什么、以及
如何请求被省略细节时,字段才能移出默认热路径。

### Roadmap-Aligned Optimization

For recurring command, recovery or retrieval costs, use the existing
[overall roadmap](../architecture/rfcs/loopx-overall-roadmap-v0.md) and owning
RFC acceptance to select the repair. Typed semantics and bridge costs belong
to the [TS migration RFC](../architecture/rfcs/typescript-control-plane-migration-v0.md);
backend capacity, retention and cutover qualification belong to the
[shared-authority RFC](../architecture/rfcs/shared-goal-authority-state-provider-v0.md).
Use the existing task/PR evidence and update its owning checkpoint when warranted;
this adds no approval, receipt or requirement to complete unrelated milestones.

- **Locate the cost before selecting an abstraction.** Separate caller repeats,
output/context expansion, process/bridge/serialization cost, shared semantic
work, and backend IO/verification/contention. A display filter does not reduce
upstream work; a short response or fast isolated query does not establish a
faster recovery loop. Name the real consumer and measure its useful outcome.
- **Share contracts, qualify implementations.** Put common selection, bounded
reads and observation reuse at the existing consumer/typed owner boundary.
Preserve completeness, current authority, receipt replay and lease/CAS checks.
File, SQLite and PostgreSQL need not share cache invalidation, indexing or
history layout. Keep backend-specific algorithms in their providers; do not
duplicate authority in Python or invent a common cache to conceal those costs.
Run the applicable real-backend validation above for every affected backend.
- **Treat migration as a hypothesis, not a cause.** Compare the same operation,
revision, data/history size, runtime configuration and concurrency where
possible. Separate cold/warm reads, alternating stores, writes and recovery
when those paths are affected. Without a controlled before/after comparison,
report measured costs and uncertainty. Successful ownership transfer does not
establish long-running latency, capacity or recovery equivalence.
- **Qualify the intended outcome.** Retrieval relevance needs independently
labeled queries, ambiguous/no-match cases and disclosed language/corpus limits;
a small development set is not production accuracy or task-success evidence.
Runtime optimization needs the original failing workload plus semantic and
scale checks. Follow the budget decisions below rather than hiding regressions
with a larger timeout, smaller fixture or truncated decision evidence.
- **Keep the delivered boundary honest.** State whether the change removes the
owning bottleneck or only mitigates its consumer impact. Reuse an existing
successor for an evidenced remaining gap and identify its acceptance. Do not
count a prompt reduction as backend qualification, or repeated repair PRs as
default-provider readiness; do not create follow-ups for hypothetical work.

对反复出现的命令、恢复或检索开销,先对应总 roadmap 和所属 RFC 的验收,再选择修复:
TS RFC 管语义 owner 与跨语言成本,shared-authority RFC 管后端容量、保留与切换验证。
沿用现有任务、PR 证据和验收记录,不增加审批、回执或无关里程碑前置条件。

- **先定位成本。** 区分重复调用、输出与上下文展开、进程与序列化、公共语义计算、
后端 IO/校验/竞争。输出过滤不减少上游计算;单次查询变快不等于恢复闭环变快。
- **共用合同,分别验证实现。** 选择、有界读取和观测复用归现有调用方或 typed owner;
保留完整性、当前权限、回执重放及 lease/CAS。缓存失效、索引和历史布局可以因后端
而异,不能在 Python 复制权威,也不强造统一缓存。受影响后端遵循上文真实路径验证。
- **迁移是待验证的原因。** 尽量控制操作、版本、数据与历史规模、配置、并发;按影响
区分冷/热读、多存储交替、写入与恢复。没有受控前后对照就披露不确定性;晋升成功
不证明长程延迟、容量和恢复等价。
- **验证实际目标。** 检索用独立标注、歧义/无命中和语言边界验证;开发小样本不代表
生产精度或任务成功率。运行时优化保留原失败负载和语义/规模检查,预算按下节处理。
- **区分缓解与闭环。** 写明消除了所属瓶颈,还是只减轻调用方影响;剩余真实缺口复用
已有后续任务并指明验收。提示词缩短不算后端验证,修复 PR 数量不算默认切换就绪。

### Budget Failure Decisions

Classify the limit by its owning contract before deciding how to repair a
Expand Down
14 changes: 6 additions & 8 deletions examples/install-local-smoke.py
Original file line number Diff line number Diff line change
Expand Up @@ -405,8 +405,8 @@ def main() -> int:
"loopx review-packet --goal-id <STABLE_GOAL_ID> --handoff-only",
"loopx --format json review-packet --goal-id",
"target project agent must not run this draft",
"This command is read-only",
"JSON output returns a minimized handoff payload with `handoff_text` instead of the full operator packet",
"This read-only command assembles agent context directly from current status",
"JSON `handoff_text` and `project_agent_handoff` always contain complete prepared text",
"--classification <PUBLIC_SAFE_PROGRESS_CLASSIFICATION>",
"--delivery-batch-scale <ACTUAL_DELIVERY_BATCH_SCALE>",
"--delivery-outcome <ACTUAL_DELIVERY_OUTCOME>",
Expand Down Expand Up @@ -512,12 +512,10 @@ def main() -> int:
self_repair_skill = codex_home / "skills" / "loopx-self-repair" / "SKILL.md"
self_repair_text = " ".join(self_repair_skill.read_text(encoding="utf-8").split())
for phrase in (
"Build a compact evidence packet",
"loopx --format json diagnose --goal-id <goal-id>",
"loopx --format json status --goal-id <goal-id> --limit 20",
"status` defaults to the registry/dashboard view, but accepts `--goal-id`",
"registry-declared active state file",
"references/repair-patterns.md",
"Reuse evidence before collecting more",
"scripts/find_pattern.py",
"references/targeted-diagnostics.md",
"references/pattern-lookup.md",
"Repair at the lowest durable layer",
"Do not solve contradictory payloads by guessing",
):
Expand Down
9 changes: 4 additions & 5 deletions loopx/doctor.py
Original file line number Diff line number Diff line change
Expand Up @@ -71,11 +71,10 @@
"For a generic library microbenchmark",
),
"loopx-self-repair": (
"Build a compact evidence packet",
"loopx --format json diagnose --goal-id <goal-id>",
"loopx --format json status --goal-id <goal-id> --limit 20",
"registry-declared active state file",
"references/repair-patterns.md",
"Reuse evidence before collecting more",
"scripts/find_pattern.py",
"references/targeted-diagnostics.md",
"references/pattern-lookup.md",
"Repair at the lowest durable layer",
),
}
Expand Down
5 changes: 5 additions & 0 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -125,8 +125,13 @@ include = ["loopx*"]
]
"share/loopx/skills/loopx-self-repair/references" = [
"skills/loopx-self-repair/references/repair-patterns.md",
"skills/loopx-self-repair/references/pattern-lookup.md",
"skills/loopx-self-repair/references/targeted-diagnostics.md",
"skills/loopx-self-repair/references/upstream-issue-escalation.md",
]
"share/loopx/skills/loopx-self-repair/scripts" = [
"skills/loopx-self-repair/scripts/find_pattern.py",
]

[tool.pytest.ini_options]
markers = ["stage2c_e2e: real process Stage 2C correctness and recovery acceptance"]
Expand Down
41 changes: 20 additions & 21 deletions skills/loopx-self-repair/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,25 +12,22 @@ not only an apology or a one-off explanation.

1. **Pause delivery selection.** Do not spend quota or continue adapter work
until the control-plane facts explain why that work is valid.
2. **Build a compact evidence packet.** Prefer structured surfaces:

```bash
git status --short --branch
loopx --format json diagnose --goal-id <goal-id>
loopx --format json status --goal-id <goal-id> --limit 20
loopx --format json quota should-run --goal-id <goal-id> [--agent-id <agent-id>]
loopx --format json history --goal-id <goal-id> --limit 5
```

`status` defaults to the registry/dashboard view, but accepts `--goal-id`
when the repair needs one goal-focused projection. Use
`diagnose --goal-id` for the richer goal-specific agent reasoning packet.
Also inspect the project-local registry and the registry-declared active
state file when relevant. Use the shared global registry for heartbeat/quota
truth.
3. **Classify the failure.** Read
`references/repair-patterns.md` and match the symptoms to a known pattern.
If no pattern fits, add one after the fix.
2. **Reuse evidence before collecting more.** Start with the current failed
command's structured response, error code and operation identity. An already
loaded packet is evidence for that observation, not permission for a later
write. Fetch fresh authority when required by its admission/lease contract.
Read [targeted diagnostics](references/targeted-diagnostics.md) when deciding
which missing fact to collect or investigating slow commands. Do not run
diagnose, status, quota and history as a fixed preflight: diagnose already
composes status and quota work. Recording an already-understood repair Todo
does not require rediscovering the incident.
3. **Look up the symptom.** Run `python3 scripts/find_pattern.py --query
'<error code or symptom terms>'` from this skill directory, or invoke its
absolute path. Use `--id <returned-id>` to read the relevant full guidance.
[Search instructions](references/pattern-lookup.md) explain pagination and
fallback. Do not load the complete catalog, paginate it into context, or
reread unchanged references already available in this task. If no pattern
fits, diagnose from current facts and add one after the fix.
4. **Assign the responsible layer.** Separate:
- agent behavior mistake;
- state projection or quota payload bug;
Expand Down Expand Up @@ -145,8 +142,10 @@ replan, or terminal closeout must return to the strict semantic checkpoint.

## Reference Routes

- For known symptom-to-repair mappings, read
`references/repair-patterns.md`.
- For known symptom-to-repair mappings, search with `scripts/find_pattern.py`;
`references/pattern-lookup.md` explains the lookup, not a required full read.
- For missing facts, slow commands and response truncation, read
`references/targeted-diagnostics.md`.
- For guarded public GitHub issue escalation, read
`references/upstream-issue-escalation.md`.
- For user/agent/state channel semantics, read
Expand Down
45 changes: 45 additions & 0 deletions skills/loopx-self-repair/references/pattern-lookup.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,45 @@
# Find a repair pattern

The catalog is a reference database, not a prerequisite reading list. Search
using the exact error code or two distinctive symptom terms, then expand the
relevant result. Run from this skill directory, or use the script's absolute
path from any working directory:

```bash
python3 scripts/find_pattern.py --query 'turn recovery'
python3 scripts/find_pattern.py --id acceptance_scope_capture
```

Search returns at most five ids with their complete symptoms, plus
`total_matches` and `next_offset`. BM25 ranks lexical matches; a complete pattern
id, then an exact backtick-delimited code token, take precedence. Results label
that precedence in `exact_match`. Identifiers split at underscores and punctuation, ignoring
case. Refine the query or pass `--offset` to see the next page. If an exact
error code has no match, try its domain and symptom words. `--list` browses ids;
`--id` returns the complete evidence, root and repair guidance without truncation.
Each result exposes `matched_terms`, and `unmatched_terms` names query words
absent from the corpus. Scores are ordering signals, not confidence or a
diagnosis. A generic shared word such as `turn` can rank unrelated incidents:
expand the symptom with the failing action and exact error before choosing a
repair. There is no synonym expansion, stemming, translation or semantic model;
for the English catalog, use English terms or literal protocol identifiers.
A match is a hypothesis: verify it against the current typed contract and facts.

The small manually labeled regression set in
`tests/fixtures/self_repair_queries.json` includes an intentionally under-specified
query to retain this limitation. Its hit rates are a local sanity check, not
measured production search accuracy. BM25 uses fixed `k1=1.2`, `b=0.75` and
positive IDF; these parameters are not fitted to that set. See the
[Lucene BM25 formula and defaults](https://lucene.apache.org/core/9_12_1/core/org/apache/lucene/search/similarities/BM25Similarity.html).

The script uses only Python's standard library, resolves resources relative to
itself, and does not invoke LoopX, read a registry, or create a Turn. It is also
delivered with installed workflow skills. If Python is unavailable, search
the source with `rg -n -F '<distinctive term>' references/repair-patterns.md`
and read only the matching rows or section. Do not recover truncated output by
paging through the entire catalog.

Maintainers add or amend the single canonical
[pattern catalog](repair-patterns.md), preserving its five-column table and
unique ids. Prose appendices are searchable as `note_*` entries. Existing
patterns are diagnosis aids, not additional authority or mandatory checklists.
Loading
Loading