From 2e1db80b127d9172d6fd97ce36f58fa53711bf45 Mon Sep 17 00:00:00 2001 From: huangruiteng <14976749+huangruiteng@users.noreply.github.com> Date: Sat, 26 Sep 2026 18:51:48 +0800 Subject: [PATCH 1/2] perf(skills): bound self-repair lookup and diagnostic reads Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com> --- .../rfcs/loopx-overall-roadmap-v0.md | 2 +- examples/install-local-smoke.py | 14 +- loopx/doctor.py | 9 +- pyproject.toml | 5 + skills/loopx-self-repair/SKILL.md | 41 +++--- .../references/pattern-lookup.md | 45 ++++++ .../references/targeted-diagnostics.md | 80 +++++++++++ .../loopx-self-repair/scripts/find_pattern.py | 126 ++++++++++++++++ tests/fixtures/self_repair_queries.json | 12 ++ tests/test_repair_pattern_lookup.py | 135 ++++++++++++++++++ 10 files changed, 434 insertions(+), 35 deletions(-) create mode 100644 skills/loopx-self-repair/references/pattern-lookup.md create mode 100644 skills/loopx-self-repair/references/targeted-diagnostics.md create mode 100644 skills/loopx-self-repair/scripts/find_pattern.py create mode 100644 tests/fixtures/self_repair_queries.json create mode 100644 tests/test_repair_pattern_lookup.py diff --git a/docs/architecture/rfcs/loopx-overall-roadmap-v0.md b/docs/architecture/rfcs/loopx-overall-roadmap-v0.md index 91e0086dc7..546c1d4786 100644 --- a/docs/architecture/rfcs/loopx-overall-roadmap-v0.md +++ b/docs/architecture/rfcs/loopx-overall-roadmap-v0.md @@ -44,7 +44,7 @@ P0 blocks correctness or continuity in the current user journey. P1 enables repe | **S7 Budget, scheduling and fleet scale · P0 observation/P1–P2 expansion** | Quota/scheduler and partial usage aggregates exist; full provider cost, distributed reservations and hundred-Agent concurrency need evidence | Separate configured budget, admission, consumption and estimates; unknown is not zero and replay cannot double-charge. R7 pagination/bounded summaries and [complete-history transport](typescript-control-plane-migration-v0.md), including refresh/replay/single-debit evidence beyond the RPC limit; provider/host limits, fairness, backpressure, event wake and isolation; report registration/activity/throughput and cost per accepted outcome separately | | **S8 Capabilities, extensions and domain integration · P1/P2** | Capability catalog, extension lifecycle, hooks, engineering/research/content/office capabilities and computer-use contracts exist | First exercise the shared control plane with existing issue-fix/PR-review and material/research callers. Every provider has readiness/version/permissions/default-off/uninstall/rollback/isolation and real-entry evidence. New domain effects start with one simulated operation, not a marketplace or workflow DSL | | **S9 Identity, authority, privacy and trust · continuous P0/P1–P2 remote** | Public/private scope, capability gates, fencing and confirmation contracts belong to existing owners | R1/R3 cover sender/audience/artifact scope and stale authority; R6 authenticates tenant/Goal/actor/host, rotation/revocation and least privilege. Qualify credential custody, untrusted tool/document inputs, dependency supply chain, audit retention/deletion and vulnerability response through real paths; roles/messages/memory mint no write authority | -| **S10 Reliability, diagnostics and operations · P0/P1** | Recovery/canary, read-only diagnostics prototype and DSH event adapter exist; C0/C1, overhead and full operations qualification are open | Failure classification→observable state→recovery drill→regression prevention; process/storage/network/delivery failures and data growth. Freeze SLO/RPO/RTO/capacity/retention boundaries and measure before qualification. Runbooks include upgrade, restore, stop and human takeover; test counts do not prove recovery | +| **S10 Reliability, diagnostics and operations · P0/P1** | Recovery/canary, read-only diagnostics prototype and DSH event adapter exist; C0/C1, overhead and full operations qualification are open | Failure classification→observable state→recovery drill→regression prevention; process/storage/network/delivery failures and data growth. Use [bounded repair lookup and targeted diagnostics](../../../skills/loopx-self-repair/references/targeted-diagnostics.md) to reduce redundant reads above the provider boundary; measure backend-specific cold/warm reads, writes and lock waits separately. Freeze SLO/RPO/RTO/capacity/retention boundaries and measure before qualification. Runbooks include upgrade, restore, stop and human takeover; test counts do not prove recovery | | **S11 Evaluation and scientific research · continuous P1/P2 research** | Benchmark toolkit, Explore, long-horizon portfolio and ten frontier-science tracks have designs/partial implementations | Pin native/passive/governed arms, model/harness/budget/task split and evaluator; report native scores, cost, failures, attention and uncertainty. Prioritize sequential evidence, continuation and stride; memory, formal kernel, curriculum/evolution, active experiments and multiscale state follow T01–T10 gates without automatic production treatment | | **S12 Release, developer experience and community governance · P0 hygiene/P1** | Install/source validation, registration, DCO/PR, test layers, contributor routes and bilingual docs exist | Qualify first work and upgrade/rollback from clean machines/release artifacts; host/OS support follows the release contract. Reduce localization/test/review effort for useful changes; preserve exact-head evidence, fixtures, compatibility, maintainer routing and contributor credit; retire duplicate protocols/stale evidence | | **S13 Adoption, ecosystem and sustainability · P1 discovery/P2 pilots** | Public adoption loop, showcases, licensing/governance and observer-first product contract exist; paid PMF is unproven | Gather independent first/repeat usage and exit reasons; reproducible cases and pilots with fixed budgets/acceptance/rollback. Retain reusable adapters/delivery guides. Account for model/compute/storage/support and maintenance costs; only repeated demand justifies commercial hosting/support/distribution decisions, with no invented SLA or open-source-term change | diff --git a/examples/install-local-smoke.py b/examples/install-local-smoke.py index f8a0446a0a..3bfcb16ef4 100644 --- a/examples/install-local-smoke.py +++ b/examples/install-local-smoke.py @@ -405,8 +405,8 @@ def main() -> int: "loopx review-packet --goal-id --handoff-only", "loopx --format json review-packet --goal-id", "target project agent must not run this draft", - "This command is read-only", - "JSON output returns a minimized handoff payload with `handoff_text` instead of the full operator packet", + "This read-only command assembles agent context directly from current status", + "JSON `handoff_text` and `project_agent_handoff` always contain complete prepared text", "--classification ", "--delivery-batch-scale ", "--delivery-outcome ", @@ -512,12 +512,10 @@ def main() -> int: self_repair_skill = codex_home / "skills" / "loopx-self-repair" / "SKILL.md" self_repair_text = " ".join(self_repair_skill.read_text(encoding="utf-8").split()) for phrase in ( - "Build a compact evidence packet", - "loopx --format json diagnose --goal-id ", - "loopx --format json status --goal-id --limit 20", - "status` defaults to the registry/dashboard view, but accepts `--goal-id`", - "registry-declared active state file", - "references/repair-patterns.md", + "Reuse evidence before collecting more", + "scripts/find_pattern.py", + "references/targeted-diagnostics.md", + "references/pattern-lookup.md", "Repair at the lowest durable layer", "Do not solve contradictory payloads by guessing", ): diff --git a/loopx/doctor.py b/loopx/doctor.py index df4fc19e2d..bffb9fc22c 100644 --- a/loopx/doctor.py +++ b/loopx/doctor.py @@ -71,11 +71,10 @@ "For a generic library microbenchmark", ), "loopx-self-repair": ( - "Build a compact evidence packet", - "loopx --format json diagnose --goal-id ", - "loopx --format json status --goal-id --limit 20", - "registry-declared active state file", - "references/repair-patterns.md", + "Reuse evidence before collecting more", + "scripts/find_pattern.py", + "references/targeted-diagnostics.md", + "references/pattern-lookup.md", "Repair at the lowest durable layer", ), } diff --git a/pyproject.toml b/pyproject.toml index 111681fae1..99b3f2fd2d 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -125,8 +125,13 @@ include = ["loopx*"] ] "share/loopx/skills/loopx-self-repair/references" = [ "skills/loopx-self-repair/references/repair-patterns.md", + "skills/loopx-self-repair/references/pattern-lookup.md", + "skills/loopx-self-repair/references/targeted-diagnostics.md", "skills/loopx-self-repair/references/upstream-issue-escalation.md", ] +"share/loopx/skills/loopx-self-repair/scripts" = [ + "skills/loopx-self-repair/scripts/find_pattern.py", +] [tool.pytest.ini_options] markers = ["stage2c_e2e: real process Stage 2C correctness and recovery acceptance"] diff --git a/skills/loopx-self-repair/SKILL.md b/skills/loopx-self-repair/SKILL.md index b9a7cba2ec..049be9ad06 100644 --- a/skills/loopx-self-repair/SKILL.md +++ b/skills/loopx-self-repair/SKILL.md @@ -12,25 +12,22 @@ not only an apology or a one-off explanation. 1. **Pause delivery selection.** Do not spend quota or continue adapter work until the control-plane facts explain why that work is valid. -2. **Build a compact evidence packet.** Prefer structured surfaces: - - ```bash - git status --short --branch - loopx --format json diagnose --goal-id - loopx --format json status --goal-id --limit 20 - loopx --format json quota should-run --goal-id [--agent-id ] - loopx --format json history --goal-id --limit 5 - ``` - - `status` defaults to the registry/dashboard view, but accepts `--goal-id` - when the repair needs one goal-focused projection. Use - `diagnose --goal-id` for the richer goal-specific agent reasoning packet. - Also inspect the project-local registry and the registry-declared active - state file when relevant. Use the shared global registry for heartbeat/quota - truth. -3. **Classify the failure.** Read - `references/repair-patterns.md` and match the symptoms to a known pattern. - If no pattern fits, add one after the fix. +2. **Reuse evidence before collecting more.** Start with the current failed + command's structured response, error code and operation identity. An already + loaded packet is evidence for that observation, not permission for a later + write. Fetch fresh authority when required by its admission/lease contract. + Read [targeted diagnostics](references/targeted-diagnostics.md) when deciding + which missing fact to collect or investigating slow commands. Do not run + diagnose, status, quota and history as a fixed preflight: diagnose already + composes status and quota work. Recording an already-understood repair Todo + does not require rediscovering the incident. +3. **Look up the symptom.** Run `python3 scripts/find_pattern.py --query + ''` from this skill directory, or invoke its + absolute path. Use `--id ` to read the relevant full guidance. + [Search instructions](references/pattern-lookup.md) explain pagination and + fallback. Do not load the complete catalog, paginate it into context, or + reread unchanged references already available in this task. If no pattern + fits, diagnose from current facts and add one after the fix. 4. **Assign the responsible layer.** Separate: - agent behavior mistake; - state projection or quota payload bug; @@ -145,8 +142,10 @@ replan, or terminal closeout must return to the strict semantic checkpoint. ## Reference Routes -- For known symptom-to-repair mappings, read - `references/repair-patterns.md`. +- For known symptom-to-repair mappings, search with `scripts/find_pattern.py`; + `references/pattern-lookup.md` explains the lookup, not a required full read. +- For missing facts, slow commands and response truncation, read + `references/targeted-diagnostics.md`. - For guarded public GitHub issue escalation, read `references/upstream-issue-escalation.md`. - For user/agent/state channel semantics, read diff --git a/skills/loopx-self-repair/references/pattern-lookup.md b/skills/loopx-self-repair/references/pattern-lookup.md new file mode 100644 index 0000000000..24d9903036 --- /dev/null +++ b/skills/loopx-self-repair/references/pattern-lookup.md @@ -0,0 +1,45 @@ +# Find a repair pattern + +The catalog is a reference database, not a prerequisite reading list. Search +using the exact error code or two distinctive symptom terms, then expand the +relevant result. Run from this skill directory, or use the script's absolute +path from any working directory: + +```bash +python3 scripts/find_pattern.py --query 'turn recovery' +python3 scripts/find_pattern.py --id acceptance_scope_capture +``` + +Search returns at most five ids with their complete symptoms, plus +`total_matches` and `next_offset`. BM25 ranks lexical matches; a complete pattern +id, then an exact backtick-delimited code token, take precedence. Results label +that precedence in `exact_match`. Identifiers split at underscores and punctuation, ignoring +case. Refine the query or pass `--offset` to see the next page. If an exact +error code has no match, try its domain and symptom words. `--list` browses ids; +`--id` returns the complete evidence, root and repair guidance without truncation. +Each result exposes `matched_terms`, and `unmatched_terms` names query words +absent from the corpus. Scores are ordering signals, not confidence or a +diagnosis. A generic shared word such as `turn` can rank unrelated incidents: +expand the symptom with the failing action and exact error before choosing a +repair. There is no synonym expansion, stemming, translation or semantic model; +for the English catalog, use English terms or literal protocol identifiers. +A match is a hypothesis: verify it against the current typed contract and facts. + +The small manually labeled regression set in +`tests/fixtures/self_repair_queries.json` includes an intentionally under-specified +query to retain this limitation. Its hit rates are a local sanity check, not +measured production search accuracy. BM25 uses fixed `k1=1.2`, `b=0.75` and +positive IDF; these parameters are not fitted to that set. See the +[Lucene BM25 formula and defaults](https://lucene.apache.org/core/9_12_1/core/org/apache/lucene/search/similarities/BM25Similarity.html). + +The script uses only Python's standard library, resolves resources relative to +itself, and does not invoke LoopX, read a registry, or create a Turn. It is also +delivered with installed workflow skills. If Python is unavailable, search +the source with `rg -n -F '' references/repair-patterns.md` +and read only the matching rows or section. Do not recover truncated output by +paging through the entire catalog. + +Maintainers add or amend the single canonical +[pattern catalog](repair-patterns.md), preserving its five-column table and +unique ids. Prose appendices are searchable as `note_*` entries. Existing +patterns are diagnosis aids, not additional authority or mandatory checklists. diff --git a/skills/loopx-self-repair/references/targeted-diagnostics.md b/skills/loopx-self-repair/references/targeted-diagnostics.md new file mode 100644 index 0000000000..382342634c --- /dev/null +++ b/skills/loopx-self-repair/references/targeted-diagnostics.md @@ -0,0 +1,80 @@ +# Targeted diagnosis and command cost + +Choose the missing fact, then the existing owner that can answer it. Preserve +the current Goal, runtime/registry route, Agent and Turn identity; diagnosis +does not authorize borrowing another session's identity or changing providers. + +| Missing fact | First useful surface | +| --- | --- | +| Why the current operation was rejected | Its existing JSON error, typed recovery and operation receipt; look up the error before opening other projections | +| Whether an ambiguous Todo create committed | `loopx --format json todo receipt --goal-id --operation-id `; inspect the receipt before any retry with the same identity | +| Current lease ownership/version | `loopx --format json task-lease inspect --goal-id --todo-id ` | +| One Todo's current state | `loopx --format json todo list --goal-id --todo-id `; use `--agent-id` when needed by the existing lane | +| Why an existing quota packet selected recovery | Inspect that packet's recovery and interaction fields first; request fresh quota when binding/settling the current Turn or after a relevant state change | +| Overall unexplained Goal health | `loopx --format json diagnose --goal-id --agent-id `; this composes status and quota reads | +| A dashboard discrepancy | `loopx --format json status --goal-id --limit 5` | +| Recent execution order | `loopx --format json history --goal-id --limit 5` | + +These are alternatives, not a sequence. Use the resolved registry/runtime +options for all calls. `quota should-run` may establish host-Turn state; it is +not a harmless latency probe. Never remove a Turn id, capability declaration, +lease proof, revision or receipt check to make a command faster. + +## Read one response more than once, not one command + +When JSON may exceed the tool output budget, capture it once in an ignored +private file with restricted permissions, or in the tool runtime's value store. +Keep the command exit status and validate that the capture is complete JSON. +Select fields from that saved response; if you need additional fields, read +the same capture. Do not rerun a costly or stateful command just to increase +`max_output_tokens`. A tool's truncated display does not mean the command +failed or did not commit. + +Use server-side selectors where they exist: `todo list --todo-id` for a known +Todo and a bounded `pr-review --limit` for queue selection. A shell `jq` filter +reduces displayed text but does not avoid constructing the full upstream +response. Keep total/completeness markers when limiting discovery; a bounded +page cannot prove the absence of other work. Do not silently truncate evidence +used for authorization, review or settlement. + +On an ambiguous write, follow its typed recovery contract. An absent receipt +does not prove that a timed-out writer stopped: retain the original operation +id and payload, wait/inspect as directed, and avoid concurrent retries. After +a relevant mutation, refresh the facts needed by the next operation; a cached +diagnostic is not current authority. + +## Measure the right latency + +Distinguish command start-to-exit duration from the tool's initial yield or +poll wait. Separately record response bytes, output truncation, read count and +time between useful operations. A fast `cat` can still consume substantial +context and reasoning time when its output is repeatedly expanded. + +Before changing runtime budgets, identify the installed source revision and +provider, compare the same workload on base/head, and separate startup, lock +wait, backend verification and response construction. Use an isolated runtime +and synthetic fixture or authorized read-only snapshot. Preserve integrity, +receipt recovery and lease/CAS semantics; do not benchmark by mutating an active +Goal. Check existing PRs before starting an overlapping store refactor. + +Searchable reference lookup has no Goal authority and can stay in the skill. +Runtime admission, recovery decisions and provider integrity remain in their +existing typed owners. This is the S10 diagnostic-efficiency boundary alongside +S2 transaction migration; shorter prompts alone do not qualify provider cutover. + +## Provider-neutral optimization boundary + +Apply query selection and response reuse above the provider interface: File, +SQLite and PostgreSQL all benefit when consumers avoid redundant requests. +Reuse within one observation only; a provider revision, current lease or CAS +precondition must still come from its authority owner. Do not add a Python +cache that bypasses the typed owner, or assume different providers share a +filesystem invalidation rule. + +Separate this from storage-specific work. File-v0's retained journal decoding +and whole-file rewrite, SQLite transactions/indexes, and PostgreSQL queries and +network round trips have different costs. Prove a shared optimization through +the common read/transaction contract, then qualify each affected real backend. +Preserve original-receipt recovery, stale-revision rejection and missing-state +fail-closed behavior. A successful promotion establishes authority ownership; +it does not establish latency or capacity equivalence to the previous path. diff --git a/skills/loopx-self-repair/scripts/find_pattern.py b/skills/loopx-self-repair/scripts/find_pattern.py new file mode 100644 index 0000000000..02acafd490 --- /dev/null +++ b/skills/loopx-self-repair/scripts/find_pattern.py @@ -0,0 +1,126 @@ +"""Search the packaged repair reference without loading LoopX or any Goal state.""" + +from __future__ import annotations + +import argparse +from collections import Counter +import json +from math import log1p +from pathlib import Path +import re + + +CATALOG = Path(__file__).resolve().parents[1] / "references" / "repair-patterns.md" +FIELDS = ("pattern", "symptoms", "evidence", "likely_root", "durable_repair") + + +def read_patterns(path: Path) -> list[dict[str, str]]: + """Keep complete table cells and prose appendices; never silently drop a row.""" + patterns: list[dict[str, str]] = [] + section: dict[str, str] | None = None + for line_number, line in enumerate(path.read_text(encoding="utf-8").splitlines(), 1): + if line.startswith("| `"): + # Delimiters have spaces; a code span such as `merged|all` is content. + cells = re.split(r"\s+\|\s+", line.removeprefix("| ").removesuffix(" |")) + if len(cells) != len(FIELDS) or not re.fullmatch(r"`[a-z0-9_]+`", cells[0]): + raise ValueError(f"invalid pattern row at line {line_number}") + row = dict(zip(FIELDS, cells)) + row["pattern"] = row["pattern"].strip("`") + patterns.append(row) + elif line.startswith(("## ", "### ")): + title = line.lstrip("# ") + section = { + "pattern": "note_" + re.sub(r"[^a-z0-9]+", "_", title.lower()).strip("_"), + "symptoms": title, + "guidance": "", + } + patterns.append(section) + elif section is not None: + section["guidance"] += line + "\n" + elif line.startswith("|") and not line.startswith(("| Pattern |", "| --- |")): + raise ValueError(f"unrecognized catalog row at line {line_number}") + ids = [row["pattern"] for row in patterns] + if not ids or len(set(ids)) != len(ids): + raise ValueError("catalog must contain unique pattern ids") + return patterns + + +def search_patterns( + patterns: list[dict[str, str]], query: str, *, offset: int, limit: int +) -> dict[str, object]: + # BM25 (k1=1.2, b=0.75), with Lucene's positive IDF. Whole identifiers + # remain addressable by --id; tokenization also exposes their components. + terms = set(re.findall(r"[^\W_]+", query.casefold())) + documents = [Counter(re.findall(r"[^\W_]+", "\n".join(row.values()).casefold())) for row in patterns] + lengths = [sum(doc.values()) for doc in documents] + average = sum(lengths) / len(lengths) if lengths else 1 + frequencies = Counter(term for doc in documents for term in doc) + scored = [] + for row, document, length in zip(patterns, documents, lengths): + matched_terms = sorted(terms & document.keys()) + score = sum( + log1p((len(patterns) - frequencies[term] + 0.5) / (frequencies[term] + 0.5)) + * document[term] * 2.2 + / (document[term] + 1.2 * (0.25 + 0.75 * length / (average or 1))) + for term in matched_terms + ) + exact_id = row["pattern"].casefold() == query.strip().casefold() + exact_code = any(query.strip().casefold() == code.casefold() + for code in re.findall(r"`([^`\n]+)`", "\n".join(row.values()))) + if score > 0 or exact_id or not query: + scored.append((exact_id, exact_code, score, row, matched_terms)) + # Stable id tie-break keeps pagination reproducible; scores are not confidence. + if query: + scored.sort(key=lambda item: (-item[0], -item[1], -item[2], item[3]["pattern"])) + matched = [{"pattern": row["pattern"], "symptoms": row["symptoms"], + "score": round(score, 6), "matched_terms": tokens, + "exact_match": "pattern_id" if exact_id else "code" if exact_code else None} + for exact_id, exact_code, score, row, tokens in scored] + page = matched[offset:offset + limit] + end = offset + len(page) + return { + "ok": True, + "query": query, + "match_mode": "bm25_exact_first" if query else "catalog_order", + "unmatched_terms": sorted(terms - frequencies.keys()), + "total_matches": len(matched), + "offset": offset, + "next_offset": end if end < len(matched) else None, + "patterns": page, + } + + +def main() -> int: + parser = argparse.ArgumentParser(description=__doc__) + mode = parser.add_mutually_exclusive_group(required=True) + mode.add_argument("--query", help="BM25 lexical search; exact id/code first, no regex or embeddings.") + mode.add_argument("--id", help="Return the complete guidance for one exact pattern id.") + mode.add_argument("--list", action="store_true", help="Browse a page of ids and symptoms.") + parser.add_argument("--limit", type=int, default=5, help="Search page size, 1–20 (default: 5).") + parser.add_argument("--offset", type=int, default=0, help="Search offset from next_offset.") + args = parser.parse_args() + if not 1 <= args.limit <= 20 or args.offset < 0: + parser.error("--limit must be 1–20 and --offset must be nonnegative") + if args.query is not None and not args.query.strip(): + parser.error("--query must contain a term; use --list to browse") + if args.query is not None and not re.search(r"[^\W_]+", args.query): + parser.error("--query must contain a word or identifier") + if args.id is not None and not args.id.strip(): + parser.error("--id must name a pattern") + try: + patterns = read_patterns(CATALOG) + if args.id is not None: + row = next((row for row in patterns if row["pattern"] == args.id), None) + payload = {"ok": row is not None, "pattern": row} + if row is None: + payload["error"] = "unknown pattern id; search with --query or --list" + else: + payload = search_patterns(patterns, args.query or "", offset=args.offset, limit=args.limit) + except (OSError, ValueError) as exc: + payload = {"ok": False, "error": str(exc)} + print(json.dumps(payload, ensure_ascii=False, indent=2)) + return 0 if payload["ok"] else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/tests/fixtures/self_repair_queries.json b/tests/fixtures/self_repair_queries.json new file mode 100644 index 0000000000..d2695d6ea0 --- /dev/null +++ b/tests/fixtures/self_repair_queries.json @@ -0,0 +1,12 @@ +[ + {"query": "prior heartbeat closeout after blocked writeback", "relevant": "host_closeout_presentation_lookup_gap"}, + {"query": "unsettled turn", "relevant": "host_closeout_presentation_lookup_gap", "limitation": "Under-specified lexical query; unsettled does not occur in the relevant pattern."}, + {"query": "quota timeout long goal history closeout", "relevant": "host_closeout_history_scan_amplification"}, + {"query": "released lease old execution key completes new generation", "relevant": "task_lease_generation_aba"}, + {"query": "validation contract unrelated work unbound", "relevant": "acceptance_scope_capture"}, + {"query": "native todo index null invalid dashboard response", "relevant": "native_todo_status_index_schema_gap"}, + {"query": "CLI still old quota behavior after merge", "relevant": "local_release_snapshot_stale_after_merge"}, + {"query": "shadow capture oversized RPC response", "relevant": "shadow_proof_transport_amplification"}, + {"query": "review retains duplicate protocol compatibility decoders", "relevant": "review_compatibility_assumption_gap"}, + {"query": "contract criteria runner failed workspace completion", "relevant": "acceptance_validation_failure_flattened"} +] diff --git a/tests/test_repair_pattern_lookup.py b/tests/test_repair_pattern_lookup.py new file mode 100644 index 0000000000..5202a26179 --- /dev/null +++ b/tests/test_repair_pattern_lookup.py @@ -0,0 +1,135 @@ +"""Bounded discovery, complete guidance and real installed skill execution.""" + +from __future__ import annotations + +import importlib.util +import json +from pathlib import Path +import subprocess +import sys +import tomllib + +import pytest + + +ROOT = Path(__file__).resolve().parents[1] +SKILL = ROOT / "skills" / "loopx-self-repair" +SCRIPT = SKILL / "scripts" / "find_pattern.py" +spec = importlib.util.spec_from_file_location("repair_lookup", SCRIPT) +assert spec is not None and spec.loader is not None +lookup = importlib.util.module_from_spec(spec) +spec.loader.exec_module(lookup) + + +def test_search_paginates_without_losing_or_expanding_guidance(): + rows = [{ + "pattern": f"pattern_{n}", "symptoms": "State [invalid]", "durable_repair": "Keep CAS proof." + } for n in range(8)] + first = lookup.search_patterns(rows, "[INVALID] proof", offset=0, limit=5) + second = lookup.search_patterns(rows, "[INVALID] proof", offset=first["next_offset"], limit=5) + assert first["total_matches"] == second["total_matches"] == 8 + assert second["next_offset"] is None + assert [r["pattern"] for r in first["patterns"] + second["patterns"]] == [r["pattern"] for r in rows] + assert all("durable_repair" not in r for r in first["patterns"]) + assert lookup.search_patterns(rows, "missing", offset=0, limit=5)["total_matches"] == 0 + + +def test_bm25_handles_partial_terms_and_keeps_exact_ids_first(): + rows = [ + {"pattern": "lease_recovery", "symptoms": "Refresh the released lease proof: `stale_proof`"}, + {"pattern": "other", "symptoms": "lease recovery lease recovery"}, + {"pattern": "display", "symptoms": "Turn display repair"}, + ] + result = lookup.search_patterns(rows, "lease_recovery", offset=0, limit=5) + assert result["patterns"][0]["pattern"] == "lease_recovery" + result = lookup.search_patterns(rows, "released unavailableword", offset=0, limit=5) + assert [r["pattern"] for r in result["patterns"]] == ["lease_recovery"] + assert result["unmatched_terms"] == ["unavailableword"] + assert result["patterns"][0]["matched_terms"] == ["released"] + result = lookup.search_patterns(rows, "stale_proof", offset=0, limit=5) + assert result["patterns"][0]["exact_match"] == "code" + assert result["patterns"][0]["pattern"] == "lease_recovery" + + +def test_labeled_repair_queries_improve_over_all_terms_filter(): + rows = lookup.read_patterns(lookup.CATALOG) + cases = json.loads((ROOT / "tests/fixtures/self_repair_queries.json").read_text()) + baseline_hits = candidate_hits = 0 + for case in cases: + assert any(row["pattern"] == case["relevant"] for row in rows) + baseline = [row["pattern"] for row in rows if all( + term in "\n".join(row.values()).casefold() for term in case["query"].casefold().split() + )][:3] + candidate = [row["pattern"] for row in lookup.search_patterns( + rows, case["query"], offset=0, limit=3)["patterns"]] + baseline_hits += case["relevant"] in baseline + candidate_hits += case["relevant"] in candidate + # Preserve a deliberately ambiguous case; this is not a broad precision claim. + assert candidate_hits > baseline_hits + + +def test_catalog_preserves_every_table_row_and_prose_appendix(): + text = lookup.CATALOG.read_text() + patterns = lookup.read_patterns(lookup.CATALOG) + by_id = {row["pattern"]: row for row in patterns} + for line in text.splitlines(): + if line.startswith("| `"): + pattern_id = line.split("`", 2)[1] + row = by_id[pattern_id] + # Whole original cells, including code pipes and safety constraints. + assert "| `" + pattern_id + "` | " + " | ".join(row[f] for f in lookup.FIELDS[1:]) + " |" == line + assert "--state merged|all" in by_id["pr_review_default_lifecycle_overreach"]["durable_repair"] + assert "wait_for_ci=false" in by_id["note_review_ignores_a_goal_s_ci_waiting_configuration"]["guidance"] + assert "caches automation rows" in by_id["note_minimal_evidence_packet"]["guidance"] + + +@pytest.mark.parametrize("body", [ + "| `broken` | missing columns |\n", + "| silently discarded row | other cells |\n", + "| `same` | s | e | r | d |\n| `same` | s | e | r | d |\n", +]) +def test_invalid_catalog_is_reported_instead_of_omitted(tmp_path, body): + path = tmp_path / "catalog.md" + path.write_text(body) + with pytest.raises(ValueError): + lookup.read_patterns(path) + + +@pytest.mark.parametrize("args", [ + ["--query", " "], ["--list", "--limit", "0"], ["--list", "--offset", "-1"], + ["--list", "--limit", "21"], ["--id", "does_not_exist"], ["--id", ""], ["--query", "[]"], +]) +def test_cli_rejects_invalid_requests(tmp_path, args): + result = subprocess.run([sys.executable, "-I", str(SCRIPT), *args], cwd=tmp_path, capture_output=True, text=True) + assert result.returncode != 0 + + +def test_real_install_delivers_lookup_and_runs_without_loopx_on_path(tmp_path): + destination = tmp_path / "host-skills" + result = subprocess.run([ + sys.executable, "-m", "loopx.cli", "--format", "json", "workflow-skills", + "--install", "--skills-dir", str(destination), + ], cwd=ROOT, capture_output=True, text=True, check=True) + assert json.loads(result.stdout)["after"]["ready"] + from loopx.doctor import installed_skill_summary + + assert installed_skill_summary((destination,))["loopx-self-repair"]["required_phrases"] + installed = destination / "loopx-self-repair" / "scripts" / "find_pattern.py" + result = subprocess.run([ + sys.executable, "-I", str(installed), "--query", "closeout recovery", + ], cwd=tmp_path, env={"PATH": str(tmp_path)}, capture_output=True, text=True, check=True) + page = json.loads(result.stdout) + assert page["total_matches"] > 0 + assert len(page["patterns"]) <= 5 + result = subprocess.run([ + sys.executable, "-I", str(installed), "--id", "host_closeout_presentation_lookup_gap", + ], cwd=tmp_path, env={"PATH": str(tmp_path)}, capture_output=True, text=True, check=True) + row = json.loads(result.stdout)["pattern"] + assert row == next(r for r in lookup.read_patterns(lookup.CATALOG) if r["pattern"] == row["pattern"]) + # Wheel data must deliver the same resources that source installation copied. + with (ROOT / "pyproject.toml").open("rb") as stream: + data = tomllib.load(stream)["tool"]["setuptools"]["data-files"] + included = {file for files in data.values() for file in files} + for path in SKILL.rglob("*"): + if path.is_file() and "__pycache__" not in path.parts: + assert str(path.relative_to(ROOT)) in included From 713fb8229c973bd240da3bcb27786eccfedfe8f5 Mon Sep 17 00:00:00 2001 From: huangruiteng <14976749+huangruiteng@users.noreply.github.com> Date: Sat, 26 Sep 2026 19:06:15 +0800 Subject: [PATCH 2/2] docs: tie operational optimization to roadmap acceptance Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com> --- AGENTS.md | 10 +++++ docs/development/testing-and-quality.md | 58 +++++++++++++++++++++++++ 2 files changed, 68 insertions(+) diff --git a/AGENTS.md b/AGENTS.md index 37d416abc0..85c25e743c 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -9,6 +9,16 @@ and relevant domain acceptance; do not make every fix wait for every RFC or invent a roadmap id. Check latest `main`, related PRs and canonical Todos so an older task description cannot override corrected direction or duplicate work. +For recurring operational or performance problems, connect the demonstrated +failure to the owning roadmap/RFC acceptance before choosing a repair. Separate +caller overhead, shared typed semantics/transport and provider-specific costs; +prefer the existing common contract where behavior is shared. Follow the +[optimization evidence guide](docs/development/testing-and-quality.md#roadmap-aligned-optimization). +Distinguish an interim mitigation from closing the owning acceptance: faster +lookup, a larger timeout or successful promotion alone does not qualify sustained +operation. Reconcile the existing checkpoint when the evidence changes it; +do not add a parallel roadmap or require unrelated RFC work for a bounded fix. + Carry one compact delivery brief from task to PR: goal/source, current gap, observable result, owning boundary and decisive acceptance evidence. Reuse the existing task/PR fields; keep private Goal state out of public artifacts. diff --git a/docs/development/testing-and-quality.md b/docs/development/testing-and-quality.md index 4c5be58bf2..d947b4e6f7 100644 --- a/docs/development/testing-and-quality.md +++ b/docs/development/testing-and-quality.md @@ -560,6 +560,64 @@ agent what to do and how to request the omitted detail. 完整诊断包保留为显式 drill-down。只有默认路径仍能告诉 agent 下一步做什么、以及 如何请求被省略细节时,字段才能移出默认热路径。 +### Roadmap-Aligned Optimization + +For recurring command, recovery or retrieval costs, use the existing +[overall roadmap](../architecture/rfcs/loopx-overall-roadmap-v0.md) and owning +RFC acceptance to select the repair. Typed semantics and bridge costs belong +to the [TS migration RFC](../architecture/rfcs/typescript-control-plane-migration-v0.md); +backend capacity, retention and cutover qualification belong to the +[shared-authority RFC](../architecture/rfcs/shared-goal-authority-state-provider-v0.md). +Use the existing task/PR evidence and update its owning checkpoint when warranted; +this adds no approval, receipt or requirement to complete unrelated milestones. + +- **Locate the cost before selecting an abstraction.** Separate caller repeats, + output/context expansion, process/bridge/serialization cost, shared semantic + work, and backend IO/verification/contention. A display filter does not reduce + upstream work; a short response or fast isolated query does not establish a + faster recovery loop. Name the real consumer and measure its useful outcome. +- **Share contracts, qualify implementations.** Put common selection, bounded + reads and observation reuse at the existing consumer/typed owner boundary. + Preserve completeness, current authority, receipt replay and lease/CAS checks. + File, SQLite and PostgreSQL need not share cache invalidation, indexing or + history layout. Keep backend-specific algorithms in their providers; do not + duplicate authority in Python or invent a common cache to conceal those costs. + Run the applicable real-backend validation above for every affected backend. +- **Treat migration as a hypothesis, not a cause.** Compare the same operation, + revision, data/history size, runtime configuration and concurrency where + possible. Separate cold/warm reads, alternating stores, writes and recovery + when those paths are affected. Without a controlled before/after comparison, + report measured costs and uncertainty. Successful ownership transfer does not + establish long-running latency, capacity or recovery equivalence. +- **Qualify the intended outcome.** Retrieval relevance needs independently + labeled queries, ambiguous/no-match cases and disclosed language/corpus limits; + a small development set is not production accuracy or task-success evidence. + Runtime optimization needs the original failing workload plus semantic and + scale checks. Follow the budget decisions below rather than hiding regressions + with a larger timeout, smaller fixture or truncated decision evidence. +- **Keep the delivered boundary honest.** State whether the change removes the + owning bottleneck or only mitigates its consumer impact. Reuse an existing + successor for an evidenced remaining gap and identify its acceptance. Do not + count a prompt reduction as backend qualification, or repeated repair PRs as + default-provider readiness; do not create follow-ups for hypothetical work. + +对反复出现的命令、恢复或检索开销,先对应总 roadmap 和所属 RFC 的验收,再选择修复: +TS RFC 管语义 owner 与跨语言成本,shared-authority RFC 管后端容量、保留与切换验证。 +沿用现有任务、PR 证据和验收记录,不增加审批、回执或无关里程碑前置条件。 + +- **先定位成本。** 区分重复调用、输出与上下文展开、进程与序列化、公共语义计算、 + 后端 IO/校验/竞争。输出过滤不减少上游计算;单次查询变快不等于恢复闭环变快。 +- **共用合同,分别验证实现。** 选择、有界读取和观测复用归现有调用方或 typed owner; + 保留完整性、当前权限、回执重放及 lease/CAS。缓存失效、索引和历史布局可以因后端 + 而异,不能在 Python 复制权威,也不强造统一缓存。受影响后端遵循上文真实路径验证。 +- **迁移是待验证的原因。** 尽量控制操作、版本、数据与历史规模、配置、并发;按影响 + 区分冷/热读、多存储交替、写入与恢复。没有受控前后对照就披露不确定性;晋升成功 + 不证明长程延迟、容量和恢复等价。 +- **验证实际目标。** 检索用独立标注、歧义/无命中和语言边界验证;开发小样本不代表 + 生产精度或任务成功率。运行时优化保留原失败负载和语义/规模检查,预算按下节处理。 +- **区分缓解与闭环。** 写明消除了所属瓶颈,还是只减轻调用方影响;剩余真实缺口复用 + 已有后续任务并指明验收。提示词缩短不算后端验证,修复 PR 数量不算默认切换就绪。 + ### Budget Failure Decisions Classify the limit by its owning contract before deciding how to repair a