diff --git a/SKILL.md b/SKILL.md index dd40a10..5b63f91 100644 --- a/SKILL.md +++ b/SKILL.md @@ -95,15 +95,12 @@ description: "Meta-research под вопрос или решение: веб-п ## После завершения — finish-up -0. **Детерминированные артефакты, не руками:** `python scripts/build_sources_csv.py --research-dir /` (единый источник колонок) · `python eval/check_citations.py --research-dir / --json --out //.verify/citations` (без `--out` файл уйдёт в `eval/output/` и gate его не найдёт). -0.2. **Компиляция в вики — `python scripts/wiki_ingest.py --research-dir /`**, детерминированно и **на любой глубине**. Пишет квитанцию `.verify/wiki_ingest.json`, без неё phase-gate красный. Изредка `python scripts/wiki_lint.py`. -0.5. **Числа — два прохода, `--research-dir / --strict`:** `check_number_provenance.py` (число без производителя; одно значение при разных корнях = ложная независимость) · `check_number_arithmetic.py` (пересчёт `derived`, доли к 100, производное число в memo без строки в `numbers.csv`). -1. **Phase-gate — БЛОКЕР:** `python scripts/validate_phases.py --research-dir / --strict`. Красный ⇒ фаза пропущена ⇒ вернись, доделай, перезапусти: не показывать путь, не писать резюме, не рапортовать «готово». -2. Пути markdown-ссылками: сначала `memo.md` (вход потребителя), затем отчёт. -3. Резюме в чат 5–8 строк: 3 ответа + главный контр-аргумент + чего не нашли + итог walkthrough из `application.md`. -4. Предложи 2–3 следующих ресёрча. -5. Есть `memory/` — предложи 1–3 кандидата (тезис + confidence + источники; авторитетный источник как `[reference]`). -6. Есть `anthropic-skills:humanizer-ru` — прогони им финальный отчёт (опционально). +0. **Одна команда, БЛОКЕР:** `python scripts/finish.py --research-dir /` — строит `sources.csv`, гоняет liveness (`--offline` без сети), компилирует прогон в вики (квитанция `.verify/wiki_ingest.json`, на любой глубине), два прохода по числам (`check_number_provenance` / `check_number_arithmetic`, `--strict`) и последним — phase-gate `validate_phases.py --strict`. Шаги не останавливаются на первом отказе, отчёт показывает все. **Красный ⇒ фаза пропущена ⇒ вернись, доделай, перезапусти:** не показывать путь, не писать резюме, не рапортовать «готово». Режим гейт берёт из `mode:` в `plan.md`, без него выводит по артефактам и предупреждает. Изредка `python scripts/wiki_lint.py`. +1. Пути markdown-ссылками: сначала `memo.md` (вход потребителя), затем отчёт. +2. Резюме в чат 5–8 строк: 3 ответа + главный контр-аргумент + чего не нашли + итог walkthrough из `application.md`. +3. Предложи 2–3 следующих ресёрча. +4. Есть `memory/` — предложи 1–3 кандидата (тезис + confidence + источники; авторитетный источник как `[reference]`). +5. Есть `anthropic-skills:humanizer-ru` — прогони им финальный отчёт (опционально). ## Что НЕ делать @@ -112,7 +109,7 @@ description: "Meta-research под вопрос или решение: веб-п - Не редактировать страницы вики руками и не заводить вики внутри проекта: слой один на все ресёрчи. Не гейтить `wiki_ingest` по глубине. - Не запускать medium/deep без единой if-then вилки Decision Spec. - Не завершать синтез финалом «it depends» без разрешённых условий. -- Не оставлять `root:` пустым и не копировать `discovery_path:` между источниками — это 3-е и 4-е условия триангуляции. +- Не оставлять `root:` пустым и не копировать `discovery_path:` между источниками — это 3-е и 4-е условия триангуляции. `claims.csv` без колонок `roots`/`paths`/`dissent`/`as_of` на medium+ гейт роняет: правила триангуляции по ним и работают. - Не разводить fetch-агентов только по подтемам — ещё и по осям поиска (EN-академия / RU + регуляторы / практики / реестры): один шаблон + одна модель + один язык = одна траектория. - Не выбрасывать дубли URL между агентами молча — считать `overlap_rate` в `plan.md` §15: совпадение это замер конформизма, а не подтверждение. - Не передавать во второй раунд находки соседей — только дыры. Не дописывать `state.md` — он перезаписывается. @@ -127,7 +124,7 @@ description: "Meta-research под вопрос или решение: веб-п - Фаза 5.5: не переписывать `sources/NN.md` (архив), не фильтровать по `total` вместо релевантности фрагмента к claim. - Фаза 6.5: не доверять наличию ссылки — проверять entailment по дословной цитате; вердикты писать в `.verify/*.json` и не пересчитывать в rubric/F10; пары брать из `evidence/`, не пересканировать `sources/`. Чинить отчёт, а не ledger. - Для fetch+save и red team — `general-purpose` с явным диапазоном номеров, не `Explore` (read-only, только разведка). Не запускать суб-агентов последовательно — только параллельно в одном сообщении. -- Не сжимать `sources/` в один файл, не выводить результат только в чат. +- Не сжимать `sources/` в один файл, не выводить результат только в чат. Не гонять шаги finish-up поодиночке вместо `finish.py` — пропуск флага (`--out`, `--strict`) и есть пропуск фазы. - Не обходить WebFetch произвольным `bash`/`curl`. Единственный санкционированный fallback — `scripts/fetch_source.py` (Фаза 4.2): он читает robots.txt, санитайзит страницу от prompt injection и проставляет `fetch_tier`. Вывод ручного `curl` в `sources/` не кладётся. - Не принимать `fetch_source.py` за средство против paywall и анти-бот-защиты: `auth-wall` и `antibot` для него — терминальный вердикт. Дальше — fallback-протокол `channels.md` или endpoint из `api_sources/`. - Не рисовать в отчёте число, которого нет в `numbers.csv`, и не подбирать палитру фигур на глаз — она валидируется скриптом. diff --git a/scripts/README.md b/scripts/README.md index 50586f0..a63f3b3 100644 --- a/scripts/README.md +++ b/scripts/README.md @@ -2,6 +2,13 @@ Automation для catalog maintenance — плюс run-time инструменты прогона (`build_sources_csv.py`, `validate_phases.py`). +## finish.py + +One finish-up command: `python scripts/finish.py --research-dir research/ [--mode medium] [--offline]`. +Runs build_sources_csv → check_citations (into `.verify/citations.json`) → wiki_ingest → +check_number_provenance --strict → check_number_arithmetic --strict → validate_phases --strict, +never stopping early; exit 1 if any step failed. `--offline` skips the liveness check. + ## build_sources_csv.py Собирает `sources.csv` прогона из `sources/NN.md` frontmatter — детерминированно, вместо diff --git a/scripts/finish.py b/scripts/finish.py new file mode 100644 index 0000000..c946ac2 --- /dev/null +++ b/scripts/finish.py @@ -0,0 +1,142 @@ +#!/usr/bin/env python3 +""" +One finish-up command for a completed run. Replaces six hand-typed invocations +(each with a flag the gate depends on — `--out` under .verify/, `--strict`, the +wiki receipt) whose omission was where phases got silently skipped. + +Runs, in order, never stopping early so one report shows everything: + 1. build_sources_csv sources/NN.md -> sources.csv + 2. check_citations liveness -> .verify/citations.json (skip: --offline) + 3. wiki_ingest run -> cross-run wiki, receipt .verify/wiki_ingest.json + 4. check_number_provenance --strict + 5. check_number_arithmetic --strict + 6. validate_phases --strict the gate — last, so it sees everything above + +Exit 1 if any step failed. Usage: + python scripts/finish.py --research-dir / [--mode medium] [--offline] +""" +from __future__ import annotations + +import argparse +import subprocess +import sys +from dataclasses import dataclass, field +from pathlib import Path +from typing import Callable + +REPO = Path(__file__).resolve().parents[1] +PY = sys.executable + + +@dataclass +class Step: + name: str + script: Path + extra: Callable[[Path, str | None, Path | None], list[str]] + network: bool = False + + +@dataclass +class StepResult: + name: str + rc: int + output: str + skipped: bool = False + argv: list[str] = field(default_factory=list) + + +def _mode(m: str | None) -> list[str]: + return ["--mode", m] if m else [] + + +STEPS: list[Step] = [ + Step("sources_csv", REPO / "scripts/build_sources_csv.py", lambda d, m, w: []), + Step( + "citations", + REPO / "eval/check_citations.py", + lambda d, m, w: ["--json", "--out", str(d / ".verify" / "citations")], + network=True, + ), + Step( + "wiki_ingest", + REPO / "scripts/wiki_ingest.py", + lambda d, m, w: ["--wiki-root", str(w)] if w else [], + ), + Step("number_provenance", REPO / "scripts/check_number_provenance.py", lambda d, m, w: ["--strict"]), + Step("number_arithmetic", REPO / "scripts/check_number_arithmetic.py", lambda d, m, w: ["--strict"]), + Step("phase_gate", REPO / "scripts/validate_phases.py", lambda d, m, w: ["--strict", *_mode(m)]), +] + +Runner = Callable[[str, list[str]], StepResult] + + +def subprocess_runner(name: str, argv: list[str]) -> StepResult: + proc = subprocess.run(argv, capture_output=True, text=True) + return StepResult(name, proc.returncode, proc.stdout + proc.stderr, argv=argv) + + +def run_finish( + d: Path, + *, + mode: str | None, + offline: bool, + wiki_root: Path | None, + runner: Runner = subprocess_runner, +) -> list[StepResult]: + (d / ".verify").mkdir(exist_ok=True) + results: list[StepResult] = [] + for step in STEPS: + if step.network and offline: + results.append(StepResult(step.name, 0, "skipped (--offline)", skipped=True)) + continue + argv = [PY, str(step.script), "--research-dir", str(d), *step.extra(d, mode, wiki_root)] + res = runner(step.name, argv) + res.argv = res.argv or argv + results.append(res) + return results + + +def exit_code(results: list[StepResult]) -> int: + return 1 if any(r.rc != 0 and not r.skipped for r in results) else 0 + + +def render(results: list[StepResult], verbose: bool) -> str: + lines = [] + for r in results: + status = "skip" if r.skipped else ("ok" if r.rc == 0 else "FAIL") + lines.append(f" {status:<5} {r.name}") + if verbose or (r.rc != 0 and not r.skipped): + out = r.output.strip().splitlines() + # Errors first: a failing gate can bury its one error under a page of warnings. + errs = [t for t in out if "ERROR" in t or "FAIL" in t] + shown = (errs or out)[-12:] if not verbose else out + lines.extend(f" {t}" for t in shown) + code = exit_code(results) + lines.append("") + lines.append( + "Finish: green — the run is complete." if code == 0 + else "Finish: RED — fix the failing step(s) above and re-run. Do not report done." + ) + return "\n".join(lines) + + +def main() -> int: + ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter) + ap.add_argument("--research-dir", required=True, type=Path) + ap.add_argument("--mode", choices=("shallow", "medium", "deep")) + ap.add_argument("--offline", action="store_true", help="skip the network liveness check") + ap.add_argument("--wiki-root", type=Path, default=None, help="override ~/.claude (tests)") + ap.add_argument("--verbose", action="store_true", help="print every step's output") + args = ap.parse_args() + d = args.research_dir + if not d.is_dir(): + print(f"ERROR: not a directory: {d}") + return 2 + print(f"Finish-up: {d}") + results = run_finish(d, mode=args.mode, offline=args.offline, wiki_root=args.wiki_root) + print(render(results, args.verbose)) + return exit_code(results) + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/scripts/validate_phases.py b/scripts/validate_phases.py index 9e32ddd..bb17e03 100644 --- a/scripts/validate_phases.py +++ b/scripts/validate_phases.py @@ -165,9 +165,18 @@ def check_phase( SOURCE_FILE_RE = re.compile(r"^(\d+)_.+\.md$") -# Columns the triangulation / dissent / provenance rules read. Warn (not error) so -# runs made before those rules existed stay validatable. +# Columns the triangulation / dissent / provenance rules read. From medium up their +# absence is an error: the ledger looks complete while the rules that make +# `triangulated` mean anything never fire. Shallow (no triangulation gate) only warns. LEDGER_COLUMNS = ("roots", "paths", "dissent", "as_of") +# Artifacts that only a medium/deep run produces — used to infer the mode when the +# run never wrote a `mode:` frontmatter (2 of 12 real runs had one). +MEDIUM_MARKERS = ( + "state.md", "outline.md", "numbers.csv", "evidence", "refresh_targets.md", + ".verify/authority.json", ".verify/wiki_pairs.json", +) +# SKILL.md depth table: shallow is 5–7 sources, medium starts at 12. +MEDIUM_SOURCE_COUNT = 10 def check_source_perimeter(d: Path, r: Report) -> None: @@ -194,7 +203,37 @@ def check_source_perimeter(d: Path, r: Report) -> None: ) -def check_ledger_columns(d: Path, r: Report) -> None: +def infer_mode(d: Path) -> tuple[str, str]: + """Best-effort depth from what the run left on disk. Never returns deep: on disk + deep and medium are the same set of files, and guessing the stricter one would + fail a run for a phase (3.5) that leaves no artifact.""" + present = [m for m in MEDIUM_MARKERS if artifact_present(d, m)] + if present: + return "medium", f"medium-only artifacts present: {', '.join(present)}" + src = d / "sources" + n = len(list(src.glob("*.md"))) if src.is_dir() else 0 + if n >= MEDIUM_SOURCE_COUNT: + return "medium", f"{n} sources (shallow is 5–7)" + return "shallow", f"no medium artifacts, {n} sources" + + +def resolve_mode(d: Path, explicit: str | None, r: Report) -> str: + """Explicit flag > frontmatter > inference. A run with no `mode:` is validated + anyway — refusing would let the least disciplined runs skip the gate entirely.""" + if explicit: + return explicit + found = detect_mode(d) + if found: + return found + mode, why = infer_mode(d) + r.warn( + f"no 'mode:' frontmatter in report or plan.md — inferred '{mode}' ({why}); " + f"add `mode: {mode}` to plan.md or pass --mode" + ) + return mode + + +def check_ledger_columns(d: Path, mode: str, r: Report) -> None: """A missing ledger column is not a formatting nit: the rule that reads it silently never fires — the 'green check, no behavior' failure mode.""" ledger = d / "claims.csv" @@ -207,11 +246,15 @@ def check_ledger_columns(d: Path, r: Report) -> None: cols = {c.strip() for c in lines[0].split(",")} missing = [c for c in LEDGER_COLUMNS if c not in cols] if missing: - r.warn( + msg = ( f"claims.csv is missing column(s) {', '.join(missing)} — the rules reading " f"them (triangulation by root/path, dissent protection, number provenance) " f"cannot fire; see references/source_scoring.md" ) + if GATE_RANK[mode] >= GATE_RANK["medium"]: + r.err(msg) + else: + r.warn(msg) OUTLINE_ROW_RE = re.compile(r"^\|(?P.+)\|\s*$") @@ -417,7 +460,7 @@ def check_wiki_ingest(d: Path, r: Report) -> None: def validate(d: Path, mode: str, phases: list[dict], r: Report) -> None: self_check(phases, r) check_source_perimeter(d, r) - check_ledger_columns(d, r) + check_ledger_columns(d, mode, r) check_state_window(d, r) check_outline_coverage(d, mode, r) check_constructs(d, r) @@ -441,7 +484,7 @@ def main() -> int: ap.add_argument( "--mode", choices=MODES, - help="run depth; auto-detected from frontmatter if omitted", + help="run depth; frontmatter, else inferred from artifacts, if omitted", ) ap.add_argument("--strict", action="store_true", help="Exit 1 if any error") ap.add_argument("--json", action="store_true") @@ -452,18 +495,12 @@ def main() -> int: print(f"ERROR: not a directory: {d}") return 2 - mode = args.mode or detect_mode(d) - if mode is None: - print( - "ERROR: could not determine run mode — pass --mode {shallow,medium,deep} " - "(no 'mode:' frontmatter found in report or plan.md)" - ) - return 2 + r = Report() + mode = resolve_mode(d, args.mode, r) phases = phases_manifest.load_phases( Path(__file__).resolve().parents[1] / "phases.yaml" ) - r = Report() validate(d, mode, phases, r) if args.json: diff --git a/tests/test_finish.py b/tests/test_finish.py new file mode 100644 index 0000000..0dd3222 --- /dev/null +++ b/tests/test_finish.py @@ -0,0 +1,81 @@ +"""Tests for scripts/finish.py — the single finish-up command.""" + +import sys +from pathlib import Path + +REPO = Path(__file__).resolve().parents[1] +sys.path.insert(0, str(REPO / "scripts")) +sys.path.insert(0, str(REPO / "tests")) + +import finish # noqa: E402 +from test_validate_phases import FULL_SET, make_run # noqa: E402 + + +def _fake_runner(fail: set[str]): + calls: list[str] = [] + + def run(name: str, argv: list[str]) -> finish.StepResult: + calls.append(name) + return finish.StepResult(name, rc=1 if name in fail else 0, output=f"{name} out") + + run.calls = calls # type: ignore[attr-defined] + return run + + +def test_step_order_is_fixed_and_gate_is_last(tmp_path): + d = make_run(tmp_path, mode="medium", phases=FULL_SET) + run = _fake_runner(fail=set()) + results = finish.run_finish(d, mode="medium", offline=False, wiki_root=None, runner=run) + assert [r.name for r in results] == [s.name for s in finish.STEPS] + assert results[-1].name == "phase_gate" + assert finish.exit_code(results) == 0 + + +def test_offline_skips_citations_but_runs_everything_else(tmp_path): + d = make_run(tmp_path, mode="medium", phases=FULL_SET) + run = _fake_runner(fail=set()) + results = finish.run_finish(d, mode="medium", offline=True, wiki_root=None, runner=run) + cit = next(r for r in results if r.name == "citations") + assert cit.skipped + assert "citations" not in run.calls + assert len(run.calls) == len(finish.STEPS) - 1 + + +def test_failed_step_does_not_stop_the_rest_and_exit_is_nonzero(tmp_path): + d = make_run(tmp_path, mode="medium", phases=FULL_SET) + run = _fake_runner(fail={"number_arithmetic"}) + results = finish.run_finish(d, mode="medium", offline=True, wiki_root=None, runner=run) + assert "phase_gate" in run.calls # the gate still ran after a failure + assert finish.exit_code(results) == 1 + + +def test_argv_carries_mode_wiki_root_and_out_dir(tmp_path): + d = make_run(tmp_path, mode="medium", phases=FULL_SET) + seen: dict[str, list[str]] = {} + + def run(name, argv): + seen[name] = [str(a) for a in argv] + return finish.StepResult(name, rc=0, output="") + + finish.run_finish(d, mode="deep", offline=False, wiki_root=tmp_path / "w", runner=run) + assert "--mode" in seen["phase_gate"] and "deep" in seen["phase_gate"] + assert "--wiki-root" in seen["wiki_ingest"] + assert str(d / ".verify" / "citations") in seen["citations"] + assert "--strict" in seen["number_provenance"] and "--strict" in seen["phase_gate"] + + +def test_end_to_end_on_synthetic_run(tmp_path): + """Real subprocesses, no network: sources.csv gets built, wiki receipt written, + number checks pass, gate is green.""" + d = make_run(tmp_path, mode="medium", phases=FULL_SET - {"wiki"}) + (d / "sources" / "01_x.md").write_text( + "---\nurl: https://example.org/x\ntitle: X\ntype: web\n---\n> quote\n", + encoding="utf-8", + ) + results = finish.run_finish( + d, mode="medium", offline=True, wiki_root=tmp_path / "wiki", runner=finish.subprocess_runner + ) + failed = [(r.name, r.output[-600:]) for r in results if r.rc != 0 and not r.skipped] + assert failed == [] + assert (d / "sources.csv").is_file() + assert (d / ".verify" / "wiki_ingest.json").is_file() diff --git a/tests/test_validate_phases.py b/tests/test_validate_phases.py index 6575119..9397560 100644 --- a/tests/test_validate_phases.py +++ b/tests/test_validate_phases.py @@ -281,9 +281,66 @@ def test_ledger_missing_new_columns_warns(tmp_path): "claim_id,status\nc1,triangulated\n", encoding="utf-8" ) r = run_validate(d, "deep") - joined = " ".join(r.warnings) + joined = " ".join(r.errors) + # A ledger without these columns means the triangulation/dissent rules silently + # never fire — the "green check, no behavior" failure. From medium up that blocks. assert "dissent" in joined and "paths" in joined - assert not any("claims.csv" in e for e in r.errors) # warning, never a blocker + assert not any("claims.csv" in w for w in r.warnings) + + +def test_ledger_missing_columns_only_warns_on_shallow(tmp_path): + d = make_run(tmp_path, mode="shallow", phases=SHALLOW_SET) + (d / "claims.csv").write_text( + "claim_id,status\nc1,triangulated\n", encoding="utf-8" + ) + r = run_validate(d, "shallow") + assert any("paths" in w for w in r.warnings) + assert not any("claims.csv" in e for e in r.errors) + + +# --- mode resolution: never refuse to validate --------------------------------- + + +def test_resolve_mode_prefers_explicit(tmp_path): + d = make_run(tmp_path, mode="medium", phases=FULL_SET) + r = vp.Report() + assert vp.resolve_mode(d, "shallow", r) == "shallow" + assert r.warnings == [] + + +def test_resolve_mode_reads_frontmatter(tmp_path): + d = make_run(tmp_path, mode="deep", phases=FULL_SET) + r = vp.Report() + assert vp.resolve_mode(d, None, r) == "deep" + assert r.warnings == [] + + +def test_resolve_mode_infers_medium_from_artifacts_and_warns(tmp_path): + d = make_run(tmp_path, mode="medium", phases=FULL_SET) + for p in (d / "plan.md", *d.glob("2026-*.md")): + p.write_text("no frontmatter\n", encoding="utf-8") + r = vp.Report() + assert vp.resolve_mode(d, None, r) == "medium" + assert any("inferred" in w and "medium" in w for w in r.warnings) + + +def test_resolve_mode_infers_shallow_when_no_medium_artifacts(tmp_path): + d = make_run(tmp_path, mode="shallow", phases=SHALLOW_SET) + for p in (d / "plan.md", *d.glob("2026-*.md")): + p.write_text("no frontmatter\n", encoding="utf-8") + r = vp.Report() + assert vp.resolve_mode(d, None, r) == "shallow" + assert any("inferred" in w for w in r.warnings) + + +def test_resolve_mode_infers_medium_from_source_count(tmp_path): + d = make_run(tmp_path, mode="shallow", phases=SHALLOW_SET) + for p in (d / "plan.md", *d.glob("2026-*.md")): + p.write_text("no frontmatter\n", encoding="utf-8") + for i in range(2, 13): + (d / "sources" / f"{i:02d}_x.md").write_text("---\nurl: http://x\n---\n") + r = vp.Report() + assert vp.resolve_mode(d, None, r) == "medium" def test_real_research_dir_flagged_incomplete_for_deep():