From 477b5b426d1e7f717a405a287803645cf3f4d137 Mon Sep 17 00:00:00 2001 From: Claude Lin & Lay Date: Thu, 20 Aug 2026 23:37:17 +0900 Subject: [PATCH 1/3] spec(evolution): hold the brake 1 fixed-axis prompt fragment and the aggregated comment preamble as copied literals [skills, docs] MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit brake 1 の計器のうち、親が作文面を持たない場所に書いていた2箇所を、保持された literal を写す形へ置き換える。#1779 ## 射程 issue の切り分けに従い **B クラス(親が作文面を持たない場所に書いた)のみ**を実装した。 A クラス(記入済みパートの多義性、7件中6件)は射程外である。wiki `presence-defect-cannot-reach-blank-label-visibility` が 「Axis statement form のパートに要求を足しても空欄ラベルの artifact visibility と 同等の構造性は持てない」と確定させており、A に要求を足す案は同じ壁に当たる。 本 PR では Axis statement form の5部分に要求を追加していない。 B が壁に当たらないのは置換先の構造が実在するからである —— 親が作文せず、 保持された literal を写す形にすればよい。写す形では書き込む欄がそもそも無い。 ## 変更内容 ### 1. 固定軸のプロンプト片(観測 #5 の形) `skills/evolution-impression-literal-detection/SKILL.md` に `## Prompt literal` を新設し、 固定軸が評価者プロンプトへ入る文言を逐語で保持した。親は写すだけで、何も足さない。 Axis statement form は固定軸を5部分の対象外と明記していたが、その除外は散文の宣言 だけで立っており、親がプロンプトの残りを書いている軸の上に欄が開いたままだった (`Scope for this axis = ...` が書き込まれた形)。保持すれば埋める空欄が無い。 literal は Positive / Negative 列挙を内側へ写さず、ファイルを名指して SHA 時点で 取得させる。同一ファイル内の1節下に第二の複製を置けば、それが drift する側になる。 文面は PR 経路にも非 PR 経路(Trigger の他項目)にも当たるよう `the draft under evaluation` で書いてあり、経路ごとの carve-out を開かない。 `skills/evolution-parallel-agent-eval/SKILL.md` 側は Procedure step 3 と Axis statement form の2箇所を、保持先を名指すポインタへ差し替えた。 ### 2. step 4 統合コメントの冒頭定型(同サイクルの統合コメント欠陥の形) `skills/evolution-parallel-agent-eval/SKILL.md` Report shape の Parent's aggregated comment に `Preamble` を追加し、冒頭 literal を保持した。 親は「比率は triage signal であって判定入力ではない」を5軸すべてにかけて書いていたが、 `Ratio is a triage signal` 節自身が固定軸を唯一の数値閾値として carve-out しており、 その一般化は固定軸の位置で偽になる。保持された文言は2節(比率の扱い + carve-out)を 両方 payload として持つため、写す限り carve-out ごと入る。 言語の継ぎ目は明示した。統合コメントの地の文は `Workspace_Language_Contract` で解決されるため、英語以外へ解決する workspace では preamble も解決後の言語で描画する。そのとき2節はどちらも落とさない。 ### 3. 追随 - `skills/evolution-impression-literal-detection/SKILL.md` の description に 「brake 1 の評価者プロンプトを組んでいて固定軸を載せるとき」を追加した。 親が保持先へ到達する経路を description 側でも開けるため。 - `docs/2.-Evolution.md` = 日本語ミラーを同一 PR 内で追随(skill 表 / 軸の書き方 / 比率 / 3成果物の形の4箇所)。 ## brake 1 の順序 **本 PR の brake 1 は変更前の計器で走る。** 新しい計器を自分自身の PR に適用すると 自己検証になるため採らない。新計器は後続 PR から効く。 ## 検査 `python -m unittest discover -s tests` = 73 件 OK。 Co-Authored-By: Claude Opus 5 --- docs/2.-Evolution.md | 8 ++++---- .../SKILL.md | 14 +++++++++++++- skills/evolution-parallel-agent-eval/SKILL.md | 11 ++++++++--- 3 files changed, 25 insertions(+), 8 deletions(-) diff --git a/docs/2.-Evolution.md b/docs/2.-Evolution.md index df584d9..0e71997 100644 --- a/docs/2.-Evolution.md +++ b/docs/2.-Evolution.md @@ -60,7 +60,7 @@ L2 Evolution layer はモデルレイヤーと同じ3種類の責務分類を使 | Persistence Tiering | `skills/evolution-persistence-tiering/SKILL.md` | 情報を memory と docs のどちらに置くか判断するとき | | Evolution Loop | `skills/evolution-loop/SKILL.md` | observe / evaluate / distill / reflect / improve / re-observe のいずれかを実行するとき | | Parallel Agent Eval | `skills/evolution-parallel-agent-eval/SKILL.md` | 自己進化 PR が CI green に到達して merge ゲートが次に来たとき(brake 1 として必須)、Li+ rules / skills / adapter の編集ドラフトが PR フローの外で収束したとき、evolution-loop の observe / evaluate が経験的 verdict を必要とするとき、N=1 self-check が positive に感じたとき、spec 改定提案に直交検証が要るとき、brake 1 の評価者報告・親の統合コメント・著者の裁定のいずれかを書こうとするとき | -| Impression-literal Detection | `skills/evolution-impression-literal-detection/SKILL.md` | 評価者が Li+ source ドラフトの impression-literal 固定軸に答えるとき、ある語句が振る舞い意味を支えているか判定するとき(brake 1 の固定軸、本体は分割で #1598 に切り出し) | +| Impression-literal Detection | `skills/evolution-impression-literal-detection/SKILL.md` | 評価者が Li+ source ドラフトの impression-literal 固定軸に答えるとき、brake 1 の評価者プロンプトを組んでいて固定軸を載せるとき、ある語句が振る舞い意味を支えているか判定するとき(brake 1 の固定軸、本体は分割で #1598 に切り出し。固定軸がプロンプトへ入る文言= Prompt literal もここが保持する) | | Evolution Full Run | `skills/evolution-full-run/SKILL.md` | 明示呼び出し時のみ(「run evolution-full-run」または scheduled-task body が名指し)。consolidate → 完全進化ループ → full refactor を順に統括する thin orchestrator。release は含まない(呼び出し側が付加) | また、`rules/evolution/*.md` のうち always-on で常在させるものは以下である。 @@ -277,10 +277,10 @@ state file は `{workspace_root}/.claude/state/last-cold-start-emit.json`(sha2 - **指摘を裁くのは親ではなく、実装した subagent を再開したもの(著者)である。** brake 1 の評価者は指摘を親へ報告として返し、PR へは何も書かない(#1732)。親は N 体の報告を 1 本の PR コメントへ統合して投稿し、著者を再開する。再開した著者はそのコメントを読み、各指摘を裁き、受け入れた分を修正し、受入 / 却下とその理由を **commit body** に記録して再び CI green で止まる(PR コメントには投稿しない)。親はその裁定を検分し、必要なら修正を名指しして著者をもう一度再開する——往復に回数上限は置かず、収束判断を保持するのは親である(`skills/model-loop-safety`。二者間ループで数える主体を決めないと、双方が相手が数えていると仮定する)。最後にセルフレビューが締める。指摘が親の文脈を通ることがこの検分を買っている——1 件も読んでいない親には、それがどう扱われたかを判定できない。親の分担は統合と投稿までであり、取捨選択はしない(受入 / 却下は著者の権限であり、投稿前に落とされた指摘はその権限を持つ主体へ届かない)。効く点は三つで、却下が往復の文脈ではなく commit body に durable に残り、単巡打ち切りの下で評価者が再反論しないため親の検分だけがその後ろに立つこと、往復は評価者ラウンドではないため single-round cap に触れないこと、および再開が実装文脈を保つため `skills/task-subagent-delegation/SKILL.md` が受容していた「親が報告から文脈を組み直す」対価が支払われなくなること(「実装は常に委譲」の判断軸である規則の単純さは不変で、条件分岐が1つ減る)。GitHub への書き込みは 1 PR あたり 5 回から 2 回(指摘なしなら 1 回)へ減り、評価者 3 体が同一 PR へ数秒以内に並列投稿するバースト形状が消える。エージェントの深さは 1 のままで、評価者を spawn するのは親であり著者は何も spawn しない。正本は `rules/evolution/initiator-autonomy.md` Two-stage brake の Adjudication actor - brake 1 の評価者へは **評価対象を書き換えるな** と委譲プロンプトで明示する。文言は `skills/evolution-parallel-agent-eval/SKILL.md` Constraint に literal として置いてあり、親はそれを写す(毎回作文しない)。上の材料に親のクローン内のパスを含めないのはこの指示と対であり、前者が書き込みの意図を断ち、後者が共有ベースラインという的そのものを外す。literal に carve-out は無い(#1732)。評価者に PR という投稿先が無くなったため、`Do not post to the PR` が報告先を否定形で述べる 1 行として literal の内側に入っており、「コメントは書き換えに当たるか」という問い自体を開かない。評価者のツール権限は絞らない(custom-agent の `tools:` は Claude Code では本文が system prompt を置換するため identity ごと変わり、probe 型の観測対象を壊す)。したがってこれは構造ではなく親が思い出す手続きであり、忘れれば効かないことを受容している。brake 2 は `adapter/claude/agents/l1-gate-eval.md` が `tools: Read`、Codex 版が `sandbox_mode = "read-only"` を持ち、材料も inline 渡しでリポジトリを名指ししないため対象外 - brake 2 は両側とも inline の形を保つ。評価者は `tools: Read` で入力も inline のため投稿先の PR 面を持たず、verdict は従来どおり親へ返る(PASS / DEVIATION はマージゲートそのものであり、マージは親のものだからである)。動くのは裁定の主体だけで、DEVIATION のとき親は修正せず、名指しされた逸脱を PR へ運んで著者を再開する -- brake 1 の**各軸の書き方**も skill 側で固定してあり、親が毎回作文しない。親は評価器具(軸文言)を、その器具の対象をリテラルに読んでいるのと同じ瞬間に書く——そしてリテラル検証は対象には届くが器具には届かない(#1692、3 日で 7 回観測、うち 5 回が「一つの軸名が複数の問いを抱える」形。可視に連結されている場合と、一語の述語へ圧縮されて一問に読める場合とがある)。対象は親が spawn 時に書き起こす**軸ごとの軸**(Trigger の `Additional axes` = `selected per draft nature`)であり、固定軸は範囲外である(固定軸の文言は毎回書かれるものではなく `skills/evolution-impression-literal-detection/SKILL.md` に置かれており、ここが塞ぐ「その場で書く」経路に乗らない)。軸は 5 つの名前つき部分として書く:**Question**(疑問文を 1 つだけ。かつ**判定を生む操作を名乗る**——評価者が材料に対して何をし、その結果のどれが所見なのか。and や読点で繋いだ 2 節は 2 操作= 2 軸である。数えるのは節ではなく操作である。評価を名乗って操作を名乗らない述語——`forced`、`consistent`、`resolves wrongly`——は数を隠して運ぶからである。読みごとに為すことが違い、連言は一語の内側へ圧縮されていて節の禁止が届かない。操作を書くことがこれを解凍する——読みは軸を書いているその場で別々の疑問文へ分かれ、そこへ節の禁止が他の対と同じように発火する。操作を名乗らない疑問文は答えられるのではなく**未記入**である。評価者は自前の読みを補わなければ着手できず、独立に補われた読みこそが不一致の顔をして現れる割れである。ただしここで成果物に載るものは上記の空欄ラベルより狭く、その差は伏せずに書いてある——空欄の部分は何も当てない読み手にも欠けて見えるが、評価的な述語は通常の文面でその部分を埋めており、区別を当てて初めて欠落として読める。得られているのは、その区別を、軸を書いているその場に在って引用可能なテキストへ当てられること(後から手順として思い出すのではなく)である。残余——著者と N 体の評価者がたまたま同じ読みへ収束し、割れが出ないため信号も出ない場合——は post-merge observation 軸で受容する(本ファイルのもう一件の実行保証なき要求と同じ扱い)。`c0020a2` の `forced`、#1763 の「整合」、#1766 の `resolve wrongly` がこの形である。このうち `c0020a2` と #1766 は 5 部分が**記入済み**でこの結果になっており、未記入の検出では防げない——だから塞ぐ先は Question の文言側である)/ **Unit**(1 つの判定が覆う単位=文・段落・ファイル・主張・出現)/ **Scope**(軸が及ぶ面=当該 PR の diff・名指しした 1 ファイル・リポジトリ・リポジトリと wiki。不在の主張はどこを掃いたかを述べる。幅には**広がりのほかに言語という次元**があり、軸は両方を述べる。本リポジトリは正規テキストの多くを二重に持つ——`rules/` と `skills/` が英語、`docs/` が日本語であり、多くは前者のミラーだが、場所によっては `docs/` 自身が正本である(`docs/5.-Notifications.md` は自らを正本と宣言しており、ミラー元となる `rules/notifications/` が存在しない)——ため、リポジトリ全体に及ぶ掃引でも英語パターンだけで走れば日本語側には一度も届いておらず、その不在の主張は見落としではなく**構造的に**掃き残している。正本側の事例のほうが鋭い——掃引が代わりに当たれる英語の対応物がそもそも無いからである(2026-08-12 観測。#1733 の brake 1 が `docs/3.-Task.md` の日本語ミラーの取り残しを 2/1 で検出し、同 PR の `fa87222` で修正済み。残渣は現存しない)。広がりだけではこれを運べない——`リポジトリ` と書かれた scope はその掃引で満たされてしまう。Report shape が復路(所見の無い軸)に掃引のパターンを要求しているのと同じ次元を、往路すなわち軸を書く側で要求する)/ **Verdict terms**(その軸での yes と no の意味を軸自身の語で。「何か落ちたか」を問う軸では所見は yes でありながら draft にとっては negative であり、極性を名乗らない軸は誤った極性を継ぐ)/ **Basis**(軸が対象・基準について述べることはすべて、名指しした SHA で解決するポインタを伴う。基準は `パス` または `パス:行` 付き逐語引用、例示は実在箇所からの引用、件数が効く軸には数値でなく数える本文を渡す。**形式だけでは満たされない**——判定を当てる先、すなわち軸が判定基準とする既存の Li+ 基準や親が依拠する論証は、必要な軸ごとに書き込む。軸は隣に書かれたものを継承せず、基準を名指さない軸は clean ではなく未記入である)。5 部分はすべて評価者のプロンプトへそのまま載る payload であり、通れば消える点検項目ではない。**書かれなかった部分は評価者が読む本文の欠けたラベルとして見える**。機構は成果物上のこの可視性であって、`rules/model/subtractive-structural-beauty.md` が「手続きを構造へ置き換えよ」と要求するときの実行保証には届かない。立っている根拠は #1630 が「落ちた手順を、憶えておくべき 1 行として別の面へ置き直す」案を却下したのと同じものである。正本は `skills/evolution-parallel-agent-eval/SKILL.md` Axis statement form +- brake 1 の**各軸の書き方**も skill 側で固定してあり、親が毎回作文しない。親は評価器具(軸文言)を、その器具の対象をリテラルに読んでいるのと同じ瞬間に書く——そしてリテラル検証は対象には届くが器具には届かない(#1692、3 日で 7 回観測、うち 5 回が「一つの軸名が複数の問いを抱える」形。可視に連結されている場合と、一語の述語へ圧縮されて一問に読める場合とがある)。対象は親が spawn 時に書き起こす**軸ごとの軸**(Trigger の `Additional axes` = `selected per draft nature`)であり、固定軸は範囲外である(固定軸の文言は毎回書かれるものではなく `skills/evolution-impression-literal-detection/SKILL.md` Prompt literal に**写すだけの literal** として保持されており、ここが塞ぐ「その場で書く」経路に乗らない。除外が保つ根拠はこの保持である——保持されていれば以下の 5 部分を書き込む空欄が固定軸に無く、散文だけの除外宣言では、親がプロンプトの残りを書いている軸の上に欄が開いたままになる)。軸は 5 つの名前つき部分として書く:**Question**(疑問文を 1 つだけ。かつ**判定を生む操作を名乗る**——評価者が材料に対して何をし、その結果のどれが所見なのか。and や読点で繋いだ 2 節は 2 操作= 2 軸である。数えるのは節ではなく操作である。評価を名乗って操作を名乗らない述語——`forced`、`consistent`、`resolves wrongly`——は数を隠して運ぶからである。読みごとに為すことが違い、連言は一語の内側へ圧縮されていて節の禁止が届かない。操作を書くことがこれを解凍する——読みは軸を書いているその場で別々の疑問文へ分かれ、そこへ節の禁止が他の対と同じように発火する。操作を名乗らない疑問文は答えられるのではなく**未記入**である。評価者は自前の読みを補わなければ着手できず、独立に補われた読みこそが不一致の顔をして現れる割れである。ただしここで成果物に載るものは上記の空欄ラベルより狭く、その差は伏せずに書いてある——空欄の部分は何も当てない読み手にも欠けて見えるが、評価的な述語は通常の文面でその部分を埋めており、区別を当てて初めて欠落として読める。得られているのは、その区別を、軸を書いているその場に在って引用可能なテキストへ当てられること(後から手順として思い出すのではなく)である。残余——著者と N 体の評価者がたまたま同じ読みへ収束し、割れが出ないため信号も出ない場合——は post-merge observation 軸で受容する(本ファイルのもう一件の実行保証なき要求と同じ扱い)。`c0020a2` の `forced`、#1763 の「整合」、#1766 の `resolve wrongly` がこの形である。このうち `c0020a2` と #1766 は 5 部分が**記入済み**でこの結果になっており、未記入の検出では防げない——だから塞ぐ先は Question の文言側である)/ **Unit**(1 つの判定が覆う単位=文・段落・ファイル・主張・出現)/ **Scope**(軸が及ぶ面=当該 PR の diff・名指しした 1 ファイル・リポジトリ・リポジトリと wiki。不在の主張はどこを掃いたかを述べる。幅には**広がりのほかに言語という次元**があり、軸は両方を述べる。本リポジトリは正規テキストの多くを二重に持つ——`rules/` と `skills/` が英語、`docs/` が日本語であり、多くは前者のミラーだが、場所によっては `docs/` 自身が正本である(`docs/5.-Notifications.md` は自らを正本と宣言しており、ミラー元となる `rules/notifications/` が存在しない)——ため、リポジトリ全体に及ぶ掃引でも英語パターンだけで走れば日本語側には一度も届いておらず、その不在の主張は見落としではなく**構造的に**掃き残している。正本側の事例のほうが鋭い——掃引が代わりに当たれる英語の対応物がそもそも無いからである(2026-08-12 観測。#1733 の brake 1 が `docs/3.-Task.md` の日本語ミラーの取り残しを 2/1 で検出し、同 PR の `fa87222` で修正済み。残渣は現存しない)。広がりだけではこれを運べない——`リポジトリ` と書かれた scope はその掃引で満たされてしまう。Report shape が復路(所見の無い軸)に掃引のパターンを要求しているのと同じ次元を、往路すなわち軸を書く側で要求する)/ **Verdict terms**(その軸での yes と no の意味を軸自身の語で。「何か落ちたか」を問う軸では所見は yes でありながら draft にとっては negative であり、極性を名乗らない軸は誤った極性を継ぐ)/ **Basis**(軸が対象・基準について述べることはすべて、名指しした SHA で解決するポインタを伴う。基準は `パス` または `パス:行` 付き逐語引用、例示は実在箇所からの引用、件数が効く軸には数値でなく数える本文を渡す。**形式だけでは満たされない**——判定を当てる先、すなわち軸が判定基準とする既存の Li+ 基準や親が依拠する論証は、必要な軸ごとに書き込む。軸は隣に書かれたものを継承せず、基準を名指さない軸は clean ではなく未記入である)。5 部分はすべて評価者のプロンプトへそのまま載る payload であり、通れば消える点検項目ではない。**書かれなかった部分は評価者が読む本文の欠けたラベルとして見える**。機構は成果物上のこの可視性であって、`rules/model/subtractive-structural-beauty.md` が「手続きを構造へ置き換えよ」と要求するときの実行保証には届かない。立っている根拠は #1630 が「落ちた手順を、憶えておくべき 1 行として別の面へ置き直す」案を却下したのと同じものである。正本は `skills/evolution-parallel-agent-eval/SKILL.md` Axis statement form - brake 1 で評価者の判定が割れたときは、①**同じ問いに答えたか**(same-question check)→ ②同じ問いなら**なぜ割れたか**(判定基準の曖昧さに遡るか、明確な基準を適用する際のゆらぎに遡るか)の順で、各一度だけ問う。①が No なら課題は基準側でなく軸の書き方(一つの軸名が複数の問いを含んでいた)にあり、その No は Axis statement form のどの部分が保たなかったかを名指す先へ着地する(従来のように「次回の申し送り」へは行かない)。事前の Question と事後の same-question check は代替関係ではない——事後側は「部分は埋まっていたが緩く埋まっていた」軸を捕まえ、本 form 以前に書かれた軸にも効き続ける。**「同じ問いに答えており、基準はこれで足りていた」は正当な帰結**であり、盲点の発見を要求しない——`rules/model/trigger-check-gate.md` と同じ形(一つでも No なら止まって取得・検証して**進む**)である。無い所に基準の隙間を書き起こすのがここでの失敗モード -- 比率(3/3 / 2/3 / 1/3)は判定入力ではなく、検証の労力をどこに先に置くかの安い事前分布である。報告してよいが判定の根拠にしてはならず、所見の可否はリテラルをソースに突き合わせて決まる(1/3 が採用され 3/3 が却下されうる)。数値閾値を持つのは固定軸(`skills/evolution-impression-literal-detection/SKILL.md` Aggregation)だけで、その数値は本節に写さない。正本は `skills/evolution-parallel-agent-eval/SKILL.md` Design Dimensions -- brake 1 が生む**3 つの成果物の形**は skill 側で固定してあり、各主体が毎回作文しない。3 つとは評価者の報告(step 3)/親の統合コメント(step 4)/著者の裁定(step 7)であり、いずれも同じ継ぎ目で非対称である——所見のある側は full、誰も争っていない側は 1 行。評価者の報告は、所見のある軸=逐語引用+`パス:行`(名指しした SHA 時点)+なぜ欠陥か/所見の無い軸=1 行(軸名+その軸の語での判定+根拠の `パス:行`。リポジトリ全体を掃く軸は指せる 1 行が無いため、再実行可能な形の掃き出し(パターンと掃いたパス)とヒット数で代える)。親の統合コメントは、届いた full の形をそのまま保つ(ここで縮めると取り直しが著者へ回り、full を要求している理由そのものを潰す。統合が稼ぐのは評価者間の重複除去であって短縮ではなく、重複は「N 体中何体が挙げたか」を持つ 1 件へ畳む)。clean な軸も 1 行のまま載せる(著者は step 6 で軸横断に集約するため、所見だけのコメントでは分母が見えない)。統合コメントで禁止するのは受入・却下・順位づけ・推奨——それは step 7 の著者のものであり、重みづけの権限を持つ主体より先に着いてしまう。著者の裁定は commit body に置き、却下=full(単巡打ち切りの下で評価者が再反論しないため、親が step 8 で読むこれだけがその後ろに立つ)/受入=1 行(何が変わったかは同じ commit の diff に既に外部化されている)。1 件も受け入れなかった場合は commit が無く commit body も無いため、裁定は著者の停止条件報告で親へ渡り、親が step 9 のセルフレビューへ運ぶ——ここが「セルフレビューだけが唯一保証された外部化先である」理由であり、他の面はすべて書く対象がある場合にのみ存在する。全側面で禁止するのは、親がプロンプトで渡した基準・閾値・軸文言の復唱である。所見側の逐語引用は削減対象ではない——著者は引用に当てて裁定するため、削るとソースの取り直しが指摘ごとに発生し、リテラルに当てずに断定する形(`rules/model/trigger-check-gate.md` Literal check / #1673 cluster)へ戻る。所見の無い軸で「見て clean と言った」を支えるのはポインタの**解決可能性**であり、名指しした SHA で開けること、開いて判定と合わないなら 1 回の参照で露見することがその働きである。掃引軸には開く先が無いため、その働きは再実行可能な形で述べた掃き出しが担う(不在の主張は開いてではなく再実行で検査され、再現しないヒット数は解決しないポインタと同じ形で落ちる)。残余(誰も開かない clean 軸のポインタ、および判定に合うよう事後選定されたポインタ)は post-merge observation 軸で受容する。**言語**の規定は評価者側から消えた(#1732)。評価者はもう契約の及ぶ面へ書かず親の文脈へ返すだけであり、契約の及ぶ面(統合コメント)を書くのは解決後の値を自分の session に持つ親だからである——親が自分で書くものの言語は既にそう解決されており、ここに残す規定は無い。残るのは解決では片づかない継ぎ目 1 つ、**逐語引用は対象外**である。根拠は都合ではなく所見が乗っている性質そのもので、著者は引用に当てて裁定するため訳された literal は literal ではない。Li+ source は英語(`rules/model/liplus-coding-rule.md` Source Language)であり、他の言語へ解決する workspace では統合コメントは構造的に混在する(引用は原文の言語のまま、地の文は解決後の言語)。明示しなければ引用まで訳されるか地の文まで引用の言語へ倒れるかのどちらかになる。brake 2 と PR 面を持たない Trigger 項目は、いずれも自前の理由で本節の外である。正本は `skills/evolution-parallel-agent-eval/SKILL.md` Report shape +- 比率(3/3 / 2/3 / 1/3)は判定入力ではなく、検証の労力をどこに先に置くかの安い事前分布である。報告してよいが判定の根拠にしてはならず、所見の可否はリテラルをソースに突き合わせて決まる(1/3 が採用され 3/3 が却下されうる)。数値閾値を持つのは固定軸(`skills/evolution-impression-literal-detection/SKILL.md` Aggregation)だけで、その数値は本節に写さない。この両半——比率が triage signal である旨と固定軸の carve-out——は統合コメント冒頭の 1 つの literal として skill 側に保持されており、親はそれを写す(この節を自分の言葉で言い直さない。言い直しが carve-out を跨いで一般化を生む)。正本は `skills/evolution-parallel-agent-eval/SKILL.md` Design Dimensions +- brake 1 が生む**3 つの成果物の形**は skill 側で固定してあり、各主体が毎回作文しない。3 つとは評価者の報告(step 3)/親の統合コメント(step 4)/著者の裁定(step 7)であり、いずれも同じ継ぎ目で非対称である——所見のある側は full、誰も争っていない側は 1 行。評価者の報告は、所見のある軸=逐語引用+`パス:行`(名指しした SHA 時点)+なぜ欠陥か/所見の無い軸=1 行(軸名+その軸の語での判定+根拠の `パス:行`。リポジトリ全体を掃く軸は指せる 1 行が無いため、再実行可能な形の掃き出し(パターンと掃いたパス)とヒット数で代える)。親の統合コメントは、**冒頭に 1 つの保持された literal**(比率は triage signal であって判定入力ではない旨と、絶対閾値を持つ固定軸だけがその外に立つ旨の 2 節)を写して開き、そのあと届いた full の形をそのまま保つ(冒頭の 2 節は両方 payload である。前者だけでは固定軸の位置で偽になり、後者だけでは軸ごとの比率が票に読める。親が自分の言葉で言い直すことがまさに carve-out を跨いだ一般化を生むため、文言を skill 側に保持する。英語以外へ解決する workspace では地の文と同じく解決後の言語で描画し、そのとき 2 節はどちらも落とさない)(ここで縮めると取り直しが著者へ回り、full を要求している理由そのものを潰す。統合が稼ぐのは評価者間の重複除去であって短縮ではなく、重複は「N 体中何体が挙げたか」を持つ 1 件へ畳む)。clean な軸も 1 行のまま載せる(著者は step 6 で軸横断に集約するため、所見だけのコメントでは分母が見えない)。統合コメントで禁止するのは受入・却下・順位づけ・推奨——それは step 7 の著者のものであり、重みづけの権限を持つ主体より先に着いてしまう。著者の裁定は commit body に置き、却下=full(単巡打ち切りの下で評価者が再反論しないため、親が step 8 で読むこれだけがその後ろに立つ)/受入=1 行(何が変わったかは同じ commit の diff に既に外部化されている)。1 件も受け入れなかった場合は commit が無く commit body も無いため、裁定は著者の停止条件報告で親へ渡り、親が step 9 のセルフレビューへ運ぶ——ここが「セルフレビューだけが唯一保証された外部化先である」理由であり、他の面はすべて書く対象がある場合にのみ存在する。全側面で禁止するのは、親がプロンプトで渡した基準・閾値・軸文言の復唱である。所見側の逐語引用は削減対象ではない——著者は引用に当てて裁定するため、削るとソースの取り直しが指摘ごとに発生し、リテラルに当てずに断定する形(`rules/model/trigger-check-gate.md` Literal check / #1673 cluster)へ戻る。所見の無い軸で「見て clean と言った」を支えるのはポインタの**解決可能性**であり、名指しした SHA で開けること、開いて判定と合わないなら 1 回の参照で露見することがその働きである。掃引軸には開く先が無いため、その働きは再実行可能な形で述べた掃き出しが担う(不在の主張は開いてではなく再実行で検査され、再現しないヒット数は解決しないポインタと同じ形で落ちる)。残余(誰も開かない clean 軸のポインタ、および判定に合うよう事後選定されたポインタ)は post-merge observation 軸で受容する。**言語**の規定は評価者側から消えた(#1732)。評価者はもう契約の及ぶ面へ書かず親の文脈へ返すだけであり、契約の及ぶ面(統合コメント)を書くのは解決後の値を自分の session に持つ親だからである——親が自分で書くものの言語は既にそう解決されており、ここに残す規定は無い。残るのは解決では片づかない継ぎ目 1 つ、**逐語引用は対象外**である。根拠は都合ではなく所見が乗っている性質そのもので、著者は引用に当てて裁定するため訳された literal は literal ではない。Li+ source は英語(`rules/model/liplus-coding-rule.md` Source Language)であり、他の言語へ解決する workspace では統合コメントは構造的に混在する(引用は原文の言語のまま、地の文は解決後の言語)。明示しなければ引用まで訳されるか地の文まで引用の言語へ倒れるかのどちらかになる。brake 2 と PR 面を持たない Trigger 項目は、いずれも自前の理由で本節の外である。正本は `skills/evolution-parallel-agent-eval/SKILL.md` Report shape - 人間ゲートはリリース / 不可逆系の軸と execution-mode の minor/major レビューに残る(人間 = 最終審判者は別軸) 段階責務: diff --git a/skills/evolution-impression-literal-detection/SKILL.md b/skills/evolution-impression-literal-detection/SKILL.md index 5dbc129..4eefe05 100644 --- a/skills/evolution-impression-literal-detection/SKILL.md +++ b/skills/evolution-impression-literal-detection/SKILL.md @@ -1,6 +1,6 @@ --- name: evolution-impression-literal-detection -description: Invoke when an evaluator is answering the fixed impression-literal axis on a Li+ source draft / a phrase in a rules or skills or adapter draft needs testing for whether it load-bears on behavior semantic / brake 1 findings on rhetorical drift are being adjudicated before merge / a Li+ source sentence is about to be kept or removed on a judgment that rests on impression rather than behavior. Provides the removal test and the refine thresholds. +description: Invoke when an evaluator is answering the fixed impression-literal axis on a Li+ source draft / a brake 1 evaluator prompt is being composed and the fixed axis has to enter it / a phrase in a rules or skills or adapter draft needs testing for whether it load-bears on behavior semantic / brake 1 findings on rhetorical drift are being adjudicated before merge / a Li+ source sentence is about to be kept or removed on a judgment that rests on impression rather than behavior. Provides the prompt literal that axis is copied from, the removal test, and the refine thresholds. layer: L2-evolution --- @@ -39,6 +39,18 @@ False-negative backstop: all-3-miss cases route to post-merge observation per `r Rationale: behavior-vs-impression boundary is context-dependent, so N=1 flag carries false-positive risk. The threshold prevents over-trimming load-bearing L1 spec phrasing. + + +## Prompt literal + +The form this axis enters an evaluator prompt in — on the brake 1 path and on the other Trigger entries alike (`skills/evolution-parallel-agent-eval/SKILL.md` Trigger). Copy it verbatim; do not re-compose it per spawn, and add nothing to it. That file's Axis statement form excludes this axis from the five parts it fixes for the axes selected per draft, and holding the wording here is what lets the exclusion hold: there is no blank for one of those parts to be filled in on this axis. The material the prompt names alongside it stays the parent's, fixed at that file's Procedure step 3. + +> **Fixed axis — impression-literal detection.** Take each phrase the draft under evaluation adds or modifies in Li+ source (`rules/*`, `skills/*`, `adapter/*`) and remove it: does the rule's behavior semantic change? Unchanged means the phrase is impression literal, and that is a finding on this axis; changed means it load-bears and is clean. Judge the added and modified lines, not the surrounding unchanged text. The Positive and Negative lists bounding this axis, including the categories that are protected and must not be flagged, are at `skills/evolution-impression-literal-detection/SKILL.md`; retrieve that file at the named SHA and apply it as written. Report each flagged phrase as a verbatim quote with its `path:line`. Do not aggregate and do not apply a threshold: this axis's thresholds are absolute and are applied to the N reports after yours arrives. + +The literal names the file instead of carrying the Positive and Negative lists inside itself: a second copy of them one section below the first is the copy that drifts, and the evaluator already holds the retrieval command that resolves a repository path at the named SHA (`skills/evolution-parallel-agent-eval/SKILL.md` Procedure step 3). + + + ## Detection signs diff --git a/skills/evolution-parallel-agent-eval/SKILL.md b/skills/evolution-parallel-agent-eval/SKILL.md index 5cd810b..0727b1a 100644 --- a/skills/evolution-parallel-agent-eval/SKILL.md +++ b/skills/evolution-parallel-agent-eval/SKILL.md @@ -74,7 +74,7 @@ Answers go where the author's adjudication goes (Report shape, Author's adjudica The ratio of evaluators reporting a finding (3/3, 2/3, 1/3) is a cheap prior on where to spend verification effort first. It is not a judgment input. The verdict on a finding comes from checking its literal against the source: a 1/3 finding that holds up is adopted, and a 3/3 finding that does not is dropped. The ratio is kept rather than discarded because verifying every finding at equal cost is not practical and the parent's own literal check is not infallible either. The parent states it on each finding when it consolidates them at Procedure step 4, and it stays a statement there — consolidation is not the place a finding is weighed. The field the ratio occupies in the self-review is fixed at Procedure step 9; the shape of the reports it is counted from is fixed at Report shape. -The one place this skill fixes a count as a threshold is the fixed axis (`skills/evolution-impression-literal-detection/SKILL.md` Aggregation), whose numbers come from the asymmetry of that judgment - over-trimming load-bearing spec phrasing is the costly error - not from counting votes. The per-judgment aggregation rule above is likewise selected from asymmetry; majority is not among its options. +The one place this skill fixes a count as a threshold is the fixed axis (`skills/evolution-impression-literal-detection/SKILL.md` Aggregation), whose numbers come from the asymmetry of that judgment - over-trimming load-bearing spec phrasing is the costly error - not from counting votes. The per-judgment aggregation rule above is likewise selected from asymmetry; majority is not among its options. Both halves — the triage signal and this carve-out — reach the author as one held literal (Report shape, Parent's aggregated comment); the parent copies it rather than restating this section in its own words at the head of that comment. @@ -95,7 +95,7 @@ The one place this skill fixes a count as a threshold is the fixed axis (`skills On the brake 1 path the material named in the prompt is the PR URL, the pushed commit SHA, and the green CI run URL — never a path inside the parent's clone, which would put every evaluator on a baseline any one of them can move. The reason that set is fixed is canonical in `rules/evolution/initiator-autonomy.md` Two-stage brake. The rule governs what the prompt *names*, so step 2's operational copy is unaffected: it reaches the evaluator as auto-injected context, not as a named path. - Each per-draft axis goes into the prompt in the form fixed at Axis statement form; the fixed axis is outside that form's scope. Five more things go in alongside the axes and the material: + Each per-draft axis goes into the prompt in the form fixed at Axis statement form; the fixed axis is outside that form's scope and enters as the held literal at `skills/evolution-impression-literal-detection/SKILL.md` Prompt literal, copied verbatim with nothing added to it. Five more things go in alongside the axes and the material: - the no-write literal verbatim (see Constraint: Evaluator does not modify the evaluation target) - the retrieval commands, so the evaluator does not assume a clone is needed: `gh pr diff --repo /` returns the diff and `gh api repos///contents/?ref=` returns any file body at that SHA - the allowance that an axis needing a repository-wide sweep clones into the evaluator's own working directory, which is off the shared surface. GitHub code search is unreliable on this repository (total hits 0), so the sweep has nowhere else to go @@ -118,7 +118,7 @@ The one place this skill fixes a count as a threshold is the fixed axis (`skills ## Axis statement form -Fixes the form each per-draft axis is written in — the set Trigger, Axis selection names with `Additional axes are selected per draft nature`, which the parent composes at spawn time. It applies at Procedure step 3, where those axes enter the prompt. The fixed axis is outside it: that axis's wording is not composed per run but held at `skills/evolution-impression-literal-detection/SKILL.md`, so the per-run authoring this section addresses does not reach it, and it enters the prompt as its own spec words it. +Fixes the form each per-draft axis is written in — the set Trigger, Axis selection names with `Additional axes are selected per draft nature`, which the parent composes at spawn time. It applies at Procedure step 3, where those axes enter the prompt. The fixed axis is outside it: that axis's wording is not composed per run but held as a copy-verbatim literal at `skills/evolution-impression-literal-detection/SKILL.md` Prompt literal, so the per-run authoring this section addresses does not reach it. The exclusion rests on the literal being held there. Held, the axis carries no blank for one of the parts below to be filled in on it; stated in prose alone, the exclusion leaves the field open on an axis the parent is writing the rest of the prompt around. The gap it closes: the parent authors the instrument in the same moment it is reading the instrument's target literally, and the literal check that reaches the target does not reach the instrument. The recurring form is one axis name carrying more than one question — joined visibly, or compressed into a single predicate that reads as one. @@ -152,6 +152,11 @@ Scope = the brake 1 path. brake 2 is out: its evaluator has its own prompt file ### Parent's aggregated comment +- **Preamble**: the comment opens with one held literal, copied rather than composed: + + > The 3/3, 2/3 and 1/3 counts below are a triage signal on where to check first, not a judgment input. Adjudicate each finding by checking its literal against the source at the named SHA, and adopt or drop it on that. One axis stands outside this line: the fixed impression-literal axis carries absolute thresholds in its own spec (`skills/evolution-impression-literal-detection/SKILL.md` Aggregation), and its counts are applied there as written. + + Both clauses are payload. The first without the second is false where the fixed axis lands; the second without the first leaves the per-draft axes' counts reading as votes (Design Dimensions, Ratio is a triage signal). The parent restating this in its own words is what generalizes the first over the carve-out the second names, which is why the wording is held here. Where the workspace resolves to a language other than English the preamble is rendered in it like the rest of the parent's prose (Language below), and rendering carries both clauses. - **A finding** keeps the full form it arrived in: the verbatim quote, its `path:line`, and why it is a defect. Compressing it here would push the re-fetch onto the author, which is the cost the full-length rule above exists to remove; consolidation earns its place by removing duplication across evaluators, not by shortening a finding. Duplicates collapse into one entry stating how many of the N raised it. - **Clean axes** carry their one line through as well. The author aggregates cross-axis at Procedure step 6, and a comment listing only findings leaves it aggregating over a denominator it cannot see. - **Prohibited**: an accept, a reject, a ranking, or a recommendation. Those are the author's at Procedure step 7, and prose that leans on a finding here arrives ahead of the actor entitled to weigh it. From 4944acce254794c5c1c22d460ed0db43ad28c4d3 Mon Sep 17 00:00:00 2001 From: Claude Lin & Lay Date: Fri, 21 Aug 2026 01:46:58 +0900 Subject: [PATCH 2/3] fix(evolution): adjudicate brake 1 findings on the two held literals MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit brake 1(評価者3体、単一ラウンド、変更前の計器)の指摘に対する裁定。#1779 ## 受入(6件、いずれも 1 行) - 所見 1-a(2/3)— `Prompt literal` の Basis が「the named SHA」で解決していた。SHA が 名指されるのは brake 1 経路だけで、同節自身が名乗る他 4 経路には指示対象が無い。 `the revision this prompt names` へ変更し、material の所在も step 3(brake 1)と step 2(それ以外)の両方を名指す形にした。 - 所見 1-b(1/3)— literal が flag 側の報告形しか述べておらず、clean 側は Report shape のどちらの分岐にも合わなかった(単一の `パス:行` でも リポジトリ全体の掃引でもない)。掃引分岐と同型の 1 行形(掃いた draft と 追加・変更行数)を literal に追加した。 - 所見 2-a(1/3)— Preamble の裁定節が無限定で、固定軸の Aggregation (1/3 は計数を理由として commit body に記録する)と衝突していた。本裁定自身が その衝突の実例である。carve-out を「両方の文の外に立つ」形へ広げた。 - 所見 2-b(1/3)— `3/3, 2/3 and 1/3` は N=3 を固定値として運んでいた。N は下限であり、 超えれば literal がそこに現れない計数を名指す。`how many of the N evaluators raised it` へ一般化した。 - 所見 2-c(1/3)— render 節が、保持が拒んでいる composition を復活させていた。 削除し、原文の言語のまま写す形にした。Language bullet の「解決では片づかない継ぎ目」 にも preamble を加えた(従来「1 つ」と数えていた)。 - 所見 4-a(1/3)— step 4 の `the surrounding prose is the parent's` が、 本 PR が「親のものではない」と定めた冒頭部まで及んでいた。冒頭 literal を除外した。 `docs/2.-Evolution.md` = 2-a / 2-b / 2-c と Language 継ぎ目の 4 点をミラーへ追随。 ## 却下(1件、固定軸 / full) **固定軸 1 of N=3 — refine しない。** flag されたリテラル、`skills/evolution-parallel-agent-eval/SKILL.md:98`: > copied verbatim with nothing added to it 理由 = **閾値未満(1/3)**である。`skills/evolution-impression-literal-detection/SKILL.md` Aggregation は 2 以上で即 refine、1 では auto-refine せず flag されたリテラルを commit body に記録し、閾値未満の計数を refine しなかった理由として名指すことを要求する。 この軸の閾値は絶対であり、比率=triage signal の一般則の外にある。したがって本件は リテラル検証で採否を決めず、計数を理由として却下する(所見 2-a が本 PR で閉じた 衝突の、まさにこの手順である)。閾値は過剰トリムを止めるために在り、 `copied verbatim with nothing added to it` は Prompt literal を開かない読み手に対して 使用時点の要求を運んでいる。 ### divergence handling(1-of-3 の割れ) - **同一問検査 = yes。** 3 体とも「追加・変更行のどの語句が impression literal か」に 答えており、clean 判定は `:98` を含む全行を見た上での「flag 無し」である。 軸の書き方の問題ではない。 - **なぜ割れたか = 明確な基準の適用のばらつき。** Negative list の `Explanatory rationale that prevents a known misinterpretation` は 「behavior semantic と revision stability の両方が不変のときのみ impression literal」と 二肢で書かれている。flag 側は第一肢(指示はポインタ先で届く)で判断し、 clean 側 2 体は第二肢(revision stability)で保護と判断した。両肢とも明文であり、 判定基準の曖昧さには遡らない。基準側の修正は行わない。 Co-Authored-By: Claude Opus 5 --- docs/2.-Evolution.md | 2 +- skills/evolution-impression-literal-detection/SKILL.md | 4 ++-- skills/evolution-parallel-agent-eval/SKILL.md | 8 ++++---- 3 files changed, 7 insertions(+), 7 deletions(-) diff --git a/docs/2.-Evolution.md b/docs/2.-Evolution.md index 0e71997..5845c20 100644 --- a/docs/2.-Evolution.md +++ b/docs/2.-Evolution.md @@ -280,7 +280,7 @@ state file は `{workspace_root}/.claude/state/last-cold-start-emit.json`(sha2 - brake 1 の**各軸の書き方**も skill 側で固定してあり、親が毎回作文しない。親は評価器具(軸文言)を、その器具の対象をリテラルに読んでいるのと同じ瞬間に書く——そしてリテラル検証は対象には届くが器具には届かない(#1692、3 日で 7 回観測、うち 5 回が「一つの軸名が複数の問いを抱える」形。可視に連結されている場合と、一語の述語へ圧縮されて一問に読める場合とがある)。対象は親が spawn 時に書き起こす**軸ごとの軸**(Trigger の `Additional axes` = `selected per draft nature`)であり、固定軸は範囲外である(固定軸の文言は毎回書かれるものではなく `skills/evolution-impression-literal-detection/SKILL.md` Prompt literal に**写すだけの literal** として保持されており、ここが塞ぐ「その場で書く」経路に乗らない。除外が保つ根拠はこの保持である——保持されていれば以下の 5 部分を書き込む空欄が固定軸に無く、散文だけの除外宣言では、親がプロンプトの残りを書いている軸の上に欄が開いたままになる)。軸は 5 つの名前つき部分として書く:**Question**(疑問文を 1 つだけ。かつ**判定を生む操作を名乗る**——評価者が材料に対して何をし、その結果のどれが所見なのか。and や読点で繋いだ 2 節は 2 操作= 2 軸である。数えるのは節ではなく操作である。評価を名乗って操作を名乗らない述語——`forced`、`consistent`、`resolves wrongly`——は数を隠して運ぶからである。読みごとに為すことが違い、連言は一語の内側へ圧縮されていて節の禁止が届かない。操作を書くことがこれを解凍する——読みは軸を書いているその場で別々の疑問文へ分かれ、そこへ節の禁止が他の対と同じように発火する。操作を名乗らない疑問文は答えられるのではなく**未記入**である。評価者は自前の読みを補わなければ着手できず、独立に補われた読みこそが不一致の顔をして現れる割れである。ただしここで成果物に載るものは上記の空欄ラベルより狭く、その差は伏せずに書いてある——空欄の部分は何も当てない読み手にも欠けて見えるが、評価的な述語は通常の文面でその部分を埋めており、区別を当てて初めて欠落として読める。得られているのは、その区別を、軸を書いているその場に在って引用可能なテキストへ当てられること(後から手順として思い出すのではなく)である。残余——著者と N 体の評価者がたまたま同じ読みへ収束し、割れが出ないため信号も出ない場合——は post-merge observation 軸で受容する(本ファイルのもう一件の実行保証なき要求と同じ扱い)。`c0020a2` の `forced`、#1763 の「整合」、#1766 の `resolve wrongly` がこの形である。このうち `c0020a2` と #1766 は 5 部分が**記入済み**でこの結果になっており、未記入の検出では防げない——だから塞ぐ先は Question の文言側である)/ **Unit**(1 つの判定が覆う単位=文・段落・ファイル・主張・出現)/ **Scope**(軸が及ぶ面=当該 PR の diff・名指しした 1 ファイル・リポジトリ・リポジトリと wiki。不在の主張はどこを掃いたかを述べる。幅には**広がりのほかに言語という次元**があり、軸は両方を述べる。本リポジトリは正規テキストの多くを二重に持つ——`rules/` と `skills/` が英語、`docs/` が日本語であり、多くは前者のミラーだが、場所によっては `docs/` 自身が正本である(`docs/5.-Notifications.md` は自らを正本と宣言しており、ミラー元となる `rules/notifications/` が存在しない)——ため、リポジトリ全体に及ぶ掃引でも英語パターンだけで走れば日本語側には一度も届いておらず、その不在の主張は見落としではなく**構造的に**掃き残している。正本側の事例のほうが鋭い——掃引が代わりに当たれる英語の対応物がそもそも無いからである(2026-08-12 観測。#1733 の brake 1 が `docs/3.-Task.md` の日本語ミラーの取り残しを 2/1 で検出し、同 PR の `fa87222` で修正済み。残渣は現存しない)。広がりだけではこれを運べない——`リポジトリ` と書かれた scope はその掃引で満たされてしまう。Report shape が復路(所見の無い軸)に掃引のパターンを要求しているのと同じ次元を、往路すなわち軸を書く側で要求する)/ **Verdict terms**(その軸での yes と no の意味を軸自身の語で。「何か落ちたか」を問う軸では所見は yes でありながら draft にとっては negative であり、極性を名乗らない軸は誤った極性を継ぐ)/ **Basis**(軸が対象・基準について述べることはすべて、名指しした SHA で解決するポインタを伴う。基準は `パス` または `パス:行` 付き逐語引用、例示は実在箇所からの引用、件数が効く軸には数値でなく数える本文を渡す。**形式だけでは満たされない**——判定を当てる先、すなわち軸が判定基準とする既存の Li+ 基準や親が依拠する論証は、必要な軸ごとに書き込む。軸は隣に書かれたものを継承せず、基準を名指さない軸は clean ではなく未記入である)。5 部分はすべて評価者のプロンプトへそのまま載る payload であり、通れば消える点検項目ではない。**書かれなかった部分は評価者が読む本文の欠けたラベルとして見える**。機構は成果物上のこの可視性であって、`rules/model/subtractive-structural-beauty.md` が「手続きを構造へ置き換えよ」と要求するときの実行保証には届かない。立っている根拠は #1630 が「落ちた手順を、憶えておくべき 1 行として別の面へ置き直す」案を却下したのと同じものである。正本は `skills/evolution-parallel-agent-eval/SKILL.md` Axis statement form - brake 1 で評価者の判定が割れたときは、①**同じ問いに答えたか**(same-question check)→ ②同じ問いなら**なぜ割れたか**(判定基準の曖昧さに遡るか、明確な基準を適用する際のゆらぎに遡るか)の順で、各一度だけ問う。①が No なら課題は基準側でなく軸の書き方(一つの軸名が複数の問いを含んでいた)にあり、その No は Axis statement form のどの部分が保たなかったかを名指す先へ着地する(従来のように「次回の申し送り」へは行かない)。事前の Question と事後の same-question check は代替関係ではない——事後側は「部分は埋まっていたが緩く埋まっていた」軸を捕まえ、本 form 以前に書かれた軸にも効き続ける。**「同じ問いに答えており、基準はこれで足りていた」は正当な帰結**であり、盲点の発見を要求しない——`rules/model/trigger-check-gate.md` と同じ形(一つでも No なら止まって取得・検証して**進む**)である。無い所に基準の隙間を書き起こすのがここでの失敗モード - 比率(3/3 / 2/3 / 1/3)は判定入力ではなく、検証の労力をどこに先に置くかの安い事前分布である。報告してよいが判定の根拠にしてはならず、所見の可否はリテラルをソースに突き合わせて決まる(1/3 が採用され 3/3 が却下されうる)。数値閾値を持つのは固定軸(`skills/evolution-impression-literal-detection/SKILL.md` Aggregation)だけで、その数値は本節に写さない。この両半——比率が triage signal である旨と固定軸の carve-out——は統合コメント冒頭の 1 つの literal として skill 側に保持されており、親はそれを写す(この節を自分の言葉で言い直さない。言い直しが carve-out を跨いで一般化を生む)。正本は `skills/evolution-parallel-agent-eval/SKILL.md` Design Dimensions -- brake 1 が生む**3 つの成果物の形**は skill 側で固定してあり、各主体が毎回作文しない。3 つとは評価者の報告(step 3)/親の統合コメント(step 4)/著者の裁定(step 7)であり、いずれも同じ継ぎ目で非対称である——所見のある側は full、誰も争っていない側は 1 行。評価者の報告は、所見のある軸=逐語引用+`パス:行`(名指しした SHA 時点)+なぜ欠陥か/所見の無い軸=1 行(軸名+その軸の語での判定+根拠の `パス:行`。リポジトリ全体を掃く軸は指せる 1 行が無いため、再実行可能な形の掃き出し(パターンと掃いたパス)とヒット数で代える)。親の統合コメントは、**冒頭に 1 つの保持された literal**(比率は triage signal であって判定入力ではない旨と、絶対閾値を持つ固定軸だけがその外に立つ旨の 2 節)を写して開き、そのあと届いた full の形をそのまま保つ(冒頭の 2 節は両方 payload である。前者だけでは固定軸の位置で偽になり、後者だけでは軸ごとの比率が票に読める。親が自分の言葉で言い直すことがまさに carve-out を跨いだ一般化を生むため、文言を skill 側に保持する。英語以外へ解決する workspace では地の文と同じく解決後の言語で描画し、そのとき 2 節はどちらも落とさない)(ここで縮めると取り直しが著者へ回り、full を要求している理由そのものを潰す。統合が稼ぐのは評価者間の重複除去であって短縮ではなく、重複は「N 体中何体が挙げたか」を持つ 1 件へ畳む)。clean な軸も 1 行のまま載せる(著者は step 6 で軸横断に集約するため、所見だけのコメントでは分母が見えない)。統合コメントで禁止するのは受入・却下・順位づけ・推奨——それは step 7 の著者のものであり、重みづけの権限を持つ主体より先に着いてしまう。著者の裁定は commit body に置き、却下=full(単巡打ち切りの下で評価者が再反論しないため、親が step 8 で読むこれだけがその後ろに立つ)/受入=1 行(何が変わったかは同じ commit の diff に既に外部化されている)。1 件も受け入れなかった場合は commit が無く commit body も無いため、裁定は著者の停止条件報告で親へ渡り、親が step 9 のセルフレビューへ運ぶ——ここが「セルフレビューだけが唯一保証された外部化先である」理由であり、他の面はすべて書く対象がある場合にのみ存在する。全側面で禁止するのは、親がプロンプトで渡した基準・閾値・軸文言の復唱である。所見側の逐語引用は削減対象ではない——著者は引用に当てて裁定するため、削るとソースの取り直しが指摘ごとに発生し、リテラルに当てずに断定する形(`rules/model/trigger-check-gate.md` Literal check / #1673 cluster)へ戻る。所見の無い軸で「見て clean と言った」を支えるのはポインタの**解決可能性**であり、名指しした SHA で開けること、開いて判定と合わないなら 1 回の参照で露見することがその働きである。掃引軸には開く先が無いため、その働きは再実行可能な形で述べた掃き出しが担う(不在の主張は開いてではなく再実行で検査され、再現しないヒット数は解決しないポインタと同じ形で落ちる)。残余(誰も開かない clean 軸のポインタ、および判定に合うよう事後選定されたポインタ)は post-merge observation 軸で受容する。**言語**の規定は評価者側から消えた(#1732)。評価者はもう契約の及ぶ面へ書かず親の文脈へ返すだけであり、契約の及ぶ面(統合コメント)を書くのは解決後の値を自分の session に持つ親だからである——親が自分で書くものの言語は既にそう解決されており、ここに残す規定は無い。残るのは解決では片づかない継ぎ目 1 つ、**逐語引用は対象外**である。根拠は都合ではなく所見が乗っている性質そのもので、著者は引用に当てて裁定するため訳された literal は literal ではない。Li+ source は英語(`rules/model/liplus-coding-rule.md` Source Language)であり、他の言語へ解決する workspace では統合コメントは構造的に混在する(引用は原文の言語のまま、地の文は解決後の言語)。明示しなければ引用まで訳されるか地の文まで引用の言語へ倒れるかのどちらかになる。brake 2 と PR 面を持たない Trigger 項目は、いずれも自前の理由で本節の外である。正本は `skills/evolution-parallel-agent-eval/SKILL.md` Report shape +- brake 1 が生む**3 つの成果物の形**は skill 側で固定してあり、各主体が毎回作文しない。3 つとは評価者の報告(step 3)/親の統合コメント(step 4)/著者の裁定(step 7)であり、いずれも同じ継ぎ目で非対称である——所見のある側は full、誰も争っていない側は 1 行。評価者の報告は、所見のある軸=逐語引用+`パス:行`(名指しした SHA 時点)+なぜ欠陥か/所見の無い軸=1 行(軸名+その軸の語での判定+根拠の `パス:行`。リポジトリ全体を掃く軸は指せる 1 行が無いため、再実行可能な形の掃き出し(パターンと掃いたパス)とヒット数で代える)。親の統合コメントは、**冒頭に 1 つの保持された literal**(計数は triage signal であって判定入力ではない旨、所見はリテラルをソースに突き合わせて裁く旨、そして固定軸だけがその両方の外に立ち——その軸では計数こそが裁定であり——絶対閾値どおりに記録される旨の 3 節)を写して開き、そのあと届いた full の形をそのまま保つ(冒頭の 3 節はすべて payload である。前 2 節だけでは固定軸の位置で偽になり、第 3 節だけでは軸ごとの計数が票に読める。親が自分の言葉で言い直すことがまさに carve-out を跨いだ一般化を生むため、文言を skill 側に保持する。したがって解決後の言語へ描画せず原文の言語のまま写す——言い直しこそが拒まれている動作であり、翻訳は言い直しである)(ここで縮めると取り直しが著者へ回り、full を要求している理由そのものを潰す。統合が稼ぐのは評価者間の重複除去であって短縮ではなく、重複は「N 体中何体が挙げたか」を持つ 1 件へ畳む)。clean な軸も 1 行のまま載せる(著者は step 6 で軸横断に集約するため、所見だけのコメントでは分母が見えない)。統合コメントで禁止するのは受入・却下・順位づけ・推奨——それは step 7 の著者のものであり、重みづけの権限を持つ主体より先に着いてしまう。著者の裁定は commit body に置き、却下=full(単巡打ち切りの下で評価者が再反論しないため、親が step 8 で読むこれだけがその後ろに立つ)/受入=1 行(何が変わったかは同じ commit の diff に既に外部化されている)。1 件も受け入れなかった場合は commit が無く commit body も無いため、裁定は著者の停止条件報告で親へ渡り、親が step 9 のセルフレビューへ運ぶ——ここが「セルフレビューだけが唯一保証された外部化先である」理由であり、他の面はすべて書く対象がある場合にのみ存在する。全側面で禁止するのは、親がプロンプトで渡した基準・閾値・軸文言の復唱である。所見側の逐語引用は削減対象ではない——著者は引用に当てて裁定するため、削るとソースの取り直しが指摘ごとに発生し、リテラルに当てずに断定する形(`rules/model/trigger-check-gate.md` Literal check / #1673 cluster)へ戻る。所見の無い軸で「見て clean と言った」を支えるのはポインタの**解決可能性**であり、名指しした SHA で開けること、開いて判定と合わないなら 1 回の参照で露見することがその働きである。掃引軸には開く先が無いため、その働きは再実行可能な形で述べた掃き出しが担う(不在の主張は開いてではなく再実行で検査され、再現しないヒット数は解決しないポインタと同じ形で落ちる)。残余(誰も開かない clean 軸のポインタ、および判定に合うよう事後選定されたポインタ)は post-merge observation 軸で受容する。**言語**の規定は評価者側から消えた(#1732)。評価者はもう契約の及ぶ面へ書かず親の文脈へ返すだけであり、契約の及ぶ面(統合コメント)を書くのは解決後の値を自分の session に持つ親だからである——親が自分で書くものの言語は既にそう解決されており、ここに残す規定は無い。残るのは解決では片づかない継ぎ目であり、**逐語引用は対象外**、**冒頭の保持 literal も対象外**である。根拠は都合ではなく所見が乗っている性質そのもので、著者は引用に当てて裁定するため訳された literal は literal ではない。Li+ source は英語(`rules/model/liplus-coding-rule.md` Source Language)であり、他の言語へ解決する workspace では統合コメントは構造的に混在する(引用は原文の言語のまま、地の文は解決後の言語)。明示しなければ引用まで訳されるか地の文まで引用の言語へ倒れるかのどちらかになる。brake 2 と PR 面を持たない Trigger 項目は、いずれも自前の理由で本節の外である。正本は `skills/evolution-parallel-agent-eval/SKILL.md` Report shape - 人間ゲートはリリース / 不可逆系の軸と execution-mode の minor/major レビューに残る(人間 = 最終審判者は別軸) 段階責務: diff --git a/skills/evolution-impression-literal-detection/SKILL.md b/skills/evolution-impression-literal-detection/SKILL.md index 4eefe05..95e5cb0 100644 --- a/skills/evolution-impression-literal-detection/SKILL.md +++ b/skills/evolution-impression-literal-detection/SKILL.md @@ -43,9 +43,9 @@ Rationale: behavior-vs-impression boundary is context-dependent, so N=1 flag car ## Prompt literal -The form this axis enters an evaluator prompt in — on the brake 1 path and on the other Trigger entries alike (`skills/evolution-parallel-agent-eval/SKILL.md` Trigger). Copy it verbatim; do not re-compose it per spawn, and add nothing to it. That file's Axis statement form excludes this axis from the five parts it fixes for the axes selected per draft, and holding the wording here is what lets the exclusion hold: there is no blank for one of those parts to be filled in on this axis. The material the prompt names alongside it stays the parent's, fixed at that file's Procedure step 3. +The form this axis enters an evaluator prompt in — on the brake 1 path and on the other Trigger entries alike (`skills/evolution-parallel-agent-eval/SKILL.md` Trigger). Copy it verbatim; do not re-compose it per spawn, and add nothing to it. That file's Axis statement form excludes this axis from the five parts it fixes for the axes selected per draft, and holding the wording here is what lets the exclusion hold: there is no blank for one of those parts to be filled in on this axis. The material the prompt names alongside it stays the parent's, fixed at that file's Procedure step 3 on the brake 1 path and at its step 2 for a draft that is not on one. The literal resolves against whichever of the two the prompt named, so it is copied unchanged on either. -> **Fixed axis — impression-literal detection.** Take each phrase the draft under evaluation adds or modifies in Li+ source (`rules/*`, `skills/*`, `adapter/*`) and remove it: does the rule's behavior semantic change? Unchanged means the phrase is impression literal, and that is a finding on this axis; changed means it load-bears and is clean. Judge the added and modified lines, not the surrounding unchanged text. The Positive and Negative lists bounding this axis, including the categories that are protected and must not be flagged, are at `skills/evolution-impression-literal-detection/SKILL.md`; retrieve that file at the named SHA and apply it as written. Report each flagged phrase as a verbatim quote with its `path:line`. Do not aggregate and do not apply a threshold: this axis's thresholds are absolute and are applied to the N reports after yours arrives. +> **Fixed axis — impression-literal detection.** Take each phrase the draft under evaluation adds or modifies in Li+ source (`rules/*`, `skills/*`, `adapter/*`) and remove it: does the rule's behavior semantic change? Unchanged means the phrase is impression literal, and that is a finding on this axis; changed means it load-bears and is clean. Judge the added and modified lines, not the surrounding unchanged text. The Positive and Negative lists bounding this axis, including the categories that are protected and must not be flagged, are at `skills/evolution-impression-literal-detection/SKILL.md`; retrieve that file at the revision this prompt names and apply it as written. Report each flagged phrase as a verbatim quote with its `path:line`. With nothing flagged, what the verdict rests on is the set you read rather than any one line: report the axis clean in one line naming the draft you swept and how many added and modified Li+ source lines it carried. Do not aggregate and do not apply a threshold: this axis's thresholds are absolute and are applied to the N reports after yours arrives. The literal names the file instead of carrying the Positive and Negative lists inside itself: a second copy of them one section below the first is the copy that drifts, and the evaluator already holds the retrieval command that resolves a repository path at the named SHA (`skills/evolution-parallel-agent-eval/SKILL.md` Procedure step 3). diff --git a/skills/evolution-parallel-agent-eval/SKILL.md b/skills/evolution-parallel-agent-eval/SKILL.md index 0727b1a..81889dd 100644 --- a/skills/evolution-parallel-agent-eval/SKILL.md +++ b/skills/evolution-parallel-agent-eval/SKILL.md @@ -103,7 +103,7 @@ The one place this skill fixes a count as a threshold is the fixed axis (`skills - the shape that report is written in: findings at full length with a verbatim quote and its `path:line`, an axis with no finding at one line, and no echo of the criteria this prompt supplies (see Report shape). That sentence is what the prompt carries; the Report shape section behind it is the parent's reference, not prompt payload, since pasting the section would put back into the prompt the bulk this contract removes. The shape is a contract on the report, so leaving it to the evaluator's discretion is what puts every axis at finding length 4. **Aggregate findings and post** - Actor = the parent. Consolidate the N evaluators' reports into one PR comment and post it (`gh pr comment --repo / --body ...`) — one comment for the whole eval, in the shape fixed at Report shape, Parent's aggregated comment. Post nothing when every axis on every report is clean: there is then nothing to adjudicate, the author is not resumed, and steps 6 to 8 have no input; the eval's record rests on the self-review at step 9. - Consolidation is summary and merge — duplicate findings raised by more than one evaluator collapse into one entry carrying its ratio, and the surrounding prose is the parent's. It stops there. **The parent does not select among findings**: it does not drop one it disagrees with, does not rank them, and does not answer one. Accept / reject is the author's authority (Constraint: Adjudication actor), and a finding dropped at this step never reaches the actor holding that authority, which is the one way this step can silently decide what it is not entitled to decide + Consolidation is summary and merge — duplicate findings raised by more than one evaluator collapse into one entry carrying its ratio, and the surrounding prose is the parent's — the held preamble the comment opens with aside (Report shape, Parent's aggregated comment). It stops there. **The parent does not select among findings**: it does not drop one it disagrees with, does not rank them, and does not answer one. Accept / reject is the author's authority (Constraint: Adjudication actor), and a finding dropped at this step never reaches the actor holding that authority, which is the one way this step can silently decide what it is not entitled to decide 5. **Runtime restore** - Restore `.claude/` to tag-match state (revert the operational copy to pre-draft). Parent-side, as the apply at step 2 was. It runs as soon as every evaluator has returned its findings, and before the author is resumed: the copy exists for the evaluators' observation surface only, so it does not wait on step 6. Skipping it carries the draft into the parent session's behavior and leaves it there for subsequent sessions. Skip only when step 2 applied nothing (skills/* direct-Read path, or permission-gate fallback): no write occurred, so there is nothing to restore 6. **Aggregate verdict** - Actor = the resumed implementation subagent, not the parent (see Constraint: Adjudication actor). It reads the parent's aggregated comment and aggregates cross-axis judgment per the Design Dimensions aggregation rule. Fixed axes may override the default per-axis (see `skills/evolution-impression-literal-detection/SKILL.md` Aggregation). Where evaluators split on an axis, run Design Dimensions Divergence handling before the verdict is written 7. **Judgment** - Actor = the resumed implementation subagent. consistent -> nothing to apply; report back. partial / negative -> adjudicate each finding against the source, apply what was accepted, commit, push, and stop again at CI green. The accept / reject verdict on each finding and its reason go in the **commit body** (shape = Report shape, Author's adjudication); they are not posted as a PR comment. Or abort. **Single round**: steps 2-6 produce one verdict per draft, and the revised draft ships without re-verification (see Non-scope: what the single-round cap gives up). The author's response-and-revision pass is the tail of that one round, not a second one. What the cap refuses is a second audit of a draft already audited, so whether a re-run is that second audit is settled by what the round audited, never by why it stopped. A re-run continues the same round when both hold: (a) the verdicts that round returned have not reached the N>=3 floor (Constraint: N=1 prohibited, minimum N=3), and (b) the baseline it ran against — the PR commit SHA — is unchanged from the first attempt. Both, and neither alone: under the pair the re-run audits a draft nothing has audited yet, so it completes round 1 instead of opening round 2. Verdicts already returned are carried into it rather than discarded — they measure the same draft. They must also share the instrument: a verdict counts toward the floor only where the axes and prompt that produced it are the ones the re-run spawns under. Repairing a prompt between attempts is permitted, and a malformed one has to be repaired before it can return anything — but the repair retires the verdicts taken under the old wording instead of adding to them, because the floor counts independent samples of one question and two wordings are two questions. Cause is not a term here: a spend limit, an evaluator crash, a malformed prompt and a timeout are one thing under this criterion, a round that returned fewer than the floor, and sorting them is an after-the-fact self-report the cap cannot check. (a) is what refuses re-buying a verdict already obtained; (b) is what refuses re-auditing the author's revision, which moves the SHA and is round 2 by definition — the cap's own target. Ceiling: a third attempt against the same baseline that still has not reached the floor stops there and escalates to human. The number and its task / debug category are `skills/model-loop-safety`'s, and this skill states none of its own; the action is not — Loop Safety prescribes stop-and-switch at its threshold, with human arriving only if the switch does not converge. Here the switch is spent before the threshold: repairing the prompt and naming a higher model tier (Constraint: Model floor) are available at every attempt, so what is left at the third is the stop @@ -154,13 +154,13 @@ Scope = the brake 1 path. brake 2 is out: its evaluator has its own prompt file - **Preamble**: the comment opens with one held literal, copied rather than composed: - > The 3/3, 2/3 and 1/3 counts below are a triage signal on where to check first, not a judgment input. Adjudicate each finding by checking its literal against the source at the named SHA, and adopt or drop it on that. One axis stands outside this line: the fixed impression-literal axis carries absolute thresholds in its own spec (`skills/evolution-impression-literal-detection/SKILL.md` Aggregation), and its counts are applied there as written. + > The count on each finding below — how many of the N evaluators raised it — is a triage signal on where to check first, not a judgment input. Adjudicate each finding by checking its literal against the source at the revision its `path:line` is given at, and adopt or drop it on that. One axis stands outside both of those sentences: on the fixed impression-literal axis the count is the adjudication, at the absolute thresholds its own spec fixes (`skills/evolution-impression-literal-detection/SKILL.md` Aggregation), and it is recorded there as that spec requires. - Both clauses are payload. The first without the second is false where the fixed axis lands; the second without the first leaves the per-draft axes' counts reading as votes (Design Dimensions, Ratio is a triage signal). The parent restating this in its own words is what generalizes the first over the carve-out the second names, which is why the wording is held here. Where the workspace resolves to a language other than English the preamble is rendered in it like the rest of the parent's prose (Language below), and rendering carries both clauses. + All three clauses are payload. The first two without the third are false where the fixed axis lands — that axis takes neither the triage reading of its count nor the literal-check route to a verdict — and the third without them leaves the per-draft axes' counts reading as votes (Design Dimensions, Ratio is a triage signal). The parent restating this in its own words is what generalizes the first two over the carve-out the third names, which is why the wording is held here. It is copied in the language the source has it in and is not rendered into the workspace's: re-wording is the move being refused, and translation is a re-wording. Language below leaves a verbatim quote where the source has it for its own reason, and the preamble sits on that side of the comment's seam rather than in the prose the parent resolves. - **A finding** keeps the full form it arrived in: the verbatim quote, its `path:line`, and why it is a defect. Compressing it here would push the re-fetch onto the author, which is the cost the full-length rule above exists to remove; consolidation earns its place by removing duplication across evaluators, not by shortening a finding. Duplicates collapse into one entry stating how many of the N raised it. - **Clean axes** carry their one line through as well. The author aggregates cross-axis at Procedure step 6, and a comment listing only findings leaves it aggregating over a denominator it cannot see. - **Prohibited**: an accept, a reject, a ranking, or a recommendation. Those are the author's at Procedure step 7, and prose that leans on a finding here arrives ahead of the actor entitled to weigh it. -- **Language**: the parent writes this comment and resolves `Workspace_Language_Contract` (`adapter/claude/CLAUDE.md` / `adapter/codex/AGENTS.md`) for it as it does for anything else it writes, so nothing about resolving it is stated here. What is stated is the one seam resolution does not settle: a verbatim quote and its `path:line` stay as the source has them. The author adjudicates against the literal and a translated literal is not the literal. Li+ source is English (`rules/model/liplus-coding-rule.md` Source Language), so in a workspace resolving to any other language this comment is mixed by construction — quotes in the source's language inside prose in the resolved one. Left unstated the pair collapses one way or the other: the quotes get translated with the prose, or the prose falls back to the quotes' language. A clean axis carries no quote, so its one line is prose and a pointer, and a pointer is not prose. +- **Language**: the parent writes this comment and resolves `Workspace_Language_Contract` (`adapter/claude/CLAUDE.md` / `adapter/codex/AGENTS.md`) for it as it does for anything else it writes, so nothing about resolving it is stated here. What is stated is what resolution does not settle: a verbatim quote and its `path:line` stay as the source has them, and so does the held preamble above. The author adjudicates against the literal and a translated literal is not the literal. Li+ source is English (`rules/model/liplus-coding-rule.md` Source Language), so in a workspace resolving to any other language this comment is mixed by construction — quotes in the source's language inside prose in the resolved one. Left unstated the pair collapses one way or the other: the quotes get translated with the prose, or the prose falls back to the quotes' language. A clean axis carries no quote, so its one line is prose and a pointer, and a pointer is not prose. ### Author's adjudication From 79bd8722fe5481300eba0da562d48bc9b520e7f5 Mon Sep 17 00:00:00 2001 From: Claude Lin & Lay Date: Fri, 21 Aug 2026 01:53:00 +0900 Subject: [PATCH 3/3] fix(evolution): scope the Prompt literal closing paragraph to both paths MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 親の検分で名指された修正 1 件と、記録漏れの明示 1 件。#1779 ## 受入(親の指摘、1 件) 所見 1-a が、修正した段落の 1 つ下で生き残っていた。`## Prompt literal` の 締め段落が `the named SHA` と Procedure step 3 のみを名乗っており、直前の段落を brake 1 経路 = step 3 / それ以外 = step 2 へ直した後は、同一節が material の 問いに二度、食い違って答える状態になっていた。しかもこの段落は Basis が ポインタで足りる理由そのものであり、非 brake-1 の spawn を組む親が読む側である。 締め段落から経路の名指しを外し、`at whichever revision the paragraph above resolved to` として直前段落へ解決させた。step-3 / step-2 の分岐は節内に 1 箇所 (46 行目)だけ残る。二重保持を作らないため、分岐そのものは写していない。 ## 記録(新規変更ではない) literal 末尾の以下の文について、親が「どの所見も要求していない追加」として 名指したため記録する。 > Do not aggregate and do not apply a threshold: this axis's thresholds are > absolute and are applied to the N reports after yours arrives. 事実関係: この文は裁定コミット `4944acc` の追加ではなく、brake 1 が実際に走った `477b5b4` の時点で literal に存在していた(`240cf0e` = main には無い)。すなわち 評価対象の draft の一部として N=3 の評価を通っており、どの評価者も flag していない。 役目: 所見 1-b で評価者に clean 行の作文を新たに求めたため、絶対閾値を評価者が 自分で適用するものと読む余地が生じる。この文はその読みを閉じ、閾値の適用主体を (親の統合後の)裁定側に固定する。1-b の実装が成り立つのはこの文が在るからであり、 1-b の受入と同じ範囲に属する。 Co-Authored-By: Claude Opus 5 --- skills/evolution-impression-literal-detection/SKILL.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/skills/evolution-impression-literal-detection/SKILL.md b/skills/evolution-impression-literal-detection/SKILL.md index 95e5cb0..872d3f3 100644 --- a/skills/evolution-impression-literal-detection/SKILL.md +++ b/skills/evolution-impression-literal-detection/SKILL.md @@ -47,7 +47,7 @@ The form this axis enters an evaluator prompt in — on the brake 1 path and on > **Fixed axis — impression-literal detection.** Take each phrase the draft under evaluation adds or modifies in Li+ source (`rules/*`, `skills/*`, `adapter/*`) and remove it: does the rule's behavior semantic change? Unchanged means the phrase is impression literal, and that is a finding on this axis; changed means it load-bears and is clean. Judge the added and modified lines, not the surrounding unchanged text. The Positive and Negative lists bounding this axis, including the categories that are protected and must not be flagged, are at `skills/evolution-impression-literal-detection/SKILL.md`; retrieve that file at the revision this prompt names and apply it as written. Report each flagged phrase as a verbatim quote with its `path:line`. With nothing flagged, what the verdict rests on is the set you read rather than any one line: report the axis clean in one line naming the draft you swept and how many added and modified Li+ source lines it carried. Do not aggregate and do not apply a threshold: this axis's thresholds are absolute and are applied to the N reports after yours arrives. -The literal names the file instead of carrying the Positive and Negative lists inside itself: a second copy of them one section below the first is the copy that drifts, and the evaluator already holds the retrieval command that resolves a repository path at the named SHA (`skills/evolution-parallel-agent-eval/SKILL.md` Procedure step 3). +The literal names the file instead of carrying the Positive and Negative lists inside itself: a second copy of them one section below the first is the copy that drifts, and the evaluator already holds a retrieval command that resolves a repository path at whichever revision the paragraph above resolved to.