Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
25 changes: 15 additions & 10 deletions .github/ISSUE_TEMPLATE/contributor-task.yml
Original file line number Diff line number Diff line change
Expand Up @@ -41,21 +41,24 @@ body:
- type: textarea
id: summary
attributes:
label: Summary
description: What should change, and why is it useful?
label: Goal and acceptance gap
description: Name the current request or contract, affected user/caller and observable result. Link an existing goal/direction when relevant; a roadmap id is not required.
placeholder: |
Goal/source:
Current gap:
Accepted outcome (before → after):
validations:
required: true
- type: textarea
id: scope
attributes:
label: Proposed scope
description: List the smallest useful slice and any known non-goals.
description: Choose a complete useful slice. State ownership, dependencies, non-goals and any staged remainder; do not split only to minimize file or PR size.
placeholder: |
In scope:
- ...

In scope / owner:
Existing related work / dependencies:
Out of scope:
- ...
If staged: useful delta, remaining gap, next owner/task, and why this boundary:
validations:
required: true
- type: input
Expand All @@ -79,10 +82,12 @@ body:
id: validation
attributes:
label: Validation plan
description: What command or review will prove the task is done?
description: What independently observable result proves completion? Include the actual entrypoint/readback and relevant failure case; passing a test count is insufficient.
placeholder: |
- python3 -m py_compile loopx/*.py
- loopx check --scan-root .
Accepted result and independent oracle:
Actual entrypoint / safe command:
Negative or recovery case:
Frontend / Lark / CLI impact or verified N/A:
validations:
required: true
- type: checkboxes
Expand Down
43 changes: 29 additions & 14 deletions .github/PULL_REQUEST_TEMPLATE.md
Original file line number Diff line number Diff line change
@@ -1,11 +1,28 @@
## Summary
## Goal And Delivered Outcome

-
<!-- Use a few concrete sentences; references are optional when the request or
regression is self-contained. State the accepted outcome, not a list of files.
For cross-cutting LoopX work, link the relevant overall-roadmap/domain acceptance
when useful. A roadmap id is not required for ordinary fixes or maintenance.
These are author facts; the review capability independently judges delivery.
-->

- Goal/source and gap:
- Observable before → after, with the validation row that proves it:
- Issue/task and intended base: <!-- Use Closes only for the issue actually completed; otherwise Related to. -->

## Scope And Continuation

## Issue Or Task
<!-- A scoped fix may be complete while the parent program remains open.
For a staged increment, explain the useful delta, remaining gap, next owner/task
and why this is an independently testable/reversible boundary. Link existing
work before creating follow-ups. Docs/research/maintenance need a concrete value,
not a fabricated runtime caller. Write "complete within this scope" when no
successor is needed. Do not grade quality from LOC, PR counts or test counts.
-->

- Closes #
- Contributor task ID:
- Completed scope and remaining work:
- Slice boundary / successor: <!-- N/A with reason when the accepted task is complete. -->

## Validation

Expand Down Expand Up @@ -88,16 +105,14 @@ even when the underlying access was authorized.

## Technical Direction

<!-- Select one. Direction labels route review; they do not imply maturity or merge authority. -->

- [ ] Core control-plane hardening
- [ ] Long-horizon benchmark evidence
- [ ] Operator surface and IM integration
- [ ] Shared Goal Authority and cross-host coordination
- [ ] Architecture and research incubator
<!-- Optional routing: Core control-plane hardening; Long-horizon benchmark evidence;
Operator surface and IM integration; Shared Goal Authority and cross-host coordination;
Architecture and research incubator. For cross-cutting work, reference an existing
roadmap S/G/R or domain acceptance id rather than copying the plan.
Routing is not maturity or implementation authority.
-->

- Target base branch:
- Direction tracker or promotion unit:
- Direction / acceptance reference, when applicable:

## Shared-authority RFC fixture impact

Expand Down
40 changes: 40 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,45 @@
# Agent Instructions

## Goal-Oriented Development

Before selecting non-trivial work, resolve the current requested outcome from
user direction, the linked issue/task, accepted contract or demonstrated bug.
For cross-cutting LoopX work, consult the [overall roadmap](docs/architecture/rfcs/loopx-overall-roadmap-v0.md)
and relevant domain acceptance; do not make every fix wait for every RFC or
invent a roadmap id. Check latest `main`, related PRs and canonical Todos so an
older task description cannot override corrected direction or duplicate work.

Carry one compact delivery brief from task to PR: goal/source, current gap,
observable result, owning boundary and decisive acceptance evidence. Reuse the
existing task/PR fields; keep private Goal state out of public artifacts.
Choose a complete, independently reviewable and reversible outcome slice.
Small diffs, fields, receipts, test counts and merged PR counts do not establish
progress. Characterization, prerequisites, research, docs and maintenance are
valid when they remove an evidenced gap or enable a named real next step.

Continue through the selected slice's implementation, integration, negative
cases and readback while authorized work remains feasible. Do not stop after
setup, a serializer, a mock or an isolated smoke when the useful outcome is
still missing. Do not expand scope merely to make a PR larger. When a staged
boundary is necessary, name the delivered delta, remaining gap, next owner/
dependency and why the boundary improves verification or rollback. Reuse or
update an existing successor; do not create ceremonial follow-up tasks for a
completed request. Real authorization, cost and operational stop gates remain.

For multi-Agent changes, qualify the relationship the user needs: dependency
artifacts, receiver adoption, claim/lease handling, independent acceptance and
result return as applicable. Sending a message or registering workers does not
prove collaboration. Missing frontend/Lark/CLI companion work makes a product
journey partial even when a backend slice is ready to merge.

Before delivery, reconcile the result with the original/current goal and
update its task and RFC checkpoint when the boundary changes. Preserve passed,
failed and untested distinctions. If user feedback exposes the same missing
outcome, repair the owning rule, active task or projection through self-repair;
do not merely append stronger instructions. PR review must execute the current
capability-owned `problem_context` delivery judgment; author declarations and
this prose do not certify it or settle a Goal.

## Commit And PR Hygiene

### Worktree And PR Gate
Expand Down
18 changes: 16 additions & 2 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -230,12 +230,26 @@ npm run smoke:demo-readiness
Before opening a pull request:

- link the issue or task ID when one exists;
- describe the behavior change and the validation you ran;
- state the requested outcome, current gap and observable before/after result;
- distinguish completion of the scoped task from a justified increment; for an
increment, name the remaining gap, next owner/dependency and why the boundary
is independently testable and reversible;
- link decisive validation to that outcome, including relevant user-entrypoint
readback and failure/recovery cases;
- keep unrelated formatting or refactors out of the PR;
- include docs or tests when changing user-visible behavior;
- confirm that no private/local runtime state was committed.

Maintainers may ask for a smaller PR if the change mixes unrelated concerns.
Use the [overall roadmap](docs/architecture/rfcs/loopx-overall-roadmap-v0.md) for
cross-cutting work, without inventing roadmap ids for ordinary fixes. Existing
issues and canonical Todos own execution; update them instead of duplicating
follow-up work. A completed task needs no invented successor. Prerequisites,
research, docs and maintenance can be useful delivered outcomes. A schema,
message, mock or passing suite alone does not complete a promised user journey.

Maintainers may request consolidation when a useful outcome was unnecessarily
split, or a smaller PR when unrelated concerns were mixed. Review evaluates the
verified goal delta and evidence, not minimum size, model identity or PR count.

### Validation disclosure

Expand Down
36 changes: 17 additions & 19 deletions docs/development/contributor-tasks.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@ into a mirror of maintainer scratch state.

| Status | Meaning |
| --- | --- |
| Available | Ready for someone to comment on the linked issue or open a small PR. |
| Available | Ready for a contributor to claim the linked outcome and deliver a cohesive PR. |
| Claimed | Someone has said they are working on it, or a maintainer assigned it. |
| Maintainer-owned | Active work is happening in maintainer/local automation; ask before touching. |
| Needs design | Discussion is welcome, but implementation needs agreement first. |
Expand All @@ -34,7 +34,7 @@ preferred review contact is not an exclusive task claim or new merge authority.
1. Prefer a linked GitHub issue. If there is no issue yet, open one with the
contributor task template.
2. Comment that you would like to work on the task. Maintainers will mark it
`claimed` or suggest a smaller slice.
`claimed` or agree a complete, independently verifiable slice.
3. For docs-only typo fixes or obviously tiny cleanups, opening a direct PR is
fine.
4. If a claimed task has no update for 14 days, maintainers may release it back
Expand All @@ -44,31 +44,29 @@ preferred review contact is not an exclusive task claim or new merge authority.

## Current Technical Directions

The canonical [Technical Directions map](../project/technical-directions.md)
explains outcomes, maturity, ownership boundaries, and promotion gates. This
board lists bounded work; it does not redefine those directions.
The [overall roadmap](../architecture/rfcs/loopx-overall-roadmap-v0.md) and
[tracking issue #4574](https://github.com/huangruiteng/loopx/issues/4574) own
cross-domain priorities and G0–G5 acceptance. The [Technical Directions map](../project/technical-directions.md)
owns contributor routing and current maturity; this board does not keep a second
copy of those stages. Before claiming a row, reconcile its linked task with
latest main, related PRs and the roadmap. Historical rows are not proof that a
missing feature remains unimplemented or a proposed slice is still useful.

| Direction | Current stage | Contributor entry | Boundary |
| --- | --- | --- | --- |
| Long-Horizon Benchmarks and Evidence | Active research | [#3243](https://github.com/huangruiteng/loopx/issues/3243) | Work on public-safe fixtures, treatment integrity, reducers, and docs; live cases and scoring remain maintainer-owned. |
| Operator Surface and IM Integration | Incubating on `frontend-control-plane-im-prototype-rfc` | [#3244](https://github.com/huangruiteng/loopx/issues/3244) | State the target base branch; UI remains a projection and promotion to `main` is staged. |
| Shared Goal Authority and Cross-host Coordination | Stage 2 slice shipped (aggregate head, file provider, `claim_work` executor); NoKV stays an unpromoted candidate | [#3245](https://github.com/huangruiteng/loopx/issues/3245) | Keep slices provider-neutral and file-backed; no second scheduler or write authority. |
| Architecture and Research Incubator | Mixed by RFC | [#3246](https://github.com/huangruiteng/loopx/issues/3246) | Read the per-exploration stage; an RFC alone does not make implementation claimable. |

Core control-plane reliability remains the shared shipped foundation. Effect
Program hardening, verified transitions, recovery, observability,
maintainability, and contributor experience continue through the focused rows
below and the existing `control-plane` label.
A claimable task names the current gap, independently useful outcome, existing
owner/caller, dependencies and decisive validation. For a staged increment,
record the remaining gap and next owner/task; do not make a field, fixture or
PR count the completion target. Preserve existing authoritative Todo/issue
identity rather than copying the whole plan here.

## Priority Queue

| Priority | Direction | Slice | Issue / PR | Status |
| --- | --- | --- | --- | --- |
| P0 | Core hardening | Exact-head review of remote execution and terminal writeback fencing: fenced journal recovery absorbed into TypeScript | #3074 | Done |
| P0 | Core hardening | Wire caller-approved `validation_command` into the remaining self-report entry points | #3082 / #3142 #3291 #3343 | Done |
| P1 | Benchmark evidence | Split one deterministic adapter-fidelity or treatment-integrity fixture | #3243 | Needs design |
| P1 | Operator surface / IM | Split one projection or session-contract characterization unit from the incubation branch | #3244 | Needs design |
| P1 | Shared coordination | Characterize the shipped file-backed `claim_work` executor with a provider-neutral parity fixture | #3700 / #3245 | Needs design |
| P1 | Benchmark evidence | Qualify a reproducible adapter-fidelity or treatment-integrity gap with existing focused fixtures | #3243 | Needs design |
| P0 | Operator surface / IM | Close the R1 confirmed-team commitment/readback gap, then qualify R2 real peer dependency handoff | #4574 / #4339 | Needs design |
| P1 | Shared coordination | Qualify the selected local authority durability and crash/replay boundary against existing D2 acceptance | #4224 / #3245 | Needs design |
| P1 | Core hardening | One budget-aware CLI output ergonomics slice | #2881 | Needs design |
| P2 | Project docs | Release docs install, activation, and recovery guidance through v0.5.4 | GH-C04 | Landed via #3982 |
| P2 | Maintainability | CLI ownership and hot-module extraction | GH-C06 | Available |
Expand Down
8 changes: 4 additions & 4 deletions examples/docs-governance-smoke.py
Original file line number Diff line number Diff line change
Expand Up @@ -513,11 +513,11 @@ def assert_technical_direction_governance_is_current() -> None:
assert required in rfc_index, required
assert "## Status matrix" not in rfc_index

# The task board routes to canonical direction/roadmap owners instead of
# duplicating their mutable maturity table.
for required in (
"Long-Horizon Benchmarks and Evidence",
"Operator Surface and IM Integration",
"Shared Goal Authority and Cross-host Coordination",
"Architecture and Research Incubator",
"../project/technical-directions.md",
"../architecture/rfcs/loopx-overall-roadmap-v0.md",
):
assert required in tasks, required

Expand Down
19 changes: 7 additions & 12 deletions examples/pr-review-command-smoke.py
Original file line number Diff line number Diff line change
Expand Up @@ -411,13 +411,7 @@ def fake_run_gh_json(args: list[str], *, cwd: Path | None = None) -> object:
assert section["word_hint"], section
assert section["agent_instruction"], section
assert "quota.py" not in section["agent_instruction"], section
assert [section["word_hint"] for section in template["sections"]] == [
"200-350字",
"300-500字",
"450-800字",
"250-500字",
"150-300字",
], template
assert all("无最低字数" in section["word_hint"] for section in template["sections"])
concrete_change = next(
section for section in template["sections"] if section["label"] == "具体改动"
)
Expand Down Expand Up @@ -962,6 +956,7 @@ def fake_run_gh_json(args: list[str], *, cwd: Path | None = None) -> object:
assert execution["completion_gate"]["metadata_only_verdict_allowed"] is False
assert execution["completion_gate"]["stale_head_verdict_allowed"] is False
assert execution["completion_gate"]["blocking_evidence_verdicts"] == {
"problem_context": ["off_goal", "fragmented", "not_yet_proven"],
"repository_reuse": ["unjustified_duplication", "not_yet_proven"],
"observable_semantics": ["unintended_drift", "not_yet_proven"],
"change_proportionality": ["disproportionate", "not_yet_proven"],
Expand Down Expand Up @@ -1100,11 +1095,11 @@ def fake_run_gh_json(args: list[str], *, cwd: Path | None = None) -> object:
assert "template below is intentionally blank" in markdown, markdown
assert "- 推荐阅读顺序:" in markdown, markdown
assert "- 五块模板(留空给 agentloop 填写):" in markdown, markdown
assert "动机(200-350字)" in markdown, markdown
assert "改动思路(300-500字)" in markdown, markdown
assert "具体改动(450-800字)" in markdown, markdown
assert "对主干的风险(250-500字)" in markdown, markdown
assert "我的整体评价(150-300字)" in markdown, markdown
assert "动机(按证据需要;无最低字数)" in markdown, markdown
assert "改动思路(按证据需要;无最低字数)" in markdown, markdown
assert "具体改动(按证据需要;无最低字数)" in markdown, markdown
assert "对主干的风险(按证据需要;无最低字数)" in markdown, markdown
assert "我的整体评价(按证据需要;无最低字数)" in markdown, markdown
assert "main regression risk:" not in markdown, markdown
assert "## Combined Review Sequence" in markdown, markdown
assert "PR #771" in markdown, markdown
Expand Down
Loading
Loading