Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
21 changes: 21 additions & 0 deletions .ci/test-impact/impact_map.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
{
"schema_version": 1,
"source_paths": {
"backend/app/core/s3_validation.py": {
"test_modules": [
"tests/test_config.py",
"tests/test_artifact_store_conformance.py",
"tests/test_s3_artifact_store.py"
],
"reason": "S3 configuration validation is exercised in the shared-foundation partition."
}
},
"path_prefixes": {
".commitrail/": {
"test_modules": [
"tests/projects/review_policy/test_semantics.py"
],
"reason": "Commitrail policy semantics are tested in the shared-foundation partition."
}
}
}
1 change: 1 addition & 0 deletions .commitrail/INDEX.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,7 @@ for current product capability.
| [WS-REV-001](initiatives/WS-REV-001/OVERVIEW.md) | Planned | Shared acceptance/source and existing fence foundations; human hidden review work remains independently dependency-gated |
| [WS-QUAL-002](initiatives/WS-QUAL-002/OVERVIEW.md) | Planned | Populate subsystem ownership before changed-line mutation work |
| [WS-QUAL-003](initiatives/WS-QUAL-003/OVERVIEW.md) | Planned | Audit and prune test proof, add missing safety cases, decompose oversized test modules |
| [WS-CI-006](initiatives/WS-CI-006/OVERVIEW.md) | Planned | Shadow-mode semantic-lane impact report beside the unchanged full required suite; gate changes require later evidence and review |
| [WS-XINT-002](initiatives/WS-XINT-002/OVERVIEW.md) | Planned | Remaining ART/AUTH activation edges only |
| [WS-XINT-003](initiatives/WS-XINT-003/OVERVIEW.md) | Planned | Resume activation only against exact merged REV behavior |
| WS-POL-002 | Superseded | Future guide inference belongs to WS-POL-003; reframe remaining executor work against current specifications |
Expand Down
51 changes: 51 additions & 0 deletions .commitrail/initiatives/WS-CI-006/OVERVIEW.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,51 @@
# WS-CI-006 — Change-impact backend verification

- Disposition: Planned
- Next usable boundary: [WS-CI-006-01](WS-CI-006-01.md)

## Intent

Reduce routine backend PR feedback by selecting tests according to demonstrated
change impact, without making incomplete test evidence authoritative. Today
the full Backend suite remains required. The first usable step is a shadow
report whose recommendation is compared with the complete suite on the same
exact PR target.

## Current evidence and limits

- `backend/scripts/test_lane_catalogue.py` assigns every discovered test module
to the complete backend run. Shared, project and task node collections are
partitioned across two, three and three jobs respectively; a module in a
partitioned group requires every shard in its group.
- `.github/workflows/backend.yml` continues to run all nine lanes, authorization
preflight, MinIO, real API proof, and evidence aggregation on every PR. The
added `impact-report` job is informational and cannot control lane execution.
- Run [36724982896](https://github.com/Flow-Research/workstream/actions/runs/36724982896)
completed 7,918 tests with zero skips/deselections in about 44 minutes. This
is one observed run, not a universal baseline.
- Initial explicit mappings are intentionally narrow: S3 validation maps to
three reviewed test modules, and Commitrail-only changes map to the policy
semantics module. Their lane owners are derived from the canonical catalogue;
changed test modules also select every lane partition that owns them. Any
unmapped path recommends all nine lanes.
- The report is generated by the PR candidate and is not independent policy
evidence. PRs that change the selector, map, catalogue or workflow cannot
validate their own changes; no lane is omitted from actual CI. The report job
is not a dependency of the required aggregate, lane fan-in or API end-to-end
proof, so its failure cannot suppress those checks. Branch protection remains
the authority over which standalone job statuses are merge requirements.

## Direction after shadow evidence

1. Observe recommendations beside complete test results on the same exact
target across representative PRs.
2. Expand mappings only after tracing production owners, downstream consumers,
shared fixtures, integration services and the lane catalogue. Unknown impact
remains full-suite.
3. Evaluate whether the accumulated evidence supports a separate bounded change
to required test execution. That decision must preserve the full suite on
`main` and broad/cross-cutting changes; this initiative does not pre-approve
a CI gate reduction or promise every PR completes in five minutes.

No external selector service, historical mutable test signal, test deletion,
coverage quota, arbitrary sharding, or product behavior change is in scope.
164 changes: 164 additions & 0 deletions .commitrail/initiatives/WS-CI-006/WS-CI-006-01.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,164 @@
# WS-CI-006-01 — Shadow-mode backend test-impact report

- Initiative: `WS-CI-006`
- Durable disposition: `Planned`
- Intended merge outcome: Backend CI reports a deterministic, exact-PR proposed semantic-lane selection and rationale while the existing complete Backend suite remains unchanged and blocking on every PR and `main` push. This change does not enable selective gating.

## Intent

Backend PRs currently pay for the complete suite even when a change has a
bounded impact. Before changing that gate, collect exact-head evidence about
which semantic lanes a conservative change-impact map would select. Use the
existing lane catalogue as the first coarse unit; later work may select within
a lane only after the shadow reports establish a sound owner-to-test closure.

The observed baseline is run
[36724982896](https://github.com/Flow-Research/workstream/actions/runs/36724982896):
7,918 tests, zero skipped/deselected, about 44 minutes wall time. This is one
run, not a universal timing estimate. The target is under five minutes for
small, well-understood changes after a later separately reviewed gating phase;
this shadow change makes no runtime reduction or timing claim.

## Current behavior and lane semantics

- `backend/scripts/test_lane_catalogue.py` assigns every discovered test module
to the complete run. Shared, project, and task groups are partitioned across
multiple jobs. A module appearing in a partitioned group requires every
shard in that group to cover all of its node IDs.
- `.github/workflows/backend.yml` runs all nine lanes, the authorization
preflight, MinIO-backed tests, real API proof and aggregate evidence on every
PR and `main` push. The required `test` result and full-suite policy remain
unchanged in this PR.
- A shadow report may recommend lanes but cannot control `if` conditions,
services, test commands, required jobs, fan-in, or merge status. Unknown,
broad, malformed, or unclassified input reports `all lanes`.

## Bounded change

### Allowed files

- `.github/workflows/backend.yml` for a PR-only classifier/report job, and
`.github/workflows/agent-gates.yml` to run repository-level selector tests;
the existing backend full-suite graph and commands stay blocking.
- `.ci/test-impact/impact_map.json` and
`scripts/backend_test_impact.py` for deterministic shadow classification and
exact-target report generation.
- `scripts/test_backend_test_impact.py` for selector behavior regressions in
the repository-level lightweight gate suite.
- `backend/scripts/test_lane_catalogue.py` only as the authoritative inventory
read by the selector; `scripts/git_delta.py` and
`scripts/test_git_delta.py` only for shared byte-safe Git delta primitives.
- `scripts/test_lightweight_agent_gates.py` for backend workflow-shape and
fan-in regressions.
- `docs/operations_backend_testing.md`, the WS-CI-006 initiative overview,
this record, and the Commitrail index for the shadow-only operating contract.
- `docs/roadmap_status.md` only if its current CI capability statement needs
correction to describe this intended merged state.

### Prohibited changes

- No changes to Backend test bodies, assertions, collection,
skip/deselect behavior, coverage policy, current lane partitioning, services,
test commands, required status checks, branch protection, or merge rules.
- The selector regression suite may be added to the existing repository-level
Agent Gates test command; this does not change Backend test collection.
- No selector-driven workflow conditions, lane omissions, workflow-level path
filters, test execution service, third-party impact product, or mutable
historical selection authority.
- No test deletion or claim that shadow output reduces CI time.

## Shadow selection contract

- Resolve changed paths from the exact PR base/head and merge base; bind a
successfully resolved report to the base SHA, head SHA, execution SHA/tree,
changed-path digest, selector/map version and digests, selected semantic
lanes and rationale. If target identity or changed paths cannot be resolved,
mark unavailable facts as unavailable and recommend all lanes.
- Use the current semantic lane catalogue to map test modules to every lane
shard that owns their nodes. A changed test module is included in the proposed
impact closure. Shared fixtures, schema/migrations, dependencies, workflow,
lane catalogue, map/selector changes, unknown source paths, or unavailable
Git evidence conservatively recommend all nine lanes.
- Initial source mapping is deliberately limited to the reviewed S3 validation
owner `backend/app/core/s3_validation.py`, mapped to
`tests/test_config.py`, `tests/test_artifact_store_conformance.py`, and
`tests/test_s3_artifact_store.py`. Resolve their owning lanes from the
canonical catalogue; do not copy shard membership into the impact map.
Changes to other application source recommend all lanes until additional
consumer closures are demonstrated and explicitly mapped.
- `.commitrail/**` remains in the changed-path report and maps to
`tests/projects/review_policy/test_semantics.py`; derive its owner lanes from
the canonical catalogue. Mixed changes union this with source selection. If
any other path is unclassified, recommend all lanes. Documentation, skills,
and agent-policy paths are not implicitly exempted.
- The report explains every selected lane and every omitted lane. Omission is
allowed in the *recommendation only* when the exact mapping explains why;
missing evidence yields all lanes. The classifier is observational: CI still
executes all nine lanes and all existing integration/preflight jobs.
- The selector and map are part of the PR candidate in this shadow phase. Their
report is candidate-produced diagnostic evidence, not a trusted test policy
or independent audit receipt. A PR that changes the selector, map, catalogue,
or workflow cannot validate those changed inputs; any later gating change
must establish trusted selection policy separately.
- The report is uploaded and linked from the PR workflow summary. It is
descriptive evidence for comparing the proposed lane set with the complete
run from the same PR head; it never attests to another commit.

## Acceptance criteria

- [ ] The current full Backend workflow still runs unchanged on every PR and
`main` push, including all nine lanes, preflight, API/integration proof, and
the required `test` aggregate.
- [ ] A successfully resolved target report includes exact base/head/execution
tree, changed paths, selector/map identity, selected lanes, omitted lanes and
per-lane reasons. If target facts cannot be established, it clearly marks
them unavailable and recommends all lanes. The artifact and summary identify
the supplied PR head without asserting unverified facts.
- [ ] The initial S3 source change recommends both shared-foundation shards;
each mapped test module resolves to every partition owning its nodes.
- [ ] Commitrail-only changes recommend the shared-foundation shards containing
the policy-semantics tests; mixed known changes union their closures.
- [ ] Unknown paths, shared test support, migrations/schema, dependency and CI
machinery changes, malformed input, stale or missing Git objects, and
selector errors recommend all nine lanes rather than a partial set.
- [ ] Adversarial tests cover malformed/empty path lists, duplicate paths,
unknown files, renames from unknown source paths, changed tests, shared
fixtures, each protected map/selector input, stale or mismatched PR targets,
and missing lane ownership.
- [ ] Workflow regression tests prove the classifier output cannot condition,
skip, replace, or weaken any full-suite job or required check; in particular,
the report job is not a dependency of the required aggregate, lane fan-in or
API end-to-end step.
- [ ] A hosted PR run shows the shadow report beside complete passing test
evidence for the same head. Subsequent naturally occurring PRs provide the
representative comparison set; do not infer safety from synthetic paths or
a single S3-only example.
- [ ] Full-suite completeness, real integration checks, authorization,
concurrency, migration and rollback proof remain blocking. Coverage remains
diagnostic only.

## Risk and review routing

- Risk class: `L1` CI/workflow integrity.
- Required tracks: `ci_integrity`, `qa`, `test_delta`, `security`,
`documentation`, and `reuse_dedup`, selected through the reviewer matrix.
- Human review focus: proof the report is observational only, partition-aware
lane selection, exact-head binding, and unchanged full-suite enforcement.

## Evidence

Local focused verification uses the selector behavior tests, the lightweight
workflow-shape and fan-in tests, Ruff on changed Python files, Commitrail
validation, Markdown-link validation and the stale-wording scan. The PR must
also complete the existing hosted Backend workflow on the exact candidate; its
nine lanes, API drill, infrastructure and evidence checks remain unchanged.

## Reconciliation

- Current source: nine complete semantic lanes and their integration/fan-in
remain required on PRs and `main`.
- Next boundary: collect shadow classifications against full runs across
representative changes, then review the map and timing evidence before
planning any separate selective-gating change.
- No contribution instructions or product capability claims are relaxed by
this change. No main-push check is removed.
1 change: 1 addition & 0 deletions .github/workflows/agent-gates.yml
Original file line number Diff line number Diff line change
Expand Up @@ -73,4 +73,5 @@ jobs:
scripts.test_commitrail_contribution_paths
scripts.test_commitrail_archive_batch
scripts.test_commitrail_markdown_structure
scripts.test_backend_test_impact
scripts.test_lightweight_agent_gates
42 changes: 42 additions & 0 deletions .github/workflows/backend.yml
Original file line number Diff line number Diff line change
Expand Up @@ -139,6 +139,48 @@ jobs:
--ledger ../.ci/auth-boundaries/TEST_STRUCTURE_DEBT.json
python -m scripts.behavior_ownership validate

impact-report:
if: ${{ github.event_name == 'pull_request' }}
runs-on: ubuntu-latest
timeout-minutes: 5
permissions:
contents: read
steps:
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5
with:
persist-credentials: false
fetch-depth: 0

- uses: actions/setup-python@a26af69be951a213d495a4c3e4e4022e16d87065
with:
python-version: "3.12"

- name: Bind and classify the exact pull request target
env:
PR_BASE_SHA: ${{ github.event.pull_request.base.sha }}
PR_HEAD_SHA: ${{ github.event.pull_request.head.sha }}
run: >-
python scripts/backend_test_impact.py
--json .ci/test-impact/report.json
--markdown .ci/test-impact/report.md

- name: Upload exact-target impact report
id: report
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02
with:
name: backend-test-impact-${{ github.sha }}-${{ github.run_attempt }}
path: |
.ci/test-impact/report.json
.ci/test-impact/report.md
if-no-files-found: error
retention-days: 7

- name: Link report artifact in run summary
env:
REPORT_URL: ${{ steps.report.outputs.artifact-url }}
run: |
printf '\n[Download exact-head test-impact report](%s)\n' "${REPORT_URL}" >> "${GITHUB_STEP_SUMMARY}"

lanes:
needs: minio-image
runs-on: ubuntu-latest
Expand Down
22 changes: 22 additions & 0 deletions docs/operations_backend_testing.md
Original file line number Diff line number Diff line change
Expand Up @@ -110,6 +110,28 @@ If provisioning fails, confirm the local PostgreSQL provisioning credential can

## Hosted semantic-lane full-suite proof

For pull requests, the workflow also publishes a test-impact shadow report in
the Backend run summary and as a seven-day artifact. When exact Git target and
changed-path evidence resolves, it records the PR base/head, merge base,
checked-out execution SHA/tree, changed paths, selector/map/catalogue digests,
and lane-level selection reasons. If that evidence cannot be established, it
marks unavailable fields and recommends all lanes rather than presenting an
unverified exact-target classification. The initial
map is deliberately narrow; changed test modules use the existing lane
catalogue, while shared fixtures, migrations/schema, dependencies, workflow or
catalogue changes and any unmapped path recommend all nine lanes. A
classification error records an all-lanes fallback. The report is observational:
all nine matrix lanes, authorization preflight, API/integration proof and
evidence fan-in remain required and run independently of its recommendation.
The report job is not a dependency of the required aggregate, lane fan-in or API
end-to-end proof; report failure cannot suppress those checks. Branch protection
remains the authority over which standalone job statuses are merge requirements.
Do not use a shadow report as evidence that omitted lanes passed. A later
change to CI selection policy requires representative same-head comparisons,
trusted selection policy, and its own review. Because the report is generated
from the PR candidate, a PR changing its selector, map, catalogue, or workflow
does not validate that changed input.

The required GitHub check remains `Backend / test`. Nine matrix jobs each own a
digest-pinned PostgreSQL service container, a pinned-source MinIO image,
and exactly one dependency lane. A step-level curl health loop admits MinIO
Expand Down
6 changes: 4 additions & 2 deletions docs/roadmap_status.md
Original file line number Diff line number Diff line change
Expand Up @@ -189,8 +189,10 @@ cannot be reused as post-submission review-gate evidence. See the
- Cross-module behavior is moving through explicit public ports under the
modular-monolith boundary. New private edges are prohibited and touched debt
is reduced incrementally.
- GitHub CI distributes the backend suite across semantic lanes, rejects
skipped/deselected tests and requires behavior, boundary and real API proof.
- GitHub CI distributes the backend suite across semantic lanes and reports a
PR-only shadow impact recommendation; all nine full-suite lanes remain
required. It rejects skipped/deselected tests and requires behavior, boundary
and real API proof.
Coverage is diagnostic only, with no percentage gate or test-count target.
Redundant coverage-only reruns are removed; their tests remain in full-suite lanes.
Its nine-lane allocation uses three project lanes, three task lanes, two
Expand Down
Loading
Loading