Skip to content

research: add pilot v1 repeat-run stability sample - #152

Merged
ErenAri merged 7 commits into
mainfrom
research/repeat-stability-v1
Sep 18, 2026
Merged

ErenAri merged 7 commits into
mainfrom
research/repeat-stability-v1

Conversation

@ErenAri

@ErenAri ErenAri commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Add a bounded, manual-only repeat-run stability protocol for the canonical
pilot v1 dataset.

This is deliberately a post-collection, purposefully stratified sample. It
does not claim preregistration or population-random sampling.

Sample

Seven canonical execution tuples are selected to cover:

  • compatible controlled baseline;
  • incompatible ring-buffer boundary;
  • compatible AlmaLinux 8 / 4.18 vendor-backport case;
  • positive and negative Falco real-loader cases;
  • positive cilium/ebpf project-loader case;
  • Oracle environment-mismatch inconclusive case.

Each tuple is repeated three times, for 21 planned attempts.

Stability semantics

The analyzer separates:

  • environment drift — repeated execution resolves to a different
    exact_environment_id than the canonical run;
  • verdict instability — same exact environment, different normalized
    compatibility verdict;
  • stable_same_environment — exact environment and verdict both match.

Environment drift is not counted as nondeterministic compatibility behavior.

Reproducibility

The repeat workflow:

  • re-materializes and verifies the same frozen v1 CLI, validator, artifacts,
    and project loaders;
  • reuses the frozen v1 profile definitions and case contracts;
  • shares the per-profile VM workdir across repeats so image bytes are cached
    within one collection;
  • records repeat provenance, raw reports, VM logs, and normalized repeat
    records;
  • uploads all evidence as a 90-day staging Actions artifact.

The PR preflight performs shell/Python syntax checks, analyzer self-tests, and
verifies that every sampled case/profile tuple resolves to exactly one row in
the canonical v1 repository dataset.

After merge

Run once from main:

gh workflow run research-repeat-v1.yml --repo Kernel-Guard/bpfcompat --ref main

The resulting stability artifact can then be locked into the research dataset
and the repeat-run item in research/ROADMAP.md can be closed.

Summary by CodeRabbit

  • New Features

    • Added a repeat-stability study for pilot v1, covering seven canonical execution scenarios with three repetitions each.
    • Added automated validation of frozen inputs, test configurations, execution results, environment identity, and compatibility verdicts.
    • Added normalized stability summaries distinguishing stable results, environment changes, and verdict instability.
    • Added downloadable reports, logs, provenance data, checksums, and generated results for each study run.
  • Documentation

    • Documented the study scope, sampling approach, execution scenarios, and interpretation of stability results.

@coderabbitai

coderabbitai Bot commented Sep 18, 2026 •

Copy link
Copy Markdown

Review Change StackReview Change Stack

Note

Currently processing new changes in this PR. This may take a few minutes, please wait...

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: b05a6980-7f97-44cf-bc1a-a733decdcb66

📥 Commits

Reviewing files that changed from the base of the PR and between 3273905 and 3bcb318.

📒 Files selected for processing (5)
  • .github/workflows/research-repeat-v1.yml
  • research/repeat/v1/README.md
  • research/repeat/v1/stability-sample.json
  • scripts/research/analyze-repeat-v1.py
  • scripts/research/run-repeat-v1.sh
 _____________________________
< I speak fluent stack trace. >
 -----------------------------
  \
   \   (\__/)
       (•ㅅ•)
       /   づ
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@ErenAri
ErenAri marked this pull request as ready for review September 18, 2026 21:48
Copilot AI lite review requested due to automatic review settings September 18, 2026 21:48

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

ErenAri commented Sep 18, 2026

Copy link
Copy Markdown
Contributor Author

@coderabbitai review

@ErenAri
ErenAri merged commit 5153cb3 into main Sep 18, 2026
13 checks passed
@coderabbitai

coderabbitai Bot commented Sep 18, 2026 •

Copy link
Copy Markdown
⚠️ Action not completed

Pull request is closed.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants