Add WAMProbe counterfactual MVP and evidence map - #1
Conversation
|
Automated progress update ( |
There was a problem hiding this comment.
Pull request overview
This PR introduces the initial WAMProbe “counterfactual MVP” as a release-candidate-quality package: a dependency-free core with analytic benchmarks/metrics, capability-aware adapter protocols (including an isolated StarWAM adapter wrapper), deterministic caching + reporting, and a documentation/reproducibility/release-evidence pipeline.
Changes:
- Adds dependency-free toy benchmarks (PointMass-2D, BlockPush-2D, Gripper-Catch), metrics/statistics utilities, and evaluation/reporting plumbing with deterministic artifacts and caches.
- Adds typed public API contracts (capabilities, adapters, robotics/manipulation types), plus an isolated StarWAM adapter contract and supporting pinned manifests.
- Adds release candidate workflows, schemas/examples, documentation site + link/schema validation, and committed evidence/reproducibility artifacts.
Reviewed changes
Copilot reviewed 131 out of 137 changed files in this pull request and generated 3 comments.
Show a summary per file
| File | Description |
|---|---|
| tests/test_video_control_study.py | Tests PSNR/global-SSIM diagnostics and the video-vs-control counterexample study + CLI report writing. |
| tests/test_statistics.py | Tests deterministic bootstrap summaries and paired metric comparisons. |
| tests/test_starwam_adapter.py | Tests StarWAM adapter capability declaration, typing, close behavior, and horizon validation. |
| tests/test_robotics_api.py | Tests robotics datamodel validation and stable digests/serialization. |
| tests/test_repository_validation.py | Ensures repo-wide schema + markdown link validation script runs in tests. |
| tests/test_release_audit.py | Tests release audit logic and CLI generation for candidate artifacts/evidence. |
| tests/test_ranking_metrics.py | Tests candidate-ranking correlation metric behavior and input validation. |
| tests/test_pointmass.py | Tests PointMass rollout correctness and suite generation invariants. |
| tests/test_metrics.py | Tests core metrics (state FDE, action dependence permutation) on toy data. |
| tests/test_manipulation_benchmarks.py | Tests BlockPush/GripperCatch dynamics, observability, scaling, and baseline profiles. |
| tests/test_libero_cf.py | Tests LIBERO-CF-Mini manifest parsing, selection semantics, and checksum validation. |
| tests/test_import.py | Updates version exposure test to 0.1.0rc1. |
| tests/test_experiments.py | Tests experiment-report analysis and CLI report emission using fixtures. |
| tests/test_evaluation.py | Tests baseline ordering and determinism for evaluation outputs. |
| tests/test_dataset_cache.py | Tests intervention JSONL IO, prediction cache behavior, and corruption detection. |
| tests/test_closed_loop_study.py | Tests closed-loop replanning study behavior, determinism, and CLI outputs. |
| tests/test_cli.py | Tests demo/report/compare/dataset CLI commands and deterministic cache hits. |
| tests/test_artifacts.py | Tests action-prediction artifact schema, cache keys, redaction of raw RGB, and validation. |
| tests/test_api.py | Tests core typed API dataclasses and validation constraints. |
| src/wamprobe/video_metrics.py | Implements dependency-free PSNR and global SSIM diagnostics + invert transform. |
| src/wamprobe/stats.py | Implements deterministic context-block bootstrap summaries and paired comparisons. |
| src/wamprobe/manipulation_evaluation.py | Adds manipulation benchmark evaluator and per-context metric computation. |
| src/wamprobe/counterfactual.py | Adds counterfactual scoring and deterministic JSON artifact writer. |
| src/wamprobe/benchmarks/rendering.py | Adds dependency-free rasterizer for toy manipulation RGB observations. |
| src/wamprobe/benchmarks/pointmass.py | Adds analytic PointMass-2D benchmark and deterministic suite generation. |
| src/wamprobe/benchmarks/gripper_catch.py | Adds analytic GripperCatch benchmark with attachment semantics and rendering. |
| src/wamprobe/benchmarks/blockpush.py | Adds analytic BlockPush benchmark with pre-contact/contact phases and rendering. |
| src/wamprobe/benchmarks/init.py | Exposes built-in benchmarks from package namespace. |
| src/wamprobe/api/types.py | Adds typed 2D context/action/trajectory/suite data structures. |
| src/wamprobe/api/robotics.py | Adds typed RGB/observation/action prediction + stable content digests. |
| src/wamprobe/api/model.py | Adds typed adapter protocols for future- and action-prediction adapters. |
| src/wamprobe/api/manipulation.py | Adds typed manipulation benchmark/action/state contracts and adapter protocol. |
| src/wamprobe/api/errors.py | Adds project exception hierarchy for validation/capability errors. |
| src/wamprobe/api/capabilities.py | Adds capability manifest datamodel and enum for future representations. |
| src/wamprobe/api/init.py | Re-exports the public API surface. |
| src/wamprobe/adapters/starwam.py | Adds StarWAM adapter wrapper and pinned release contract validation. |
| src/wamprobe/adapters/manipulation.py | Adds manipulation baseline adapters (oracle/noisy/copy-last/wrong/action-agnostic). |
| src/wamprobe/adapters/baselines.py | Adds PointMass state baseline adapters (oracle/noisy/copy-last/wrong/action-agnostic). |
| src/wamprobe/adapters/init.py | Exposes built-in adapters from package namespace. |
| src/wamprobe/main.py | Enables python -m wamprobe CLI invocation. |
| src/wamprobe/init.py | Updates package description and version to 0.1.0rc1. |
| SECURITY.md | Adds a security policy and vulnerability reporting guidance. |
| schemas/result-v0.1.schema.json | Adds Draft 2020-12 schema for evaluation summary outputs. |
| schemas/intervention-v0.1.schema.json | Adds Draft 2020-12 schema for intervention JSONL records. |
| schemas/examples/intervention-pointmass-v0.1.json | Adds canonical example instance for pointmass intervention record. |
| schemas/examples/intervention-manipulation-v0.1.json | Adds canonical example instance for manipulation intervention record. |
| schemas/capability-v0.1.schema.json | Adds Draft 2020-12 schema for model capability manifests. |
| schemas/.gitkeep | Removes placeholder now that schemas directory is populated. |
| release/README.md | Documents release-candidate procedure and public release boundary. |
| release/evidence-manifest-v0.1.schema.json | Adds schema for the release evidence manifest. |
| release/evidence-manifest-v0.1.json | Adds pinned evidence manifest referencing committed and external artifacts. |
| pyproject.toml | Updates package metadata, scripts, dev deps, build include/exclude, and strict tool config. |
| paper/references.bib | Adds technical report bibliography entries. |
| paper/README.md | Documents how to compile the Overleaf-ready technical report. |
| paper/CLAIMS_CHECKLIST.md | Adds claim↔evidence checklist to keep paper claims aligned with artifacts. |
| mkdocs.yml | Adds strict MkDocs site configuration and navigation. |
| examples/video-control-study/video-control-study.md | Adds committed markdown report for video fidelity vs control study. |
| examples/video-control-study/video-control-study.json | Adds committed JSON report for video fidelity vs control study. |
| examples/closed-loop-study/closed-loop-study.md | Adds committed markdown report for toy closed-loop replanning study. |
| environments/starwam/README.md | Documents isolated StarWAM environment setup and verified preflight/matrix steps. |
| environments/starwam/patches/0001-safe-inference-loads.patch | Adds upstream patch to restrict PyTorch loads to weights-only where verified safe. |
| environments/libero/README.md | Documents LIBERO-CF-Mini generator workflow and restore/validation invariants. |
| docs/rfcs/0002-counterfactual-metrics.md | Adds RFC for counterfactual-first metric design and sanity checks. |
| docs/rfcs/0001-scope-and-capabilities.md | Adds RFC for scope and capability-aware adapter routing. |
| docs/reproducibility/REPRODUCIBILITY.md | Adds tiered reproduction guide and canonical hashes for committed studies. |
| docs/reproducibility/GPU_NIGHTLY.md | Adds guide for provisioning/running the disabled-by-default GPU nightly runner. |
| docs/reproducibility/EXTERNAL_REPRODUCTION_TEMPLATE.md | Adds template for independent reproduction reports. |
| docs/models/STARWAM.md | Adds StarWAM adapter model card, capabilities, evaluation summary, and limitations. |
| docs/metrics/CORE_METRICS.md | Adds metric cards, applicability rules, and baseline expectations. |
| docs/index.md | Adds documentation landing page describing WAMProbe and pointers to key docs. |
| docs/experiments/TOY_CLOSED_LOOP_V0.1.md | Adds experiment card describing the toy closed-loop study protocol/results. |
| docs/experiments/STARWAM_LIBERO_CF_MINI_V0.1.md | Adds experiment card summarizing StarWAM × LIBERO-CF-Mini matrix results and skips. |
| docs/benchmarks/TOY_BENCHMARKS.md | Adds benchmark card for toy analytic benchmarks and protocols. |
| docs/benchmarks/LIBERO_CF_MINI.md | Adds benchmark card for LIBERO-CF-Mini pilot and verification procedure. |
| docs/benchmarks/libero_cf_mini_verified_v0.1.json | Adds structured verification record for LIBERO-CF-Mini pilot run. |
| CONTRIBUTING.md | Adds contributor setup + validation checklist and contribution requirements. |
| configs/models/upstream_models.json | Adds pinned upstream model store manifest (paths/sizes/hashes/licensing/status). |
| configs/models/starwam-capabilities-v0.1.json | Adds pinned capability manifest for StarWAM adapter. |
| configs/models/.gitkeep | Removes placeholder now that model configs exist. |
| configs/benchmarks/libero_cf_mini_v0.1.json | Adds pinned LIBERO-CF-Mini task manifest with checksums and metadata. |
| CODE_OF_CONDUCT.md | Adds code of conduct for repo interactions. |
| CITATION.cff | Adds citation metadata for the software release. |
| checkpoints/README.md | Documents local model store rules and hash verification via wamprobe doctor. |
| CHANGELOG.md | Adds changelog and documents the 0.1.0rc1 candidate contents. |
| .gitignore | Updates ignores for docs build output and local artifacts (models/runs/vendor/paper outputs). |
| .github/workflows/release-candidate.yml | Adds manual release-candidate build/audit/smoke workflow with attestations. |
| .github/workflows/publish-pypi.yml | Adds guarded manual PyPI publish workflow using Trusted Publishing and tag confirmation. |
| .github/workflows/gpu-nightly.yml | Adds disabled-by-default GPU nightly workflow for self-hosted runner evidence generation. |
| .github/workflows/docs.yml | Adds strict docs build + GitHub Pages deployment workflow. |
| .github/workflows/ci.yml | Adds CI quality matrix (ruff/mypy/tests/docs/schema+links) and candidate build job. |
| .github/PULL_REQUEST_TEMPLATE.md | Adds PR template with validation checklist and metric/adapter/benchmark checklist. |
| .github/ISSUE_TEMPLATE/metric_proposal.yml | Adds structured issue template for proposing new metrics. |
| .github/ISSUE_TEMPLATE/config.yml | Adds issue template config with a security contact link. |
| .github/ISSUE_TEMPLATE/bug_report.yml | Adds structured bug report template. |
| .github/ISSUE_TEMPLATE/benchmark_proposal.yml | Adds structured benchmark proposal template. |
| .github/ISSUE_TEMPLATE/adapter_proposal.yml | Adds structured adapter proposal template. |
| .github/dependabot.yml | Adds monthly Dependabot updates for pip + GitHub Actions. |
| .gitattributes | Adds whitespace handling for patch files. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| "vec2": { | ||
| "type": "array", | ||
| "prefixItems": [{"type": "number"}, {"type": "number"}], | ||
| "items": false | ||
| }, |
| title: "WAMProbe: Counterfactual Evaluation for World Action Models" | ||
| type: software | ||
| version: 0.1.0-rc.1 | ||
| license: Apache-2.0 |
| "type": "object", | ||
| "required": [ |
Summary
0.1.0rc1package candidate and an Overleaf-ready five-page technical reportWhy
A World Action Model can produce plausible outputs while ignoring an action, responding in
the wrong direction, violating contact/grasp semantics, or providing no control value.
WAMProbe evaluates shared-context action branches and keeps action dependence, physical
accuracy, control utility, and compute cost as separate, capability-gated measurements.
Verified evidence
Analytic and paired simulator tiers
0.0; repeat run: 4/4 verified cache hitsStarWAM matrix
61fd988ea648922652ef2ab23f08b14a9849d37353d2e4894eb839fe3dfdb12d3bb9d5bdab9437cf07b0ed6567a8ddb146191447931a59a098b853e2fd3e9846Candidate-action mask/shuffle is explicitly skipped because StarWAM's released interface
neither accepts candidate actions nor returns action-conditioned futures.
Video fidelity versus control value
0, CRC1, regret0, but PSNR about0.59 dB-0.161de2009fa09a025db1f16ff55b07220ee679ba228f01971e08a9808e717b6cf4This is a controlled metric counterexample. The dependency-free
global_ssimdiagnosticis explicitly not standard windowed SSIM.
Minimal closed loop
0.9855/1.000019d2a0108fc19580da4e81f15226a4706a40c9d2c0513857d6bc8dc9000189b3These are descriptive analytic sanity checks across five deliberately constructed profiles,
not evidence of real-model closed-loop task improvement.
Validation
ruff format --check .ruff check .mypy: 57 source files0.1.0rc1candidate00104868a1122efa4a29b445eb48b29a7d4591cf6f6899742a50d4cc939076db5c13313b9e418e680af16f3b4416cbee2dd1376365055cc51ebe144cd5d6f329f00d245e5b9e3582bb59d9560eed7fdf063ee2cd9aaf14049a2a686295fb4032d0cb0fb874a27385eff2bd74ee311ea581b3d2b82760a553d2860e20388b422cae366d93a89e2ed93d4f938ee138337d2b409648Formal GitHub/PyPI publication is intentionally not triggered by this draft PR.
Scope and limitations
Model weights, simulator datasets, and generated runs remain outside Git. StarWAM currently
validates action prediction and executed short-horizon behavior, not action-conditioned
video generation. Sparse LIBERO success is zero in the reported matrix, so this PR makes
no task-solving performance claim. Public PyPI installation and third-party reproduction
remain release gates that require maintainer publication and an independent external user.