Skip to content

Add WAMProbe counterfactual MVP and evidence map - #1

Merged
myheart521 merged 20 commits into
mainfrom
agent/initial-mvp
Jul 15, 2026
Merged

myheart521 merged 20 commits into
mainfrom
agent/initial-mvp

Conversation

@myheart521

@myheart521 myheart521 commented Jul 15, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • typed, capability-aware WAM/action/manipulation adapter protocols
  • paired PointMass, BlockPush, Gripper-Catch, and four-family LIBERO-CF-Mini interventions
  • causal, dynamics, ranking, regret, efficiency, uncertainty, and closed-loop utility profiles without a composite score
  • deterministic JSONL data, verified caches, paired comparison, and JSON/Markdown/HTML reports
  • pinned StarWAM observation-to-action inference and real LIBERO action execution
  • committed video-fidelity/control-value counterexample and score-execute-observe replanning studies
  • reproducible 0.1.0rc1 package candidate and an Overleaf-ready five-page technical report
  • strict MkDocs site, canonical JSON Schema instances, repository-local link checks, and GitHub Pages deployment

Why

A World Action Model can produce plausible outputs while ignoring an action, responding in
the wrong direction, violating contact/grasp semantics, or providing no control value.
WAMProbe evaluates shared-context action branches and keeps action dependence, physical
accuracy, control utility, and compute cost as separate, capability-gated measurements.

Verified evidence

Analytic and paired simulator tiers

  • three dependency-free toy benchmarks with exact state and RGB observations
  • four LIBERO task families × four branches × eight steps
  • two independent restores, repeated branches, and forward/reverse branch order all exact
  • maximum LIBERO integration-state error: 0.0; repeat run: 4/4 verified cache hits

StarWAM matrix

  • 4 tasks × 3 seeds × NFE 1/4/8 = 36/36 predictions and 36/36 executed action chunks
  • repeat run: 36/36 output-SHA-verified cache hits
  • NFE 1/4/8 mean latency: 0.780 / 0.972 / 1.216 seconds; peak GPU allocation about 11.39 GiB
  • horizon 8/16/32 mean EEF displacement: 0.1716 / 0.1121 / 0.0000
  • all short-horizon sparse returns/successes were zero and remain reported negative results
  • prediction matrix SHA256: 61fd988ea648922652ef2ab23f08b14a9849d37353d2e4894eb839fe3dfdb12d
  • execution index SHA256: 3bb9d5bdab9437cf07b0ed6567a8ddb146191447931a59a098b853e2fd3e9846

Candidate-action mask/shuffle is explicitly skipped because StarWAM's released interface
neither accepts candidate actions nor returns action-conditioned futures.

Video fidelity versus control value

  • 12 contexts on BlockPush-2D and Gripper-Catch
  • appearance-corrupted oracle: FDE 0, CRC 1, regret 0, but PSNR about 0.59 dB
  • PSNR/regret ordering conflicts on 3/9 and 5/9 comparable profile pairs
  • PSNR versus regret Pearson correlation about -0.16
  • study SHA256: 1de2009fa09a025db1f16ff55b07220ee679ba228f01971e08a9808e717b6cf4

This is a controlled metric counterexample. The dependency-free global_ssim diagnostic
is explicitly not standard windowed SSIM.

Minimal closed loop

  • score every legal candidate, execute one true-dynamics step, observe, and replan
  • oracle success is 100% on BlockPush and Gripper-Catch; noisy scorer success is 100% / 91.7%
  • copy-last, wrong-direction, and action-agnostic future scorers have zero success on both tasks
  • offline CRC versus mean closed-loop return Pearson is 0.9855 / 1.0000
  • study SHA256: 19d2a0108fc19580da4e81f15226a4706a40c9d2c0513857d6bc8dc9000189b3

These are descriptive analytic sanity checks across five deliberately constructed profiles,
not evidence of real-model closed-loop task improvement.

Validation

  • ruff format --check .
  • ruff check .
  • strict mypy: 57 source files
  • 76 tests passed with 88.35% coverage
  • Python 3.11–3.13 GitHub Actions quality matrix
  • seven canonical instances validated against four Draft 2020-12 public schemas
  • 42 repository-local links validated across 33 Markdown files
  • strict MkDocs documentation build
  • local reproducible double build, archive/metadata audit, committed-evidence hash audit
  • dependency-free, network-disabled clean-wheel install plus CLI/two-context demo smoke
  • three-pass LaTeX/BibTeX compilation: five pages, no unresolved references or layout warnings

0.1.0rc1 candidate

  • source commit: 00104868a1122efa4a29b445eb48b29a7d4591cf
  • wheel SHA256: 6f6899742a50d4cc939076db5c13313b9e418e680af16f3b4416cbee2dd13763
  • sdist SHA256: 65055cc51ebe144cd5d6f329f00d245e5b9e3582bb59d9560eed7fdf063ee2cd
  • release manifest SHA256: 9aaf14049a2a686295fb4032d0cb0fb874a27385eff2bd74ee311ea581b3d2b8
  • committed evidence manifest SHA256: 2760a553d2860e20388b422cae366d93a89e2ed93d4f938ee138337d2b409648

Formal GitHub/PyPI publication is intentionally not triggered by this draft PR.

Scope and limitations

Model weights, simulator datasets, and generated runs remain outside Git. StarWAM currently
validates action prediction and executed short-horizon behavior, not action-conditioned
video generation. Sparse LIBERO success is zero in the reported matrix, so this PR makes
no task-solving performance claim. Public PyPI installation and third-party reproduction
remain release gates that require maintainer publication and an independent external user.

@myheart521 myheart521 changed the title Add counterfactual PointMass MVP Add WAMProbe counterfactual MVP and evidence map Jul 15, 2026
@myheart521

Copy link
Copy Markdown
Owner Author

Automated progress update (c8eea56): added deterministic checksummed intervention JSONL round trips, a corruption-detecting content-addressed prediction cache, strict result reloads, JSONL reporting, and standalone report, compare, dataset-export, and dataset-validate commands. Validation: 61 tests passed, 88.30% coverage, strict mypy/Ruff passed, sdist+wheel built, and the wheel passed an isolated install/export/validate smoke test. The PR remains Draft while LIBERO-CF-Mini and real-WAM experiments continue.

@myheart521
myheart521 marked this pull request as ready for review July 15, 2026 15:38
Copilot AI review requested due to automatic review settings July 15, 2026 15:38
@myheart521
myheart521 merged commit be5374e into main Jul 15, 2026
4 checks passed

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR introduces the initial WAMProbe “counterfactual MVP” as a release-candidate-quality package: a dependency-free core with analytic benchmarks/metrics, capability-aware adapter protocols (including an isolated StarWAM adapter wrapper), deterministic caching + reporting, and a documentation/reproducibility/release-evidence pipeline.

Changes:

  • Adds dependency-free toy benchmarks (PointMass-2D, BlockPush-2D, Gripper-Catch), metrics/statistics utilities, and evaluation/reporting plumbing with deterministic artifacts and caches.
  • Adds typed public API contracts (capabilities, adapters, robotics/manipulation types), plus an isolated StarWAM adapter contract and supporting pinned manifests.
  • Adds release candidate workflows, schemas/examples, documentation site + link/schema validation, and committed evidence/reproducibility artifacts.

Reviewed changes

Copilot reviewed 131 out of 137 changed files in this pull request and generated 3 comments.

Show a summary per file
File Description
tests/test_video_control_study.py Tests PSNR/global-SSIM diagnostics and the video-vs-control counterexample study + CLI report writing.
tests/test_statistics.py Tests deterministic bootstrap summaries and paired metric comparisons.
tests/test_starwam_adapter.py Tests StarWAM adapter capability declaration, typing, close behavior, and horizon validation.
tests/test_robotics_api.py Tests robotics datamodel validation and stable digests/serialization.
tests/test_repository_validation.py Ensures repo-wide schema + markdown link validation script runs in tests.
tests/test_release_audit.py Tests release audit logic and CLI generation for candidate artifacts/evidence.
tests/test_ranking_metrics.py Tests candidate-ranking correlation metric behavior and input validation.
tests/test_pointmass.py Tests PointMass rollout correctness and suite generation invariants.
tests/test_metrics.py Tests core metrics (state FDE, action dependence permutation) on toy data.
tests/test_manipulation_benchmarks.py Tests BlockPush/GripperCatch dynamics, observability, scaling, and baseline profiles.
tests/test_libero_cf.py Tests LIBERO-CF-Mini manifest parsing, selection semantics, and checksum validation.
tests/test_import.py Updates version exposure test to 0.1.0rc1.
tests/test_experiments.py Tests experiment-report analysis and CLI report emission using fixtures.
tests/test_evaluation.py Tests baseline ordering and determinism for evaluation outputs.
tests/test_dataset_cache.py Tests intervention JSONL IO, prediction cache behavior, and corruption detection.
tests/test_closed_loop_study.py Tests closed-loop replanning study behavior, determinism, and CLI outputs.
tests/test_cli.py Tests demo/report/compare/dataset CLI commands and deterministic cache hits.
tests/test_artifacts.py Tests action-prediction artifact schema, cache keys, redaction of raw RGB, and validation.
tests/test_api.py Tests core typed API dataclasses and validation constraints.
src/wamprobe/video_metrics.py Implements dependency-free PSNR and global SSIM diagnostics + invert transform.
src/wamprobe/stats.py Implements deterministic context-block bootstrap summaries and paired comparisons.
src/wamprobe/manipulation_evaluation.py Adds manipulation benchmark evaluator and per-context metric computation.
src/wamprobe/counterfactual.py Adds counterfactual scoring and deterministic JSON artifact writer.
src/wamprobe/benchmarks/rendering.py Adds dependency-free rasterizer for toy manipulation RGB observations.
src/wamprobe/benchmarks/pointmass.py Adds analytic PointMass-2D benchmark and deterministic suite generation.
src/wamprobe/benchmarks/gripper_catch.py Adds analytic GripperCatch benchmark with attachment semantics and rendering.
src/wamprobe/benchmarks/blockpush.py Adds analytic BlockPush benchmark with pre-contact/contact phases and rendering.
src/wamprobe/benchmarks/init.py Exposes built-in benchmarks from package namespace.
src/wamprobe/api/types.py Adds typed 2D context/action/trajectory/suite data structures.
src/wamprobe/api/robotics.py Adds typed RGB/observation/action prediction + stable content digests.
src/wamprobe/api/model.py Adds typed adapter protocols for future- and action-prediction adapters.
src/wamprobe/api/manipulation.py Adds typed manipulation benchmark/action/state contracts and adapter protocol.
src/wamprobe/api/errors.py Adds project exception hierarchy for validation/capability errors.
src/wamprobe/api/capabilities.py Adds capability manifest datamodel and enum for future representations.
src/wamprobe/api/init.py Re-exports the public API surface.
src/wamprobe/adapters/starwam.py Adds StarWAM adapter wrapper and pinned release contract validation.
src/wamprobe/adapters/manipulation.py Adds manipulation baseline adapters (oracle/noisy/copy-last/wrong/action-agnostic).
src/wamprobe/adapters/baselines.py Adds PointMass state baseline adapters (oracle/noisy/copy-last/wrong/action-agnostic).
src/wamprobe/adapters/init.py Exposes built-in adapters from package namespace.
src/wamprobe/main.py Enables python -m wamprobe CLI invocation.
src/wamprobe/init.py Updates package description and version to 0.1.0rc1.
SECURITY.md Adds a security policy and vulnerability reporting guidance.
schemas/result-v0.1.schema.json Adds Draft 2020-12 schema for evaluation summary outputs.
schemas/intervention-v0.1.schema.json Adds Draft 2020-12 schema for intervention JSONL records.
schemas/examples/intervention-pointmass-v0.1.json Adds canonical example instance for pointmass intervention record.
schemas/examples/intervention-manipulation-v0.1.json Adds canonical example instance for manipulation intervention record.
schemas/capability-v0.1.schema.json Adds Draft 2020-12 schema for model capability manifests.
schemas/.gitkeep Removes placeholder now that schemas directory is populated.
release/README.md Documents release-candidate procedure and public release boundary.
release/evidence-manifest-v0.1.schema.json Adds schema for the release evidence manifest.
release/evidence-manifest-v0.1.json Adds pinned evidence manifest referencing committed and external artifacts.
pyproject.toml Updates package metadata, scripts, dev deps, build include/exclude, and strict tool config.
paper/references.bib Adds technical report bibliography entries.
paper/README.md Documents how to compile the Overleaf-ready technical report.
paper/CLAIMS_CHECKLIST.md Adds claim↔evidence checklist to keep paper claims aligned with artifacts.
mkdocs.yml Adds strict MkDocs site configuration and navigation.
examples/video-control-study/video-control-study.md Adds committed markdown report for video fidelity vs control study.
examples/video-control-study/video-control-study.json Adds committed JSON report for video fidelity vs control study.
examples/closed-loop-study/closed-loop-study.md Adds committed markdown report for toy closed-loop replanning study.
environments/starwam/README.md Documents isolated StarWAM environment setup and verified preflight/matrix steps.
environments/starwam/patches/0001-safe-inference-loads.patch Adds upstream patch to restrict PyTorch loads to weights-only where verified safe.
environments/libero/README.md Documents LIBERO-CF-Mini generator workflow and restore/validation invariants.
docs/rfcs/0002-counterfactual-metrics.md Adds RFC for counterfactual-first metric design and sanity checks.
docs/rfcs/0001-scope-and-capabilities.md Adds RFC for scope and capability-aware adapter routing.
docs/reproducibility/REPRODUCIBILITY.md Adds tiered reproduction guide and canonical hashes for committed studies.
docs/reproducibility/GPU_NIGHTLY.md Adds guide for provisioning/running the disabled-by-default GPU nightly runner.
docs/reproducibility/EXTERNAL_REPRODUCTION_TEMPLATE.md Adds template for independent reproduction reports.
docs/models/STARWAM.md Adds StarWAM adapter model card, capabilities, evaluation summary, and limitations.
docs/metrics/CORE_METRICS.md Adds metric cards, applicability rules, and baseline expectations.
docs/index.md Adds documentation landing page describing WAMProbe and pointers to key docs.
docs/experiments/TOY_CLOSED_LOOP_V0.1.md Adds experiment card describing the toy closed-loop study protocol/results.
docs/experiments/STARWAM_LIBERO_CF_MINI_V0.1.md Adds experiment card summarizing StarWAM × LIBERO-CF-Mini matrix results and skips.
docs/benchmarks/TOY_BENCHMARKS.md Adds benchmark card for toy analytic benchmarks and protocols.
docs/benchmarks/LIBERO_CF_MINI.md Adds benchmark card for LIBERO-CF-Mini pilot and verification procedure.
docs/benchmarks/libero_cf_mini_verified_v0.1.json Adds structured verification record for LIBERO-CF-Mini pilot run.
CONTRIBUTING.md Adds contributor setup + validation checklist and contribution requirements.
configs/models/upstream_models.json Adds pinned upstream model store manifest (paths/sizes/hashes/licensing/status).
configs/models/starwam-capabilities-v0.1.json Adds pinned capability manifest for StarWAM adapter.
configs/models/.gitkeep Removes placeholder now that model configs exist.
configs/benchmarks/libero_cf_mini_v0.1.json Adds pinned LIBERO-CF-Mini task manifest with checksums and metadata.
CODE_OF_CONDUCT.md Adds code of conduct for repo interactions.
CITATION.cff Adds citation metadata for the software release.
checkpoints/README.md Documents local model store rules and hash verification via wamprobe doctor.
CHANGELOG.md Adds changelog and documents the 0.1.0rc1 candidate contents.
.gitignore Updates ignores for docs build output and local artifacts (models/runs/vendor/paper outputs).
.github/workflows/release-candidate.yml Adds manual release-candidate build/audit/smoke workflow with attestations.
.github/workflows/publish-pypi.yml Adds guarded manual PyPI publish workflow using Trusted Publishing and tag confirmation.
.github/workflows/gpu-nightly.yml Adds disabled-by-default GPU nightly workflow for self-hosted runner evidence generation.
.github/workflows/docs.yml Adds strict docs build + GitHub Pages deployment workflow.
.github/workflows/ci.yml Adds CI quality matrix (ruff/mypy/tests/docs/schema+links) and candidate build job.
.github/PULL_REQUEST_TEMPLATE.md Adds PR template with validation checklist and metric/adapter/benchmark checklist.
.github/ISSUE_TEMPLATE/metric_proposal.yml Adds structured issue template for proposing new metrics.
.github/ISSUE_TEMPLATE/config.yml Adds issue template config with a security contact link.
.github/ISSUE_TEMPLATE/bug_report.yml Adds structured bug report template.
.github/ISSUE_TEMPLATE/benchmark_proposal.yml Adds structured benchmark proposal template.
.github/ISSUE_TEMPLATE/adapter_proposal.yml Adds structured adapter proposal template.
.github/dependabot.yml Adds monthly Dependabot updates for pip + GitHub Actions.
.gitattributes Adds whitespace handling for patch files.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +10 to +14
"vec2": {
"type": "array",
"prefixItems": [{"type": "number"}, {"type": "number"}],
"items": false
},
Comment thread CITATION.cff
Comment on lines +3 to +6
title: "WAMProbe: Counterfactual Evaluation for World Action Models"
type: software
version: 0.1.0-rc.1
license: Apache-2.0
Comment on lines +14 to +15
"type": "object",
"required": [
@myheart521
myheart521 deleted the agent/initial-mvp branch July 15, 2026 16:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants