Objective
Test one genuinely independent real provider on complete physical object/session groups, separating provider competence from downstream guarded BayesianPhysTwin value. Controlled synthetic graph and Prob4D-to-BayesianPhysTwin results are mechanism evidence only; they do not establish real-provider competence or unseen-object benefit.
Stage 0: support feasibility before residual outcomes
Before opening source residuals or any target outcome, publish a content-addressed ProviderSupportFeasibilityV1 decision for the exact provider/model/runtime identity. It must bind:
- complete development, calibration, target, and excluded object/session rosters;
- representation, units, metric-frame prior, causal cutoff, windows, seeds, and immutable model set;
- required streams, identities, masks, timestamps, camera geometry, and physical-query mapping;
- minimum complete-stream support and technical-failure policy; and
- exact unsupported-case fallback.
If complete-stream support fails, stop that provider version and report the negative feasibility result. Do not fit covariance, inspect target outcomes, substitute partial streams, or rescue it with a different graph/anchor/model under the same identity.
Frozen comparison arms
Select at most one Prob4D candidate using development/calibration units only. Score exactly these main arms on the same held-out units:
- unchanged physical fallback;
- simple direct visual observation or
last_residual comparator;
- one source-selected Prob4D candidate with explicit complete joint gauge uncertainty;
- exact physical fallback for every unsupported, invalid, or rejected candidate.
Metric-assisted variants must be labelled sensor-assisted and may not be compared as though they used the same information as an unanchored arm. Additional graph, association, semantic, or covariance variants are out of scope unless this frozen result localizes a specific failure that requires a new protocol.
Freeze before target access
Bind exact Prob4D, BayesianPhysTwin, observation-provider, model, environment, and numerical identities; source/calibration-only reliability and covariance calibration; the BayesianPhysTwin guard; horizons; physical queries; group weighting; bootstrap/randomization seeds; technical failures; and exact fallback. Frames, points, tracks, views, and taxels are nested observations, not independent groups.
Separate endpoints
Provider competence
Report by independent object/session and horizon:
- support and technical-failure accounting;
- point/endpoint or valid identity-based track error;
- seam/drift diagnostics where applicable;
- 50/90/95% coverage, proper score, normalized NEES, and full covariance width;
- identity retention/precision and selective risk; and
- worst-group coverage shortfall.
Downstream guarded query
Report separately:
- deployed physical-query proper score and RMSE;
- object/session-clustered paired interval versus both comparators;
- accepted, rejected, and exact-fallback counts;
- harmful accepted updates and worst-group regret;
- accepted-update coverage and width; and
- provider competence versus downstream benefit without allowing either to rescue the other.
Completion rule
Close with one complete frozen outcome:
- provider competence and guarded BayesianPhysTwin criteria both pass on the independent cohort; or
- a valid negative/bounded result localizes failure to support, means/identities, gauge/dependence, calibration shift, object/session transfer, or query identifiability.
A tie retains the simple comparator and physical fallback. Do not retune on the opened target cohort. A downstream result cannot rescue failed support or provider competence, and optional Causal4D results cannot rescue either upstream claim.
Objective
Test one genuinely independent real provider on complete physical object/session groups, separating provider competence from downstream guarded BayesianPhysTwin value. Controlled synthetic graph and Prob4D-to-BayesianPhysTwin results are mechanism evidence only; they do not establish real-provider competence or unseen-object benefit.
Stage 0: support feasibility before residual outcomes
Before opening source residuals or any target outcome, publish a content-addressed
ProviderSupportFeasibilityV1decision for the exact provider/model/runtime identity. It must bind:If complete-stream support fails, stop that provider version and report the negative feasibility result. Do not fit covariance, inspect target outcomes, substitute partial streams, or rescue it with a different graph/anchor/model under the same identity.
Frozen comparison arms
Select at most one Prob4D candidate using development/calibration units only. Score exactly these main arms on the same held-out units:
last_residualcomparator;Metric-assisted variants must be labelled sensor-assisted and may not be compared as though they used the same information as an unanchored arm. Additional graph, association, semantic, or covariance variants are out of scope unless this frozen result localizes a specific failure that requires a new protocol.
Freeze before target access
Bind exact Prob4D, BayesianPhysTwin, observation-provider, model, environment, and numerical identities; source/calibration-only reliability and covariance calibration; the BayesianPhysTwin guard; horizons; physical queries; group weighting; bootstrap/randomization seeds; technical failures; and exact fallback. Frames, points, tracks, views, and taxels are nested observations, not independent groups.
Separate endpoints
Provider competence
Report by independent object/session and horizon:
Downstream guarded query
Report separately:
Completion rule
Close with one complete frozen outcome:
A tie retains the simple comparator and physical fallback. Do not retune on the opened target cohort. A downstream result cannot rescue failed support or provider competence, and optional Causal4D results cannot rescue either upstream claim.