Evidence-first calibration for robotics sensors.
Calibrate solid-state LiDAR and multi-sensor rigs with holdouts, observability, and provenance you can reproduce.
Calibrex is an open-source, ROS-independent Python toolkit that turns candidate extrinsics, time offsets, and trajectories into evidence backed by holdout metrics, known-bad controls, observability checks, and reproducible provenance.
Use it when a transform must be more than a plausible number: evaluate native or adapter-produced LiDAR, camera, IMU, radar, RGB-D, hand-eye, and robot-world calibration without reducing the verdict to optimizer convergence or a single training residual.
Real TIERS Indoor02 moving-platform replay: a Velodyne VLP-16 motion map supports online Ouster OS1 calibration, with 106 of 108 batches accepted by holdout gates.
Every gallery asset is generated from public raw data. Dataset source,
protocol, parameters, and digests are recorded in
readme-gif-gallery.json.
Start with the small Livox sample if you want a fast, no-ROS check. It writes a
calibrated result.yaml, report.html, evidence sidecars, and a verified
provenance bundle:
python -m pip install .
calibrex demo livox-evidence \
--output-dir outputs/solid-state-livox-demo \
--json
calibrex validate outputs/solid-state-livox-demo/result.yaml
calibrex verify outputs/solid-state-livox-demo/bundle.jsonFor real per-point timing and clock-offset profiling, use the public TIERS
VLP-16 ↔ Livox Horizon bag. The bag is about 7.18 GB and is never downloaded
automatically; place it at data/public/tiers_lidars_cali/LidarsCali.bag as
described in the public-dataset tutorial.
No ROS installation is required:
python -m pip install -e ".[rosbag1-lz4]"
calibrex public-datasets show tiers_livox_lidars_cali --json
calibrex inspect data/public/tiers_lidars_cali/LidarsCali.bag \
--type rosbag1 --json
calibrex continuous-time-lidar-pair \
examples/public_datasets/tiers_livox_lidars_cali/online_continuous_time_config.yaml \
--output outputs/solid-state-tiers/continuous_time_lidar_pair.yaml \
--json
calibrex validate \
outputs/solid-state-tiers/continuous_time_lidar_pair.yaml \
--kind continuous-time-lidar-pairThe timed artifact records the config and bag SHA-256, transform convention, candidate offsets, train/holdout RMSE, and fixed-odometry provenance. On the checked public run, the selected offset was +40 ms and train RMSE changed from 0.0957 m to 0.0240 m, with 0.0242 m holdout RMSE. This is an algorithmic estimate under the declared holdout: the public sequence has no independent clock ground truth.
For an absolute solver-correctness gate, run the deterministic synthetic truth benchmark. It recovers a known extrinsic and +30 ms clock offset, then rejects a fixed-clock known-bad control:
python tools/run_solid_state_synthetic_benchmark.py \
--output outputs/solid-state-synthetic-benchmark.yaml \
--markdown-output outputs/solid-state-synthetic-benchmark.md \
--enforce
calibrex validate outputs/solid-state-synthetic-benchmark.yaml \
--kind solid-state-synthetic-benchmarkThis synthetic gate verifies solver mechanics, not real-sensor accuracy. An absolute real-sensor accuracy claim still requires independently measured extrinsics and clock truth. The checked result is available as a schema-valid YAML artifact with a compact Markdown report.
The current solid-state evaluation scope is public-data-only. Use the public benchmark runbook to compare paired solver variants on the same capture windows, temporal holdouts, sampling seeds, and known-bad/control fixtures:
calibrex validate \
examples/public_datasets/solid_state_cross_dataset_benchmark.yaml \
--kind solid-state-cross-dataset-benchmark-config
python tools/run_solid_state_benchmark_replicates.py \
examples/public_datasets/solid_state_cross_dataset_benchmark.yaml \
--output-root outputs/solid_state_benchmark_v02_replicates \
--output-spec examples/public_datasets/solid_state_cross_dataset_benchmark_v02.yaml \
--split-id middle_holdout --split-id late_holdout \
--seed 0 --seed 17
python tools/run_solid_state_benchmark_replicates.py \
examples/public_datasets/solid_state_cross_dataset_benchmark.yaml \
--output-root outputs/solid_state_benchmark_v03_replicates \
--output-spec examples/public_datasets/solid_state_cross_dataset_benchmark_v03.yaml \
--append-spec examples/public_datasets/solid_state_cross_dataset_benchmark_v02.yaml \
--split-id early_holdout --split-id middle_holdout --split-id late_holdout \
--seed 0 --seed 17 --seed 42
python tools/run_solid_state_cross_dataset_benchmark.py \
examples/public_datasets/solid_state_cross_dataset_benchmark_v03.yaml \
--output outputs/solid_state_cross_dataset_benchmark_v03.yaml \
--markdown-output outputs/solid_state_cross_dataset_benchmark_v03.md \
--html-output outputs/solid_state_cross_dataset_benchmark_v03.html
calibrex validate outputs/solid_state_cross_dataset_benchmark_v03.yaml \
--kind solid-state-cross-dataset-benchmarkThis gate can support reproducible temporal-holdout comparisons, observability, point-time/deskew evidence, convergence/failure analysis, and known-bad detection. It cannot establish absolute extrinsic or clock accuracy because the public recordings do not provide independent metrology.
The independent physical-metrology packet remains an optional future path for users who can obtain surveyed references. Its checked artifact is intentionally planned, not a result: YAML template and Markdown report.
Documentation · Five-minute quickstart · Solid-state quickstart · Calibration methods · Public-data demos
Render and validate a committed result without ROS or a dataset download:
git clone https://github.com/rsasaki0109/Calibrex.git
cd Calibrex
python -m pip install .
calibrex validate examples/precomputed/result.yaml --json
calibrex render examples/precomputed/result.yaml \
--format evidence-card \
--output outputs/quickstart/evidence-card.svg
calibrex render examples/precomputed/result.yaml \
--output-dir outputs/quickstartOpen outputs/quickstart/report.html, then inspect or verify the
schema-valid sidecars. To recompute evidence from a declared public sample:
calibrex demo livox-evidence \
--output-dir outputs/livox_horizon_horizon_pcd_sampleBefore configuring a solve, diagnose a recording and save the result for review or CI:
calibrex doctor recording.mcap \
--output outputs/doctor.json \
--json
calibrex validate outputs/doctor.jsondoctor infers supported dataset types, reports missing optional dependencies,
data coverage and degeneracy warnings, and suggests compatible evidence
workflows. Running it without a path retains the lightweight environment check.
Run the same evidence gates on every calibration change:
- uses: actions/checkout@v4
- uses: rsasaki0109/Calibrex@v0.4.0
with:
candidate: calibration/candidate.yaml
baseline: calibration/baseline.yamlThe action writes a GitHub Step Summary, fails on FAIL or INCONCLUSIVE by
default, and exposes schema-valid evidence, comparison, SVG, and
calibration-ci.json artifacts. See Calibration CI.
| Calibration path | Native solve | Evidence | Public example | Maturity |
|---|---|---|---|---|
| LiDAR ↔ LiDAR | ✅ | ✅ | ✅ | 🟢 Alpha |
| Camera ↔ LiDAR | ✅ | ✅ | ✅ | 🟡 Experimental |
Hand-eye AX=XB |
✅ | ✅ | ✅ | 🟢 Alpha |
Robot-world AX=YB |
✅ | ✅ | ✅ | 🟢 Alpha |
| LiDAR ↔ IMU | — | ✅ | ✅ | 🟡 Evidence |
| Radar extrinsic | ✅ | ✅ | ✅ | 🟡 Experimental |
| RGB-D joint SLAC | ✅ | ✅ | ✅ | 🟡 Experimental |
| Open3D / external tools | Adapter | ✅ | ✅ | 🔵 Adapter |
Calibrex evaluates all 13 native hand-eye methods and seven OpenCV 4 variants
on the ETHZ ASL real robot-arm dataset.
Every method sees the same five digest-locked absolute-pose splits. AX=XB and
AX=YB are separate equation families and are never ranked together.
| Method | Holdout rotation RMSE deg ↓ | Holdout translation RMSE mm ↓ | Failure rate | Runtime s |
|---|---|---|---|---|
| Calibrex Park-Martin | 0.867227 [0.759962, 0.958155] | 13.8759 [11.7678, 15.8266] | 0.0% | 0.153451 |
| Calibrex Tsai-Lenz | 0.875291 [0.77004, 0.967317] | 13.8652 [11.7326, 15.8293] | 0.0% | 0.126921 |
| Calibrex Daniilidis | 0.868969 [0.761845, 0.960775] | 13.8212 [11.9121, 15.6887] | 0.0% | 0.322403 |
| Calibrex Andreff | 0.867199 [0.759958, 0.958223] | 13.879 [11.7674, 15.8309] | 0.0% | 0.334728 |
| Calibrex Shiu-Ahmad | 1.6191 [0.787831, 3.10277] | 31.9219 [12.7334, 52.761] | 0.0% | 14.1537 |
| Calibrex Chou-Kamel | 0.867205 [0.759953, 0.958241] | 13.8795 [11.7684, 15.8315] | 0.0% | 0.129661 |
| Calibrex Horaud-Dornaika | 0.890804 [0.782116, 0.996268] | 14.1434 [11.8641, 16.1088] | 0.0% | 0.139859 |
| Calibrex H-D nonlinear | 0.885789 [0.777979, 0.991045] | 14.0634 [11.8408, 16.0027] | 0.0% | 19.4512 |
| OpenCV Tsai | 0.871872 [0.769051, 0.961725] | 13.8681 [11.8255, 15.6992] | 0.0% | 0.0368567 |
| OpenCV Park | 0.867225 [0.759969, 0.958144] | 13.8754 [11.7667, 15.8261] | 0.0% | 0.0280751 |
| OpenCV Horaud | 0.867203 [0.759961, 0.958229] | 13.879 [11.7674, 15.831] | 0.0% | 0.0248639 |
| OpenCV Andreff | 0.868623 [0.763653, 0.959819] | 15.4855 [13.6068, 17.338] | 0.0% | 0.0327896 |
| OpenCV Daniilidis | 0.871005 [0.764385, 0.963618] | 13.9265 [11.9635, 15.6961] | 0.0% | 0.0284915 |
| Method | Holdout rotation RMSE deg ↓ | Holdout translation RMSE mm ↓ | Failure rate | Runtime s |
|---|---|---|---|---|
| Calibrex Shah | 0.624739 [0.559308, 0.688861] | 10.7793 [9.53397, 11.9952] | 0.0% | 0.0206115 |
| Calibrex Li-Wang-Wu | 0.627131 [0.561467, 0.692796] | 19.651 [15.2668, 23.3423] | 0.0% | 0.0321106 |
| Calibrex Dornaika-Horaud | 0.624742 [0.559311, 0.688865] | 10.7793 [9.53399, 11.9952] | 0.0% | 0.0286939 |
| Calibrex Zhuang-Roth-Sudhakar | 0.651861 [0.569557, 0.733996] | 10.961 [9.65144, 12.1956] | 0.0% | 0.0213057 |
| Calibrex D-H nonlinear | 0.6265 [0.560987, 0.692012] | 10.7752 [9.52577, 12.0161] | 0.0% | 0.217955 |
| OpenCV Shah | 0.624739 [0.559308, 0.688861] | 10.7793 [9.53397, 11.9952] | 0.0% | 0.00343018 |
| OpenCV Li | 0.627131 [0.561467, 0.692796] | 19.651 [15.2668, 23.3423] | 0.0% | 0.00570322 |
These are scoped consistency results, not blanket accuracy claims: the archive does not provide an accepted ground-truth extrinsic. The tables report closure RMSE on untouched holdout poses, retain failures in the denominator, and show 95% bootstrap intervals across splits. See the full protocol, citations, limitations, and reproducible provenance.
The checked v0.3 solid-state protocol runs nine paired replicates per dataset across three temporal holdout boundaries and three deterministic sampling seeds. Lower holdout RMSE is better; positive improvement means adaptive is better than the uniform baseline.
| Public pair/control | Replicates | Adaptive win rate | Mean improvement (95% CI) | Outcome |
|---|---|---|---|---|
| AgRob Modular-e Livox MID-70 ↔ RS-LiDAR | 9 | 0.556 | −1.34% (−10.16, 6.68) | mixed; adaptive 5/9 |
| TIERS LidarsCali VLP-16 ↔ Livox Horizon | 9 | 1.00 | +78.76% (+76.00, 81.73) | adaptive 9/9 |
| AIST GLIM identity control | 9 | 1.00 | +65.67% (+64.18, 66.95) | adaptive 9/9 |
Across 27 scored replicates, adaptive wins 23/27 with mean improvement of 47.70% and bootstrap 95% CI [34.43, 59.99]%. AgRob is retained as a useful counterexample rather than hidden: its result varies across split and seed. The full checked summary is the v0.3 benchmark report. These are ground-truth-free temporal-holdout results, not a universal SOTA claim. Reproduce the runs and inspect the public-dataset protocol.
What is implemented in the current alpha?
- typed config, result, comparison, protocol, policy, and evidence schemas
- offline and online/streaming LiDAR calibration for rosbag1, rosbag2, and MCAP
- motion compensation, per-point deskew, and trajectory evidence
- targetless camera-LiDAR mutual information and online monitoring
- native point-to-point, point-to-plane, hand-eye, and robot-world baselines
- backend-neutral joint SLAC with typed Schur pose elimination
- radar velocity, LiDAR-IMU rotation, and temporal-offset evidence
- typed external-run artifacts, a Kalibr camchain importer, and Koide/Open3D, ROS, and Autoware adapter boundaries
- a frozen, SHA-bound full-scale KITTI reference-vs-known-bad falsification runner
See the calibration methods and changelog for solver-level detail and limitations.
The card below is not a hand-authored mockup. It is rendered from a
schema-valid result.yaml, bound to the
source file by SHA-256, and committed with machine-readable provenance inside
the SVG.
Generate the same artifact from any Calibrex result:
calibrex render result.yaml \
--format evidence-card \
--output evidence-card.svgMost calibration tools stop after producing a transform. Calibrex asks the next question: what evidence would falsify this transform?
| A matrix gives you | Calibrex adds |
|---|---|
| One estimated transform | Candidate, reference, and selected estimates kept separate |
| One training residual | Train/holdout metrics and temporal stability |
| Optimizer convergence | Known-bad perturbation challenges |
| A covariance matrix | Rank, weak directions, and degeneracy warnings |
| A screenshot | Schema-valid artifacts with input and producer provenance |
| A result file | A digest-locked evidence bundle that can be verified later |
This generated table compares the validated quickstart result with a declared 5° yaw / 0.15 m known-bad control. The amber protocol state is intentional: these compact example results do not declare a shared evidence protocol, so Calibrex shows the metric deltas but refuses to call the comparison compatible.
calibrex compare \
examples/precomputed/result.yaml \
examples/precomputed/known_bad_result.yaml \
--format evidence-table \
--left-label accepted \
--right-label "known bad" \
--output comparison-table.svgflowchart LR
A[Sensor data] --> B[Candidate calibration]
B --> C{Evidence gates}
C -->|Holdout| D[Generalization]
C -->|Known-bad| E[Falsification]
C -->|Observability| F[Weak directions]
D --> G[PASS / WARN / FAIL]
E --> G
F --> G
G --> H[HTML + schema-valid sidecars]
Calibrex does not turn every run green. Weak excitation and failed controls are reported as evidence, not hidden as demo noise.
| v0.3 evidence on TIERS Indoor02 | Observation | Verdict |
|---|---|---|
| Native-deskew identity selftest | translation error 9.5 cm → 4.3 cm | Improved |
| Cross-segment trajectory drift proxy | 0.143 m over 42 s | ✅ PASS |
| Two-pass odometry trajectory gate | 0.366 m | ❌ FAIL |
| Anchored temporal +50 ms injection | offset tracked exactly | ✅ Detected |
| LiDAR-IMU rotation evidence | 3.56 deg/s holdout RMSE |
The failure is part of the product: gates refuse to certify what the available data cannot falsify.
Calibrex is currently distributed from source. The quickstart above installs the core package; development and optional backends remain explicit:
python -m pip install -e ".[dev]"
python -m pip install -e ".[open3d]"The core package stays ROS-independent. ROS bags are read through typed data adapters, and GPL or ecosystem-specific tools stay behind optional adapter or subprocess boundaries.
result.yaml calibrated values + run provenance
├── report.html portable human-readable report
├── metrics.json train / holdout measurements
├── evidence.json protocols, summaries, and known-bad cases
├── assessment.json PASS / WARN / FAIL policy decision
├── observability.json rank and weak directions
├── degeneracy.json failure modes and limitations
├── protocol.json declared evaluation contract
├── policy.json falsification thresholds
├── transforms.json typed transform provenance
├── bundle.json artifact SHA-256 manifest
└── verification.json materialized bundle verification
calibrex evidence result.yaml --output evidence.json
calibrex assess evidence.json
calibrex verify bundle.json
calibrex compare reference.yaml candidate.yaml --enforce-compatiblerender never claims to recompute metrics. Cached inputs remain marked as
cached, raw-input verification remains visible, and incompatible protocols are
not silently ranked together.
The evidence card is deterministic and bound to its source result:
calibrex render examples/precomputed/result.yaml \
--format evidence-card \
--output docs/assets/readme-evidence-card.svg
calibrex compare \
examples/precomputed/result.yaml \
examples/precomputed/known_bad_result.yaml \
--format evidence-table \
--left-label accepted \
--right-label "known bad" \
--output docs/assets/readme-comparison-table.svgThe public-data GIF gallery has its own schema-valid manifest:
python tools/generate_calibration_evidence_gif.py --readme-galleryTests fail if the committed evidence card drifts from its validated source.
- Keep the sensor-agnostic core ROS-independent.
- Treat dataset calibration as reference evidence, not absolute truth.
- Require evaluation for calibration behavior changes.
- Record provenance for every generated result.
- Keep GPL and ecosystem tools behind adapters or subprocess boundaries.
- Preserve schema stability and evaluation quality over short-term convenience.
- Documentation site
- SLAC concept
- Calibration methods
- Solid-state LiDAR calibration
- Frame conventions
- Public datasets
- Open3D adapter
- LiDAR-camera adapter
- License boundaries
- Support · Contributing · Security
Apache-2.0. If you use Calibrex in research, see
CITATION.cff.







