CI guard for routing graph rebuilds. You define a suite of fixed test routes; after every rebuild the tool queries your routing endpoint, compares the results against a committed baseline, and fails the build when distance, duration, or geometry drifts beyond tolerance.
Ships as both a GitHub Action and a standalone CLI. Supports OSRM and Valhalla.
A routing graph rebuild is a silent deploy. You pull a fresh OSM extract, run
osrm-extract/osrm-customize or valhalla_build_tiles, and ship the result. Nothing
fails. Then, a week later, dispatch notices that trucks are being sent over a
weight-limited bridge, or that the depot-to-depot trunk run got eleven minutes slower.
The usual causes are all invisible at build time:
- A profile edit changed a speed class and every motorway route shifted.
- A restriction (weight limit, LEZ, turn restriction, access tag) stopped being imported and the router happily plans through it.
- The OSM extract was truncated or the bounding box moved, so part of the network is now unreachable and the router takes a 40 km detour.
- A road was retagged upstream and the graph fragmented at that point.
None of these produce an error. They produce a different route. This tool turns "different route" into a red build.
For each case in your suite, against your committed baseline:
| Check | Catches |
|---|---|
| Distance drift | Detours, fragmentation, lost shortcuts |
| Duration drift | Speed profile changes, penalty changes |
| Geometry drift | A route with the same distance and duration that now takes completely different roads |
| Routability flip | A route that appeared or disappeared entirely |
| Absolute bounds | max_duration, min_distance, … — SLAs, not just history |
| Required segments | "This run must use the A14" |
| Forbidden segments | "This truck must never touch Fen Causeway" |
| Turn count | Manoeuvre churn from a fragmented graph |
| Must-NOT-be-routable | A restriction that is supposed to block a route still does |
That last one deserves emphasis. expect: {routable: false} is how you test that a
restriction is being honoured. If a graph rebuild silently drops a weight limit, every
drift-based check still passes — the route just becomes available. Only an explicit
unroutable expectation catches it.
Geometry drift is the other one that is hard to get any other way. A rebuild can reroute a leg onto a parallel road with almost identical length and travel time. Distance and duration both pass; the route is wrong. Comparing shapes catches it.
routes.yml ──┐
├─► runner ──► engine client ──HTTP──► OSRM / Valhalla
baseline.json┘ │ (or fake, for tests)
▼
compare ──► tolerances ─┐
│ ├─► checks ──► severity
└─► geometry metric ──┘
│
┌──────────────────────┼─────────────────────┐
▼ ▼ ▼ ▼ ▼
terminal step summary annotations JUnit JSON
The pipeline is deliberately split so that everything worth testing is pure:
suite.py— parses and validates the YAML suite. Errors name the offending case.baseline.py— loads, saves, and diffs the committed baseline. Schema-versioned.engines/— the only code that speaks HTTP. Everything above it works on a normalisedRouteResult, which is why the entire test suite runs offline against a fake client.geometry.py— pure-Python polyline5/6 codec plus discrete Fréchet and Hausdorff distance.tolerance.py— two-tier (warn / fail) thresholds, absolute or relative.compare.py— pure decision logic. No I/O, no clock, no network.runner.py— orchestration; owns the only two impure things (engine calls, time).reporting/— renderers for terminal, markdown, annotations, JUnit, JSON.
Comparing two routes needs a shape metric, not an index-by-index one — a rebuilt graph can return the same road with a different number of vertices, and zipping the coordinate lists together would call that a regression.
Both metrics are implemented in pure Python (no shapely, no scipy — nothing that could lack a wheel on one of the supported Python versions) and report metres via an equirectangular projection around the mean latitude of the two curves. Over the length of a test route that approximation is sub-metre accurate, and geometry tolerances are set in the tens of metres.
- Discrete Fréchet (default) — the "dog walking" distance: the shortest leash that lets a walker traverse one curve and a dog the other, both only moving forward. It respects traversal order, so a route that uses the same roads in a different sequence is caught. Because the discrete variant couples vertices, it is somewhat sensitive to re-sampling density.
- Hausdorff — each vertex measured to the nearest point on a segment of the other
curve, so it is genuinely insensitive to re-sampling. But it ignores order: a
reversed route scores zero. Pick it (
--geometry-metric hausdorff) if your engine varies geometry simplification between builds and you only care which roads were used.
Routing graphs move a little on every rebuild. A one-metre distance change is noise.
A single threshold forces you to choose between a noisy build and a blind one, so every
tolerance has two lines: warn ("someone should glance at this") and fail ("do not
ship this graph"). By default only fail breaks the build; --fail-on warn makes
warnings blocking once you have tuned them, and --fail-on never reports without
blocking while you are still calibrating.
This project is not published to PyPI. Clone it and install from the local checkout.
git clone https://github.com/geospatialrouting/route-regression-check.git
cd route-regression-check
python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"Requires Python 3.11 or newer. Runtime dependencies are httpx, click, pyyaml,
pydantic, and rich — all pure-Python.
Both of these work after installing:
route-regression-check --help # console script
python -m route_regression_check --help # module entry pointThe repository ships a complete example: a suite, a JSON fixture standing in for a routing engine, and a baseline. No server required.
$ route-regression-check validate --suite examples/routes.yml --baseline examples/route-baseline.json --strict
suite 'cambridge-freight' is valid: 5 case(s), 4 active.
baseline 'examples/route-baseline.json' is valid: 4 entries.$ route-regression-check update-baseline \
--suite examples/routes.yml \
--baseline examples/route-baseline.json \
--fixture examples/fixtures/cambridge.json
4 case(s) would change in examples/route-baseline.json:
added a14-corridor-with-via: 10233.8m / 705.0s / 2 turns
added depot-to-city-centre: 5421.3m / 612.0s / 2 turns
added hgv-avoids-weight-limited-bridge: 3180.5m / 428.0s / 2 turns
added pedestrian-only-street-is-not-drivable: unroutable
Write these changes? [y/N]: y
wrote examples/route-baseline.jsonCommit that file. It is a reviewed artefact — the pull request that changes it is exactly where a human decides whether a routing change was intended.
A clean run:
$ route-regression-check check --engine osrm --endpoint http://localhost:5000 \
--suite examples/routes.yml --baseline examples/route-baseline.json --warmup 120
╭──────────────── route regression check — osrm ─────────────────╮
│ status case profile Δ dist Δ dur ... │
│ PASS a14-corridor-with-via driving +0m +0s │
│ PASS depot-to-city-centre driving +0m +0s │
│ PASS hgv-avoids-weight-limi... truck +0m +0s │
│ PASS pedestrian-only-street... driving - - │
│ SKIP seasonal-ferry-link driving - - │
│ 5 cases · 4 passed · 0 warned · 0 failed · 1 skipped · 0.41s │
╰────────────────────────────────────────────────────────────────╯
$ echo $?
0And a rebuild that broke something — the A14 leg got rerouted through a residential street, and the city-centre run got slower:
$ route-regression-check check --suite examples/routes.yml \
--baseline examples/route-baseline.json --fixture drifted.json
╭─────────────────────────── route regression check — fake ───────────────────────────╮
│ status case profile Δ dist Δ dur geom detail │
│ FAIL a14-corridor-w driving +216m (+2.1%) +385s (+54.6%) 1423m required │
│ ith-via segment 'A14' │
│ is missing │
│ from the │
│ route; │
│ forbidden │
│ segment 'Green │
│ End Road' │
│ appears in the │
│ route (+3 more)│
│ WARN depot-to-city- driving +0m (+0.0%) +16s (+2.6%) 0m duration │
│ centre drifted +16.0s │
│ (+2.61%) │
│ PASS hgv-avoids-wei truck +0m (+0.0%) +0s (+0.0%) 0m matches │
│ ght-limited-br baseline │
│ PASS pedestrian-onl driving - - - matches │
│ y-street-is-no baseline │
│ SKIP seasonal-ferry driving - - - Ferry is out │
│ -link of service │
│ 5 cases · 2 passed · 1 warned · 1 failed · 1 skipped · 0.00s │
╰─────────────────────────────────────────────────────────────────────────────────────╯
::error title=Route regression%3A a14-corridor-with-via::required segment 'A14' is missing from the route; forbidden segment 'Green End Road' appears in the route; distance drifted +216.2m (+2.11%25) (baseline 10233.8m, allowed +/-204.7m); duration drifted +385.0s (+54.61%25) (baseline 705.0s, allowed +/-70.5s); geometry diverged: frechet distance 1423.3m exceeds 150.0m — the route is following different roads
::warning title=Route regression%3A depot-to-city-centre::duration drifted +16.0s (+2.61%25) (baseline 612.0s, allowed +/-12.2s)
$ echo $?
1The matching step summary:
## Route regression check
**FAIL** — one or more routes regressed beyond tolerance.
Engine `fake` · suite `cambridge-freight` · 0.00s
| total | passed | warned | failed | skipped |
| ----: | -----: | -----: | -----: | ------: |
| 5 | 2 | 1 | 1 | 1 |
| | case | profile | Δ distance | Δ duration | geometry | detail |
| :-: | --- | --- | ---: | ---: | ---: | --- |
| ❌ | `a14-corridor-with-via` | driving | +216m (+2.1%) | +385s (+54.6%) | 1423m | required segment 'A14' is missing from the route; … (+2 more) |
| ⚠️ | `depot-to-city-centre` | driving | +0m (+0.0%) | +16s (+2.6%) | 0m | duration drifted +16.0s (+2.61%) (baseline 612.0s, allowed +/-12.2s) |
| ✅ | `hgv-avoids-weight-limited-bridge` | truck | +0m (+0.0%) | +0s (+0.0%) | 0m | matches baseline |Note the failing case: distance drifted only 2.1%, which on its own is under the fail line. What made it a failure was the required segment vanishing, the forbidden street appearing, and a 1.4 km geometry divergence. A distance-only check would have shipped it.
- name: Route regression check
id: routes
uses: geospatialrouting/route-regression-check@v1
with:
engine: osrm
endpoint: http://localhost:5000
suite: routes/suite.yml
baseline: routes/baseline.json
junit-path: reports/routes.junit.xml
report-path: reports/routes.json
warmup: "120" # the graph is still loading when this step starts
- name: Publish
if: steps.routes.outputs.status != 'fail'
run: ./scripts/publish-graph.shThree ready-to-copy workflows live in examples/workflows/: rebuild-then-check, a
nightly non-blocking staging check, and a matrix over vehicle profiles.
| Input | Default | Description |
|---|---|---|
endpoint |
"" |
Base URL of the routing engine. Required unless fixture is set. |
engine |
osrm |
osrm or valhalla. |
suite |
routes.yml |
Path to the route suite YAML. |
baseline |
route-baseline.json |
Path to the committed baseline JSON. |
fixture |
"" |
Serve canned responses from a JSON file; no network call. |
tolerance-distance |
"" |
Override the distance FAIL threshold. |
tolerance-duration |
"" |
Override the duration FAIL threshold. |
tolerance-geometry |
"" |
Override the geometry FAIL threshold, in metres. |
tolerance-mode |
relative |
How distance/duration overrides are read. |
geometry-metric |
frechet |
frechet or hausdorff. |
fail-on |
fail |
warn, fail, or never. |
annotate |
true |
Emit ::error / ::warning annotations. |
summary |
true |
Write the markdown table to the job step summary. |
junit-path |
"" |
Where to write JUnit XML. Empty disables. |
report-path |
route-regression-report.json |
Where to write the JSON artifact. |
retries |
3 |
Attempts per route request. |
warmup |
0 |
Seconds to wait for the engine to report healthy. |
timeout |
20 |
Per-request timeout in seconds. |
polyline-precision |
5 |
OSRM geometry precision (5 or 6). Valhalla always uses 6. |
units |
kilometers |
Valhalla request units. OSRM ignores this. |
python-version |
3.12 |
Python used to run the check. |
| Output | Description |
|---|---|
status |
pass, warn, or fail. |
failed-count |
Number of cases at FAIL severity. |
warned-count |
Number of cases at WARN severity. |
report-path |
Path to the JSON report, or empty. |
version: 1 # schema version; only 1 is supported
name: cambridge-freight
description: Core freight routes.
profile: driving # default profile for cases that omit one
tolerances: # suite-wide defaults
distance:
mode: relative # 'relative' (fraction) or 'absolute' (metres)
warn: 0.01
fail: 0.05
duration:
mode: relative # 'relative' (fraction) or 'absolute' (seconds)
warn: 0.02
fail: 0.10
geometry:
mode: absolute # geometry is ALWAYS absolute, in metres
warn: 25
fail: 150
cases:
- name: depot-to-city-centre # required; keys the baseline, must be unique
description: Daily trunk run.
profile: truck # overrides the suite default
coordinates: # two or more; extras are via points
- { lat: 52.2100, lon: 0.1200 }
- { lat: 52.2053, lon: 0.1218 }
options: # engine-specific request extras
exclude: motorway
expect:
routable: true # set false to assert a restriction holds
max_distance: 12000 # metres
min_distance: 4000
max_duration: 1800 # seconds
min_duration: 300
required_segments: ["A14"] # case-insensitive substring, name or ref
forbidden_segments: ["Fen Causeway"]
turn_count: 6 # exact
max_turns: 10
tolerances: # per-case overlay; unset fields inherit
distance: { fail: 0.02 }
skip: false
skip_reason: null # required when skip is true| Key | Type | Notes |
|---|---|---|
version |
int | Must be 1. |
name |
string | Shown in reports. |
description |
string | Free text. |
profile |
string | Default profile pushed into cases that omit one. |
tolerances |
mapping | distance, duration, geometry; each {mode, warn, fail}. |
cases |
list | At least one. |
| Key | Type | Notes |
|---|---|---|
name |
string | Required, unique — this is the baseline key. |
description |
string | Free text. |
profile |
string | Engine profile / costing model. |
coordinates |
list | ≥ 2 {lat, lon} objects. Named fields, so you cannot swap them by accident. |
options |
mapping | Merged into the outgoing request (OSRM query params, Valhalla body). |
expect |
mapping | See below. |
tolerances |
mapping | Overlays the suite tolerances field by field. |
skip |
bool | Excludes the case from the run. |
skip_reason |
string | Required whenever skip is true. |
| Key | Type | Notes |
|---|---|---|
routable |
bool | Default true. false asserts the engine finds NO route. |
max_distance / min_distance |
float | Metres. Inclusive bounds. |
max_duration / min_duration |
float | Seconds. Inclusive bounds. |
required_segments |
list of string | Case-insensitive substring match against street names and refs. |
forbidden_segments |
list of string | Same matching; presence is a failure. |
turn_count |
int | Exact turn count (excludes depart/arrive/continue). |
max_turns |
int | Upper bound on turns. |
A case with routable: false may not also assert route properties — validation rejects
that combination rather than silently ignoring half of it.
Replays the suite and compares against the baseline.
| Flag | Default | Description |
|---|---|---|
--engine |
osrm |
osrm or valhalla. |
--endpoint |
"" |
Engine base URL. Env: ROUTE_CHECK_ENDPOINT. |
--suite |
routes.yml |
Suite YAML path. |
--baseline |
route-baseline.json |
Baseline JSON path. |
--fixture |
– | Serve canned responses from a JSON file; no network. |
--timeout |
20.0 |
Per-request timeout, seconds. |
--retries |
3 |
Attempts per route request. |
--warmup |
0.0 |
Seconds to poll the engine's health endpoint before starting. |
--polyline-precision |
5 |
OSRM geometry precision (5 or 6). |
--units |
kilometers |
Valhalla units (kilometers or miles). |
--tolerance-distance |
– | Override the distance FAIL threshold. |
--tolerance-duration |
– | Override the duration FAIL threshold. |
--tolerance-geometry |
– | Override the geometry FAIL threshold, metres. |
--tolerance-mode |
relative |
How distance/duration overrides are read. |
--geometry-metric |
frechet |
frechet or hausdorff. |
--fail-on |
fail |
warn, fail, or never. |
--only |
– | Run only the named case. Repeatable. |
--junit-path |
– | Write JUnit XML here. |
--json-path |
– | Write the JSON artifact here. |
--summary / --no-summary |
on | Write to $GITHUB_STEP_SUMMARY. |
--annotate / --no-annotate |
on | Emit workflow annotations. |
--quiet |
off | Suppress the terminal table. |
Exit codes: 0 clean (or below --fail-on), 1 at or past --fail-on, 2 a usage
or configuration error (bad suite, missing baseline, unknown case).
Re-records the baseline from a graph you trust, after showing a diff preview.
| Flag | Default | Description |
|---|---|---|
--only |
– | Update only the named case; others are carried over unchanged. Repeatable. |
--yes / -y |
off | Write without the confirmation prompt. |
--dry-run |
off | Show the diff and exit without writing. |
Shares the engine, suite, baseline, fixture, timeout, retry, precision and units flags
with check.
Checks the suite — and optionally the baseline — with no engine involved. Fast enough to run as a pre-commit hook.
| Flag | Default | Description |
|---|---|---|
--suite |
routes.yml |
Suite YAML path. |
--baseline |
route-baseline.json |
Baseline JSON path, if present. |
--strict |
off | Require a baseline that covers every active case, with no orphans. |
Four outputs, each for a different reader:
- Terminal table (Rich) — for local runs. Long detail collapses to
(+N more)so one broken case cannot push the others off screen. - Step summary (
$GITHUB_STEP_SUMMARY) — verdict, counts, per-case table. This is what people actually read. - Annotations (
::error/::warning) — one per failing case, not per failing check; five annotations for one route buries everything else. - JUnit XML (
--junit-path) — so existing test reporters display routing regressions alongside unit tests. Warnings become passing tests withsystem-out, since JUnit has no warning state. - JSON (
--json-path) — the lossless artifact. Every check with its baseline, actual, delta, relative delta and threshold. Chart drift over time from this.
The test suite makes zero network calls — that is a hard design constraint, not an aspiration. Two seams make it work:
RoutingEngineis aProtocol.FakeEnginesatisfies it and serves cannedRouteResults keyed by case name, with knobs for simulated transport failures (fail_first) and slow warm-up (healthy_after).HttpTransportaccepts an injectedhttpx.Client, so the real OSRM and Valhalla clients are exercised end to end againsthttpx.MockTransport.
You can use the same seam yourself: --fixture some.json runs your whole suite,
reports and exit codes included, with no server anywhere.
This repository dogfoods its own action in CI (.github/workflows/action.yml): a tiny
stub HTTP server replays recorded OSRM responses, the real action.yml runs against
it, and a deliberately-perturbed baseline proves the failure path works too. The whole
job takes seconds.
- Only OSRM and Valhalla. GraphHopper, ORS and others are not implemented. The client protocol is small if you want to add one.
- Geometry distances use an equirectangular projection, accurate to well under a metre over a single test route but not appropriate for continent-scale geometries.
- Discrete Fréchet is O(n·m) in vertex counts. A route with tens of thousands of
vertices will be slow; use
overview=simplifiedon very long routes, or Hausdorff. - Segment matching is substring-based and case-insensitive.
"High Street"will match"High Street West". That is deliberate (engines decorate names inconsistently) but it means very short segment names can match unintentionally. - Turn counting is heuristic — depart, arrive, continue and notification manoeuvres are excluded. It is stable enough to assert on across rebuilds of the same engine, but turn counts are not comparable between OSRM and Valhalla.
- No alternatives, matrix, isochrone or map-matching endpoints. Route only.
- The baseline is single-engine. Recording with OSRM and checking with Valhalla will produce noise; keep one baseline per engine and profile.
- A relative tolerance against a zero baseline degenerates to an exact-match requirement. That is intentional, but worth knowing if you have degenerate cases.
Part of a small set of routing-infrastructure tools:
- osrm-quickstart — get an OSRM instance running from an OSM extract without fighting the toolchain.
- speed-profile-builder — build and calibrate vehicle speed profiles.
Background on the problems this tool guards against:
- Benchmarking routing engine latency and accuracy — on what "accuracy" means for a routing engine and how to measure it repeatably, which is the question a baseline is an answer to.
- OSRM vs Valhalla vs GraphHopper for freight routing — how the engines differ in restriction handling, which explains why the same suite needs a separate baseline per engine.
- Graph fragmentation prevention in OSM data — the failure mode behind most large geometry divergences after a rebuild.
- Speed profile calibration for heavy vehicles — useful when duration drifts but distance does not, which is almost always a profile change rather than a graph change.
- geospatialrouting.com — the rest of the site.
See CONTRIBUTING.md.
MIT — see LICENSE.
Maintained by geospatialrouting.com.