Rust PR Bench compares Rust benchmark results between a base and head revision using Gungraun, Criterion, or both. It produces a job summary, can maintain a sticky pull-request report, and can fail CI when a regression exceeds a configured threshold.
This repository is the successor to
terjekv/github-action-iai-callgrind.
The original repository remains available for existing v1–v3 reusable-workflow consumers.
Rust PR Bench provides two interfaces backed by the same comparison and reporting code:
- The root action is the simplest interface and is suitable for GitHub Marketplace. It runs all selected benchmark cases sequentially in one caller-owned job.
- The reusable workflow precompiles and fans benchmark cases out across a dynamic job matrix. Use it for larger suites where parallel execution is worth the additional jobs and artifacts.
Input syntax differs between the interfaces:
- Root composite-action inputs are strings. Quote booleans and numbers, such as
"true"and"3". - Reusable-workflow inputs use the types declared by
workflow_call. Write booleans and numbers without quotes, such astrueand3.
The action requires an Ubuntu runner and a full checkout so both revisions are available. The
caller controls job permissions; pull-requests: write is needed only when PR comments are enabled.
Comment updates are best-effort: fork pull requests normally receive a read-only token, so the
report remains available in the job summary when GitHub does not permit the action to update the PR.
name: Rust PR Bench
on:
pull_request:
permissions:
contents: read
pull-requests: write
jobs:
bench:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
with:
fetch-depth: 0
- name: Compare benchmarks
id: bench
uses: terjekv/rust-pr-bench@v1
with:
# Composite-action inputs are strings.
backend: all
auto_discover: "true"
regression_threshold_pct_gungraun: "3"
regression_threshold_pct_criterion: "10"
fail_on_regression: "true"Call the reusable workflow directly from a job. The workflow owns its Ubuntu jobs, matrix, precompiled benchmark artifacts, report artifact, and PR comment.
name: Rust PR Bench
on:
pull_request:
permissions:
contents: read
pull-requests: write
jobs:
bench:
uses: terjekv/rust-pr-bench/.github/workflows/rust-pr-bench.yml@v1
with:
# Reusable-workflow booleans and numbers are typed values.
backend: all
auto_discover: true
feature_sets_json: >-
[
{"name":"default","features":""},
{"name":"simd","features":"simd"}
]
regression_threshold_pct_gungraun: 3
regression_threshold_pct_criterion: 10
fail_on_regression: true- Compares pull-request head and base revisions using isolated Git worktrees.
- Supports Gungraun instruction/event counts, Criterion wall-clock measurements, or both.
- Discovers standalone and workspace benchmark targets or accepts explicit commands.
- Tests multiple Cargo feature sets and honors benchmark target
required-features. - Forms target-level deltas only from metric identities present in both revisions; added or removed benchmark functions remain visible as unknown metrics without skewing the aggregate.
- Manages per-execution setup, readiness, and guaranteed teardown for service-backed benchmarks.
- Supports benchmarks moved between workspace members.
- Selects the exact Gungraun runner required by each benchmark executable.
- Can execute an
iai-callgrind 0.16.1benchmark from an older base revision during migration. - Publishes Markdown summaries and sticky PR comments with bounded history.
- Exposes raw and unaccepted regression signals separately.
- Supports label-approved, auditable exceptions for intentional regressions.
The action and reusable workflow share the inputs below. As noted above, action values are strings; reusable-workflow booleans and numbers use their native YAML types.
| Input | Default | Description |
|---|---|---|
backend |
gungraun |
gungraun, criterion, or all |
benchmarks_json |
[] |
Explicit benchmark specifications |
auto_discover |
true |
Discover benches/*.rs when no explicit specifications are supplied |
auto_detect_moved_benchmarks |
false |
Pair uniquely moved workspace benchmarks across revisions |
feature_sets_json |
default feature set | Feature combinations to test |
working_directory |
. |
Cargo project or workspace directory |
toolchain |
stable |
Rust toolchain |
cargo_args |
empty | Extra Cargo arguments |
setup_command |
empty | Runtime setup run separately for head and base |
readiness_command |
empty | Readiness probe retried after setup |
teardown_command |
empty | Cleanup always attempted after each execution |
readiness_timeout_seconds |
60 |
Per-execution readiness timeout |
criterion_cli_args |
--noplot |
Criterion bench-binary arguments |
criterion_statistic |
mean |
mean or median |
base_sha |
PR base SHA | Explicit base revision |
head_sha |
PR head SHA | Root-action-only explicit head revision |
regression_threshold_pct |
3 |
Generic regression threshold |
regression_threshold_pct_gungraun |
-1 |
Gungraun override; -1 uses the generic threshold |
regression_threshold_pct_criterion |
-1 |
Criterion override; -1 uses the generic threshold |
fail_on_regression |
false |
Fail for an unaccepted regression above its threshold |
regression_override_label |
empty | Label required to approve PR-body exceptions |
comment_mode |
always |
always, on-regression, or never |
The reusable workflow resolves its helper scripts from the exact called-workflow commit. It also
exposes action_repository and action_ref overrides for testing a fork or pull-request revision
of the workflow implementation.
Service-backed benchmarks can use runtime lifecycle commands without putting Docker or process management in the benchmark binary:
setup_command: |
NAME="bench-postgres-${RUST_PR_BENCH_EXECUTION_ID}"
docker run --detach --name "$NAME" --publish-all \
--env POSTGRES_PASSWORD=bench \
postgres:16.14-alpine3.24@sha256:57c72fd2a128e416c7fcc499958864df5301e940bca0a56f58fddf30ffc07777
PORT="$(docker port "$NAME" 5432/tcp | sed 's/.*://')"
printf 'PG_CONTAINER=%s\n' "$NAME" >> "$RUST_PR_BENCH_ENV_FILE"
printf 'DATABASE_URL=postgres://postgres:bench@127.0.0.1:%s/postgres\n' "$PORT" \
>> "$RUST_PR_BENCH_ENV_FILE"
readiness_command: |
# The image briefly starts a temporary server while initializing. Wait until
# the final PostgreSQL server has replaced the entrypoint as PID 1.
docker exec "$PG_CONTAINER" sh -ceu \
'test "$(cat /proc/1/comm)" = postgres'
docker exec "$PG_CONTAINER" pg_isready --username postgres
teardown_command: |
docker rm --force "$PG_CONTAINER"
readiness_timeout_seconds: 90The lifecycle is setup → readiness → benchmark → teardown for head, followed by the same
fresh sequence for base. Each command runs from working_directory after the relevant revision is
checked out. Readiness is retried once per second until it succeeds or its timeout expires. Teardown
is attempted even when setup, readiness, compilation, or measurement fails; only runner termination
or job cancellation can prevent it.
The reusable workflow runs every benchmark/feature/backend matrix case in its own job. The root
action runs cases sequentially, but still gives every head and base execution a separate lifecycle.
Rust PR Bench does not carry lifecycle environment state between executions. Use
RUST_PR_BENCH_EXECUTION_ID for unique container, database, network, or volume names when matrix
jobs may run concurrently. A caller that deliberately wants shared state must make its commands
connect to a caller-managed external service explicitly.
Lifecycle commands receive these environment variables:
RUST_PR_BENCH_SIDE:headorbase.RUST_PR_BENCH_EXECUTION_ID: unique, shell-safe identifier for this execution.RUST_PR_BENCH_ENV_FILE: fresh per-execution file for passing values to later stages.RUST_PR_BENCH_REPOSITORY: absolute caller-repository checkout path.RUST_PR_BENCH_WORKING_DIRECTORY: absolute command working directory.CARGO_TARGET_DIR: isolated target directory used by the benchmark.
The setup command can append NAME=value lines to RUST_PR_BENCH_ENV_FILE. An optional export
prefix is accepted; values are treated literally, without shell evaluation. These values are loaded
for readiness, the benchmark, and teardown, then the file is deleted. CARGO_TARGET_DIR and
RUST_PR_BENCH_* variables cannot be overridden through the file. Do not print secrets: command
output is captured in lifecycle logs.
The reusable workflow uploads head.setup.log, head.readiness.log, head.teardown.log, and their
base equivalents with the case result whenever the corresponding commands are configured. Failed
stage output is also included in the benchmark error diagnostics and report stage label.
Both interfaces expose:
has_regressions: at least one measurement exceeded its configured threshold, including an accepted regression.has_unaccepted_regressions: at least one threshold regression lacked an approved exception.
The root action additionally exposes:
had_errors: at least one benchmark command failed.report_path: absolute path to the generated Markdown report for later workflow steps.
When benchmarks_json is empty and auto_discover is enabled, benchmark filenames route cases:
- A name containing
gungraun,iai_callgrind, orcallgrind, but notcriterion, selects Gungraun. - A name containing
criterion, but no Callgrind-family marker, selects Criterion. - Other names run for every selected backend.
Explicit entries may be strings or objects:
benchmarks_json: >-
[
{"name":"parser-events","bench":"parser_gungraun","backend":"gungraun"},
{
"name":"parser-time",
"bench":"parser_criterion",
"backend":"criterion",
"criterion_args":"--noplot --sample-size 80 --measurement-time 6"
}
]Object fields include:
name: report label.bench: Cargo benchmark target.backend:gungraunorcriterion.command: complete command override.manifest_path,package, andargs: Cargo command helpers.required_features: benchmark-specific Cargo features that are always enabled.criterion_args: per-benchmark Criterion arguments.head,base,head_command, andbase_command: explicit mappings when a benchmark moved or changed names.
For example, a Gungraun benchmark can compare against an older IAI-Callgrind target without making IAI-Callgrind part of the new public backend API:
[
{
"name": "parser-migration",
"bench": "parser_gungraun",
"backend": "gungraun",
"base": {"bench": "parser_iai_callgrind"}
}
]feature_sets_json is an array of names or objects:
[
{"name":"default","features":""},
{"name":"simd","features":"simd"},
{"name":"minimal","features":"serde","no_default_features":true}
]Autodiscovery reads required-features from each Cargo [[bench]] target and combines them with
the selected feature set. Declare target-specific requirements there instead of placing them in a
global feature set, which applies to every discovered benchmark. Workspace members without those
target requirements are then left unchanged. Explicit benchmark objects can provide the equivalent
required_features string or array.
Set regression_override_label to require an explicit approval label. A PR can then declare
bounded exceptions in one fenced block:
```rust-pr-bench
{
"accept_regressions": [
{
"benchmark": "verify_password",
"backend": "gungraun",
"feature": "constant-time",
"max_regression_pct": 25,
"reason": "Constant-time verification"
}
]
}
```The approval label must have been applied after the most recent PR-body edit. Editing the directive invalidates an older approval. Reports always show the measured regression; approval only changes whether it is considered unaccepted by the CI gate.
Existing callers may remain on the original repository and its v1–v3 tags. Migration is
explicit because GitHub Actions does not follow repository-rename redirects.
For the parallel workflow, change only the repository and major version first:
- uses: terjekv/github-action-iai-callgrind/.github/workflows/rust-pr-bench.yml@v3
+ uses: terjekv/rust-pr-bench/.github/workflows/rust-pr-bench.yml@v1Then replace deprecated public values:
backend: iai-callgrind,iai, orcallgrind→backend: gungraunregression_threshold_pct_iai_callgrind→regression_threshold_pct_gungraun
The original compatibility workflow path does not exist in this repository. Use
.github/workflows/rust-pr-bench.yml or migrate to the root action.
Run the local tests with:
actionlint
python3 -m unittest discover -s tests -v
node --test tests/test_github_pr_comment.js
cargo clippy --manifest-path examples/sample-rust-app/Cargo.toml \
--all-targets --all-features -- -D warningsThe sample CI workflow exercises Gungraun, Criterion, mixed runner versions, both public interfaces,
regression outputs, and the service lifecycle hooks against a real PostgreSQL SELECT 1 query.
Rust PR Bench is licensed under the MIT License.