Skip to content

aggregate blocked impl - #24928

Draft
rluvaton wants to merge 95 commits into
apache:mainfrom
rluvaton:add-blocks-impl-from-scratch
Draft

rluvaton wants to merge 95 commits into
apache:mainfrom
rluvaton:add-blocks-impl-from-scratch

Conversation

@rluvaton

@rluvaton rluvaton commented Sep 3, 2026 •

Copy link
Copy Markdown
Member

Huge blocked impl that is backwards compatible
And just see the performance cost

Currently it only contain blocked impl for single group by and some aggregate expression

Which issue does this PR close?

  • Closes #.

Rationale for this change

What changes are included in this PR?

What is the testing strategy for this PR?

Are there any user-facing changes?

# Conflicts:
#	datafusion/expr-common/src/groups_accumulator.rs
@adriangbot

Copy link
Copy Markdown

Benchmark for this request failed before finishing (Kubernetes reason: BackoffLimitExceeded).

Benchmarks requested: h2o_medium

Runner log (last 40 lines)
2026-09-09T18:45:22.676291Z  INFO runner starting benchmark runner bench_type=Datafusion, pr_url=https://github.com/apache/datafusion/pull/24928, benchmarks=h2o_medium
2026-09-09T18:45:22.771467Z  INFO benchmark_controller::runner::bench_datafusion === Cloning PR branch ===
2026-09-09T18:45:22.771578Z  INFO benchmark_controller::runner::shell running command cmd=git, args=["clone", "--depth=200", "https://github.com/apache/datafusion.git", "/workspace/datafusion-branch"], cwd="/"
2026-09-09T18:45:27.780079Z  INFO benchmark_controller::runner::shell running command cmd=git, args=["fetch", "origin", "refs/pull/24928/head:add-blocks-impl-from-scratch", "main"], cwd="/workspace/datafusion-branch"
2026-09-09T18:45:32.782703Z  INFO benchmark_controller::runner::shell running command cmd=git, args=["checkout", "add-blocks-impl-from-scratch"], cwd="/workspace/datafusion-branch"
2026-09-09T18:45:37.783989Z  INFO benchmark_controller::runner::shell running command cmd=git, args=["merge-base", "HEAD", "origin/main"], cwd="/workspace/datafusion-branch"
2026-09-09T18:45:42.786842Z  INFO benchmark_controller::runner::bench_datafusion === Cloning merge-base ===
2026-09-09T18:45:42.786858Z  INFO benchmark_controller::runner::shell running command cmd=git, args=["clone", "--depth=200", "https://github.com/apache/datafusion.git", "/workspace/datafusion-base"], cwd="/"
2026-09-09T18:45:47.787929Z  INFO benchmark_controller::runner::shell running command cmd=git, args=["-c", "advice.detachedHead=false", "checkout", "40488988ad596c9b093ad60e1453430d803ce33c"], cwd="/workspace/datafusion-base"
2026-09-09T18:45:52.789146Z  INFO benchmark_controller::runner::shell running command cmd=rustc, args=["--version"], cwd="/"
2026-09-09T18:45:57.791473Z  INFO benchmark_controller::runner::shell running command cmd=cargo, args=["metadata", "--no-deps", "--format-version", "1"], cwd="/workspace/datafusion-branch/benchmarks"
2026-09-09T18:46:02.795549Z  INFO benchmark_controller::runner::bench_datafusion === Compiling dfbench for PR branch and merge-base in parallel ===
2026-09-09T18:46:02.835389Z  INFO benchmark_controller::runner::shell running command cmd=git, args=["rev-parse", "HEAD"], cwd="/workspace/datafusion-branch"
2026-09-09T18:46:07.855026Z  INFO benchmark_controller::runner::shell running command cmd=git, args=["rev-parse", "HEAD"], cwd="/workspace/datafusion-base"
2026-09-09T18:46:12.858039Z  INFO benchmark_controller::runner::controller_client post_comment _repo=apache/datafusion, _pr_number=24928, job_id=2276
2026-09-09T18:47:28.101519Z ERROR runner benchmark failed error=send request: error sending request for url (http://benchmark-controller.benchmarking.svc.cluster.local:8080/jobs/2276/comment): client error (Connect): tcp connect error: Connection refused (os error 111)
Kubernetes message
Job has reached the specified backoff limit

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c5606968809-2278-4xgxl 6.12.94+ #1 SMP Tue Aug 4 08:44:15 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing add-blocks-impl-from-scratch (78834e5) to 4048898 (merge-base) diff

Run configuration
run benchmark clickbench_partitioned

Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c5606967120-2275-ll72d 6.12.94+ #1 SMP Fri Jul 17 09:42:57 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing add-blocks-impl-from-scratch (78834e5) to 4048898 (merge-base) diff

Run configuration
run benchmark external_aggr

Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c5606968809-2279-vs6tz 6.12.94+ #1 SMP Tue Aug 4 08:44:15 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing add-blocks-impl-from-scratch (78834e5) to 4048898 (merge-base) diff

Run configuration
run benchmark tpcds

Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c5606968809-2280-jlpz6 6.12.94+ #1 SMP Tue Aug 4 08:44:15 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing add-blocks-impl-from-scratch (78834e5) to 4048898 (merge-base) diff

Run configuration
run benchmark tpch

Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c5606967120-2277-pnd78 6.12.94+ #1 SMP Tue Aug 4 08:44:15 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing add-blocks-impl-from-scratch (78834e5) to 4048898 (merge-base) diff

Run configuration
run benchmark tpch10

Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing add-blocks-impl-from-scratch (78834e5) to 4048898 (merge-base) diff

Run configuration
run benchmark tpch
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and add-blocks-impl-from-scratch
--------------------
Benchmark tpch_sf1.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━┓
┃ Query     ┃     HEAD ┃ add-blocks-impl-from-scratch ┃       Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━┩
│ QQuery 1  │ 39.08 ms │                     42.01 ms │ 1.07x slower │
│ QQuery 2  │ 19.13 ms │                     19.83 ms │    no change │
│ QQuery 3  │ 29.75 ms │                     29.62 ms │    no change │
│ QQuery 4  │ 18.21 ms │                     18.50 ms │    no change │
│ QQuery 5  │ 36.93 ms │                     37.58 ms │    no change │
│ QQuery 6  │ 16.44 ms │                     16.85 ms │    no change │
│ QQuery 7  │ 42.20 ms │                     42.36 ms │    no change │
│ QQuery 8  │ 41.78 ms │                     42.15 ms │    no change │
│ QQuery 9  │ 51.30 ms │                     50.12 ms │    no change │
│ QQuery 10 │ 43.57 ms │                     43.80 ms │    no change │
│ QQuery 11 │ 14.24 ms │                     14.14 ms │    no change │
│ QQuery 12 │ 24.56 ms │                     24.44 ms │    no change │
│ QQuery 13 │ 42.09 ms │                     46.43 ms │ 1.10x slower │
│ QQuery 14 │ 25.26 ms │                     25.50 ms │    no change │
│ QQuery 15 │ 31.78 ms │                     32.72 ms │    no change │
│ QQuery 16 │ 14.17 ms │                     14.68 ms │    no change │
│ QQuery 17 │ 75.07 ms │                     89.77 ms │ 1.20x slower │
│ QQuery 18 │ 63.00 ms │                     63.42 ms │    no change │
│ QQuery 19 │ 33.63 ms │                     33.75 ms │    no change │
│ QQuery 20 │ 33.14 ms │                     33.07 ms │    no change │
│ QQuery 21 │ 58.70 ms │                     57.88 ms │    no change │
│ QQuery 22 │ 14.71 ms │                     14.98 ms │    no change │
└───────────┴──────────┴──────────────────────────────┴──────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━┓
┃ Benchmark Summary                           ┃          ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━┩
│ Total Time (HEAD)                           │ 768.75ms │
│ Total Time (add-blocks-impl-from-scratch)   │ 793.59ms │
│ Average Time (HEAD)                         │  34.94ms │
│ Average Time (add-blocks-impl-from-scratch) │  36.07ms │
│ Queries Faster                              │        0 │
│ Queries Slower                              │        3 │
│ Queries with No Change                      │       19 │
│ Queries with Failure                        │        0 │
└─────────────────────────────────────────────┴──────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and add-blocks-impl-from-scratch
--------------------
Benchmark tpch_sf1.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━┓
┃ Query     ┃                           HEAD ┃   add-blocks-impl-from-scratch ┃       Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━┩
│ QQuery 1  │ 39.08 / 39.82 ±0.94 / 41.65 ms │ 42.01 / 43.34 ±1.28 / 44.93 ms │ 1.09x slower │
│ QQuery 2  │ 19.13 / 19.93 ±0.78 / 21.41 ms │ 19.83 / 19.89 ±0.07 / 20.02 ms │    no change │
│ QQuery 3  │ 29.75 / 30.07 ±0.19 / 30.32 ms │ 29.62 / 29.94 ±0.21 / 30.15 ms │    no change │
│ QQuery 4  │ 18.21 / 19.80 ±1.92 / 23.36 ms │ 18.50 / 19.08 ±0.94 / 20.96 ms │    no change │
│ QQuery 5  │ 36.93 / 37.99 ±1.11 / 40.00 ms │ 37.58 / 37.72 ±0.17 / 38.06 ms │    no change │
│ QQuery 6  │ 16.44 / 16.58 ±0.19 / 16.95 ms │ 16.85 / 16.98 ±0.08 / 17.10 ms │    no change │
│ QQuery 7  │ 42.20 / 44.20 ±3.11 / 50.40 ms │ 42.36 / 44.00 ±1.20 / 45.46 ms │    no change │
│ QQuery 8  │ 41.78 / 42.01 ±0.24 / 42.47 ms │ 42.15 / 43.26 ±1.45 / 46.06 ms │    no change │
│ QQuery 9  │ 51.30 / 52.82 ±1.38 / 55.19 ms │ 50.12 / 51.52 ±1.21 / 53.30 ms │    no change │
│ QQuery 10 │ 43.57 / 44.82 ±1.27 / 46.38 ms │ 43.80 / 44.03 ±0.17 / 44.29 ms │    no change │
│ QQuery 11 │ 14.24 / 14.31 ±0.09 / 14.49 ms │ 14.14 / 14.31 ±0.16 / 14.53 ms │    no change │
│ QQuery 12 │ 24.56 / 25.67 ±1.05 / 27.66 ms │ 24.44 / 25.98 ±1.83 / 29.47 ms │    no change │
│ QQuery 13 │ 42.09 / 42.49 ±0.47 / 43.20 ms │ 46.43 / 47.39 ±1.01 / 49.06 ms │ 1.12x slower │
│ QQuery 14 │ 25.26 / 25.42 ±0.15 / 25.65 ms │ 25.50 / 25.69 ±0.14 / 25.86 ms │    no change │
│ QQuery 15 │ 31.78 / 32.40 ±0.53 / 33.34 ms │ 32.72 / 32.85 ±0.17 / 33.17 ms │    no change │
│ QQuery 16 │ 14.17 / 14.36 ±0.12 / 14.55 ms │ 14.68 / 14.91 ±0.18 / 15.13 ms │    no change │
│ QQuery 17 │ 75.07 / 76.95 ±2.39 / 81.51 ms │ 89.77 / 91.51 ±1.35 / 93.61 ms │ 1.19x slower │
│ QQuery 18 │ 63.00 / 63.84 ±0.62 / 64.59 ms │ 63.42 / 65.87 ±1.76 / 68.89 ms │    no change │
│ QQuery 19 │ 33.63 / 33.99 ±0.27 / 34.29 ms │ 33.75 / 34.34 ±0.77 / 35.83 ms │    no change │
│ QQuery 20 │ 33.14 / 33.29 ±0.12 / 33.42 ms │ 33.07 / 33.80 ±0.42 / 34.30 ms │    no change │
│ QQuery 21 │ 58.70 / 59.61 ±1.10 / 61.77 ms │ 57.88 / 58.35 ±0.32 / 58.82 ms │    no change │
│ QQuery 22 │ 14.71 / 15.31 ±0.50 / 16.06 ms │ 14.98 / 16.60 ±1.70 / 19.48 ms │ 1.08x slower │
└───────────┴────────────────────────────────┴────────────────────────────────┴──────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━┓
┃ Benchmark Summary                           ┃          ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━┩
│ Total Time (HEAD)                           │ 785.67ms │
│ Total Time (add-blocks-impl-from-scratch)   │ 811.36ms │
│ Average Time (HEAD)                         │  35.71ms │
│ Average Time (add-blocks-impl-from-scratch) │  36.88ms │
│ Queries Faster                              │        0 │
│ Queries Slower                              │        4 │
│ Queries with No Change                      │       18 │
│ Queries with Failure                        │        0 │
└─────────────────────────────────────────────┴──────────┘

Resource Usage

tpch — base (merge-base)

Metric Value
Wall time 5.0s
Peak memory 1.2 GiB
Avg memory 502.4 MiB
CPU user 22.2s
CPU sys 1.8s
Peak spill 0 B

tpch — branch

Metric Value
Wall time 5.0s
Peak memory 1.3 GiB
Avg memory 687.7 MiB
CPU user 23.2s
CPU sys 1.8s
Peak spill 0 B

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing add-blocks-impl-from-scratch (78834e5) to 4048898 (merge-base) diff

Run configuration
run benchmark tpch10
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and add-blocks-impl-from-scratch
--------------------
Benchmark tpch_sf10.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━┓
┃ Query     ┃      HEAD ┃ add-blocks-impl-from-scratch ┃       Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━┩
│ QQuery 1  │ 309.57 ms │                    335.93 ms │ 1.09x slower │
│ QQuery 2  │  90.48 ms │                     92.49 ms │    no change │
│ QQuery 3  │ 216.41 ms │                    218.59 ms │    no change │
│ QQuery 4  │ 113.14 ms │                    113.91 ms │    no change │
│ QQuery 5  │ 350.59 ms │                    351.57 ms │    no change │
│ QQuery 6  │ 123.14 ms │                    125.04 ms │    no change │
│ QQuery 7  │ 448.85 ms │                    454.60 ms │    no change │
│ QQuery 8  │ 360.08 ms │                    356.79 ms │    no change │
│ QQuery 9  │ 540.98 ms │                    519.62 ms │    no change │
│ QQuery 10 │ 309.89 ms │                    302.95 ms │    no change │
│ QQuery 11 │  64.55 ms │                     64.08 ms │    no change │
│ QQuery 12 │ 182.51 ms │                    180.21 ms │    no change │
│ QQuery 13 │ 307.48 ms │                    322.92 ms │ 1.05x slower │
│ QQuery 14 │ 172.16 ms │                    170.09 ms │    no change │
│ QQuery 15 │ 301.70 ms │                    310.07 ms │    no change │
│ QQuery 16 │  64.35 ms │                     66.23 ms │    no change │
│ QQuery 17 │ 569.74 ms │                    684.02 ms │ 1.20x slower │
│ QQuery 18 │ 703.89 ms │                    747.30 ms │ 1.06x slower │
│ QQuery 19 │ 244.96 ms │                    250.50 ms │    no change │
│ QQuery 20 │ 256.49 ms │                    285.20 ms │ 1.11x slower │
│ QQuery 21 │ 661.01 ms │                    691.69 ms │    no change │
│ QQuery 22 │  61.30 ms │                     64.71 ms │ 1.06x slower │
└───────────┴───────────┴──────────────────────────────┴──────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Benchmark Summary                           ┃           ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ Total Time (HEAD)                           │ 6453.26ms │
│ Total Time (add-blocks-impl-from-scratch)   │ 6708.50ms │
│ Average Time (HEAD)                         │  293.33ms │
│ Average Time (add-blocks-impl-from-scratch) │  304.93ms │
│ Queries Faster                              │         0 │
│ Queries Slower                              │         6 │
│ Queries with No Change                      │        16 │
│ Queries with Failure                        │         0 │
└─────────────────────────────────────────────┴───────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and add-blocks-impl-from-scratch
--------------------
Benchmark tpch_sf10.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━┓
┃ Query     ┃                               HEAD ┃       add-blocks-impl-from-scratch ┃       Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━┩
│ QQuery 1  │  309.57 / 311.89 ±1.97 / 315.04 ms │  335.93 / 340.59 ±6.08 / 352.55 ms │ 1.09x slower │
│ QQuery 2  │     90.48 / 91.97 ±1.46 / 94.37 ms │     92.49 / 94.10 ±1.24 / 96.22 ms │    no change │
│ QQuery 3  │  216.41 / 219.46 ±3.49 / 226.25 ms │  218.59 / 222.47 ±3.26 / 227.10 ms │    no change │
│ QQuery 4  │  113.14 / 115.12 ±1.76 / 118.24 ms │  113.91 / 115.85 ±1.40 / 117.59 ms │    no change │
│ QQuery 5  │  350.59 / 355.83 ±4.42 / 363.79 ms │  351.57 / 362.45 ±5.82 / 366.73 ms │    no change │
│ QQuery 6  │  123.14 / 124.73 ±1.77 / 128.05 ms │  125.04 / 128.56 ±3.41 / 134.38 ms │    no change │
│ QQuery 7  │  448.85 / 456.21 ±5.73 / 466.47 ms │ 454.60 / 469.62 ±10.91 / 485.24 ms │    no change │
│ QQuery 8  │  360.08 / 363.44 ±2.85 / 367.84 ms │  356.79 / 362.37 ±4.05 / 367.67 ms │    no change │
│ QQuery 9  │ 540.98 / 554.64 ±14.34 / 582.03 ms │  519.62 / 527.81 ±4.66 / 532.59 ms │    no change │
│ QQuery 10 │ 309.89 / 321.08 ±10.02 / 338.90 ms │  302.95 / 306.18 ±3.50 / 312.41 ms │    no change │
│ QQuery 11 │     64.55 / 66.07 ±1.89 / 69.65 ms │     64.08 / 67.76 ±6.44 / 80.63 ms │    no change │
│ QQuery 12 │  182.51 / 193.21 ±8.69 / 204.83 ms │  180.21 / 184.56 ±3.97 / 189.50 ms │    no change │
│ QQuery 13 │  307.48 / 319.78 ±9.45 / 335.27 ms │  322.92 / 328.55 ±4.43 / 333.62 ms │    no change │
│ QQuery 14 │  172.16 / 178.32 ±9.02 / 196.07 ms │  170.09 / 176.75 ±6.76 / 187.07 ms │    no change │
│ QQuery 15 │  301.70 / 309.82 ±5.14 / 317.58 ms │  310.07 / 314.12 ±4.36 / 320.61 ms │    no change │
│ QQuery 16 │     64.35 / 68.59 ±3.73 / 73.40 ms │     66.23 / 67.48 ±1.53 / 70.46 ms │    no change │
│ QQuery 17 │ 569.74 / 590.36 ±14.44 / 613.42 ms │ 684.02 / 697.57 ±11.73 / 717.30 ms │ 1.18x slower │
│ QQuery 18 │  703.89 / 713.60 ±8.91 / 730.16 ms │ 747.30 / 766.87 ±15.68 / 785.53 ms │ 1.07x slower │
│ QQuery 19 │  244.96 / 256.77 ±7.29 / 267.14 ms │ 250.50 / 265.74 ±16.83 / 293.46 ms │    no change │
│ QQuery 20 │ 256.49 / 273.02 ±10.31 / 288.92 ms │  285.20 / 291.77 ±4.57 / 297.34 ms │ 1.07x slower │
│ QQuery 21 │ 661.01 / 676.76 ±10.14 / 686.15 ms │ 691.69 / 710.59 ±15.82 / 735.94 ms │    no change │
│ QQuery 22 │     61.30 / 65.39 ±3.84 / 70.84 ms │     64.71 / 69.24 ±4.50 / 76.63 ms │ 1.06x slower │
└───────────┴────────────────────────────────────┴────────────────────────────────────┴──────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Benchmark Summary                           ┃           ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ Total Time (HEAD)                           │ 6626.02ms │
│ Total Time (add-blocks-impl-from-scratch)   │ 6870.98ms │
│ Average Time (HEAD)                         │  301.18ms │
│ Average Time (add-blocks-impl-from-scratch) │  312.32ms │
│ Queries Faster                              │         0 │
│ Queries Slower                              │         5 │
│ Queries with No Change                      │        17 │
│ Queries with Failure                        │         0 │
└─────────────────────────────────────────────┴───────────┘

Resource Usage

tpch10 — base (merge-base)

Metric Value
Wall time 35.0s
Peak memory 5.3 GiB
Avg memory 1.7 GiB
CPU user 337.5s
CPU sys 19.4s
Peak spill 0 B

tpch10 — branch

Metric Value
Wall time 35.0s
Peak memory 5.3 GiB
Avg memory 1.6 GiB
CPU user 347.9s
CPU sys 19.1s
Peak spill 0 B

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing add-blocks-impl-from-scratch (78834e5) to 4048898 (merge-base) diff

Run configuration
run benchmark tpcds
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and add-blocks-impl-from-scratch
--------------------
Benchmark tpcds_sf1.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃       HEAD ┃ add-blocks-impl-from-scratch ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 1  │    5.52 ms │                      6.34 ms │  1.15x slower │
│ QQuery 2  │   79.20 ms │                     85.30 ms │  1.08x slower │
│ QQuery 3  │   28.48 ms │                     29.77 ms │     no change │
│ QQuery 4  │  465.00 ms │                    491.42 ms │  1.06x slower │
│ QQuery 5  │   51.61 ms │                     51.88 ms │     no change │
│ QQuery 6  │   35.09 ms │                     35.28 ms │     no change │
│ QQuery 7  │   74.88 ms │                     74.10 ms │     no change │
│ QQuery 8  │   35.51 ms │                     36.20 ms │     no change │
│ QQuery 9  │   53.89 ms │                     53.33 ms │     no change │
│ QQuery 10 │   62.49 ms │                     62.17 ms │     no change │
│ QQuery 11 │  287.07 ms │                    347.73 ms │  1.21x slower │
│ QQuery 12 │   28.14 ms │                     29.79 ms │  1.06x slower │
│ QQuery 13 │  117.36 ms │                    121.85 ms │     no change │
│ QQuery 14 │  417.43 ms │                    432.34 ms │     no change │
│ QQuery 15 │   56.49 ms │                     62.66 ms │  1.11x slower │
│ QQuery 16 │    6.65 ms │                      7.53 ms │  1.13x slower │
│ QQuery 17 │   79.28 ms │                     82.27 ms │     no change │
│ QQuery 18 │  103.05 ms │                    110.87 ms │  1.08x slower │
│ QQuery 19 │   40.83 ms │                     43.48 ms │  1.06x slower │
│ QQuery 20 │   34.97 ms │                     38.69 ms │  1.11x slower │
│ QQuery 21 │   16.81 ms │                     17.54 ms │     no change │
│ QQuery 22 │   62.47 ms │                     70.53 ms │  1.13x slower │
│ QQuery 23 │  303.52 ms │                    359.86 ms │  1.19x slower │
│ QQuery 24 │  190.20 ms │                    214.66 ms │  1.13x slower │
│ QQuery 25 │  108.41 ms │                    115.48 ms │  1.07x slower │
│ QQuery 26 │   48.60 ms │                     50.86 ms │     no change │
│ QQuery 27 │    6.02 ms │                      6.84 ms │  1.14x slower │
│ QQuery 28 │   60.40 ms │                     58.83 ms │     no change │
│ QQuery 29 │   95.58 ms │                    101.72 ms │  1.06x slower │
│ QQuery 30 │   31.89 ms │                     34.62 ms │  1.09x slower │
│ QQuery 31 │  109.27 ms │                    118.05 ms │  1.08x slower │
│ QQuery 32 │   19.88 ms │                     22.22 ms │  1.12x slower │
│ QQuery 33 │   37.07 ms │                     39.87 ms │  1.08x slower │
│ QQuery 34 │    9.72 ms │                     10.79 ms │  1.11x slower │
│ QQuery 35 │   71.92 ms │                     78.66 ms │  1.09x slower │
│ QQuery 36 │    5.68 ms │                      6.16 ms │  1.08x slower │
│ QQuery 37 │    6.67 ms │                      7.05 ms │  1.06x slower │
│ QQuery 38 │   60.91 ms │                     69.11 ms │  1.13x slower │
│ QQuery 39 │   88.54 ms │                    103.08 ms │  1.16x slower │
│ QQuery 40 │   23.17 ms │                     26.02 ms │  1.12x slower │
│ QQuery 41 │   11.24 ms │                     12.23 ms │  1.09x slower │
│ QQuery 42 │   23.33 ms │                     24.85 ms │  1.07x slower │
│ QQuery 43 │    5.19 ms │                      5.55 ms │  1.07x slower │
│ QQuery 44 │    9.49 ms │                      9.96 ms │  1.05x slower │
│ QQuery 45 │   38.03 ms │                     42.72 ms │  1.12x slower │
│ QQuery 46 │   11.77 ms │                     12.36 ms │     no change │
│ QQuery 47 │  219.66 ms │                    259.75 ms │  1.18x slower │
│ QQuery 48 │   94.70 ms │                     98.57 ms │     no change │
│ QQuery 49 │   69.91 ms │                     74.15 ms │  1.06x slower │
│ QQuery 50 │   60.18 ms │                     61.19 ms │     no change │
│ QQuery 51 │   92.98 ms │                     93.76 ms │     no change │
│ QQuery 52 │   24.71 ms │                     25.13 ms │     no change │
│ QQuery 53 │   30.00 ms │                     29.65 ms │     no change │
│ QQuery 54 │   55.87 ms │                     57.47 ms │     no change │
│ QQuery 55 │   24.17 ms │                     24.44 ms │     no change │
│ QQuery 56 │   40.09 ms │                     40.98 ms │     no change │
│ QQuery 57 │  181.28 ms │                    190.49 ms │  1.05x slower │
│ QQuery 58 │  113.44 ms │                    115.67 ms │     no change │
│ QQuery 59 │  117.23 ms │                    127.48 ms │  1.09x slower │
│ QQuery 60 │   40.37 ms │                     40.82 ms │     no change │
│ QQuery 61 │   12.84 ms │                     13.52 ms │  1.05x slower │
│ QQuery 62 │   47.77 ms │                     49.08 ms │     no change │
│ QQuery 63 │   30.11 ms │                     30.20 ms │     no change │
│ QQuery 64 │  385.31 ms │                    386.87 ms │     no change │
│ QQuery 65 │  124.91 ms │                    134.80 ms │  1.08x slower │
│ QQuery 66 │   81.03 ms │                     84.80 ms │     no change │
│ QQuery 67 │  264.96 ms │                    275.52 ms │     no change │
│ QQuery 68 │   12.56 ms │                     13.13 ms │     no change │
│ QQuery 69 │   58.62 ms │                     59.02 ms │     no change │
│ QQuery 70 │  109.38 ms │                    112.43 ms │     no change │
│ QQuery 71 │   36.01 ms │                     36.68 ms │     no change │
│ QQuery 72 │ 1907.51 ms │                   1887.64 ms │     no change │
│ QQuery 73 │   10.54 ms │                      9.82 ms │ +1.07x faster │
│ QQuery 74 │  196.12 ms │                    175.46 ms │ +1.12x faster │
│ QQuery 75 │  150.88 ms │                    148.59 ms │     no change │
│ QQuery 76 │   36.86 ms │                     35.33 ms │     no change │
│ QQuery 77 │   62.67 ms │                     61.63 ms │     no change │
│ QQuery 78 │  242.10 ms │                    207.49 ms │ +1.17x faster │
│ QQuery 79 │   69.15 ms │                     66.20 ms │     no change │
│ QQuery 80 │  102.75 ms │                     99.36 ms │     no change │
│ QQuery 81 │   27.45 ms │                     25.90 ms │ +1.06x faster │
│ QQuery 82 │   17.04 ms │                     16.26 ms │     no change │
│ QQuery 83 │   34.97 ms │                     34.09 ms │     no change │
│ QQuery 84 │   30.42 ms │                     29.42 ms │     no change │
│ QQuery 85 │  105.07 ms │                    102.13 ms │     no change │
│ QQuery 86 │   27.37 ms │                     24.91 ms │ +1.10x faster │
│ QQuery 87 │   67.66 ms │                     63.83 ms │ +1.06x faster │
│ QQuery 88 │   64.65 ms │                     63.34 ms │     no change │
│ QQuery 89 │   36.83 ms │                     35.40 ms │     no change │
│ QQuery 90 │   17.91 ms │                     17.14 ms │     no change │
│ QQuery 91 │   46.09 ms │                     44.50 ms │     no change │
│ QQuery 92 │   30.14 ms │                     28.90 ms │     no change │
│ QQuery 93 │   50.71 ms │                     49.85 ms │     no change │
│ QQuery 94 │   38.68 ms │                     37.78 ms │     no change │
│ QQuery 95 │   83.94 ms │                     80.79 ms │     no change │
│ QQuery 96 │   24.37 ms │                     23.76 ms │     no change │
│ QQuery 97 │   53.74 ms │                     54.93 ms │     no change │
│ QQuery 98 │   43.04 ms │                     42.53 ms │     no change │
│ QQuery 99 │   70.59 ms │                     70.99 ms │     no change │
└───────────┴────────────┴──────────────────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Benchmark Summary                           ┃           ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ Total Time (HEAD)                           │ 9496.11ms │
│ Total Time (add-blocks-impl-from-scratch)   │ 9796.18ms │
│ Average Time (HEAD)                         │   95.92ms │
│ Average Time (add-blocks-impl-from-scratch) │   98.95ms │
│ Queries Faster                              │         6 │
│ Queries Slower                              │        38 │
│ Queries with No Change                      │        55 │
│ Queries with Failure                        │         0 │
└─────────────────────────────────────────────┴───────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and add-blocks-impl-from-scratch
--------------------
Benchmark tpcds_sf1.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃                                  HEAD ┃          add-blocks-impl-from-scratch ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 1  │           5.52 / 6.08 ±0.95 / 7.97 ms │           6.34 / 6.84 ±0.90 / 8.65 ms │  1.12x slower │
│ QQuery 2  │        79.20 / 80.19 ±1.03 / 82.15 ms │        85.30 / 85.94 ±0.42 / 86.46 ms │  1.07x slower │
│ QQuery 3  │        28.48 / 28.89 ±0.26 / 29.23 ms │        29.77 / 30.15 ±0.24 / 30.50 ms │     no change │
│ QQuery 4  │     465.00 / 474.60 ±9.54 / 492.85 ms │    491.42 / 535.23 ±22.78 / 554.14 ms │  1.13x slower │
│ QQuery 5  │        51.61 / 54.83 ±4.36 / 63.34 ms │        51.88 / 52.28 ±0.32 / 52.72 ms │     no change │
│ QQuery 6  │        35.09 / 35.96 ±0.57 / 36.86 ms │        35.28 / 35.93 ±0.46 / 36.63 ms │     no change │
│ QQuery 7  │        74.88 / 75.31 ±0.31 / 75.73 ms │        74.10 / 75.59 ±1.75 / 78.85 ms │     no change │
│ QQuery 8  │        35.51 / 35.92 ±0.30 / 36.39 ms │        36.20 / 36.67 ±0.50 / 37.62 ms │     no change │
│ QQuery 9  │        53.89 / 56.55 ±3.31 / 62.97 ms │        53.33 / 54.04 ±0.71 / 55.37 ms │     no change │
│ QQuery 10 │        62.49 / 62.70 ±0.27 / 63.24 ms │        62.17 / 63.74 ±1.22 / 65.13 ms │     no change │
│ QQuery 11 │     287.07 / 289.34 ±1.88 / 292.45 ms │     347.73 / 353.22 ±3.83 / 358.46 ms │  1.22x slower │
│ QQuery 12 │        28.14 / 30.42 ±2.22 / 34.61 ms │        29.79 / 30.52 ±0.61 / 31.38 ms │     no change │
│ QQuery 13 │     117.36 / 118.20 ±0.95 / 119.92 ms │     121.85 / 123.75 ±2.06 / 127.56 ms │     no change │
│ QQuery 14 │     417.43 / 422.01 ±5.93 / 433.59 ms │     432.34 / 437.48 ±4.29 / 444.95 ms │     no change │
│ QQuery 15 │        56.49 / 57.38 ±1.53 / 60.42 ms │        62.66 / 65.16 ±1.63 / 67.42 ms │  1.14x slower │
│ QQuery 16 │           6.65 / 6.75 ±0.08 / 6.84 ms │          7.53 / 8.98 ±2.61 / 14.17 ms │  1.33x slower │
│ QQuery 17 │        79.28 / 80.02 ±0.70 / 81.22 ms │        82.27 / 84.62 ±2.10 / 88.52 ms │  1.06x slower │
│ QQuery 18 │     103.05 / 104.21 ±0.88 / 105.23 ms │     110.87 / 111.60 ±0.83 / 113.22 ms │  1.07x slower │
│ QQuery 19 │        40.83 / 41.07 ±0.27 / 41.55 ms │        43.48 / 45.22 ±1.53 / 47.68 ms │  1.10x slower │
│ QQuery 20 │        34.97 / 35.49 ±0.30 / 35.80 ms │        38.69 / 39.58 ±0.71 / 40.78 ms │  1.12x slower │
│ QQuery 21 │        16.81 / 17.01 ±0.12 / 17.14 ms │        17.54 / 18.27 ±0.45 / 18.83 ms │  1.07x slower │
│ QQuery 22 │        62.47 / 62.96 ±0.55 / 63.96 ms │        70.53 / 72.07 ±0.98 / 73.63 ms │  1.14x slower │
│ QQuery 23 │     303.52 / 308.68 ±3.32 / 313.90 ms │     359.86 / 370.55 ±6.14 / 378.92 ms │  1.20x slower │
│ QQuery 24 │     190.20 / 193.26 ±3.65 / 200.20 ms │     214.66 / 219.31 ±5.32 / 228.85 ms │  1.13x slower │
│ QQuery 25 │     108.41 / 110.27 ±1.72 / 113.44 ms │     115.48 / 117.82 ±2.20 / 121.91 ms │  1.07x slower │
│ QQuery 26 │        48.60 / 49.15 ±0.58 / 50.25 ms │        50.86 / 51.14 ±0.24 / 51.51 ms │     no change │
│ QQuery 27 │           6.02 / 6.14 ±0.12 / 6.36 ms │           6.84 / 7.06 ±0.16 / 7.31 ms │  1.15x slower │
│ QQuery 28 │        60.40 / 60.66 ±0.23 / 61.03 ms │        58.83 / 61.47 ±1.88 / 63.28 ms │     no change │
│ QQuery 29 │       95.58 / 98.55 ±2.74 / 103.30 ms │     101.72 / 106.62 ±7.23 / 120.91 ms │  1.08x slower │
│ QQuery 30 │        31.89 / 32.65 ±0.59 / 33.38 ms │        34.62 / 35.24 ±0.50 / 36.12 ms │  1.08x slower │
│ QQuery 31 │     109.27 / 110.40 ±1.51 / 113.37 ms │     118.05 / 121.44 ±3.44 / 127.53 ms │  1.10x slower │
│ QQuery 32 │        19.88 / 20.26 ±0.43 / 21.09 ms │        22.22 / 22.65 ±0.49 / 23.30 ms │  1.12x slower │
│ QQuery 33 │        37.07 / 37.77 ±0.41 / 38.20 ms │        39.87 / 40.13 ±0.35 / 40.82 ms │  1.06x slower │
│ QQuery 34 │         9.72 / 10.07 ±0.33 / 10.62 ms │        10.79 / 11.15 ±0.38 / 11.74 ms │  1.11x slower │
│ QQuery 35 │        71.92 / 72.44 ±0.35 / 73.00 ms │        78.66 / 81.46 ±4.52 / 90.47 ms │  1.12x slower │
│ QQuery 36 │           5.68 / 5.81 ±0.14 / 6.09 ms │           6.16 / 6.50 ±0.19 / 6.72 ms │  1.12x slower │
│ QQuery 37 │           6.67 / 6.76 ±0.08 / 6.89 ms │           7.05 / 7.41 ±0.23 / 7.79 ms │  1.10x slower │
│ QQuery 38 │        60.91 / 62.91 ±2.11 / 66.48 ms │        69.11 / 70.70 ±1.47 / 73.42 ms │  1.12x slower │
│ QQuery 39 │        88.54 / 89.77 ±1.27 / 91.94 ms │     103.08 / 105.52 ±1.80 / 107.82 ms │  1.18x slower │
│ QQuery 40 │        23.17 / 23.49 ±0.21 / 23.77 ms │        26.02 / 29.02 ±3.44 / 35.73 ms │  1.24x slower │
│ QQuery 41 │        11.24 / 11.38 ±0.11 / 11.50 ms │        12.23 / 12.36 ±0.07 / 12.42 ms │  1.09x slower │
│ QQuery 42 │        23.33 / 23.65 ±0.34 / 24.31 ms │        24.85 / 25.67 ±0.65 / 26.42 ms │  1.09x slower │
│ QQuery 43 │           5.19 / 5.29 ±0.12 / 5.53 ms │           5.55 / 5.71 ±0.15 / 5.98 ms │  1.08x slower │
│ QQuery 44 │           9.49 / 9.54 ±0.05 / 9.64 ms │         9.96 / 10.23 ±0.19 / 10.50 ms │  1.07x slower │
│ QQuery 45 │        38.03 / 39.23 ±1.35 / 41.81 ms │        42.72 / 43.59 ±0.80 / 44.60 ms │  1.11x slower │
│ QQuery 46 │        11.77 / 12.39 ±0.38 / 12.90 ms │        12.36 / 12.70 ±0.25 / 13.12 ms │     no change │
│ QQuery 47 │     219.66 / 223.17 ±2.52 / 227.38 ms │     259.75 / 266.01 ±5.69 / 273.35 ms │  1.19x slower │
│ QQuery 48 │        94.70 / 95.30 ±0.63 / 96.22 ms │      98.57 / 101.66 ±4.41 / 110.38 ms │  1.07x slower │
│ QQuery 49 │        69.91 / 72.82 ±2.94 / 78.21 ms │        74.15 / 75.19 ±0.74 / 76.07 ms │     no change │
│ QQuery 50 │        60.18 / 60.48 ±0.30 / 61.01 ms │        61.19 / 64.22 ±3.04 / 69.76 ms │  1.06x slower │
│ QQuery 51 │        92.98 / 94.51 ±0.91 / 95.34 ms │       93.76 / 97.42 ±2.15 / 100.53 ms │     no change │
│ QQuery 52 │        24.71 / 24.94 ±0.18 / 25.22 ms │        25.13 / 25.51 ±0.34 / 26.06 ms │     no change │
│ QQuery 53 │        30.00 / 32.54 ±4.20 / 40.90 ms │        29.65 / 30.08 ±0.24 / 30.27 ms │ +1.08x faster │
│ QQuery 54 │        55.87 / 56.77 ±0.74 / 58.10 ms │        57.47 / 60.42 ±5.29 / 70.99 ms │  1.06x slower │
│ QQuery 55 │        24.17 / 24.75 ±0.72 / 26.16 ms │        24.44 / 25.02 ±0.54 / 25.97 ms │     no change │
│ QQuery 56 │        40.09 / 40.59 ±0.52 / 41.52 ms │        40.98 / 41.72 ±0.55 / 42.47 ms │     no change │
│ QQuery 57 │     181.28 / 183.66 ±2.83 / 189.18 ms │     190.49 / 193.86 ±2.55 / 198.21 ms │  1.06x slower │
│ QQuery 58 │     113.44 / 115.00 ±1.23 / 117.19 ms │     115.67 / 118.17 ±3.01 / 123.63 ms │     no change │
│ QQuery 59 │     117.23 / 118.93 ±1.13 / 120.72 ms │     127.48 / 128.74 ±1.07 / 130.68 ms │  1.08x slower │
│ QQuery 60 │        40.37 / 42.24 ±3.17 / 48.56 ms │        40.82 / 41.59 ±0.43 / 42.15 ms │     no change │
│ QQuery 61 │        12.84 / 13.07 ±0.12 / 13.18 ms │        13.52 / 15.22 ±2.74 / 20.64 ms │  1.17x slower │
│ QQuery 62 │        47.77 / 47.90 ±0.11 / 48.04 ms │        49.08 / 49.60 ±0.65 / 50.87 ms │     no change │
│ QQuery 63 │        30.11 / 30.33 ±0.23 / 30.72 ms │        30.20 / 30.67 ±0.56 / 31.37 ms │     no change │
│ QQuery 64 │     385.31 / 388.72 ±2.94 / 393.66 ms │     386.87 / 395.79 ±6.74 / 407.06 ms │     no change │
│ QQuery 65 │     124.91 / 129.67 ±3.72 / 135.16 ms │     134.80 / 138.59 ±3.21 / 142.69 ms │  1.07x slower │
│ QQuery 66 │        81.03 / 84.51 ±1.76 / 85.82 ms │        84.80 / 85.65 ±0.59 / 86.28 ms │     no change │
│ QQuery 67 │     264.96 / 271.30 ±6.08 / 282.10 ms │     275.52 / 280.56 ±4.24 / 288.09 ms │     no change │
│ QQuery 68 │        12.56 / 12.69 ±0.18 / 13.04 ms │        13.13 / 13.75 ±0.84 / 15.40 ms │  1.08x slower │
│ QQuery 69 │        58.62 / 60.49 ±2.04 / 64.45 ms │        59.02 / 63.30 ±6.58 / 76.35 ms │     no change │
│ QQuery 70 │     109.38 / 111.13 ±1.55 / 112.88 ms │     112.43 / 114.75 ±3.26 / 121.12 ms │     no change │
│ QQuery 71 │        36.01 / 36.26 ±0.23 / 36.64 ms │        36.68 / 38.11 ±1.33 / 40.25 ms │  1.05x slower │
│ QQuery 72 │ 1907.51 / 1941.37 ±48.71 / 2037.27 ms │ 1887.64 / 2008.16 ±89.14 / 2127.37 ms │     no change │
│ QQuery 73 │        10.54 / 11.32 ±0.67 / 12.29 ms │         9.82 / 12.90 ±4.04 / 20.85 ms │  1.14x slower │
│ QQuery 74 │     196.12 / 199.34 ±2.99 / 204.15 ms │     175.46 / 177.00 ±1.62 / 180.01 ms │ +1.13x faster │
│ QQuery 75 │     150.88 / 155.36 ±4.85 / 164.60 ms │     148.59 / 150.99 ±3.52 / 157.86 ms │     no change │
│ QQuery 76 │        36.86 / 37.51 ±0.53 / 38.27 ms │        35.33 / 36.15 ±0.99 / 38.08 ms │     no change │
│ QQuery 77 │        62.67 / 63.46 ±0.75 / 64.45 ms │        61.63 / 62.00 ±0.32 / 62.40 ms │     no change │
│ QQuery 78 │    242.10 / 255.03 ±14.16 / 279.15 ms │    207.49 / 213.91 ±10.33 / 234.43 ms │ +1.19x faster │
│ QQuery 79 │        69.15 / 71.02 ±0.98 / 72.07 ms │        66.20 / 66.70 ±0.40 / 67.31 ms │ +1.06x faster │
│ QQuery 80 │     102.75 / 108.85 ±8.54 / 125.79 ms │      99.36 / 101.80 ±3.36 / 108.33 ms │ +1.07x faster │
│ QQuery 81 │        27.45 / 28.53 ±1.75 / 32.03 ms │        25.90 / 26.42 ±0.54 / 27.14 ms │ +1.08x faster │
│ QQuery 82 │        17.04 / 17.42 ±0.38 / 18.08 ms │        16.26 / 16.66 ±0.34 / 17.23 ms │     no change │
│ QQuery 83 │        34.97 / 35.41 ±0.31 / 35.84 ms │        34.09 / 34.29 ±0.14 / 34.47 ms │     no change │
│ QQuery 84 │        30.42 / 30.90 ±0.39 / 31.51 ms │        29.42 / 29.56 ±0.19 / 29.91 ms │     no change │
│ QQuery 85 │     105.07 / 107.92 ±3.74 / 115.24 ms │     102.13 / 105.37 ±2.93 / 110.58 ms │     no change │
│ QQuery 86 │        27.37 / 28.69 ±1.27 / 31.06 ms │        24.91 / 25.42 ±0.34 / 25.83 ms │ +1.13x faster │
│ QQuery 87 │        67.66 / 69.18 ±1.52 / 71.77 ms │        63.83 / 64.72 ±0.47 / 65.21 ms │ +1.07x faster │
│ QQuery 88 │        64.65 / 65.28 ±0.75 / 66.70 ms │        63.34 / 64.52 ±1.65 / 67.77 ms │     no change │
│ QQuery 89 │        36.83 / 38.80 ±3.15 / 45.05 ms │        35.40 / 38.61 ±4.84 / 48.25 ms │     no change │
│ QQuery 90 │        17.91 / 19.02 ±1.14 / 21.16 ms │        17.14 / 17.46 ±0.20 / 17.71 ms │ +1.09x faster │
│ QQuery 91 │        46.09 / 46.47 ±0.26 / 46.77 ms │        44.50 / 45.39 ±0.84 / 46.55 ms │     no change │
│ QQuery 92 │        30.14 / 30.95 ±0.58 / 31.86 ms │        28.90 / 29.74 ±1.14 / 31.95 ms │     no change │
│ QQuery 93 │        50.71 / 51.93 ±0.78 / 53.18 ms │        49.85 / 50.39 ±0.33 / 50.83 ms │     no change │
│ QQuery 94 │        38.68 / 40.91 ±1.58 / 43.29 ms │        37.78 / 39.34 ±1.52 / 41.56 ms │     no change │
│ QQuery 95 │        83.94 / 84.90 ±0.88 / 86.22 ms │        80.79 / 83.13 ±3.22 / 89.52 ms │     no change │
│ QQuery 96 │        24.37 / 24.65 ±0.22 / 24.96 ms │        23.76 / 24.13 ±0.29 / 24.58 ms │     no change │
│ QQuery 97 │        53.74 / 55.14 ±1.10 / 56.64 ms │        54.93 / 56.13 ±0.91 / 57.59 ms │     no change │
│ QQuery 98 │        43.04 / 45.53 ±1.31 / 46.78 ms │        42.53 / 42.93 ±0.30 / 43.21 ms │ +1.06x faster │
│ QQuery 99 │        70.59 / 71.76 ±1.50 / 74.64 ms │        70.99 / 73.06 ±3.90 / 80.85 ms │     no change │
└───────────┴───────────────────────────────────────┴───────────────────────────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┓
┃ Benchmark Summary                           ┃            ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━┩
│ Total Time (HEAD)                           │  9684.78ms │
│ Total Time (add-blocks-impl-from-scratch)   │ 10134.84ms │
│ Average Time (HEAD)                         │    97.83ms │
│ Average Time (add-blocks-impl-from-scratch) │   102.37ms │
│ Queries Faster                              │         10 │
│ Queries Slower                              │         44 │
│ Queries with No Change                      │         45 │
│ Queries with Failure                        │          0 │
└─────────────────────────────────────────────┴────────────┘

Resource Usage

tpcds — base (merge-base)

Metric Value
Wall time 50.0s
Peak memory 2.2 GiB
Avg memory 1.5 GiB
CPU user 213.7s
CPU sys 5.6s
Peak spill 0 B

tpcds — branch

Metric Value
Wall time 55.0s
Peak memory 2.0 GiB
Avg memory 1.3 GiB
CPU user 221.6s
CPU sys 6.2s
Peak spill 0 B

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing add-blocks-impl-from-scratch (78834e5) to 4048898 (merge-base) diff

Run configuration
run benchmark clickbench_partitioned
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and add-blocks-impl-from-scratch
--------------------
Benchmark clickbench_partitioned.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃       HEAD ┃ add-blocks-impl-from-scratch ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 0  │    1.24 ms │                      1.30 ms │     no change │
│ QQuery 1  │   11.80 ms │                     11.03 ms │ +1.07x faster │
│ QQuery 2  │   36.81 ms │                     35.57 ms │     no change │
│ QQuery 3  │   31.68 ms │                     30.84 ms │     no change │
│ QQuery 4  │  238.76 ms │                    312.66 ms │  1.31x slower │
│ QQuery 5  │  288.97 ms │                    314.38 ms │  1.09x slower │
│ QQuery 6  │    1.28 ms │                      1.35 ms │     no change │
│ QQuery 7  │   13.88 ms │                     12.10 ms │ +1.15x faster │
│ QQuery 8  │  355.21 ms │                    362.16 ms │     no change │
│ QQuery 9  │  479.60 ms │                    544.55 ms │  1.14x slower │
│ QQuery 10 │   71.68 ms │                     69.60 ms │     no change │
│ QQuery 11 │   83.23 ms │                     80.59 ms │     no change │
│ QQuery 12 │  283.99 ms │                    303.25 ms │  1.07x slower │
│ QQuery 13 │  368.85 ms │                    411.87 ms │  1.12x slower │
│ QQuery 14 │  288.40 ms │                    325.15 ms │  1.13x slower │
│ QQuery 15 │  273.47 ms │                    381.47 ms │  1.39x slower │
│ QQuery 16 │  622.77 ms │                    686.75 ms │  1.10x slower │
│ QQuery 17 │  617.21 ms │                    700.73 ms │  1.14x slower │
│ QQuery 18 │ 1266.00 ms │                   1356.13 ms │  1.07x slower │
│ QQuery 19 │   27.46 ms │                     26.35 ms │     no change │
│ QQuery 20 │  521.67 ms │                    523.41 ms │     no change │
│ QQuery 21 │  526.48 ms │                    519.98 ms │     no change │
│ QQuery 22 │ 1000.65 ms │                    993.77 ms │     no change │
│ QQuery 23 │ 3136.79 ms │                   3137.60 ms │     no change │
│ QQuery 24 │   41.52 ms │                     40.10 ms │     no change │
│ QQuery 25 │  114.02 ms │                    109.35 ms │     no change │
│ QQuery 26 │   41.67 ms │                     40.59 ms │     no change │
│ QQuery 27 │  518.60 ms │                    533.47 ms │     no change │
│ QQuery 28 │ 2989.39 ms │                   3009.32 ms │     no change │
│ QQuery 29 │   42.54 ms │                     41.94 ms │     no change │
│ QQuery 30 │  311.57 ms │                    336.47 ms │  1.08x slower │
│ QQuery 31 │  285.21 ms │                    286.45 ms │     no change │
│ QQuery 32 │  956.20 ms │                    981.79 ms │     no change │
│ QQuery 33 │ 1471.69 ms │                   1682.47 ms │  1.14x slower │
│ QQuery 34 │ 1593.22 ms │                   1703.10 ms │  1.07x slower │
│ QQuery 35 │  291.11 ms │                    382.71 ms │  1.31x slower │
│ QQuery 36 │   67.23 ms │                     71.78 ms │  1.07x slower │
│ QQuery 37 │   38.32 ms │                     37.07 ms │     no change │
│ QQuery 38 │   44.02 ms │                     40.78 ms │ +1.08x faster │
│ QQuery 39 │  142.99 ms │                    159.40 ms │  1.11x slower │
│ QQuery 40 │   14.54 ms │                     13.13 ms │ +1.11x faster │
│ QQuery 41 │   14.18 ms │                     12.42 ms │ +1.14x faster │
│ QQuery 42 │   13.83 ms │                     12.92 ms │ +1.07x faster │
└───────────┴────────────┴──────────────────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┓
┃ Benchmark Summary                           ┃            ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━┩
│ Total Time (HEAD)                           │ 19539.75ms │
│ Total Time (add-blocks-impl-from-scratch)   │ 20637.80ms │
│ Average Time (HEAD)                         │   454.41ms │
│ Average Time (add-blocks-impl-from-scratch) │   479.95ms │
│ Queries Faster                              │          6 │
│ Queries Slower                              │         16 │
│ Queries with No Change                      │         21 │
│ Queries with Failure                        │          0 │
└─────────────────────────────────────────────┴────────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and add-blocks-impl-from-scratch
--------------------
Benchmark clickbench_partitioned.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃                                  HEAD ┃          add-blocks-impl-from-scratch ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 0  │          1.24 / 4.14 ±5.62 / 15.38 ms │          1.30 / 4.25 ±5.70 / 15.64 ms │     no change │
│ QQuery 1  │        11.80 / 12.37 ±0.33 / 12.65 ms │        11.03 / 11.15 ±0.08 / 11.26 ms │ +1.11x faster │
│ QQuery 2  │        36.81 / 37.16 ±0.31 / 37.71 ms │        35.57 / 35.86 ±0.30 / 36.40 ms │     no change │
│ QQuery 3  │        31.68 / 32.60 ±0.50 / 33.15 ms │        30.84 / 31.88 ±1.10 / 33.48 ms │     no change │
│ QQuery 4  │     238.76 / 242.91 ±3.60 / 247.72 ms │     312.66 / 317.95 ±3.84 / 322.57 ms │  1.31x slower │
│ QQuery 5  │     288.97 / 291.94 ±2.29 / 295.34 ms │     314.38 / 320.41 ±5.25 / 326.99 ms │  1.10x slower │
│ QQuery 6  │           1.28 / 1.42 ±0.22 / 1.85 ms │           1.35 / 1.49 ±0.22 / 1.93 ms │     no change │
│ QQuery 7  │        13.88 / 14.92 ±1.68 / 18.28 ms │        12.10 / 13.31 ±2.00 / 17.29 ms │ +1.12x faster │
│ QQuery 8  │     355.21 / 356.92 ±2.35 / 361.50 ms │     362.16 / 369.29 ±5.41 / 377.20 ms │     no change │
│ QQuery 9  │     479.60 / 484.96 ±6.14 / 494.34 ms │     544.55 / 559.75 ±8.46 / 567.49 ms │  1.15x slower │
│ QQuery 10 │        71.68 / 76.04 ±6.51 / 88.95 ms │        69.60 / 75.13 ±5.92 / 82.84 ms │     no change │
│ QQuery 11 │        83.23 / 83.98 ±0.61 / 85.04 ms │        80.59 / 81.14 ±0.41 / 81.62 ms │     no change │
│ QQuery 12 │    283.99 / 295.02 ±12.41 / 318.37 ms │     303.25 / 312.02 ±7.89 / 321.81 ms │  1.06x slower │
│ QQuery 13 │     368.85 / 381.56 ±7.67 / 391.32 ms │    411.87 / 424.48 ±12.96 / 447.13 ms │  1.11x slower │
│ QQuery 14 │     288.40 / 294.52 ±4.31 / 300.87 ms │     325.15 / 331.46 ±5.19 / 339.09 ms │  1.13x slower │
│ QQuery 15 │     273.47 / 280.73 ±8.99 / 297.96 ms │    381.47 / 405.07 ±19.83 / 428.84 ms │  1.44x slower │
│ QQuery 16 │     622.77 / 629.55 ±3.74 / 633.60 ms │     686.75 / 691.43 ±3.13 / 695.79 ms │  1.10x slower │
│ QQuery 17 │     617.21 / 630.52 ±9.48 / 646.05 ms │    700.73 / 709.56 ±11.90 / 732.07 ms │  1.13x slower │
│ QQuery 18 │ 1266.00 / 1287.85 ±14.16 / 1305.53 ms │ 1356.13 / 1460.18 ±52.50 / 1498.91 ms │  1.13x slower │
│ QQuery 19 │        27.46 / 27.93 ±0.37 / 28.45 ms │        26.35 / 33.86 ±9.01 / 47.78 ms │  1.21x slower │
│ QQuery 20 │    521.67 / 531.05 ±11.99 / 552.73 ms │    523.41 / 541.68 ±11.96 / 556.34 ms │     no change │
│ QQuery 21 │     526.48 / 536.73 ±9.62 / 553.01 ms │    519.98 / 529.36 ±10.57 / 549.28 ms │     no change │
│ QQuery 22 │ 1000.65 / 1020.79 ±13.57 / 1040.98 ms │  993.77 / 1012.67 ±18.06 / 1046.45 ms │     no change │
│ QQuery 23 │ 3136.79 / 3194.10 ±37.65 / 3245.70 ms │ 3137.60 / 3167.68 ±21.99 / 3199.78 ms │     no change │
│ QQuery 24 │        41.52 / 42.58 ±1.38 / 45.19 ms │        40.10 / 42.84 ±3.46 / 49.61 ms │     no change │
│ QQuery 25 │     114.02 / 118.66 ±4.02 / 125.95 ms │     109.35 / 111.15 ±1.66 / 113.70 ms │ +1.07x faster │
│ QQuery 26 │        41.67 / 42.95 ±0.79 / 44.01 ms │        40.59 / 41.86 ±1.13 / 43.34 ms │     no change │
│ QQuery 27 │     518.60 / 529.34 ±6.53 / 538.10 ms │     533.47 / 545.42 ±7.44 / 555.30 ms │     no change │
│ QQuery 28 │  2989.39 / 2998.00 ±8.93 / 3013.56 ms │ 3009.32 / 3037.05 ±23.91 / 3071.72 ms │     no change │
│ QQuery 29 │       42.54 / 63.03 ±18.73 / 91.47 ms │       41.94 / 58.99 ±21.93 / 95.91 ms │ +1.07x faster │
│ QQuery 30 │     311.57 / 318.82 ±8.16 / 333.79 ms │     336.47 / 345.26 ±6.95 / 356.63 ms │  1.08x slower │
│ QQuery 31 │     285.21 / 292.98 ±6.45 / 304.77 ms │     286.45 / 293.95 ±5.59 / 300.51 ms │     no change │
│ QQuery 32 │  956.20 / 1000.27 ±36.96 / 1067.24 ms │  981.79 / 1007.80 ±28.93 / 1061.44 ms │     no change │
│ QQuery 33 │ 1471.69 / 1517.71 ±46.25 / 1604.72 ms │ 1682.47 / 1719.67 ±39.58 / 1795.40 ms │  1.13x slower │
│ QQuery 34 │ 1593.22 / 1608.44 ±11.93 / 1628.97 ms │ 1703.10 / 1800.87 ±60.69 / 1876.95 ms │  1.12x slower │
│ QQuery 35 │    291.11 / 323.85 ±30.02 / 380.45 ms │    382.71 / 446.93 ±59.31 / 536.92 ms │  1.38x slower │
│ QQuery 36 │        67.23 / 78.35 ±6.58 / 85.16 ms │        71.78 / 81.13 ±8.01 / 92.46 ms │     no change │
│ QQuery 37 │        38.32 / 40.38 ±2.57 / 45.42 ms │        37.07 / 42.02 ±3.90 / 48.02 ms │     no change │
│ QQuery 38 │        44.02 / 45.83 ±1.81 / 49.32 ms │        40.78 / 47.30 ±7.14 / 60.92 ms │     no change │
│ QQuery 39 │    142.99 / 157.71 ±10.98 / 170.96 ms │     159.40 / 164.43 ±4.62 / 172.80 ms │     no change │
│ QQuery 40 │        14.54 / 14.85 ±0.37 / 15.56 ms │        13.13 / 13.46 ±0.38 / 14.20 ms │ +1.10x faster │
│ QQuery 41 │        14.18 / 15.76 ±2.62 / 20.95 ms │        12.42 / 12.80 ±0.36 / 13.24 ms │ +1.23x faster │
│ QQuery 42 │        13.83 / 14.07 ±0.17 / 14.22 ms │        12.92 / 14.58 ±2.99 / 20.55 ms │     no change │
└───────────┴───────────────────────────────────────┴───────────────────────────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┓
┃ Benchmark Summary                           ┃            ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━┩
│ Total Time (HEAD)                           │ 19973.29ms │
│ Total Time (add-blocks-impl-from-scratch)   │ 21268.55ms │
│ Average Time (HEAD)                         │   464.50ms │
│ Average Time (add-blocks-impl-from-scratch) │   494.62ms │
│ Queries Faster                              │          6 │
│ Queries Slower                              │         15 │
│ Queries with No Change                      │         22 │
│ Queries with Failure                        │          0 │
└─────────────────────────────────────────────┴────────────┘

Resource Usage

clickbench_partitioned — base (merge-base)

Metric Value
Wall time 105.0s
Peak memory 11.7 GiB
Avg memory 4.5 GiB
CPU user 1022.1s
CPU sys 74.1s
Peak spill 0 B

clickbench_partitioned — branch

Metric Value
Wall time 110.0s
Peak memory 11.4 GiB
Avg memory 4.6 GiB
CPU user 1094.0s
CPU sys 77.9s
Peak spill 0 B

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing add-blocks-impl-from-scratch (78834e5) to 4048898 (merge-base) diff

Run configuration
run benchmark external_aggr
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and add-blocks-impl-from-scratch
--------------------
Benchmark external_aggr.json
--------------------
┏━━━━━━━━━━━━━━┳━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━┓
┃ Query        ┃      HEAD ┃ add-blocks-impl-from-scratch ┃       Change ┃
┡━━━━━━━━━━━━━━╇━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━┩
│ Q1(64.0 MB)  │  51.57 ms │                     64.53 ms │ 1.25x slower │
│ Q1(32.0 MB)  │  51.31 ms │                     55.53 ms │ 1.08x slower │
│ Q1(16.0 MB)  │  47.90 ms │                     52.53 ms │ 1.10x slower │
│ Q2(512.0 MB) │ 277.39 ms │                    291.35 ms │ 1.05x slower │
│ Q2(256.0 MB) │ 254.93 ms │                    266.64 ms │    no change │
│ Q2(128.0 MB) │ 239.40 ms │                    264.40 ms │ 1.10x slower │
│ Q2(64.0 MB)  │ 237.10 ms │                    260.94 ms │ 1.10x slower │
│ Q2(32.0 MB)  │ 293.12 ms │                    319.73 ms │ 1.09x slower │
└──────────────┴───────────┴──────────────────────────────┴──────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Benchmark Summary                           ┃           ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ Total Time (HEAD)                           │ 1452.72ms │
│ Total Time (add-blocks-impl-from-scratch)   │ 1575.64ms │
│ Average Time (HEAD)                         │  181.59ms │
│ Average Time (add-blocks-impl-from-scratch) │  196.96ms │
│ Queries Faster                              │         0 │
│ Queries Slower                              │         7 │
│ Queries with No Change                      │         1 │
│ Queries with Failure                        │         0 │
└─────────────────────────────────────────────┴───────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and add-blocks-impl-from-scratch
--------------------
Benchmark external_aggr.json
--------------------
┏━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━┓
┃ Query        ┃                               HEAD ┃       add-blocks-impl-from-scratch ┃       Change ┃
┡━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━┩
│ Q1(64.0 MB)  │     51.57 / 54.29 ±2.57 / 59.17 ms │     64.53 / 68.85 ±5.10 / 78.69 ms │ 1.27x slower │
│ Q1(32.0 MB)  │     51.31 / 52.10 ±0.45 / 52.53 ms │     55.53 / 56.45 ±1.35 / 59.13 ms │ 1.08x slower │
│ Q1(16.0 MB)  │     47.90 / 50.39 ±1.39 / 51.90 ms │     52.53 / 54.26 ±2.06 / 58.19 ms │ 1.08x slower │
│ Q2(512.0 MB) │ 277.39 / 288.72 ±10.82 / 305.67 ms │ 291.35 / 303.78 ±12.84 / 324.98 ms │ 1.05x slower │
│ Q2(256.0 MB) │ 254.93 / 285.34 ±23.12 / 315.67 ms │  266.64 / 271.33 ±2.70 / 274.60 ms │    no change │
│ Q2(128.0 MB) │  239.40 / 245.30 ±7.30 / 257.94 ms │  264.40 / 268.77 ±5.06 / 278.33 ms │ 1.10x slower │
│ Q2(64.0 MB)  │  237.10 / 238.44 ±1.11 / 239.79 ms │  260.94 / 264.56 ±4.00 / 272.16 ms │ 1.11x slower │
│ Q2(32.0 MB)  │  293.12 / 298.12 ±3.79 / 302.35 ms │  319.73 / 323.91 ±3.99 / 329.31 ms │ 1.09x slower │
└──────────────┴────────────────────────────────────┴────────────────────────────────────┴──────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Benchmark Summary                           ┃           ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ Total Time (HEAD)                           │ 1512.69ms │
│ Total Time (add-blocks-impl-from-scratch)   │ 1611.90ms │
│ Average Time (HEAD)                         │  189.09ms │
│ Average Time (add-blocks-impl-from-scratch) │  201.49ms │
│ Queries Faster                              │         0 │
│ Queries Slower                              │         7 │
│ Queries with No Change                      │         1 │
│ Queries with Failure                        │         0 │
└─────────────────────────────────────────────┴───────────┘

Memory Pool Peaks

Peak MemoryPool reservation per query — what DataFusion's accounting believes it reserved. Recorded only when the benchmark runs with DATAFUSION_RUNTIME_MEMORY_LIMIT set.

Base: 4048898 (merge-base) | Changed: add-blocks-impl-from-scratch

external_aggr

Query Base Changed Change
1(64.0 MB) 35.0 MiB 29.2 MiB -16.5%
1(32.0 MB) 18.6 MiB 18.6 MiB -0.1%
1(16.0 MB) 11.4 MiB 13.4 MiB +17.5%
2(512.0 MB) 136.2 MiB 186.3 MiB +36.8%
2(256.0 MB) 97.9 MiB 100.1 MiB +2.2%
2(128.0 MB) 49.4 MiB 52.7 MiB +6.8%
2(64.0 MB) 30.5 MiB 29.0 MiB -5.1%
2(32.0 MB) 30.0 MiB 30.0 MiB +0.0%

Pool accounting vs. process RSS

Max pool peak is the largest reservation any single query in the run reached; peak RSS covers the whole invocation, including data loading and allocator retention, and the two high-water marks need not coincide in time. The gap is therefore an upper bound on what the pool did not account for, not a measurement of it.

Benchmark Side Max pool peak Peak RSS Gap RSS / pool
external_aggr base (4048898 (merge-base)) 136.2 MiB 435.7 MiB 299.5 MiB 3.2×
external_aggr changed (add-blocks-impl-from-scratch) 186.3 MiB 531.6 MiB 345.3 MiB 2.9×
Resource Usage

external_aggr — base (merge-base)

Metric Value
Wall time 525.1s
Peak memory 435.7 MiB
Avg memory 9.2 MiB
CPU user 25.7s
CPU sys 3.7s
Peak spill 0 B

external_aggr — branch

Metric Value
Wall time 560.1s
Peak memory 531.6 MiB
Avg memory 10.1 MiB
CPU user 27.1s
CPU sys 3.2s
Peak spill 0 B

File an issue against this benchmark runner

# Conflicts:
#	datafusion/functions-aggregate-common/src/aggregate/groups_accumulator/accumulate.rs
#	datafusion/functions-aggregate/src/average.rs
@rluvaton
rluvaton force-pushed the add-blocks-impl-from-scratch branch from f935f74 to db6903c Compare September 16, 2026 17:03
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

auto detected api change Auto detected API change common Related to common crate core Core DataFusion crate functions Changes to functions implementation logical-expr Logical plan and expressions optimizer Optimizer rules physical-expr Changes to the physical-expr crates physical-plan Changes to the physical-plan crate proto Related to proto crate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants