Skip to content

feat(perf): Improve filter performance with per word bit filtering (and BMI when supported) - #10136

Merged
alamb merged 28 commits into
apache:mainfrom
devanbenz:db/10098/bmi-null-bool-filter
Sep 25, 2026
Merged

alamb merged 28 commits into
apache:mainfrom
devanbenz:db/10098/bmi-null-bool-filter

Conversation

@devanbenz

@devanbenz devanbenz commented Jun 12, 2026 •

Copy link
Copy Markdown
Contributor

This commit adds the ability for bit filtering to be done using the _pext_u64 BMI with a scalar fallback.

We cannot run the BMI2 feature on the benchmark host machine, here are the benchmarks from my machine:

Specs:

Architecture:                x86_64
  CPU op-mode(s):            32-bit, 64-bit
  Address sizes:             46 bits physical, 48 bits virtual
  Byte Order:                Little Endian
Vendor ID:                   GenuineIntel
  Model name:                12th Gen Intel(R) Core(TM) i7-12700K
Case Build main #10136 #10136 vs main
filter context i32 w NULLs (kept 1/2) scalar 47.54 µs 28.52 µs -40.0%
bmi2 53.28 µs 10.63 µs -80.1%
filter context u8 w NULLs (kept 1/2) scalar 39.32 µs 26.41 µs -32.9%
bmi2 37.90 µs 9.10 µs -76.0%
filter context string dictionary w NULLs (kept 1/2) scalar 36.39 µs 28.40 µs -21.9%
bmi2 44.88 µs 10.69 µs -76.2%
filter f32 (kept 1/2) scalar 61.87 µs 39.97 µs -35.4%
bmi2 61.96 µs 22.74 µs -63.3%
filter context f32 (kept 1/2) scalar 44.43 µs 28.40 µs -36.1%
bmi2 36.93 µs 10.80 µs -70.8%
filter context short string view (kept 1/2) scalar 51.30 µs 44.48 µs -13.3%
bmi2 50.56 µs 29.11 µs -42.4%
filter context mixed string view (kept 1/2) scalar 54.59 µs 43.77 µs -19.8%
bmi2 56.13 µs 26.35 µs -53.1%
boolean 65 536 bits, kept 1/2 (filter context fsb …, 9 rows) scalar 28.74–39.63 µs 19.15–19.39 µs -51…-33%
bmi2 32.59–41.92 µs 1.71–1.77 µs -96…-95%

With these changes we see the following improvements for filter kernels

Build Avg improvement Including boolean row
scalar ~28.5% faster ~30.2% faster
bmi2 ~66.0% faster ~69.7% faster
overall ~47.2% faster ~49.9% faster

@github-actions github-actions Bot added the arrow Changes to the arrow crate label Jun 12, 2026
Comment thread arrow-buffer/src/util/bit_util.rs
@devanbenz

Copy link
Copy Markdown
Contributor Author

run benchmark arrow-select

@adriangbot

Copy link
Copy Markdown

Hi @devanbenz, thanks for the request (#10136 (comment)). Only whitelisted users can trigger benchmarks. Allowed users: Dandandan, Fokko, Jefffrey, Omega359, adriangb, alamb, asubiotto, brunal, buraksenn, cetra3, codephage2020, coderfender, comphead, erenavsarogullari, etseidl, friendlymatthew, gabotechs, geoffreyclaude, grtlr, haohuaijin, jonathanc-n, kevinjqliu, klion26, kosiew, kumarUjjawal, kunalsinghdadhwal, liamzwbao, mbutrovich, mkleen, mzabaluev, neilconway, rluvaton, sdf-jkl, timsaucer, xudong963, zhuqi-lucas.


File an issue against this benchmark runner

@alamb

alamb commented Jun 23, 2026

Copy link
Copy Markdown
Contributor

run benchmark arrow-select

@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c4782118681-633-b8j47 6.12.68+ #1 SMP Sat May 2 07:49:07 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing db/10098/bmi-null-bool-filter (76c7efc) to 1ba5d48 (merge-base) diff
BENCH_NAME=arrow-select
BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench arrow-select
BENCH_FILTER=
Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

Benchmark for this request failed.

Last 20 lines of output:

Click to expand
    record_batch
    regexp_kernels
    row_format
    row_group_index_reader
    row_selection_cursor
    row_selector
    serde
    sort_kernel
    string_dictionary_builder
    string_run_builder
    string_run_iterator
    substring_kernels
    take_kernels
    union_array
    variant_builder
    variant_kernels
    variant_validation
    view_types
    writer_overhead
    zip_kernels

File an issue against this benchmark runner

@alamb

alamb commented Jun 23, 2026

Copy link
Copy Markdown
Contributor

run benchmark arrow_select

@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c4783458128-638-vs74c 6.12.68+ #1 SMP Sat May 2 07:49:07 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing db/10098/bmi-null-bool-filter (76c7efc) to 1ba5d48 (merge-base) diff
BENCH_NAME=arrow_select
BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench arrow_select
BENCH_FILTER=
Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

Benchmark for this request failed.

Last 20 lines of output:

Click to expand
    record_batch
    regexp_kernels
    row_format
    row_group_index_reader
    row_selection_cursor
    row_selector
    serde
    sort_kernel
    string_dictionary_builder
    string_run_builder
    string_run_iterator
    substring_kernels
    take_kernels
    union_array
    variant_builder
    variant_kernels
    variant_validation
    view_types
    writer_overhead
    zip_kernels

File an issue against this benchmark runner

@devanbenz

Copy link
Copy Markdown
Contributor Author

run benchmark arrow-select

@alamb my apologies, the bench is called filter_bits :P

@alamb

alamb commented Jun 24, 2026

Copy link
Copy Markdown
Contributor

run benchmark filter_bits

@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c4792893778-664-tn2q8 6.12.68+ #1 SMP Sat May 2 07:49:07 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing db/10098/bmi-null-bool-filter (76c7efc) to 1ba5d48 (merge-base) diff
BENCH_NAME=filter_bits
BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench filter_bits
BENCH_FILTER=
Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

New benchmark — branch-only results (no baseline comparison)

Details

group                                             db_10098_bmi-null-bool-filter
-----                                             -----------------------------
filter_bits (indices, kept 1/10)                  1.00     13.5±0.18µs        ? ?/sec
filter_bits (indices, kept 1/1024)                1.00   1014.5±3.55ns        ? ?/sec
filter_bits (indices, kept 1/2)                   1.00     77.3±0.43µs        ? ?/sec
filter_bits (slices, kept 1023/1024)              1.00      2.5±0.10µs        ? ?/sec
filter_bits (slices, kept 9/10)                   1.00     88.1±0.35µs        ? ?/sec
filter_bits optimized (indices, kept 1/10)        1.00     14.7±1.27µs        ? ?/sec
filter_bits optimized (indices, kept 1/1024)      1.00   225.8±12.39ns        ? ?/sec
filter_bits optimized (indices, kept 1/2)         1.00     73.9±6.47µs        ? ?/sec
filter_bits optimized (slices, kept 1023/1024)    1.00   1706.1±3.71ns        ? ?/sec
filter_bits optimized (slices, kept 9/10)         1.00     83.6±1.10µs        ? ?/sec
filter_bits sliced (indices, kept 1/10)           1.00     13.5±0.18µs        ? ?/sec
filter_bits sliced (indices, kept 1/1024)         1.00   1014.6±2.75ns        ? ?/sec
filter_bits sliced (indices, kept 1/2)            1.00     77.1±0.38µs        ? ?/sec
filter_bits sliced (slices, kept 1023/1024)       1.00      2.4±0.01µs        ? ?/sec
filter_bits sliced (slices, kept 9/10)            1.00     88.9±0.27µs        ? ?/sec

Resource Usage

branch

Metric Value
Wall time 155.0s
Peak memory 6.8 MiB
Avg memory 5.2 MiB
CPU user 148.7s
CPU sys 0.0s
Peak spill 0 B

File an issue against this benchmark runner

@devanbenz

Copy link
Copy Markdown
Contributor Author

@alamb Could you please run the benchmark again but setting RUSTFLAGS="-C target-feature=-bmi2" to disable BMI on the host device? 🙏

@alamb

This comment was marked as outdated.

@adriangbot

This comment was marked as outdated.

@alamb

alamb commented Jun 29, 2026

Copy link
Copy Markdown
Contributor

run benchmark filter_bits

env:
RUSTFLAGS: "-C target-feature=-bmi2"

@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c4833873588-734-dldxz 6.12.85+ #1 SMP Mon May 11 08:17:35 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing db/10098/bmi-null-bool-filter (8e326b4) to 9f37683 (merge-base) diff
BENCH_NAME=filter_bits
BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench filter_bits
BENCH_FILTER=
Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

New benchmark — branch-only results (no baseline comparison)

Details

group                                             db_10098_bmi-null-bool-filter
-----                                             -----------------------------
filter_bits (indices, kept 1/10)                  1.00     13.5±0.17µs        ? ?/sec
filter_bits (indices, kept 1/1024)                1.00   1011.4±0.74ns        ? ?/sec
filter_bits (indices, kept 1/2)                   1.00     76.8±0.32µs        ? ?/sec
filter_bits (slices, kept 1023/1024)              1.00      2.5±0.00µs        ? ?/sec
filter_bits (slices, kept 9/10)                   1.00     95.1±0.30µs        ? ?/sec
filter_bits optimized (indices, kept 1/10)        1.00     14.3±1.07µs        ? ?/sec
filter_bits optimized (indices, kept 1/1024)      1.00    220.3±9.41ns        ? ?/sec
filter_bits optimized (indices, kept 1/2)         1.00     72.2±5.38µs        ? ?/sec
filter_bits optimized (slices, kept 1023/1024)    1.00   1707.5±3.10ns        ? ?/sec
filter_bits optimized (slices, kept 9/10)         1.00     62.6±0.70µs        ? ?/sec
filter_bits sliced (indices, kept 1/10)           1.00     13.5±0.18µs        ? ?/sec
filter_bits sliced (indices, kept 1/1024)         1.00   1011.5±0.76ns        ? ?/sec
filter_bits sliced (indices, kept 1/2)            1.00     76.7±0.38µs        ? ?/sec
filter_bits sliced (slices, kept 1023/1024)       1.00      2.4±0.00µs        ? ?/sec
filter_bits sliced (slices, kept 9/10)            1.00     96.1±0.31µs        ? ?/sec

Resource Usage

branch

Metric Value
Wall time 145.0s
Peak memory 9.1 MiB
Avg memory 5.0 MiB
CPU user 141.8s
CPU sys 0.0s
Peak spill 0 B

File an issue against this benchmark runner

@devanbenz

Copy link
Copy Markdown
Contributor Author

@alamb Taking a look at the host machines. It doesn't look like the processor supports BMI2 intrinsics. So the benches are basically equal.

@devanbenz

devanbenz commented Jul 2, 2026 •

Copy link
Copy Markdown
Contributor Author

Now that #9848 Is merged in, I'm going to merge this PR with it and move the compress BMI method in to a shared arrow package. i.e. arrow-buffer.

alamb added a commit that referenced this pull request Sep 24, 2026
# Which issue does this PR close?

- follow up to review discussion on
#10136

# Rationale for this change

The "Test Release Mode" job in `arrow.yml` compiles with the default
`x86_64` target, which only enables SSE2:

```
$ rustc | grep target_feature
target_feature="fxsr"
target_feature="sse"
target_feature="sse2"
```

So any code gated on `cfg(target_feature = ...)`, such as the AVX / AVX2
paths in `arrow-arith/src/aggregate.rs` and
`arrow-array/src/array/union_array.rs`, or the BMI2 `pext` path proposed
in #10136, is never compiled or tested on CI.

# What changes are included in this PR?

Add `-C target-cpu=native` to `RUSTFLAGS` for the `linux-release-test`
job, so we test with every instruction set extension of the runner CPU.

# Are these changes tested?

By CI: this job should now build and run the tests with the target
features enabled.

Example:
https://github.com/apache/arrow-rs/actions/runs/35922205002/job/107388798752?pr=11192

<img width="1160" height="624" alt="Screenshot 2026-09-23 at 5 29 40 PM"
src="https://github.com/user-attachments/assets/dac161bf-a918-4e86-88c3-e3584ecc0cf2"
/>


# Are there any user-facing changes?

No

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
/// bits of the result, in their original order:
///
/// ```text
/// bit: 7 6 5 4 3 2 1 0

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Apparently there is no CI coverage for this yet (the existing release job doesn't use all the fancy architectural features). I made a PR to propose doing this

@alamb

alamb commented Sep 24, 2026

Copy link
Copy Markdown
Contributor

What do you think about having the fallback walk whichever set of bits is smaller? The current loop runs once per set bit in mask, so at 9/10 density it takes about 58 iterations per word. Each iteration is 9 instructions on aarch64:

In my opinion, we should merge this PR and then improve the fallback code as a follow on issue / PR

It seems like this PR is already faster even with the somewhat basic scalar fallback. I am sure we can all then geek out trying to improve the performance of the fallback with more crazy bithacks (this is a good thing)

I looked over comments, and it looks to me like these are the only remaining outstanding comments about adding some additional asserts

For fun, I will re-run my benchmark run with the latest fixes

@devanbenz

Copy link
Copy Markdown
Contributor Author

I am sure we can all then geek out trying to improve the performance of the fallback with more crazy bithacks (this is a good thing)

Reading hackers delight as we speak

@devanbenz

Copy link
Copy Markdown
Contributor Author

What do you think about having the fallback walk whichever set of bits is smaller? The current loop runs once per set bit in mask, so at 9/10 density it takes about 58 iterations per word. Each iteration is 9 instructions on aarch64:

In my opinion, we should merge this PR and then improve the fallback code as a follow on issue / PR

It seems like this PR is already faster even with the somewhat basic scalar fallback. I am sure we can all then geek out trying to improve the performance of the fallback with more crazy bithacks (this is a good thing)

I looked over comments, and it looks to me like these are the only remaining outstanding comments about adding some additional asserts

* [feat(perf): Improve filter performance with per word bit filtering (and `BMI` when supported) #10136 (comment)](https://github.com/apache/arrow-rs/pull/10136#discussion_r4087289204) from me and @mbutrovich

* [feat(perf): Improve filter performance with per word bit filtering (and `BMI` when supported) #10136 (comment)](https://github.com/apache/arrow-rs/pull/10136#discussion_r4094224975) from @mbutrovich

For fun, I will re-run my benchmark run with the latest fixes

I've gone ahead and added the suggested debug_asserts. I've also added a new benchmark for FSB which includes nulls to stress the code path in filter_bits.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@alamb

alamb commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

I merged up from main to resolve a merge conflict. I also re-ran this code on

==> Host CPU
Vendor ID:                               GenuineIntel
Model name:                              Intel(R) Xeon(R) CPU @ 3.10GHz

==> Target features enabled by -C target-cpu=native
target_feature="adx" target_feature="aes" target_feature="avx" target_feature="avx2" target_feature="avx512bw" target_feature="avx512cd" target_feature="avx512dq" target_feature="avx512f" target_feature="avx512vl" target_feature="avx512vnni" target_feature="bmi1" target_feature="bmi2" target_feature="cmpxchg16b" target_feature="f16c" target_feature="fma" target_feature="fxsr" target_feature="lzcnt" target_feature="movbe" target_feature="pclmulqdq" target_feature="popcnt" target_feature="rdrand" target_feature="rdseed" target_feature="sse" target_feature="sse2" target_feature="sse3" target_feature="sse4.1" target_feature="sse4.2" target_feature="ssse3" target_feature="xsave" target_feature="xsavec" target_feature="xsaveopt" target_feature="xsaves"
bmi2: ENABLED (pext path will be compiled)

There is quite a bit of noise, but I think overall this looks like a win to me

Details
==> Paired runs, only rows differing by more than 5%
--- base-1 vs pr-1
group                                                                         base-1                                 pr-1
-----                                                                         ------                                 ----
filter context decimal128 (kept 1/2)                                          1.00     46.4±0.92µs        ? ?/sec    1.39     64.4±0.45µs        ? ?/sec
filter context decimal128 high selectivity (kept 1023/1024)                   1.00     48.9±1.40µs        ? ?/sec    1.12     54.8±3.52µs        ? ?/sec
filter context f32 (kept 1/2)                                                 1.77    117.8±0.19µs        ? ?/sec    1.00     66.7±0.65µs        ? ?/sec
filter context fsb with value length 20 (kept 1/2)                            1.00     73.1±3.50µs        ? ?/sec    1.19     86.9±2.59µs        ? ?/sec
filter context fsb with value length 50 (kept 1/2)                            1.00   163.0±12.11µs        ? ?/sec    1.19    194.0±3.51µs        ? ?/sec
filter context fsb with value length 50 high selectivity (kept 1023/1024)     1.00    187.4±3.81µs        ? ?/sec    1.28    239.0±0.97µs        ? ?/sec
filter context fsb with value length 50 low selectivity (kept 1/1024)         1.14    341.1±4.34ns        ? ?/sec    1.00    298.6±2.94ns        ? ?/sec
filter context i32 high selectivity (kept 1023/1024)                          1.17      7.4±0.29µs        ? ?/sec    1.00      6.3±0.23µs        ? ?/sec
filter context i32 w NULLs (kept 1/2)                                         1.77    118.2±1.12µs        ? ?/sec    1.00     66.7±0.07µs        ? ?/sec
filter context i32 w NULLs high selectivity (kept 1023/1024)                  1.06      8.6±0.36µs        ? ?/sec    1.00      8.1±0.03µs        ? ?/sec
filter context mixed string view (kept 1/2)                                   1.72    104.7±1.01µs        ? ?/sec    1.00     60.9±4.73µs        ? ?/sec
filter context mixed string view high selectivity (kept 1023/1024)            1.00     54.0±0.91µs        ? ?/sec    1.19     64.3±0.18µs        ? ?/sec
filter context short string view (kept 1/2)                                   1.73    109.5±2.18µs        ? ?/sec    1.00     63.3±5.50µs        ? ?/sec
filter context string (kept 1/2)                                              1.14   563.7±10.80µs        ? ?/sec    1.00    496.5±2.11µs        ? ?/sec
filter context string dictionary w NULLs (kept 1/2)                           1.77    118.3±1.44µs        ? ?/sec    1.00     66.9±0.68µs        ? ?/sec
filter context string dictionary w NULLs high selectivity (kept 1023/1024)    1.00      8.3±0.28µs        ? ?/sec    1.07      8.9±0.26µs        ? ?/sec
filter context u8 high selectivity (kept 1023/1024)                           1.00  1727.5±31.04ns        ? ?/sec    1.06  1824.7±10.44ns        ? ?/sec
filter context u8 w NULLs (kept 1/2)                                          4.44     66.2±1.23µs        ? ?/sec    1.00     14.9±0.15µs        ? ?/sec
filter decimal128 (kept 1/2)                                                  1.00     39.2±4.29µs        ? ?/sec    1.20     47.2±6.34µs        ? ?/sec
filter decimal128 high selectivity (kept 1023/1024)                           1.00     51.2±1.46µs        ? ?/sec    1.16     59.5±5.40µs        ? ?/sec
filter f32 (kept 1/2)                                                         3.61    124.1±0.42µs        ? ?/sec    1.00     34.4±0.31µs        ? ?/sec
filter fsb with value length 20 high selectivity (kept 1023/1024)             1.00     70.1±2.07µs        ? ?/sec    1.08     75.7±3.48µs        ? ?/sec
filter fsb with value length 20 low selectivity (kept 1/1024)                 1.00      2.2±0.02µs        ? ?/sec    1.31      2.9±0.00µs        ? ?/sec
filter fsb with value length 5 low selectivity (kept 1/1024)                  1.00      2.1±0.02µs        ? ?/sec    1.29      2.7±0.03µs        ? ?/sec
filter fsb with value length 50 (kept 1/2)                                    1.00    140.6±6.68µs        ? ?/sec    1.25   176.1±17.51µs        ? ?/sec
filter fsb with value length 50 high selectivity (kept 1023/1024)             1.00   198.6±15.64µs        ? ?/sec    1.15    227.9±3.12µs        ? ?/sec
filter fsb with value length 50 low selectivity (kept 1/1024)                 1.00      2.2±0.01µs        ? ?/sec    1.31      2.8±0.01µs        ? ?/sec
--- base-2 vs pr-2
group                                                                        base-2                                 pr-2
-----                                                                        ------                                 ----
filter context decimal128 (kept 1/2)                                         1.16     67.2±3.46µs        ? ?/sec    1.00     57.7±5.57µs        ? ?/sec
filter context decimal128 high selectivity (kept 1023/1024)                  1.00     49.9±0.62µs        ? ?/sec    1.16     57.6±3.54µs        ? ?/sec
filter context decimal128 low selectivity (kept 1/1024)                      1.09    189.7±3.57ns        ? ?/sec    1.00    174.3±1.89ns        ? ?/sec
filter context f32 (kept 1/2)                                                1.77    118.0±1.10µs        ? ?/sec    1.00     66.8±0.63µs        ? ?/sec
filter context fsb with value length 20 (kept 1/2)                           1.16     84.8±3.18µs        ? ?/sec    1.00     72.8±1.97µs        ? ?/sec
filter context fsb with value length 20 high selectivity (kept 1023/1024)    1.17     78.6±0.86µs        ? ?/sec    1.00     66.9±2.79µs        ? ?/sec
filter context fsb with value length 20 low selectivity (kept 1/1024)        1.00    300.6±4.32ns        ? ?/sec    1.11    333.1±3.88ns        ? ?/sec
filter context fsb with value length 5 high selectivity (kept 1023/1024)     1.25      9.9±0.22µs        ? ?/sec    1.00      7.9±0.25µs        ? ?/sec
filter context fsb with value length 50 (kept 1/2)                           1.29    195.4±3.61µs        ? ?/sec    1.00    151.6±8.88µs        ? ?/sec
filter context fsb with value length 50 high selectivity (kept 1023/1024)    1.08    204.5±6.19µs        ? ?/sec    1.00    189.3±6.46µs        ? ?/sec
filter context fsb with value length 50 low selectivity (kept 1/1024)        1.00    312.3±4.42ns        ? ?/sec    1.07    332.8±3.96ns        ? ?/sec
filter context i32 high selectivity (kept 1023/1024)                         1.00      5.9±0.17µs        ? ?/sec    1.05      6.2±0.23µs        ? ?/sec
filter context i32 w NULLs (kept 1/2)                                        1.77    117.9±1.56µs        ? ?/sec    1.00     66.8±0.09µs        ? ?/sec
filter context mixed string view (kept 1/2)                                  1.96    112.9±2.79µs        ? ?/sec    1.00     57.6±2.08µs        ? ?/sec
filter context short string view (kept 1/2)                                  1.98    115.5±3.96µs        ? ?/sec    1.00     58.3±0.83µs        ? ?/sec
filter context short string view high selectivity (kept 1023/1024)           1.22     63.9±3.03µs        ? ?/sec    1.00     52.4±1.78µs        ? ?/sec
filter context string (kept 1/2)                                             1.14    545.3±7.43µs        ? ?/sec    1.00    479.4±6.52µs        ? ?/sec
filter context string dictionary high selectivity (kept 1023/1024)           1.00      6.0±0.07µs        ? ?/sec    1.06      6.4±0.25µs        ? ?/sec
filter context string dictionary w NULLs (kept 1/2)                          1.77    117.9±0.32µs        ? ?/sec    1.00     66.7±0.10µs        ? ?/sec
filter context string high selectivity (kept 1023/1024)                      1.05   671.6±17.79µs        ? ?/sec    1.00    638.8±9.68µs        ? ?/sec
filter context u8 w NULLs (kept 1/2)                                         4.41     65.9±0.60µs        ? ?/sec    1.00     15.0±0.18µs        ? ?/sec
filter decimal128 (kept 1/2)                                                 1.00     38.8±0.15µs        ? ?/sec    1.14     44.2±5.61µs        ? ?/sec
filter decimal128 high selectivity (kept 1023/1024)                          1.07     60.4±0.55µs        ? ?/sec    1.00     56.2±0.56µs        ? ?/sec
filter f32 (kept 1/2)                                                        3.61    123.9±0.36µs        ? ?/sec    1.00     34.4±0.35µs        ? ?/sec
filter fsb with value length 20 low selectivity (kept 1/1024)                1.00      2.2±0.03µs        ? ?/sec    1.30      2.8±0.03µs        ? ?/sec
filter fsb with value length 5 high selectivity (kept 1023/1024)             1.14     11.6±0.13µs        ? ?/sec    1.00     10.2±0.54µs        ? ?/sec
filter fsb with value length 5 low selectivity (kept 1/1024)                 1.00      2.1±0.03µs        ? ?/sec    1.29      2.7±0.03µs        ? ?/sec
filter fsb with value length 50 (kept 1/2)                                   1.29    185.9±5.06µs        ? ?/sec    1.00   144.1±10.29µs        ? ?/sec
filter fsb with value length 50 high selectivity (kept 1023/1024)            1.24    245.3±2.78µs        ? ?/sec    1.00    197.3±7.03µs        ? ?/sec
filter fsb with value length 50 low selectivity (kept 1/1024)                1.00      2.1±0.03µs        ? ?/sec    1.34      2.9±0.03µs        ? ?/sec
--- base-3 vs pr-3
group                                                                        base-3                                 pr-3
-----                                                                        ------                                 ----
filter context decimal128 (kept 1/2)                                         1.00     54.3±0.54µs        ? ?/sec    1.15     62.5±0.25µs        ? ?/sec
filter context f32 (kept 1/2)                                                1.76    117.8±0.16µs        ? ?/sec    1.00     67.0±0.63µs        ? ?/sec
filter context fsb with value length 20 (kept 1/2)                           1.08     86.7±1.17µs        ? ?/sec    1.00     80.5±5.62µs        ? ?/sec
filter context fsb with value length 20 high selectivity (kept 1023/1024)    1.09     87.1±0.62µs        ? ?/sec    1.00     79.7±0.38µs        ? ?/sec
filter context fsb with value length 5 high selectivity (kept 1023/1024)     1.00      8.6±0.18µs        ? ?/sec    1.13      9.7±0.63µs        ? ?/sec
filter context fsb with value length 50 (kept 1/2)                           1.10    193.2±4.92µs        ? ?/sec    1.00    175.2±9.39µs        ? ?/sec
filter context fsb with value length 50 low selectivity (kept 1/1024)        1.00    295.1±3.44ns        ? ?/sec    1.08    318.4±3.69ns        ? ?/sec
filter context i32 w NULLs (kept 1/2)                                        1.76    117.9±1.19µs        ? ?/sec    1.00     67.0±0.78µs        ? ?/sec
filter context i32 w NULLs high selectivity (kept 1023/1024)                 1.00      8.3±0.11µs        ? ?/sec    1.09      9.1±0.43µs        ? ?/sec
filter context mixed string view (kept 1/2)                                  1.76    117.6±1.38µs        ? ?/sec    1.00     67.0±3.74µs        ? ?/sec
filter context mixed string view high selectivity (kept 1023/1024)           1.06     57.8±1.52µs        ? ?/sec    1.00     54.6±1.85µs        ? ?/sec
filter context short string view (kept 1/2)                                  1.81    120.5±0.29µs        ? ?/sec    1.00     66.5±0.38µs        ? ?/sec
filter context short string view high selectivity (kept 1023/1024)           1.00     52.6±2.25µs        ? ?/sec    1.26     66.4±0.28µs        ? ?/sec
filter context string (kept 1/2)                                             1.08    541.5±6.25µs        ? ?/sec    1.00    503.3±4.90µs        ? ?/sec
filter context string dictionary w NULLs (kept 1/2)                          1.76    118.1±1.16µs        ? ?/sec    1.00     67.1±0.74µs        ? ?/sec
filter context u8 high selectivity (kept 1023/1024)                          1.00  1714.0±17.42ns        ? ?/sec    1.19      2.0±0.02µs        ? ?/sec
filter context u8 w NULLs (kept 1/2)                                         4.41     65.9±0.65µs        ? ?/sec    1.00     14.9±0.05µs        ? ?/sec
filter context u8 w NULLs high selectivity (kept 1023/1024)                  1.06      4.1±0.04µs        ? ?/sec    1.00      3.9±0.05µs        ? ?/sec
filter decimal128 (kept 1/2)                                                 1.00     39.5±0.32µs        ? ?/sec    1.34     52.8±0.14µs        ? ?/sec
filter decimal128 high selectivity (kept 1023/1024)                          1.00     54.1±3.46µs        ? ?/sec    1.16     62.7±0.57µs        ? ?/sec
filter f32 (kept 1/2)                                                        3.61    124.0±1.35µs        ? ?/sec    1.00     34.4±0.41µs        ? ?/sec
filter fsb with value length 20 high selectivity (kept 1023/1024)            1.19     82.7±4.80µs        ? ?/sec    1.00     69.6±1.47µs        ? ?/sec
filter fsb with value length 20 low selectivity (kept 1/1024)                1.00      2.2±0.02µs        ? ?/sec    1.27      2.8±0.01µs        ? ?/sec
filter fsb with value length 5 high selectivity (kept 1023/1024)             1.00     10.7±0.28µs        ? ?/sec    1.07     11.4±0.28µs        ? ?/sec
filter fsb with value length 5 low selectivity (kept 1/1024)                 1.00      2.2±0.03µs        ? ?/sec    1.26      2.7±0.00µs        ? ?/sec
filter fsb with value length 50 (kept 1/2)                                   1.06    183.6±1.38µs        ? ?/sec    1.00    173.3±7.81µs        ? ?/sec
filter fsb with value length 50 high selectivity (kept 1023/1024)            1.20    265.4±1.94µs        ? ?/sec    1.00    220.3±3.60µs        ? ?/sec
filter fsb with value length 50 low selectivity (kept 1/1024)                1.00      2.2±0.02µs        ? ?/sec    1.32      2.8±0.01µs        ? ?/sec
--- base-4 vs pr-4
group                                                                        base-4                                 pr-4
-----                                                                        ------                                 ----
filter context decimal128 (kept 1/2)                                         1.25     61.5±3.11µs        ? ?/sec    1.00     49.3±0.86µs        ? ?/sec
filter context decimal128 high selectivity (kept 1023/1024)                  1.00     56.1±0.57µs        ? ?/sec    1.07     59.8±0.22µs        ? ?/sec
filter context f32 (kept 1/2)                                                1.76    118.0±0.31µs        ? ?/sec    1.00     66.9±0.71µs        ? ?/sec
filter context f32 high selectivity (kept 1023/1024)                         1.05      8.8±0.05µs        ? ?/sec    1.00      8.3±0.07µs        ? ?/sec
filter context fsb with value length 20 high selectivity (kept 1023/1024)    1.20     86.6±8.46µs        ? ?/sec    1.00     72.3±0.57µs        ? ?/sec
filter context fsb with value length 50 high selectivity (kept 1023/1024)    1.12    245.7±0.36µs        ? ?/sec    1.00    220.2±3.03µs        ? ?/sec
filter context i32 w NULLs (kept 1/2)                                        1.77    117.9±1.12µs        ? ?/sec    1.00     66.8±0.08µs        ? ?/sec
filter context i32 w NULLs high selectivity (kept 1023/1024)                 1.07      8.6±0.16µs        ? ?/sec    1.00      8.1±0.02µs        ? ?/sec
filter context mixed string view (kept 1/2)                                  1.71    117.9±0.88µs        ? ?/sec    1.00     69.0±0.30µs        ? ?/sec
filter context short string view (kept 1/2)                                  1.84    113.8±3.78µs        ? ?/sec    1.00     61.8±5.17µs        ? ?/sec
filter context string (kept 1/2)                                             1.10    560.3±5.19µs        ? ?/sec    1.00    511.6±9.17µs        ? ?/sec
filter context string dictionary w NULLs (kept 1/2)                          1.76    117.9±0.25µs        ? ?/sec    1.00     67.0±0.70µs        ? ?/sec
filter context u8 high selectivity (kept 1023/1024)                          1.16      2.0±0.02µs        ? ?/sec    1.00  1737.1±19.60ns        ? ?/sec
filter context u8 w NULLs (kept 1/2)                                         4.40     65.8±0.10µs        ? ?/sec    1.00     14.9±0.04µs        ? ?/sec
filter decimal128 high selectivity (kept 1023/1024)                          1.26     65.2±0.34µs        ? ?/sec    1.00     51.6±0.92µs        ? ?/sec
filter f32 (kept 1/2)                                                        3.60    124.0±1.25µs        ? ?/sec    1.00     34.4±0.38µs        ? ?/sec
filter fsb with value length 20 high selectivity (kept 1023/1024)            1.05     81.5±0.65µs        ? ?/sec    1.00     77.4±0.45µs        ? ?/sec
filter fsb with value length 20 low selectivity (kept 1/1024)                1.00      2.2±0.03µs        ? ?/sec    1.26      2.8±0.01µs        ? ?/sec
filter fsb with value length 5 low selectivity (kept 1/1024)                 1.00      2.1±0.00µs        ? ?/sec    1.29      2.7±0.00µs        ? ?/sec
filter fsb with value length 50 (kept 1/2)                                   1.00   164.7±12.08µs        ? ?/sec    1.08    178.2±0.71µs        ? ?/sec
filter fsb with value length 50 high selectivity (kept 1023/1024)            1.09    264.7±1.74µs        ? ?/sec    1.00    243.4±9.09µs        ? ?/sec
filter fsb with value length 50 low selectivity (kept 1/1024)                1.00      2.2±0.00µs        ? ?/sec    1.29      2.8±0.03µs        ? ?/sec
filter i32 high selectivity (kept 1023/1024)                                 1.06      7.7±0.39µs        ? ?/sec    1.00      7.3±0.08µs        ? ?/sec
--- base-5 vs pr-5
group                                                                       base-5                                 pr-5
-----                                                                       ------                                 ----
filter context decimal128 high selectivity (kept 1023/1024)                 1.09     54.9±3.98µs        ? ?/sec    1.00     50.3±0.68µs        ? ?/sec
filter context f32 (kept 1/2)                                               1.76    117.9±1.14µs        ? ?/sec    1.00     66.9±0.65µs        ? ?/sec
filter context f32 high selectivity (kept 1023/1024)                        1.06      8.9±0.11µs        ? ?/sec    1.00      8.4±0.35µs        ? ?/sec
filter context fsb with value length 5 high selectivity (kept 1023/1024)    1.37     11.6±0.18µs        ? ?/sec    1.00      8.5±0.40µs        ? ?/sec
filter context fsb with value length 50 low selectivity (kept 1/1024)       1.00    287.1±3.03ns        ? ?/sec    1.15    329.7±0.34ns        ? ?/sec
filter context i32 w NULLs (kept 1/2)                                       1.76    117.9±1.21µs        ? ?/sec    1.00     66.8±0.29µs        ? ?/sec
filter context i32 w NULLs high selectivity (kept 1023/1024)                1.07      9.0±0.26µs        ? ?/sec    1.00      8.4±0.16µs        ? ?/sec
filter context mixed string view (kept 1/2)                                 2.13    122.0±0.89µs        ? ?/sec    1.00     57.3±0.84µs        ? ?/sec
filter context mixed string view high selectivity (kept 1023/1024)          1.05     55.5±1.05µs        ? ?/sec    1.00     52.8±0.61µs        ? ?/sec
filter context short string view (kept 1/2)                                 1.97    121.8±0.87µs        ? ?/sec    1.00     61.7±0.99µs        ? ?/sec
filter context short string view high selectivity (kept 1023/1024)          1.09     57.9±2.44µs        ? ?/sec    1.00     53.2±0.16µs        ? ?/sec
filter context string (kept 1/2)                                            1.10    526.6±8.26µs        ? ?/sec    1.00    479.2±3.04µs        ? ?/sec
filter context string dictionary w NULLs (kept 1/2)                         1.76    118.2±0.92µs        ? ?/sec    1.00     67.0±0.71µs        ? ?/sec
filter context u8 w NULLs (kept 1/2)                                        4.42     65.9±0.60µs        ? ?/sec    1.00     14.9±0.10µs        ? ?/sec
filter context u8 w NULLs high selectivity (kept 1023/1024)                 1.07      4.3±0.04µs        ? ?/sec    1.00      4.0±0.08µs        ? ?/sec
filter decimal128 high selectivity (kept 1023/1024)                         1.16     61.9±0.19µs        ? ?/sec    1.00     53.4±0.28µs        ? ?/sec
filter f32 (kept 1/2)                                                       3.61    124.0±0.79µs        ? ?/sec    1.00     34.3±0.12µs        ? ?/sec
filter fsb with value length 20 low selectivity (kept 1/1024)               1.00      2.2±0.03µs        ? ?/sec    1.25      2.8±0.01µs        ? ?/sec
filter fsb with value length 5 high selectivity (kept 1023/1024)            1.27     13.2±0.24µs        ? ?/sec    1.00     10.4±0.41µs        ? ?/sec
filter fsb with value length 5 low selectivity (kept 1/1024)                1.00      2.2±0.03µs        ? ?/sec    1.25      2.8±0.00µs        ? ?/sec
filter fsb with value length 50 high selectivity (kept 1023/1024)           1.00    194.7±9.60µs        ? ?/sec    1.13   220.7±15.98µs        ? ?/sec
filter fsb with value length 50 low selectivity (kept 1/1024)               1.00      2.2±0.03µs        ? ?/sec    1.30      2.8±0.03µs        ? ?/sec

@alamb

alamb commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

run benchmark filter_kernels

@alamb

alamb commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

run benchmark filter_kernels

env:
BENCH_FILTER: NULL

@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c5830825981-2722-6lvx7 6.12.94+ #1 SMP Wed Aug 19 07:47:20 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing db/10098/bmi-null-bool-filter (c7118c6) to c77f08d (merge-base) diff

Run configuration
run benchmark filter_kernels

BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench filter_kernels
Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c5830848503-2723-r4rf7 6.12.94+ #1 SMP Wed Aug 19 07:47:20 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing db/10098/bmi-null-bool-filter (c7118c6) to c77f08d (merge-base) diff

Run configuration
run benchmark filter_kernels
env:
  BENCH_FILTER: "NULL"

BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench filter_kernels
Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing db/10098/bmi-null-bool-filter (c7118c6) to c77f08d (merge-base) diff

Run configuration
run benchmark filter_kernels
env:
  BENCH_FILTER: "NULL"
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

group                                                                                db_10098_bmi-null-bool-filter          main
-----                                                                                -----------------------------          ----
filter context fsb with value length 20 w NULLs (kept 1/2)                           1.00     76.6±0.54µs        ? ?/sec  
filter context fsb with value length 20 w NULLs high selectivity (kept 1023/1024)    1.00     26.5±0.48µs        ? ?/sec  
filter context fsb with value length 20 w NULLs low selectivity (kept 1/1024)        1.00    476.3±2.64ns        ? ?/sec  
filter context fsb with value length 5 w NULLs (kept 1/2)                            1.00     74.0±0.12µs        ? ?/sec  
filter context fsb with value length 5 w NULLs high selectivity (kept 1023/1024)     1.00      6.6±0.02µs        ? ?/sec  
filter context fsb with value length 5 w NULLs low selectivity (kept 1/1024)         1.00    433.5±2.78ns        ? ?/sec  
filter context fsb with value length 50 w NULLs (kept 1/2)                           1.00    115.9±3.47µs        ? ?/sec  
filter context fsb with value length 50 w NULLs high selectivity (kept 1023/1024)    1.00     65.0±2.70µs        ? ?/sec  
filter context fsb with value length 50 w NULLs low selectivity (kept 1/1024)        1.00    508.1±2.26ns        ? ?/sec  
filter context i32 w NULLs (kept 1/2)                                                1.00     42.7±0.11µs        ? ?/sec    1.91     81.5±0.21µs        ? ?/sec
filter context i32 w NULLs high selectivity (kept 1023/1024)                         1.00      5.6±0.01µs        ? ?/sec    1.00      5.6±0.02µs        ? ?/sec
filter context i32 w NULLs low selectivity (kept 1/1024)                             1.00    326.7±1.93ns        ? ?/sec    1.00    325.5±1.90ns        ? ?/sec
filter context string dictionary w NULLs (kept 1/2)                                  1.00     42.5±0.09µs        ? ?/sec    1.92     81.8±0.19µs        ? ?/sec
filter context string dictionary w NULLs high selectivity (kept 1023/1024)           1.01      5.7±0.01µs        ? ?/sec    1.00      5.7±0.01µs        ? ?/sec
filter context string dictionary w NULLs low selectivity (kept 1/1024)               1.03    402.6±3.22ns        ? ?/sec    1.00    390.0±2.31ns        ? ?/sec
filter context u8 w NULLs (kept 1/2)                                                 1.00     42.1±0.10µs        ? ?/sec    1.93     81.3±0.14µs        ? ?/sec
filter context u8 w NULLs high selectivity (kept 1023/1024)                          1.04      2.9±0.01µs        ? ?/sec    1.00      2.8±0.01µs        ? ?/sec
filter context u8 w NULLs low selectivity (kept 1/1024)                              1.01    309.8±2.19ns        ? ?/sec    1.00    307.1±1.07ns        ? ?/sec
filter fsb with value length 20 w NULLs (kept 1/2)                                   1.00     88.6±0.54µs        ? ?/sec  
filter fsb with value length 20 w NULLs high selectivity (kept 1023/1024)            1.00     26.7±0.38µs        ? ?/sec  
filter fsb with value length 20 w NULLs low selectivity (kept 1/1024)                1.00   1751.1±2.20ns        ? ?/sec  
filter fsb with value length 5 w NULLs (kept 1/2)                                    1.00     85.4±0.14µs        ? ?/sec  
filter fsb with value length 5 w NULLs high selectivity (kept 1023/1024)             1.00      8.1±0.02µs        ? ?/sec  
filter fsb with value length 5 w NULLs low selectivity (kept 1/1024)                 1.00   1700.9±5.69ns        ? ?/sec  
filter fsb with value length 50 w NULLs (kept 1/2)                                   1.00    124.7±3.80µs        ? ?/sec  
filter fsb with value length 50 w NULLs high selectivity (kept 1023/1024)            1.00     75.8±4.25µs        ? ?/sec  
filter fsb with value length 50 w NULLs low selectivity (kept 1/1024)                1.00   1763.7±1.75ns        ? ?/sec  

Resource Usage

base (merge-base)

Metric Value
Wall time 100.0s
Peak memory 23.0 MiB
Avg memory 14.0 MiB
CPU user 94.8s
CPU sys 0.0s
Peak spill 0 B

branch

Metric Value
Wall time 265.1s
Peak memory 28.6 MiB
Avg memory 21.5 MiB
CPU user 261.3s
CPU sys 0.0s
Peak spill 0 B

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing db/10098/bmi-null-bool-filter (c7118c6) to c77f08d (merge-base) diff

Run configuration
run benchmark filter_kernels
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

group                                                                                db_10098_bmi-null-bool-filter          main
-----                                                                                -----------------------------          ----
filter context decimal128 (kept 1/2)                                                 1.03     21.1±0.10µs        ? ?/sec    1.00     20.5±0.15µs        ? ?/sec
filter context decimal128 high selectivity (kept 1023/1024)                          1.04     19.7±0.17µs        ? ?/sec    1.00     18.9±0.24µs        ? ?/sec
filter context decimal128 low selectivity (kept 1/1024)                              1.00    144.3±0.89ns        ? ?/sec    1.00    144.1±1.08ns        ? ?/sec
filter context f32 (kept 1/2)                                                        1.00     41.6±0.07µs        ? ?/sec    1.99     82.7±0.76µs        ? ?/sec
filter context f32 high selectivity (kept 1023/1024)                                 1.02      5.7±0.02µs        ? ?/sec    1.00      5.6±0.02µs        ? ?/sec
filter context f32 low selectivity (kept 1/1024)                                     1.00    326.9±2.23ns        ? ?/sec    1.00    327.4±3.29ns        ? ?/sec
filter context fsb with value length 20 (kept 1/2)                                   1.00     46.1±0.43µs        ? ?/sec    1.02     47.0±0.56µs        ? ?/sec
filter context fsb with value length 20 high selectivity (kept 1023/1024)            1.00     24.2±0.19µs        ? ?/sec    1.06     25.6±0.92µs        ? ?/sec
filter context fsb with value length 20 low selectivity (kept 1/1024)                1.00    261.6±0.84ns        ? ?/sec    1.07    280.7±1.12ns        ? ?/sec
filter context fsb with value length 20 w NULLs (kept 1/2)                           1.00     77.3±0.92µs        ? ?/sec  
filter context fsb with value length 20 w NULLs high selectivity (kept 1023/1024)    1.00     26.2±0.33µs        ? ?/sec  
filter context fsb with value length 20 w NULLs low selectivity (kept 1/1024)        1.00    457.0±4.27ns        ? ?/sec  
filter context fsb with value length 5 (kept 1/2)                                    1.00     45.6±0.62µs        ? ?/sec    1.01     45.9±0.22µs        ? ?/sec
filter context fsb with value length 5 high selectivity (kept 1023/1024)             1.00      4.7±0.01µs        ? ?/sec    1.01      4.7±0.00µs        ? ?/sec
filter context fsb with value length 5 low selectivity (kept 1/1024)                 1.00    235.4±0.73ns        ? ?/sec    1.03    241.3±1.00ns        ? ?/sec
filter context fsb with value length 5 w NULLs (kept 1/2)                            1.00     74.9±0.80µs        ? ?/sec  
filter context fsb with value length 5 w NULLs high selectivity (kept 1023/1024)     1.00      6.5±0.01µs        ? ?/sec  
filter context fsb with value length 5 w NULLs low selectivity (kept 1/1024)         1.00    418.5±2.56ns        ? ?/sec  
filter context fsb with value length 50 (kept 1/2)                                   1.00     88.7±2.05µs        ? ?/sec    1.04     92.4±1.09µs        ? ?/sec
filter context fsb with value length 50 high selectivity (kept 1023/1024)            1.00     68.8±0.73µs        ? ?/sec    1.03     71.1±2.08µs        ? ?/sec
filter context fsb with value length 50 low selectivity (kept 1/1024)                1.00    291.5±1.17ns        ? ?/sec    1.01    294.5±1.02ns        ? ?/sec
filter context fsb with value length 50 w NULLs (kept 1/2)                           1.00    122.0±3.88µs        ? ?/sec  
filter context fsb with value length 50 w NULLs high selectivity (kept 1023/1024)    1.00     67.1±3.31µs        ? ?/sec  
filter context fsb with value length 50 w NULLs low selectivity (kept 1/1024)        1.00    481.8±3.40ns        ? ?/sec  
filter context i32 (kept 1/2)                                                        1.00     12.3±0.03µs        ? ?/sec    1.00     12.3±0.02µs        ? ?/sec
filter context i32 high selectivity (kept 1023/1024)                                 1.00      3.7±0.00µs        ? ?/sec    1.01      3.7±0.01µs        ? ?/sec
filter context i32 low selectivity (kept 1/1024)                                     1.06    144.2±1.49ns        ? ?/sec    1.00    136.5±0.95ns        ? ?/sec
filter context i32 w NULLs (kept 1/2)                                                1.00     42.4±0.11µs        ? ?/sec    1.94     82.0±0.18µs        ? ?/sec
filter context i32 w NULLs high selectivity (kept 1023/1024)                         1.00      5.5±0.01µs        ? ?/sec    1.01      5.5±0.01µs        ? ?/sec
filter context i32 w NULLs low selectivity (kept 1/1024)                             1.00    331.7±2.34ns        ? ?/sec    1.00    330.6±3.41ns        ? ?/sec
filter context mixed string view (kept 1/2)                                          1.00     50.9±0.08µs        ? ?/sec    1.79     91.4±0.53µs        ? ?/sec
filter context mixed string view high selectivity (kept 1023/1024)                   1.00     20.4±0.12µs        ? ?/sec    1.06     21.6±0.15µs        ? ?/sec
filter context mixed string view low selectivity (kept 1/1024)                       1.00    322.3±2.23ns        ? ?/sec    1.00    322.5±2.42ns        ? ?/sec
filter context short string view (kept 1/2)                                          1.00     50.1±0.15µs        ? ?/sec    1.79     89.9±0.60µs        ? ?/sec
filter context short string view high selectivity (kept 1023/1024)                   1.04     22.1±0.12µs        ? ?/sec    1.00     21.3±0.19µs        ? ?/sec
filter context short string view low selectivity (kept 1/1024)                       1.00    322.9±2.09ns        ? ?/sec    1.01    324.8±2.85ns        ? ?/sec
filter context string (kept 1/2)                                                     1.00   387.3±10.24µs        ? ?/sec    1.10    424.2±9.84µs        ? ?/sec
filter context string dictionary (kept 1/2)                                          1.01     12.5±0.03µs        ? ?/sec    1.00     12.4±0.03µs        ? ?/sec
filter context string dictionary high selectivity (kept 1023/1024)                   1.01      3.7±0.01µs        ? ?/sec    1.00      3.7±0.00µs        ? ?/sec
filter context string dictionary low selectivity (kept 1/1024)                       1.00    198.2±0.85ns        ? ?/sec    1.00    199.1±2.10ns        ? ?/sec
filter context string dictionary w NULLs (kept 1/2)                                  1.00     42.6±0.77µs        ? ?/sec    1.91     81.5±0.15µs        ? ?/sec
filter context string dictionary w NULLs high selectivity (kept 1023/1024)           1.00      5.6±0.02µs        ? ?/sec    1.01      5.6±0.01µs        ? ?/sec
filter context string dictionary w NULLs low selectivity (kept 1/1024)               1.01    400.2±8.31ns        ? ?/sec    1.00    394.6±2.22ns        ? ?/sec
filter context string high selectivity (kept 1023/1024)                              1.00   323.0±26.58µs        ? ?/sec    1.02   328.7±23.59µs        ? ?/sec
filter context string low selectivity (kept 1/1024)                                  1.00    782.3±2.30ns        ? ?/sec    1.08    847.4±2.80ns        ? ?/sec
filter context u8 (kept 1/2)                                                         1.00     12.1±0.04µs        ? ?/sec    1.15     14.0±0.22µs        ? ?/sec
filter context u8 high selectivity (kept 1023/1024)                                  1.07   1122.1±3.69ns        ? ?/sec    1.00   1048.4±4.28ns        ? ?/sec
filter context u8 low selectivity (kept 1/1024)                                      1.01    127.1±0.66ns        ? ?/sec    1.00    125.4±1.22ns        ? ?/sec
filter context u8 w NULLs (kept 1/2)                                                 1.00     41.7±0.07µs        ? ?/sec    1.95     81.4±0.06µs        ? ?/sec
filter context u8 w NULLs high selectivity (kept 1023/1024)                          1.00      2.8±0.01µs        ? ?/sec    1.02      2.9±0.01µs        ? ?/sec
filter context u8 w NULLs low selectivity (kept 1/1024)                              1.00    307.8±1.23ns        ? ?/sec    1.00    306.6±3.12ns        ? ?/sec
filter decimal128 (kept 1/2)                                                         1.00     31.5±0.07µs        ? ?/sec    1.05     32.9±0.09µs        ? ?/sec
filter decimal128 high selectivity (kept 1023/1024)                                  1.01     20.8±0.13µs        ? ?/sec    1.00     20.5±0.19µs        ? ?/sec
filter decimal128 low selectivity (kept 1/1024)                                      1.00   1206.5±2.64ns        ? ?/sec    1.00   1205.3±2.58ns        ? ?/sec
filter f32 (kept 1/2)                                                                1.00     58.3±0.09µs        ? ?/sec    1.85    107.7±0.42µs        ? ?/sec
filter fsb with value length 20 (kept 1/2)                                           1.00     57.9±0.07µs        ? ?/sec    1.10     63.5±0.26µs        ? ?/sec
filter fsb with value length 20 high selectivity (kept 1023/1024)                    1.01     24.3±0.26µs        ? ?/sec    1.00     24.0±0.15µs        ? ?/sec
filter fsb with value length 20 low selectivity (kept 1/1024)                        1.00   1246.5±3.11ns        ? ?/sec    1.00   1242.4±2.08ns        ? ?/sec
filter fsb with value length 20 w NULLs (kept 1/2)                                   1.00     87.4±0.10µs        ? ?/sec  
filter fsb with value length 20 w NULLs high selectivity (kept 1023/1024)            1.00     26.8±0.20µs        ? ?/sec  
filter fsb with value length 20 w NULLs low selectivity (kept 1/1024)                1.00   1738.0±3.81ns        ? ?/sec  
filter fsb with value length 5 (kept 1/2)                                            1.00     55.9±0.14µs        ? ?/sec    1.09     61.0±0.05µs        ? ?/sec
filter fsb with value length 5 high selectivity (kept 1023/1024)                     1.00      5.5±0.01µs        ? ?/sec    1.01      5.5±0.01µs        ? ?/sec
filter fsb with value length 5 low selectivity (kept 1/1024)                         1.00   1186.3±6.00ns        ? ?/sec    1.00   1191.6±1.72ns        ? ?/sec
filter fsb with value length 5 w NULLs (kept 1/2)                                    1.00     85.8±0.06µs        ? ?/sec  
filter fsb with value length 5 w NULLs high selectivity (kept 1023/1024)             1.00      8.1±0.02µs        ? ?/sec  
filter fsb with value length 5 w NULLs low selectivity (kept 1/1024)                 1.00   1689.4±4.70ns        ? ?/sec  
filter fsb with value length 50 (kept 1/2)                                           1.01     96.8±2.72µs        ? ?/sec    1.00     96.1±0.29µs        ? ?/sec
filter fsb with value length 50 high selectivity (kept 1023/1024)                    1.00     71.0±2.03µs        ? ?/sec    1.05     74.2±4.41µs        ? ?/sec
filter fsb with value length 50 low selectivity (kept 1/1024)                        1.00   1237.1±1.85ns        ? ?/sec    1.01   1249.9±1.22ns        ? ?/sec
filter fsb with value length 50 w NULLs (kept 1/2)                                   1.00    126.1±0.82µs        ? ?/sec  
filter fsb with value length 50 w NULLs high selectivity (kept 1023/1024)            1.00     71.3±4.13µs        ? ?/sec  
filter fsb with value length 50 w NULLs low selectivity (kept 1/1024)                1.00   1747.5±4.18ns        ? ?/sec  
filter i32 (kept 1/2)                                                                1.09     29.2±0.02µs        ? ?/sec    1.00     26.9±0.05µs        ? ?/sec
filter i32 high selectivity (kept 1023/1024)                                         1.02      4.4±0.01µs        ? ?/sec    1.00      4.3±0.01µs        ? ?/sec
filter i32 low selectivity (kept 1/1024)                                             1.01   1155.9±1.88ns        ? ?/sec    1.00   1141.7±2.99ns        ? ?/sec
filter optimize (kept 1/2)                                                           1.01     27.0±0.15µs        ? ?/sec    1.00     26.7±0.12µs        ? ?/sec
filter optimize high selectivity (kept 1023/1024)                                    1.00   1497.2±1.24ns        ? ?/sec    1.00   1500.9±1.59ns        ? ?/sec
filter optimize low selectivity (kept 1/1024)                                        1.00    807.1±1.05ns        ? ?/sec    1.00    803.5±0.52ns        ? ?/sec
filter run array (kept 1/2)                                                          1.00    282.6±1.89µs        ? ?/sec    1.01    285.6±0.99µs        ? ?/sec
filter run array high selectivity (kept 1023/1024)                                   1.00    285.4±3.82µs        ? ?/sec    1.02    291.4±1.59µs        ? ?/sec
filter run array low selectivity (kept 1/1024)                                       1.00    235.5±0.82µs        ? ?/sec    1.00    235.9±0.96µs        ? ?/sec
filter single record batch                                                           1.00     28.2±0.06µs        ? ?/sec    1.01     28.6±0.08µs        ? ?/sec
filter u8 (kept 1/2)                                                                 1.00     29.3±0.03µs        ? ?/sec    1.00     29.4±0.10µs        ? ?/sec
filter u8 high selectivity (kept 1023/1024)                                          1.00   1824.5±5.39ns        ? ?/sec    1.00   1820.5±7.01ns        ? ?/sec
filter u8 low selectivity (kept 1/1024)                                              1.00  1113.4±11.87ns        ? ?/sec    1.00  1113.4±14.56ns        ? ?/sec

Resource Usage

base (merge-base)

Metric Value
Wall time 675.1s
Peak memory 37.8 MiB
Avg memory 18.8 MiB
CPU user 669.4s
CPU sys 0.1s
Peak spill 0 B

branch

Metric Value
Wall time 845.2s
Peak memory 36.0 MiB
Avg memory 21.4 MiB
CPU user 839.0s
CPU sys 0.1s
Peak spill 0 B

File an issue against this benchmark runner

@alamb

alamb commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

run benchmark filter_kernels

@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c5831250775-2724-crxcc 6.12.94+ #1 SMP Wed Aug 19 07:47:20 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing db/10098/bmi-null-bool-filter (c7118c6) to c77f08d (merge-base) diff

Run configuration
run benchmark filter_kernels

BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench filter_kernels
Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing db/10098/bmi-null-bool-filter (c7118c6) to c77f08d (merge-base) diff

Run configuration
run benchmark filter_kernels
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

group                                                                                db_10098_bmi-null-bool-filter          main
-----                                                                                -----------------------------          ----
filter context decimal128 (kept 1/2)                                                 1.00     20.0±0.14µs        ? ?/sec    1.00     20.1±0.09µs        ? ?/sec
filter context decimal128 high selectivity (kept 1023/1024)                          1.02     19.4±0.13µs        ? ?/sec    1.00     19.1±0.39µs        ? ?/sec
filter context decimal128 low selectivity (kept 1/1024)                              1.03    148.0±2.19ns        ? ?/sec    1.00    143.8±1.10ns        ? ?/sec
filter context f32 (kept 1/2)                                                        1.00     43.6±0.73µs        ? ?/sec    1.85     80.7±0.06µs        ? ?/sec
filter context f32 high selectivity (kept 1023/1024)                                 1.00      5.5±0.02µs        ? ?/sec    1.05      5.7±0.02µs        ? ?/sec
filter context f32 low selectivity (kept 1/1024)                                     1.02    333.4±1.53ns        ? ?/sec    1.00    328.4±2.09ns        ? ?/sec
filter context fsb with value length 20 (kept 1/2)                                   1.00     45.8±0.16µs        ? ?/sec    1.04     47.5±0.35µs        ? ?/sec
filter context fsb with value length 20 high selectivity (kept 1023/1024)            1.00     23.7±0.29µs        ? ?/sec    1.01     24.0±0.30µs        ? ?/sec
filter context fsb with value length 20 low selectivity (kept 1/1024)                1.00    262.0±0.76ns        ? ?/sec    1.06    277.2±2.17ns        ? ?/sec
filter context fsb with value length 20 w NULLs (kept 1/2)                           1.00     76.5±0.44µs        ? ?/sec  
filter context fsb with value length 20 w NULLs high selectivity (kept 1023/1024)    1.00     26.7±0.57µs        ? ?/sec  
filter context fsb with value length 20 w NULLs low selectivity (kept 1/1024)        1.00    463.2±0.89ns        ? ?/sec  
filter context fsb with value length 5 (kept 1/2)                                    1.02     45.1±0.06µs        ? ?/sec    1.00     44.3±0.04µs        ? ?/sec
filter context fsb with value length 5 high selectivity (kept 1023/1024)             1.00      4.7±0.00µs        ? ?/sec    1.00      4.7±0.00µs        ? ?/sec
filter context fsb with value length 5 low selectivity (kept 1/1024)                 1.00    235.6±1.00ns        ? ?/sec    1.02    241.0±1.94ns        ? ?/sec
filter context fsb with value length 5 w NULLs (kept 1/2)                            1.00     74.7±0.21µs        ? ?/sec  
filter context fsb with value length 5 w NULLs high selectivity (kept 1023/1024)     1.00      6.5±0.01µs        ? ?/sec  
filter context fsb with value length 5 w NULLs low selectivity (kept 1/1024)         1.00    418.7±1.55ns        ? ?/sec  
filter context fsb with value length 50 (kept 1/2)                                   1.00     85.2±0.76µs        ? ?/sec    1.00     85.0±1.95µs        ? ?/sec
filter context fsb with value length 50 high selectivity (kept 1023/1024)            1.05     74.1±2.49µs        ? ?/sec    1.00     70.6±0.92µs        ? ?/sec
filter context fsb with value length 50 low selectivity (kept 1/1024)                1.00    294.5±0.90ns        ? ?/sec    1.04    307.2±1.84ns        ? ?/sec
filter context fsb with value length 50 w NULLs (kept 1/2)                           1.00    118.8±1.46µs        ? ?/sec  
filter context fsb with value length 50 w NULLs high selectivity (kept 1023/1024)    1.00     72.8±5.70µs        ? ?/sec  
filter context fsb with value length 50 w NULLs low selectivity (kept 1/1024)        1.00    488.3±0.85ns        ? ?/sec  
filter context i32 (kept 1/2)                                                        1.00     12.4±0.01µs        ? ?/sec    1.00     12.3±0.04µs        ? ?/sec
filter context i32 high selectivity (kept 1023/1024)                                 1.00      3.7±0.00µs        ? ?/sec    1.01      3.7±0.01µs        ? ?/sec
filter context i32 low selectivity (kept 1/1024)                                     1.08    146.9±2.29ns        ? ?/sec    1.00    136.1±1.39ns        ? ?/sec
filter context i32 w NULLs (kept 1/2)                                                1.00     42.2±0.14µs        ? ?/sec    1.94     81.7±0.14µs        ? ?/sec
filter context i32 w NULLs high selectivity (kept 1023/1024)                         1.00      5.5±0.01µs        ? ?/sec    1.01      5.5±0.01µs        ? ?/sec
filter context i32 w NULLs low selectivity (kept 1/1024)                             1.03    339.1±1.91ns        ? ?/sec    1.00    330.1±1.90ns        ? ?/sec
filter context mixed string view (kept 1/2)                                          1.00     51.5±0.62µs        ? ?/sec    1.74     89.7±0.19µs        ? ?/sec
filter context mixed string view high selectivity (kept 1023/1024)                   1.04     21.9±0.54µs        ? ?/sec    1.00     21.0±0.18µs        ? ?/sec
filter context mixed string view low selectivity (kept 1/1024)                       1.03    331.0±3.10ns        ? ?/sec    1.00    321.8±2.21ns        ? ?/sec
filter context short string view (kept 1/2)                                          1.00     50.4±0.64µs        ? ?/sec    1.79     90.2±0.17µs        ? ?/sec
filter context short string view high selectivity (kept 1023/1024)                   1.01     21.8±0.66µs        ? ?/sec    1.00     21.6±0.11µs        ? ?/sec
filter context short string view low selectivity (kept 1/1024)                       1.03    331.8±3.16ns        ? ?/sec    1.00    322.2±2.40ns        ? ?/sec
filter context string (kept 1/2)                                                     1.00    381.0±1.97µs        ? ?/sec    1.10    420.8±3.58µs        ? ?/sec
filter context string dictionary (kept 1/2)                                          1.13     14.1±0.04µs        ? ?/sec    1.00     12.5±0.02µs        ? ?/sec
filter context string dictionary high selectivity (kept 1023/1024)                   1.00      3.7±0.01µs        ? ?/sec    1.00      3.8±0.01µs        ? ?/sec
filter context string dictionary low selectivity (kept 1/1024)                       1.02    202.2±3.09ns        ? ?/sec    1.00    199.2±5.43ns        ? ?/sec
filter context string dictionary w NULLs (kept 1/2)                                  1.00     42.8±0.10µs        ? ?/sec    1.91     81.5±0.06µs        ? ?/sec
filter context string dictionary w NULLs high selectivity (kept 1023/1024)           1.00      5.6±0.01µs        ? ?/sec    1.00      5.6±0.01µs        ? ?/sec
filter context string dictionary w NULLs low selectivity (kept 1/1024)               1.04    404.1±3.38ns        ? ?/sec    1.00    389.8±3.53ns        ? ?/sec
filter context string high selectivity (kept 1023/1024)                              1.00    314.4±6.43µs        ? ?/sec    1.00    315.9±6.26µs        ? ?/sec
filter context string low selectivity (kept 1/1024)                                  1.00    788.6±2.60ns        ? ?/sec    1.08    849.1±1.51ns        ? ?/sec
filter context u8 (kept 1/2)                                                         1.00     12.1±0.02µs        ? ?/sec    1.15     14.0±0.21µs        ? ?/sec
filter context u8 high selectivity (kept 1023/1024)                                  1.07   1117.6±3.96ns        ? ?/sec    1.00   1048.2±2.02ns        ? ?/sec
filter context u8 low selectivity (kept 1/1024)                                      1.05    130.4±2.01ns        ? ?/sec    1.00    123.7±1.07ns        ? ?/sec
filter context u8 w NULLs (kept 1/2)                                                 1.00     42.0±0.10µs        ? ?/sec    1.94     81.4±0.10µs        ? ?/sec
filter context u8 w NULLs high selectivity (kept 1023/1024)                          1.00      2.8±0.01µs        ? ?/sec    1.02      2.9±0.01µs        ? ?/sec
filter context u8 w NULLs low selectivity (kept 1/1024)                              1.03    316.3±1.70ns        ? ?/sec    1.00    307.2±1.66ns        ? ?/sec
filter decimal128 (kept 1/2)                                                         1.00     30.7±0.05µs        ? ?/sec    1.05     32.2±0.08µs        ? ?/sec
filter decimal128 high selectivity (kept 1023/1024)                                  1.00     20.4±0.16µs        ? ?/sec    1.00     20.5±0.27µs        ? ?/sec
filter decimal128 low selectivity (kept 1/1024)                                      1.00   1209.8±3.26ns        ? ?/sec    1.00   1206.4±2.72ns        ? ?/sec
filter f32 (kept 1/2)                                                                1.00     58.4±0.20µs        ? ?/sec    1.85    108.1±0.46µs        ? ?/sec
filter fsb with value length 20 (kept 1/2)                                           1.00     57.7±0.07µs        ? ?/sec    1.09     62.9±0.83µs        ? ?/sec
filter fsb with value length 20 high selectivity (kept 1023/1024)                    1.00     23.1±0.13µs        ? ?/sec    1.01     23.4±0.20µs        ? ?/sec
filter fsb with value length 20 low selectivity (kept 1/1024)                        1.00   1239.4±1.47ns        ? ?/sec    1.00   1239.6±1.21ns        ? ?/sec
filter fsb with value length 20 w NULLs (kept 1/2)                                   1.00     89.2±0.20µs        ? ?/sec  
filter fsb with value length 20 w NULLs high selectivity (kept 1023/1024)            1.00     25.6±0.15µs        ? ?/sec  
filter fsb with value length 20 w NULLs low selectivity (kept 1/1024)                1.00   1745.7±3.91ns        ? ?/sec  
filter fsb with value length 5 (kept 1/2)                                            1.00     55.7±0.03µs        ? ?/sec    1.09     61.0±0.04µs        ? ?/sec
filter fsb with value length 5 high selectivity (kept 1023/1024)                     1.00      5.5±0.00µs        ? ?/sec    1.00      5.5±0.01µs        ? ?/sec
filter fsb with value length 5 low selectivity (kept 1/1024)                         1.00   1184.6±6.17ns        ? ?/sec    1.01   1193.2±2.57ns        ? ?/sec
filter fsb with value length 5 w NULLs (kept 1/2)                                    1.00     85.9±0.19µs        ? ?/sec  
filter fsb with value length 5 w NULLs high selectivity (kept 1023/1024)             1.00      8.0±0.01µs        ? ?/sec  
filter fsb with value length 5 w NULLs low selectivity (kept 1/1024)                 1.00   1689.7±3.96ns        ? ?/sec  
filter fsb with value length 50 (kept 1/2)                                           1.00     93.3±0.51µs        ? ?/sec    1.00     93.2±0.70µs        ? ?/sec
filter fsb with value length 50 high selectivity (kept 1023/1024)                    1.02     71.8±0.84µs        ? ?/sec    1.00     70.6±0.64µs        ? ?/sec
filter fsb with value length 50 low selectivity (kept 1/1024)                        1.00   1235.2±1.18ns        ? ?/sec    1.02   1259.2±3.64ns        ? ?/sec
filter fsb with value length 50 w NULLs (kept 1/2)                                   1.00    125.1±1.23µs        ? ?/sec  
filter fsb with value length 50 w NULLs high selectivity (kept 1023/1024)            1.00     70.3±1.88µs        ? ?/sec  
filter fsb with value length 50 w NULLs low selectivity (kept 1/1024)                1.00   1742.7±1.84ns        ? ?/sec  
filter i32 (kept 1/2)                                                                1.10     29.2±0.04µs        ? ?/sec    1.00     26.5±0.05µs        ? ?/sec
filter i32 high selectivity (kept 1023/1024)                                         1.01      4.4±0.01µs        ? ?/sec    1.00      4.3±0.01µs        ? ?/sec
filter i32 low selectivity (kept 1/1024)                                             1.01   1160.7±1.89ns        ? ?/sec    1.00   1147.9±4.61ns        ? ?/sec
filter optimize (kept 1/2)                                                           1.02     27.1±0.10µs        ? ?/sec    1.00     26.5±0.15µs        ? ?/sec
filter optimize high selectivity (kept 1023/1024)                                    1.02   1517.1±2.57ns        ? ?/sec    1.00   1480.3±1.16ns        ? ?/sec
filter optimize low selectivity (kept 1/1024)                                        1.00    807.0±0.84ns        ? ?/sec    1.00    807.0±0.57ns        ? ?/sec
filter run array (kept 1/2)                                                          1.00    286.3±2.69µs        ? ?/sec    1.00    287.5±0.90µs        ? ?/sec
filter run array high selectivity (kept 1023/1024)                                   1.00    286.1±4.49µs        ? ?/sec    1.02    290.7±0.89µs        ? ?/sec
filter run array low selectivity (kept 1/1024)                                       1.00    235.8±0.89µs        ? ?/sec    1.00    235.8±0.92µs        ? ?/sec
filter single record batch                                                           1.01     28.8±0.05µs        ? ?/sec    1.00     28.5±0.04µs        ? ?/sec
filter u8 (kept 1/2)                                                                 1.00     29.5±0.01µs        ? ?/sec    1.00     29.4±0.03µs        ? ?/sec
filter u8 high selectivity (kept 1023/1024)                                          1.00   1809.0±5.25ns        ? ?/sec    1.00   1814.6±6.70ns        ? ?/sec
filter u8 low selectivity (kept 1/1024)                                              1.00  1116.4±12.42ns        ? ?/sec    1.01  1122.3±21.30ns        ? ?/sec

Resource Usage

base (merge-base)

Metric Value
Wall time 675.1s
Peak memory 40.5 MiB
Avg memory 19.4 MiB
CPU user 668.5s
CPU sys 0.1s
Peak spill 0 B

branch

Metric Value
Wall time 850.2s
Peak memory 36.0 MiB
Avg memory 21.4 MiB
CPU user 844.0s
CPU sys 0.1s
Peak spill 0 B

File an issue against this benchmark runner

@alamb

alamb commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

I think this one has seen enough now; let's merge it in and improve things as a follow on PR

@alamb
alamb merged commit 5d2a28c into apache:main Sep 25, 2026
44 checks passed
@alamb

alamb commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

Nice work @devanbenz @Rich-T-kid and @mbutrovich

alamb added a commit to alamb/arrow-rs that referenced this pull request Sep 25, 2026
Applies review suggestion from apache#10136 (comment)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
haohuaijin pushed a commit to haohuaijin/arrow-rs that referenced this pull request Sep 25, 2026
…nd `BMI` when supported) (apache#10136)

- close apache#11060

This commit adds the ability for bit filtering to be done using the
[`_pext_u64`](https://www.intel.com/content/www/us/en/docs/intrinsics-guide/index.html#text=_pext_u64)
BMI with a scalar fallback.

We cannot run the BMI2 feature on the benchmark host machine, here are
the benchmarks from my machine:

Specs:
```
Architecture:                x86_64
  CPU op-mode(s):            32-bit, 64-bit
  Address sizes:             46 bits physical, 48 bits virtual
  Byte Order:                Little Endian
Vendor ID:                   GenuineIntel
  Model name:                12th Gen Intel(R) Core(TM) i7-12700K
```

| Case | Build | main | apache#10136 | apache#10136 vs main |
|---|---|---:|---:|---:|
| filter context i32 w NULLs (kept 1/2) | scalar | 47.54 µs | 28.52 µs |
-40.0% |
| | bmi2 | 53.28 µs | 10.63 µs | -80.1% |
| filter context u8 w NULLs (kept 1/2) | scalar | 39.32 µs | 26.41 µs |
-32.9% |
| | bmi2 | 37.90 µs | 9.10 µs | -76.0% |
| filter context string dictionary w NULLs (kept 1/2) | scalar | 36.39
µs | 28.40 µs | -21.9% |
| | bmi2 | 44.88 µs | 10.69 µs | -76.2% |
| filter f32 (kept 1/2) | scalar | 61.87 µs | 39.97 µs | -35.4% |
| | bmi2 | 61.96 µs | 22.74 µs | -63.3% |
| filter context f32 (kept 1/2) | scalar | 44.43 µs | 28.40 µs | -36.1%
|
| | bmi2 | 36.93 µs | 10.80 µs | -70.8% |
| filter context short string view (kept 1/2) | scalar | 51.30 µs |
44.48 µs | -13.3% |
| | bmi2 | 50.56 µs | 29.11 µs | -42.4% |
| filter context mixed string view (kept 1/2) | scalar | 54.59 µs |
43.77 µs | -19.8% |
| | bmi2 | 56.13 µs | 26.35 µs | -53.1% |
| boolean 65 536 bits, kept 1/2 (filter context fsb …, 9 rows) | scalar
| 28.74–39.63 µs | 19.15–19.39 µs | -51…-33% |
| | bmi2 | 32.59–41.92 µs | 1.71–1.77 µs | -96…-95% |

With these changes we see the following improvements for filter kernels

| Build | Avg improvement | Including boolean row |
|---|---:|---:|
| scalar | ~28.5% faster | ~30.2% faster |
| bmi2 | ~66.0% faster | ~69.7% faster |
| overall | ~47.2% faster | ~49.9% faster |

---------

Co-authored-by: Andrew Lamb <andrew@nerdnetworks.org>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
alamb added a commit that referenced this pull request Sep 27, 2026
# Which issue does this PR close?

- Follow on to #10136

# Rationale for this change

Applies the review suggestion in
#10136 (comment)
from @mbutrovich to add a doc example for the new public
`BitChunks::chunk` API.

# What changes are included in this PR?

Adds a doctest to `BitChunks::chunk`.

# Are these changes tested?

Yes, by the new doctest.

# Are there any user-facing changes?

Docs only.

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

arrow Changes to the arrow crate arrow-buffer arrow-select parquet Changes to the parquet crate performance

Projects

None yet

Development

Successfully merging this pull request may close these issues.

replace bit-at-a-time null bitmap filtering with word-level

8 participants