Skip to content

perf(filter): replace bit-at-a-time null bitmap filtering - #11055

Open
Rich-T-kid wants to merge 2 commits into
apache:mainfrom
Rich-T-kid:rich-t-kid/filter-bits-gather-optimization
Open

Rich-T-kid wants to merge 2 commits into
apache:mainfrom
Rich-T-kid:rich-t-kid/filter-bits-gather-optimization

Conversation

@Rich-T-kid

@Rich-T-kid Rich-T-kid commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

Which issue does this PR close?

Rationale for this change

When filtering an array that has a null bitmap, Arrow walks the null bitmap
one bit at a time, one branch and one memory access per selected row. For a
65 536-row array at 50% filter density that is ~32 000 sequential bit reads.
On ARM (Neoverse-V2, Apple Silicon) where there is no hardware PEXT
instruction, the software fallback also serialises all 64 bit-positions in a
word into a loop-carried dependency chain, leaving gains on the table

What changes are included in this PR?

  • gather_bits replaces the bit-at-a-time loops in filter_bits_strategy.
    Instead of one get_bit_raw call per selected row, it zips 64-bit chunks of
    the filter and source bitmaps, calls compress() once per word, and streams
    the packed result into the output. The Indices branch retains the old
    index-lookup path only for filters with < 1/64 bits set, where scanning the
    full bitmap costs more than jumping to the few set positions.

  • Sparse path (compress_sparse, ≤ 8 bits set): keeps the same
    bit-by-bit loop but it now terminates in at most 8 iterations. For 1/1024
    selectivity, most 64-bit mask words have 0–1 bits set, so this costs 0–1
    iterations instead of 0–64.

  • Dense path (compress_dense, > 8 bits set): uses NIBBLE_PEXT, a
    256-byte compile-time lookup table covering all 4-bit mask/value pairs. The
    64-bit word is processed as 8 bytes × 2 nibbles = 16 independent table
    lookups. Per-byte output offsets are precomputed so all 8 writes into
    result are data-independent (hopfully CPU can
    issue them in parallel)

Are these changes tested?

yes, existing test

Are there any user-facing changes?

no

@github-actions github-actions Bot added arrow Changes to the arrow crate arrow-select labels Sep 10, 2026
@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

run benchmark filter_kernels
env:
BENCH_FILTER: NULL

@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c5624007162-2312-89jl4 6.12.94+ #1 SMP Tue Aug 4 08:44:15 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing rich-t-kid/filter-bits-gather-optimization (75281be) to 2078680 (merge-base) diff

Run configuration
run benchmark filter_kernels
env:
  BENCH_FILTER: "NULL"

BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench filter_kernels
Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing rich-t-kid/filter-bits-gather-optimization (75281be) to 2078680 (merge-base) diff

Run configuration
run benchmark filter_kernels
env:
  BENCH_FILTER: "NULL"
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

group                                                                         main                                   rich-t-kid_filter-bits-gather-optimization
-----                                                                         ----                                   ------------------------------------------
filter context i32 w NULLs (kept 1/2)                                         1.39     81.5±0.06µs        ? ?/sec    1.00     58.7±0.04µs        ? ?/sec
filter context i32 w NULLs high selectivity (kept 1023/1024)                  1.00      5.6±0.01µs        ? ?/sec    1.00      5.6±0.01µs        ? ?/sec
filter context i32 w NULLs low selectivity (kept 1/1024)                      1.01    321.5±0.97ns        ? ?/sec    1.00    316.8±1.26ns        ? ?/sec
filter context string dictionary w NULLs (kept 1/2)                           1.40     81.6±0.05µs        ? ?/sec    1.00     58.1±0.05µs        ? ?/sec
filter context string dictionary w NULLs high selectivity (kept 1023/1024)    1.02      5.8±0.01µs        ? ?/sec    1.00      5.6±0.01µs        ? ?/sec
filter context string dictionary w NULLs low selectivity (kept 1/1024)        1.00    376.7±2.31ns        ? ?/sec    1.02    382.5±2.52ns        ? ?/sec
filter context u8 w NULLs (kept 1/2)                                          1.40     81.5±0.19µs        ? ?/sec    1.00     58.1±0.05µs        ? ?/sec
filter context u8 w NULLs high selectivity (kept 1023/1024)                   1.00      2.8±0.01µs        ? ?/sec    1.00      2.8±0.01µs        ? ?/sec
filter context u8 w NULLs low selectivity (kept 1/1024)                       1.00    306.2±1.66ns        ? ?/sec    1.02    313.5±3.56ns        ? ?/sec

Resource Usage

base (merge-base)

Metric Value
Wall time 95.0s
Peak memory 23.1 MiB
Avg memory 14.0 MiB
CPU user 91.8s
CPU sys 0.0s
Peak spill 0 B

branch

Metric Value
Wall time 95.0s
Peak memory 22.5 MiB
Avg memory 13.7 MiB
CPU user 89.1s
CPU sys 0.0s
Peak spill 0 B

File an issue against this benchmark runner

@Rich-T-kid Rich-T-kid left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

trim

Comment thread arrow-select/src/filter.rs Outdated
Comment thread arrow-select/src/filter.rs Outdated
@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

run benchmark filter_kernels
env:
BENCH_FILTER: NULL

1 similar comment
@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

run benchmark filter_kernels
env:
BENCH_FILTER: NULL

@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c5639714015-2320-55t89 6.12.94+ #1 SMP Tue Aug 4 08:44:15 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing rich-t-kid/filter-bits-gather-optimization (200598b) to 2078680 (merge-base) diff

Run configuration
run benchmark filter_kernels
env:
  BENCH_FILTER: "NULL"

BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench filter_kernels
Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c5639714741-2321-fh8nd 6.12.94+ #1 SMP Tue Aug 4 08:44:15 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing rich-t-kid/filter-bits-gather-optimization (200598b) to 2078680 (merge-base) diff

Run configuration
run benchmark filter_kernels
env:
  BENCH_FILTER: "NULL"

BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench filter_kernels
Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing rich-t-kid/filter-bits-gather-optimization (200598b) to 2078680 (merge-base) diff

Run configuration
run benchmark filter_kernels
env:
  BENCH_FILTER: "NULL"
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

group                                                                         main                                   rich-t-kid_filter-bits-gather-optimization
-----                                                                         ----                                   ------------------------------------------
filter context i32 w NULLs (kept 1/2)                                         1.77     82.2±0.66µs        ? ?/sec    1.00     46.6±0.09µs        ? ?/sec
filter context i32 w NULLs high selectivity (kept 1023/1024)                  1.00      5.6±0.01µs        ? ?/sec    1.00      5.6±0.01µs        ? ?/sec
filter context i32 w NULLs low selectivity (kept 1/1024)                      1.00    321.9±1.51ns        ? ?/sec    1.00    320.9±2.30ns        ? ?/sec
filter context string dictionary w NULLs (kept 1/2)                           1.67     81.5±0.07µs        ? ?/sec    1.00     48.7±0.88µs        ? ?/sec
filter context string dictionary w NULLs high selectivity (kept 1023/1024)    1.00      5.6±0.01µs        ? ?/sec    1.02      5.7±0.01µs        ? ?/sec
filter context string dictionary w NULLs low selectivity (kept 1/1024)        1.00    376.2±2.84ns        ? ?/sec    1.02    382.8±1.67ns        ? ?/sec
filter context u8 w NULLs (kept 1/2)                                          1.71     81.2±0.03µs        ? ?/sec    1.00     47.6±1.32µs        ? ?/sec
filter context u8 w NULLs high selectivity (kept 1023/1024)                   1.01      2.8±0.01µs        ? ?/sec    1.00      2.8±0.01µs        ? ?/sec
filter context u8 w NULLs low selectivity (kept 1/1024)                       1.00    306.7±2.05ns        ? ?/sec    1.02    313.2±2.33ns        ? ?/sec

Resource Usage

base (merge-base)

Metric Value
Wall time 100.0s
Peak memory 23.1 MiB
Avg memory 14.1 MiB
CPU user 94.8s
CPU sys 0.0s
Peak spill 0 B

branch

Metric Value
Wall time 90.0s
Peak memory 22.5 MiB
Avg memory 13.5 MiB
CPU user 84.1s
CPU sys 0.0s
Peak spill 0 B

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing rich-t-kid/filter-bits-gather-optimization (200598b) to 2078680 (merge-base) diff

Run configuration
run benchmark filter_kernels
env:
  BENCH_FILTER: "NULL"
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

group                                                                         main                                   rich-t-kid_filter-bits-gather-optimization
-----                                                                         ----                                   ------------------------------------------
filter context i32 w NULLs (kept 1/2)                                         1.70     82.1±1.93µs        ? ?/sec    1.00     48.3±0.36µs        ? ?/sec
filter context i32 w NULLs high selectivity (kept 1023/1024)                  1.00      5.6±0.02µs        ? ?/sec    1.00      5.6±0.01µs        ? ?/sec
filter context i32 w NULLs low selectivity (kept 1/1024)                      1.02    323.8±3.80ns        ? ?/sec    1.00    318.8±0.92ns        ? ?/sec
filter context string dictionary w NULLs (kept 1/2)                           1.79     82.1±0.23µs        ? ?/sec    1.00     45.8±0.14µs        ? ?/sec
filter context string dictionary w NULLs high selectivity (kept 1023/1024)    1.00      5.6±0.01µs        ? ?/sec    1.00      5.6±0.01µs        ? ?/sec
filter context string dictionary w NULLs low selectivity (kept 1/1024)        1.00    376.2±2.63ns        ? ?/sec    1.03    387.4±3.06ns        ? ?/sec
filter context u8 w NULLs (kept 1/2)                                          1.75     81.3±0.37µs        ? ?/sec    1.00     46.4±0.23µs        ? ?/sec
filter context u8 w NULLs high selectivity (kept 1023/1024)                   1.00      2.8±0.02µs        ? ?/sec    1.00      2.8±0.01µs        ? ?/sec
filter context u8 w NULLs low selectivity (kept 1/1024)                       1.00    308.7±5.47ns        ? ?/sec    1.01    312.9±4.28ns        ? ?/sec

Resource Usage

base (merge-base)

Metric Value
Wall time 100.0s
Peak memory 23.1 MiB
Avg memory 14.0 MiB
CPU user 94.8s
CPU sys 0.0s
Peak spill 0 B

branch

Metric Value
Wall time 90.0s
Peak memory 22.5 MiB
Avg memory 14.2 MiB
CPU user 87.1s
CPU sys 0.0s
Peak spill 0 B

File an issue against this benchmark runner

@Rich-T-kid
Rich-T-kid force-pushed the rich-t-kid/filter-bits-gather-optimization branch from c275451 to 4727e6a Compare September 11, 2026 19:54
@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

run benchmark filter_kernels
env:
BENCH_FILTER: NULL

1 similar comment
@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

run benchmark filter_kernels
env:
BENCH_FILTER: NULL

@Rich-T-kid
Rich-T-kid marked this pull request as ready for review September 11, 2026 19:55
@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

@sdf-jkl could you take a look 🚀

@Rich-T-kid Rich-T-kid changed the title WIP perf(filter): replace bit-at-a-time null bitmap filtering perf(filter): replace bit-at-a-time null bitmap filtering Sep 11, 2026
@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c5639890125-2322-crmbd 6.12.94+ #1 SMP Tue Aug 4 08:44:15 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing rich-t-kid/filter-bits-gather-optimization (4727e6a) to 2078680 (merge-base) diff

Run configuration
run benchmark filter_kernels
env:
  BENCH_FILTER: "NULL"

BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench filter_kernels
Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c5639890747-2323-rrvcc 6.12.94+ #1 SMP Tue Aug 4 08:44:15 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing rich-t-kid/filter-bits-gather-optimization (4727e6a) to 2078680 (merge-base) diff

Run configuration
run benchmark filter_kernels
env:
  BENCH_FILTER: "NULL"

BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench filter_kernels
Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing rich-t-kid/filter-bits-gather-optimization (4727e6a) to 2078680 (merge-base) diff

Run configuration
run benchmark filter_kernels
env:
  BENCH_FILTER: "NULL"
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

group                                                                         main                                   rich-t-kid_filter-bits-gather-optimization
-----                                                                         ----                                   ------------------------------------------
filter context i32 w NULLs (kept 1/2)                                         1.73     81.4±0.07µs        ? ?/sec    1.00     47.0±0.62µs        ? ?/sec
filter context i32 w NULLs high selectivity (kept 1023/1024)                  1.00      5.6±0.01µs        ? ?/sec    1.00      5.6±0.02µs        ? ?/sec
filter context i32 w NULLs low selectivity (kept 1/1024)                      1.01    323.2±1.53ns        ? ?/sec    1.00    320.8±6.27ns        ? ?/sec
filter context string dictionary w NULLs (kept 1/2)                           1.69     81.6±0.12µs        ? ?/sec    1.00     48.2±0.15µs        ? ?/sec
filter context string dictionary w NULLs high selectivity (kept 1023/1024)    1.00      5.6±0.01µs        ? ?/sec    1.00      5.6±0.01µs        ? ?/sec
filter context string dictionary w NULLs low selectivity (kept 1/1024)        1.01    382.0±2.70ns        ? ?/sec    1.00    377.3±2.04ns        ? ?/sec
filter context u8 w NULLs (kept 1/2)                                          1.75     81.3±0.15µs        ? ?/sec    1.00     46.3±0.14µs        ? ?/sec
filter context u8 w NULLs high selectivity (kept 1023/1024)                   1.00      2.8±0.01µs        ? ?/sec    1.01      2.9±0.02µs        ? ?/sec
filter context u8 w NULLs low selectivity (kept 1/1024)                       1.00    308.4±2.22ns        ? ?/sec    1.00    308.1±2.60ns        ? ?/sec

Resource Usage

base (merge-base)

Metric Value
Wall time 100.0s
Peak memory 23.1 MiB
Avg memory 14.2 MiB
CPU user 95.6s
CPU sys 0.0s
Peak spill 0 B

branch

Metric Value
Wall time 90.0s
Peak memory 22.5 MiB
Avg memory 14.3 MiB
CPU user 87.1s
CPU sys 0.0s
Peak spill 0 B

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing rich-t-kid/filter-bits-gather-optimization (4727e6a) to 2078680 (merge-base) diff

Run configuration
run benchmark filter_kernels
env:
  BENCH_FILTER: "NULL"
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

group                                                                         main                                   rich-t-kid_filter-bits-gather-optimization
-----                                                                         ----                                   ------------------------------------------
filter context i32 w NULLs (kept 1/2)                                         1.76     81.6±0.62µs        ? ?/sec    1.00     46.5±0.15µs        ? ?/sec
filter context i32 w NULLs high selectivity (kept 1023/1024)                  1.00      5.6±0.02µs        ? ?/sec    1.00      5.5±0.01µs        ? ?/sec
filter context i32 w NULLs low selectivity (kept 1/1024)                      1.01    319.5±1.19ns        ? ?/sec    1.00    316.3±0.85ns        ? ?/sec
filter context string dictionary w NULLs (kept 1/2)                           1.76     81.5±0.05µs        ? ?/sec    1.00     46.4±0.17µs        ? ?/sec
filter context string dictionary w NULLs high selectivity (kept 1023/1024)    1.00      5.7±0.01µs        ? ?/sec    1.02      5.7±0.01µs        ? ?/sec
filter context string dictionary w NULLs low selectivity (kept 1/1024)        1.00    377.6±1.91ns        ? ?/sec    1.01    381.2±1.61ns        ? ?/sec
filter context u8 w NULLs (kept 1/2)                                          1.75     81.2±0.04µs        ? ?/sec    1.00     46.4±0.55µs        ? ?/sec
filter context u8 w NULLs high selectivity (kept 1023/1024)                   1.00      2.8±0.01µs        ? ?/sec    1.00      2.8±0.01µs        ? ?/sec
filter context u8 w NULLs low selectivity (kept 1/1024)                       1.00    306.4±1.73ns        ? ?/sec    1.00    308.0±2.56ns        ? ?/sec

Resource Usage

base (merge-base)

Metric Value
Wall time 100.0s
Peak memory 23.1 MiB
Avg memory 14.0 MiB
CPU user 94.7s
CPU sys 0.0s
Peak spill 0 B

branch

Metric Value
Wall time 90.0s
Peak memory 22.5 MiB
Avg memory 13.5 MiB
CPU user 84.1s
CPU sys 0.0s
Peak spill 0 B

File an issue against this benchmark runner

Comment thread arrow-select/src/filter.rs
@sdf-jkl

sdf-jkl commented Sep 11, 2026

Copy link
Copy Markdown
Member

I'll take a look this weekend

@sdf-jkl sdf-jkl left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@Rich-T-kid looking good

Comment thread arrow-select/src/filter.rs Outdated

/// Collects the bits of `val` wherever `mask` is 1, packed into the low bits of the result.
#[inline(always)]
fn pext64(val: u64, mut mask: u64) -> u64 {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There already is an implementation of PEXT in the codebase --

pub(crate) fn compress(value: u64, mask: u64) -> u64 {
#[cfg(all(target_arch = "x86_64", target_feature = "bmi2"))]
{
// SAFETY: the `bmi2` target feature is statically enabled for this
// build, so the `pext` instruction is guaranteed to be available.
unsafe { std::arch::x86_64::_pext_u64(value, mask) }
}
#[cfg(not(all(target_arch = "x86_64", target_feature = "bmi2")))]
{
let mut mask = mask;
let mut result: u64 = 0;
let mut dest_bit: u64 = 1;
while mask != 0 {
let lowest = mask & mask.wrapping_neg();
if value & lowest != 0 {
result |= dest_bit;
}
dest_bit <<= 1;
mask ^= lowest;
}
result
}
}
#[cfg(test)]

Yours seems to perform better though (on my machine 🤓 ) If anything we can drop drop the parquet one and make it import your implementation.

Rust has a nightly impl-- doc.rust-lang.org/std/primitive.u64.html#method.extract_bits

The issue here - rust-lang/rust#149069 - explains that the implementation is waiting on supporting the new LLVM intrinsics that support automatically using the PEXT/PDEP machine instructions if machine supports them. When the instructions are available the perf is pretty epic.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yea I think moving this out so it can be re-used is a good idea. Would be nice to see how much of a speed up parquet gets from this as well

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

arrow-buffer/src/util/bit_util.rs is where @devanbenz put it

}
};

for (filter_word, src_word) in filter_chunks.iter().zip(src_chunks.iter()) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Processing each word is independent from each other so this could be paralellizable.

I killed some time on it today. You can take a look here -- #11086

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

a great follow on task perhaps

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree, can open up a follow on PR

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Once we merge #10136, would you be willing to try and revive this code (to see if it is faster than what went in to #10136 ?)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yea sure, I'll keep a look out

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@sdf-jkl mhm I cant seem to get similar results to what you did on your branch. everything was pretty much within noise. maybe you can take a closer look

@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

@sdf-jkl your PR seems to be faster than this one, should I close this in favor of that one?

@sdf-jkl

sdf-jkl commented Sep 15, 2026

Copy link
Copy Markdown
Member

Mine was just pathfinding and only better when SIMD enabled, not on scalar. You can port smth / take inspiration from my PR here.

@alamb alamb left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is very neat -- thank you @Rich-T-kid @sdf-jkl and @devanbenz

It will be pretty amazing to get 75% faster on some filter kernels 🤯

I'll keep an eye on this one

FYI @jhorstmann and @hhhizzz you may be interested too

Comment thread arrow-select/src/filter.rs Outdated
}
};

for (filter_word, src_word) in filter_chunks.iter().zip(src_chunks.iter()) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

a great follow on task perhaps

Comment thread arrow-select/src/filter.rs Outdated

/// Collects the bits of `val` wherever `mask` is 1, packed into the low bits of the result.
#[inline(always)]
fn pext64(val: u64, mut mask: u64) -> u64 {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

arrow-buffer/src/util/bit_util.rs is where @devanbenz put it

@alamb alamb left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it might also make sense to ensure the benchmarks cover this case as well (clearly some do but maybe not all)

Comment thread arrow-select/src/filter.rs
@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c5873115614-2891-dksdt 6.12.94+ #1 SMP Wed Aug 19 07:47:20 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing rich-t-kid/filter-bits-gather-optimization (b64eae3) to 1c2c390 (merge-base) diff

Run configuration
run benchmark filter_kernels
env:
  BENCH_FILTER: "NULL"

BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench filter_kernels
Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing rich-t-kid/filter-bits-gather-optimization (b64eae3) to 1c2c390 (merge-base) diff

Run configuration
run benchmark filter_kernels
env:
  BENCH_FILTER: "NULL"
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

group                                                                                main                                   rich-t-kid_filter-bits-gather-optimization
-----                                                                                ----                                   ------------------------------------------
filter context fsb with value length 20 w NULLs (kept 1/2)                           1.22     77.8±1.18µs        ? ?/sec    1.00     63.7±0.67µs        ? ?/sec
filter context fsb with value length 20 w NULLs high selectivity (kept 1023/1024)    1.00     24.8±0.42µs        ? ?/sec    1.13     28.0±0.71µs        ? ?/sec
filter context fsb with value length 20 w NULLs low selectivity (kept 1/1024)        1.01    475.2±2.88ns        ? ?/sec    1.00    468.3±5.84ns        ? ?/sec
filter context fsb with value length 5 w NULLs (kept 1/2)                            1.22     74.7±0.13µs        ? ?/sec    1.00     61.4±0.16µs        ? ?/sec
filter context fsb with value length 5 w NULLs high selectivity (kept 1023/1024)     1.00      6.6±0.02µs        ? ?/sec    1.00      6.6±0.03µs        ? ?/sec
filter context fsb with value length 5 w NULLs low selectivity (kept 1/1024)         1.00    425.2±4.52ns        ? ?/sec    1.01    429.0±4.26ns        ? ?/sec
filter context fsb with value length 50 w NULLs (kept 1/2)                           1.14    129.5±2.59µs        ? ?/sec    1.00    113.9±3.89µs        ? ?/sec
filter context fsb with value length 50 w NULLs high selectivity (kept 1023/1024)    1.07     83.4±7.64µs        ? ?/sec    1.00     77.7±4.21µs        ? ?/sec
filter context fsb with value length 50 w NULLs low selectivity (kept 1/1024)        1.00    484.6±1.50ns        ? ?/sec    1.02    494.5±3.71ns        ? ?/sec
filter context i32 w NULLs (kept 1/2)                                                1.46     43.1±0.36µs        ? ?/sec    1.00     29.5±0.14µs        ? ?/sec
filter context i32 w NULLs high selectivity (kept 1023/1024)                         1.00      5.6±0.02µs        ? ?/sec    1.00      5.6±0.02µs        ? ?/sec
filter context i32 w NULLs low selectivity (kept 1/1024)                             1.05    339.3±4.04ns        ? ?/sec    1.00    324.4±1.56ns        ? ?/sec
filter context string dictionary w NULLs (kept 1/2)                                  1.45     42.6±0.08µs        ? ?/sec    1.00     29.4±0.09µs        ? ?/sec
filter context string dictionary w NULLs high selectivity (kept 1023/1024)           1.00      5.7±0.02µs        ? ?/sec    1.00      5.7±0.02µs        ? ?/sec
filter context string dictionary w NULLs low selectivity (kept 1/1024)               1.00    393.4±2.56ns        ? ?/sec    1.00    392.9±2.16ns        ? ?/sec
filter context u8 w NULLs (kept 1/2)                                                 1.41     43.9±0.12µs        ? ?/sec    1.00     31.2±0.09µs        ? ?/sec
filter context u8 w NULLs high selectivity (kept 1023/1024)                          1.00      2.8±0.01µs        ? ?/sec    1.00      2.8±0.01µs        ? ?/sec
filter context u8 w NULLs low selectivity (kept 1/1024)                              1.02    313.4±1.16ns        ? ?/sec    1.00    306.5±1.44ns        ? ?/sec
filter fsb with value length 20 w NULLs (kept 1/2)                                   1.21     90.8±0.63µs        ? ?/sec    1.00     75.3±0.38µs        ? ?/sec
filter fsb with value length 20 w NULLs high selectivity (kept 1023/1024)            1.01     26.4±0.35µs        ? ?/sec    1.00     26.0±0.22µs        ? ?/sec
filter fsb with value length 20 w NULLs low selectivity (kept 1/1024)                1.00   1758.3±7.13ns        ? ?/sec    1.00   1766.4±2.16ns        ? ?/sec
filter fsb with value length 5 w NULLs (kept 1/2)                                    1.19     86.5±0.18µs        ? ?/sec    1.00     72.7±0.11µs        ? ?/sec
filter fsb with value length 5 w NULLs high selectivity (kept 1023/1024)             1.00      8.0±0.01µs        ? ?/sec    1.01      8.1±0.08µs        ? ?/sec
filter fsb with value length 5 w NULLs low selectivity (kept 1/1024)                 1.00  1688.7±14.23ns        ? ?/sec    1.02   1723.6±1.89ns        ? ?/sec
filter fsb with value length 50 w NULLs (kept 1/2)                                   1.19    139.8±8.04µs        ? ?/sec    1.00    117.1±1.69µs        ? ?/sec
filter fsb with value length 50 w NULLs high selectivity (kept 1023/1024)            1.15     85.7±5.26µs        ? ?/sec    1.00     74.8±2.60µs        ? ?/sec
filter fsb with value length 50 w NULLs low selectivity (kept 1/1024)                1.00   1746.6±3.63ns        ? ?/sec    1.02   1784.8±1.92ns        ? ?/sec

Resource Usage

base (merge-base)

Metric Value
Wall time 275.1s
Peak memory 32.6 MiB
Avg memory 21.7 MiB
CPU user 270.0s
CPU sys 0.0s
Peak spill 0 B

branch

Metric Value
Wall time 260.1s
Peak memory 27.0 MiB
Avg memory 20.7 MiB
CPU user 256.3s
CPU sys 0.0s
Peak spill 0 B

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Arrow criterion benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing rich-t-kid/filter-bits-gather-optimization (b64eae3) to 1c2c390 (merge-base) diff

Run configuration
run benchmark filter_kernels
env:
  BENCH_FILTER: "NULL"
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

group                                                                                main                                   rich-t-kid_filter-bits-gather-optimization
-----                                                                                ----                                   ------------------------------------------
filter context fsb with value length 20 w NULLs (kept 1/2)                           1.24     78.3±0.72µs        ? ?/sec    1.00     62.9±0.23µs        ? ?/sec
filter context fsb with value length 20 w NULLs high selectivity (kept 1023/1024)    1.00     26.8±0.48µs        ? ?/sec    1.06     28.4±0.50µs        ? ?/sec
filter context fsb with value length 20 w NULLs low selectivity (kept 1/1024)        1.00    475.5±2.77ns        ? ?/sec    1.00    477.0±4.17ns        ? ?/sec
filter context fsb with value length 5 w NULLs (kept 1/2)                            1.23     75.2±1.05µs        ? ?/sec    1.00     61.1±0.11µs        ? ?/sec
filter context fsb with value length 5 w NULLs high selectivity (kept 1023/1024)     1.00      6.6±0.01µs        ? ?/sec    1.01      6.6±0.01µs        ? ?/sec
filter context fsb with value length 5 w NULLs low selectivity (kept 1/1024)         1.00    418.4±4.62ns        ? ?/sec    1.05    439.4±4.03ns        ? ?/sec
filter context fsb with value length 50 w NULLs (kept 1/2)                           1.12    122.9±4.43µs        ? ?/sec    1.00    109.4±2.38µs        ? ?/sec
filter context fsb with value length 50 w NULLs high selectivity (kept 1023/1024)    1.00     74.0±4.37µs        ? ?/sec    1.02     75.6±4.58µs        ? ?/sec
filter context fsb with value length 50 w NULLs low selectivity (kept 1/1024)        1.00    482.6±1.53ns        ? ?/sec    1.05    506.0±2.61ns        ? ?/sec
filter context i32 w NULLs (kept 1/2)                                                1.41     42.8±0.13µs        ? ?/sec    1.00     30.4±0.35µs        ? ?/sec
filter context i32 w NULLs high selectivity (kept 1023/1024)                         1.00      5.6±0.01µs        ? ?/sec    1.00      5.6±0.01µs        ? ?/sec
filter context i32 w NULLs low selectivity (kept 1/1024)                             1.02    332.0±1.38ns        ? ?/sec    1.00    325.7±2.27ns        ? ?/sec
filter context string dictionary w NULLs (kept 1/2)                                  1.47     43.1±0.07µs        ? ?/sec    1.00     29.2±0.06µs        ? ?/sec
filter context string dictionary w NULLs high selectivity (kept 1023/1024)           1.00      5.6±0.01µs        ? ?/sec    1.01      5.7±0.01µs        ? ?/sec
filter context string dictionary w NULLs low selectivity (kept 1/1024)               1.00    391.7±2.17ns        ? ?/sec    1.01    397.2±2.84ns        ? ?/sec
filter context u8 w NULLs (kept 1/2)                                                 1.39     42.3±0.06µs        ? ?/sec    1.00     30.5±0.03µs        ? ?/sec
filter context u8 w NULLs high selectivity (kept 1023/1024)                          1.01      2.8±0.02µs        ? ?/sec    1.00      2.8±0.01µs        ? ?/sec
filter context u8 w NULLs low selectivity (kept 1/1024)                              1.01    315.0±1.89ns        ? ?/sec    1.00    310.8±2.05ns        ? ?/sec
filter fsb with value length 20 w NULLs (kept 1/2)                                   1.18     89.9±0.64µs        ? ?/sec    1.00     76.0±0.53µs        ? ?/sec
filter fsb with value length 20 w NULLs high selectivity (kept 1023/1024)            1.00     26.7±0.39µs        ? ?/sec    1.00     26.8±0.18µs        ? ?/sec
filter fsb with value length 20 w NULLs low selectivity (kept 1/1024)                1.00   1751.1±2.24ns        ? ?/sec    1.01   1770.6±2.62ns        ? ?/sec
filter fsb with value length 5 w NULLs (kept 1/2)                                    1.19     86.2±0.18µs        ? ?/sec    1.00     72.6±0.11µs        ? ?/sec
filter fsb with value length 5 w NULLs high selectivity (kept 1023/1024)             1.01      8.1±0.01µs        ? ?/sec    1.00      8.1±0.02µs        ? ?/sec
filter fsb with value length 5 w NULLs low selectivity (kept 1/1024)                 1.00  1685.0±14.56ns        ? ?/sec    1.02   1726.7±2.03ns        ? ?/sec
filter fsb with value length 50 w NULLs (kept 1/2)                                   1.09    129.6±2.75µs        ? ?/sec    1.00    118.6±4.03µs        ? ?/sec
filter fsb with value length 50 w NULLs high selectivity (kept 1023/1024)            1.00     74.2±4.40µs        ? ?/sec    1.00     74.1±5.81µs        ? ?/sec
filter fsb with value length 50 w NULLs low selectivity (kept 1/1024)                1.00   1741.9±2.28ns        ? ?/sec    1.02   1784.0±1.78ns        ? ?/sec

Resource Usage

base (merge-base)

Metric Value
Wall time 275.1s
Peak memory 32.6 MiB
Avg memory 21.6 MiB
CPU user 269.1s
CPU sys 0.0s
Peak spill 0 B

branch

Metric Value
Wall time 265.1s
Peak memory 27.0 MiB
Avg memory 20.5 MiB
CPU user 258.3s
CPU sys 0.0s
Peak spill 0 B

File an issue against this benchmark runner

mightsleep added a commit to mightsleep/arrow-rs that referenced this pull request Sep 28, 2026
@mightsleep

Copy link
Copy Markdown
Contributor

Yes hello,
@Rich-T-kid some numbers, since you asked for a second pair of eyes. For #11213 I have a different portable compress (not a PR yet) and timed it with this PR and #11086 against main on GitHub runners: filter_bits from main, the large masks from #11271, and 22 masks from real TPC-H and NYC taxi predicates. main / variant, above 1 is faster, 71 cases:

N2 geomean N2 min Zen 4 geomean Zen 4 min
this PR 1.31 0.62 1.16 0.57
#11086 0.90 0.27 0.83 0.20
#11213 draft 2.24 0.92 1.94 0.89

The draft is w/o any LUT table for nibbles; it dispatches on the kept count k, with the bit loop for k ≤ 16 and a direct formula when k ≤ 2 or k ≥ 62. In between, each kept bit moves down by the number of dropped bits below it within its byte, done as shifts by 1, 2 and 4 on all eight bytes at once. Byte i then lands at the sum of the counts of bytes below it, and one multiply by 0x0101..01 gives all eight of those sums. Before opening a PR I am still on its tests, which cover every path of the dispatch, and on Zen 4, where it is still slower on sparse masks.

The benchmarks in #11271 might be useful here too: the existing filter_bits cases repeat a 1024-word mask, which the branch predictor learns, so they hide most of the difference between fallbacks. The large random and clustered masks there do not.

gather_bits in the strategy paths is a clear win (N2: indices 1/2 1.44, slices 9/10 2.58). The nibble table loses to main at about 9 to 20 kept bits per word, which is common in real predicates (TPC-H Q6 date range 0.62).

Would it work to keep gather_bits here and take the portable compress into a separate PR for #11213? Happy to rebase on whichever lands first.

Runs: https://github.com/mightsleep/arrow-rs/actions/runs/36493097865

Is this ok?

@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

interesting, would be nice to have the benchmarks from #11271

@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

run benchmark arrow-select
env:
BENCH_FILTER: filter_bits batches

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark starting (GKE) | trigger

Target: arrow-select

Sharding factor: 1 (1 workers).

Comparing rich-t-kid/filter-bits-gather-optimization (b64eae3) to 1c2c390 (merge-base) diff

Run configuration
run benchmark arrow-select
env:
  BENCH_FILTER: filter_bits batches
shards: 1

Results will be posted when all workers finish.


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark failed or incomplete (GKE) | trigger

Target: arrow-select

Comparing rich-t-kid/filter-bits-gather-optimization (b64eae3) to 1c2c390 (merge-base) diff

Run configuration
run benchmark arrow-select
env:
  BENCH_FILTER: filter_bits batches
shards: 1
Errors / missing shards
arrow-select — shard 1/1: Benchmark for [this request](https://github.com/apache/arrow-rs/pull/11055#issuecomment-6024027681) failed before finishing (Kubernetes reason: `BackoffLimitExceeded`).

Benchmarks requested: `arrow-select`

<details><summary>Runner log (last 40 lines)</summary>

```
    mutable_array
    mutable_buffer_repeat_slice
    null_dict
    nullif_kernel
    occupancy
    offset
    parquet_round_trip
    parse_date
    parse_decimal
    parse_time
    parse_timestamp
    partition_kernels
    primitive_array
    primitive_run_accessor
    primitive_run_take
    project_record
    push_decoder
    record_batch
    regexp_kernels
    row_format
    row_group_index_reader
    row_selection_cursor
    row_selector
    row_selector_boolean_buffer
    serde
    sort_kernel
    string_dictionary_builder
    string_run_builder
    string_run_iterator
    substring_kernels
    take_kernels
    union_array
    variant_builder
    variant_kernels
    variant_validation
    view_types
    writer_overhead
    writer_page_windows
    zip_kernels
```

</details>

<details><summary>Kubernetes message</summary>

```
Job has reached the specified backoff limit
```

</details>

---
[File an issue](https://github.com/adriangb/datafusion-benchmarking/issues) against this benchmark runner
Per-runner information
arrow-select — shard 1/1

Node: gk3-benchmark-cluster-nap-5e8o8x4q-b79dfcaa-hpwl

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

uname:

Linux bench-c6024027681-3100-wpvnw 6.12.94+ #1 SMP Fri Aug 21 08:00:16 UTC 2026 aarch64 GNU/Linux

BENCH_COMMAND:

cargo bench --features=arrow,async,test_common,experimental,object_store --bench arrow-select
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Resource Usage

No completed measurement resource samples.


File an issue against this benchmark runner

@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

@mightsleep do you remember the name of the benchmarks you added?

@mightsleep

Copy link
Copy Markdown
Contributor

It's filter_bits (arrow-select/benches/filter_bits.rs). The batched cases are filter_bits batches (...):

run benchmark filter_bits
env:
  BENCH_FILTER: batches

The portable compress I mentioned is #11322 now; I'll run the same there, so the two are on the same machine.

@mightsleep

Copy link
Copy Markdown
Contributor

the bot doesn't take the command from me; could you trigger the same filter_bits / batches run on my PR? ty.

@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

run benchmark filter_bits
env:
BENCH_FILTER: batches

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark starting (GKE) | trigger

Target: filter_bits

Sharding factor: 1 (1 workers).

Comparing rich-t-kid/filter-bits-gather-optimization (b64eae3) to 1c2c390 (merge-base) diff

Run configuration
run benchmark filter_bits
env:
  BENCH_FILTER: batches
shards: 1

Results will be posted when all workers finish.


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Target: filter_bits

Comparing rich-t-kid/filter-bits-gather-optimization (b64eae3) to 1c2c390 (merge-base) diff

Run configuration
run benchmark filter_bits
env:
  BENCH_FILTER: batches
shards: 1
Details
No matching cases; no measurements run.

Per-runner information
filter_bits — shard 1/1

Node: gk3-benchmark-cluster-nap-wik5zfr4-ca7d2def-nrtv

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

uname:

Linux bench-c6026007676-3101-ppxlm 6.12.94+ #1 SMP Fri Aug 21 08:00:16 UTC 2026 aarch64 GNU/Linux

BENCH_COMMAND:

cargo bench --features=arrow,async,test_common,experimental,object_store --bench filter_bits
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Resource Usage

base (merge-base)

Metric Value
Wall time 5.0s
Peak memory 3.2 MiB
Avg memory 544.0 KiB
CPU user 0.0s
CPU sys 0.0s
Peak spill 0 B

branch

Metric Value
Wall time 5.0s
Peak memory 0 B
Avg memory 0 B
CPU user 0.0s
CPU sys 0.0s
Peak spill 0 B

File an issue against this benchmark runner

@mightsleep

Copy link
Copy Markdown
Contributor

your branch is from before the batched benches landed, merge main for fix.

@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

run benchmark filter_bits
env:
BENCH_FILTER: batches

1 similar comment
@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

run benchmark filter_bits
env:
BENCH_FILTER: batches

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark starting (GKE) | trigger

Target: filter_bits

Sharding factor: 1 (1 workers).

Comparing rich-t-kid/filter-bits-gather-optimization (946cfbc) to 6b34163 (merge-base) diff

Run configuration
run benchmark filter_bits
env:
  BENCH_FILTER: batches
shards: 1

Results will be posted when all workers finish.


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark starting (GKE) | trigger

Target: filter_bits

Sharding factor: 1 (1 workers).

Comparing rich-t-kid/filter-bits-gather-optimization (946cfbc) to 6b34163 (merge-base) diff

Run configuration
run benchmark filter_bits
env:
  BENCH_FILTER: batches
shards: 1

Results will be posted when all workers finish.


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Target: filter_bits

Comparing rich-t-kid/filter-bits-gather-optimization (946cfbc) to 6b34163 (merge-base) diff

Run configuration
run benchmark filter_bits
env:
  BENCH_FILTER: batches
shards: 1
Details
group                                          base                                   changed
-----                                          ----                                   -------
filter_bits batches (random, kept 1/1024)      1.00     90.2±0.52µs        ? ?/sec    1.05     95.0±0.51µs        ? ?/sec
filter_bits batches (random, kept 1/16)        1.00    648.2±0.65µs        ? ?/sec    1.03    666.5±1.50µs        ? ?/sec
filter_bits batches (random, kept 1/2)         2.01      2.3±0.00ms        ? ?/sec    1.00   1137.1±0.61µs        ? ?/sec
filter_bits batches (random, kept 1/256)       1.00    149.2±0.64µs        ? ?/sec    1.12    167.6±0.51µs        ? ?/sec
filter_bits batches (random, kept 1/4)         1.03   1416.5±3.14µs        ? ?/sec    1.00   1371.0±0.74µs        ? ?/sec
filter_bits batches (random, kept 1/64)        1.00    359.0±3.93µs        ? ?/sec    1.07    385.5±1.06µs        ? ?/sec
filter_bits batches (random, kept 15/16)       3.32      3.8±0.00ms        ? ?/sec    1.00   1142.3±0.98µs        ? ?/sec
filter_bits batches (random, kept 3/4)         2.85      3.3±0.00ms        ? ?/sec    1.00   1144.3±0.62µs        ? ?/sec
filter_bits batches (runs of 512, kept 1/2)    2.77   1888.3±0.74µs        ? ?/sec    1.00    681.9±0.67µs        ? ?/sec
filter_bits batches (runs of 512, kept 1/8)    2.36    538.8±0.35µs        ? ?/sec    1.00    228.1±0.54µs        ? ?/sec
filter_bits batches (runs of 512, kept 7/8)    2.93      3.3±0.00ms        ? ?/sec    1.00   1116.1±0.82µs        ? ?/sec
filter_bits batches (runs of 64, kept 1/2)     1.97      2.3±0.00ms        ? ?/sec    1.00   1151.2±1.53µs        ? ?/sec
filter_bits batches (runs of 64, kept 1/8)     1.70    656.5±0.46µs        ? ?/sec    1.00    386.5±0.46µs        ? ?/sec
filter_bits batches (runs of 64, kept 7/8)     2.77      3.5±0.00ms        ? ?/sec    1.00   1265.0±0.53µs        ? ?/sec

Per-runner information
filter_bits — shard 1/1

Node: gk3-benchmark-cluster-nap-1w7yp11l-147db3f1-zq6k

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

uname:

Linux bench-c6028733926-3103-ntsn2 6.12.94+ #1 SMP Fri Aug 21 08:00:16 UTC 2026 aarch64 GNU/Linux

BENCH_COMMAND:

cargo bench --features=arrow,async,test_common,experimental,object_store --bench filter_bits
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Resource Usage

base (merge-base)

Metric Value
Wall time 150.0s
Peak memory 21.5 MiB
Avg memory 20.5 MiB
CPU user 147.5s
CPU sys 0.0s
Peak spill 0 B

branch

Metric Value
Wall time 155.0s
Peak memory 21.5 MiB
Avg memory 20.3 MiB
CPU user 150.2s
CPU sys 0.0s
Peak spill 0 B

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Target: filter_bits

Comparing rich-t-kid/filter-bits-gather-optimization (946cfbc) to 6b34163 (merge-base) diff

Run configuration
run benchmark filter_bits
env:
  BENCH_FILTER: batches
shards: 1
Details
group                                          base                                   changed
-----                                          ----                                   -------
filter_bits batches (random, kept 1/1024)      1.00     89.0±0.52µs        ? ?/sec    1.05     93.8±1.44µs        ? ?/sec
filter_bits batches (random, kept 1/16)        1.00    648.4±0.79µs        ? ?/sec    1.03    665.0±1.94µs        ? ?/sec
filter_bits batches (random, kept 1/2)         2.01      2.3±0.00ms        ? ?/sec    1.00   1137.1±1.60µs        ? ?/sec
filter_bits batches (random, kept 1/256)       1.00    152.0±0.48µs        ? ?/sec    1.09    166.4±1.57µs        ? ?/sec
filter_bits batches (random, kept 1/4)         1.03   1409.7±0.67µs        ? ?/sec    1.00   1367.2±1.84µs        ? ?/sec
filter_bits batches (random, kept 1/64)        1.00    357.6±0.59µs        ? ?/sec    1.08    384.7±2.52µs        ? ?/sec
filter_bits batches (random, kept 15/16)       3.32      3.8±0.00ms        ? ?/sec    1.00   1141.9±1.73µs        ? ?/sec
filter_bits batches (random, kept 3/4)         2.85      3.3±0.00ms        ? ?/sec    1.00   1144.4±2.11µs        ? ?/sec
filter_bits batches (runs of 512, kept 1/2)    2.77   1887.4±0.81µs        ? ?/sec    1.00    680.5±1.59µs        ? ?/sec
filter_bits batches (runs of 512, kept 1/8)    2.38    538.9±0.65µs        ? ?/sec    1.00    226.8±1.68µs        ? ?/sec
filter_bits batches (runs of 512, kept 7/8)    2.93      3.3±0.00ms        ? ?/sec    1.00   1114.4±1.71µs        ? ?/sec
filter_bits batches (runs of 64, kept 1/2)     1.97      2.3±0.00ms        ? ?/sec    1.00   1147.1±2.36µs        ? ?/sec
filter_bits batches (runs of 64, kept 1/8)     1.70    655.8±0.54µs        ? ?/sec    1.00    385.1±1.60µs        ? ?/sec
filter_bits batches (runs of 64, kept 7/8)     2.77      3.5±0.00ms        ? ?/sec    1.00   1262.2±3.35µs        ? ?/sec

Per-runner information
filter_bits — shard 1/1

Node: gk3-benchmark-cluster-nap-1w7yp11l-7d112169-cc8b

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

uname:

Linux bench-c6028733114-3102-s2p2r 6.12.94+ #1 SMP Fri Aug 21 08:00:16 UTC 2026 aarch64 GNU/Linux

BENCH_COMMAND:

cargo bench --features=arrow,async,test_common,experimental,object_store --bench filter_bits
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Resource Usage

base (merge-base)

Metric Value
Wall time 155.0s
Peak memory 21.5 MiB
Avg memory 20.2 MiB
CPU user 150.5s
CPU sys 0.0s
Peak spill 0 B

branch

Metric Value
Wall time 155.0s
Peak memory 21.5 MiB
Avg memory 20.3 MiB
CPU user 150.2s
CPU sys 0.0s
Peak spill 0 B

File an issue against this benchmark runner

@Rich-T-kid

Copy link
Copy Markdown
Contributor Author

pretty nice results

@alamb

alamb commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

run benchmark filter_kernels

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark starting (GKE) | trigger

Target: filter_kernels

Sharding factor: 1 (1 workers).

Comparing rich-t-kid/filter-bits-gather-optimization (946cfbc) to 6b34163 (merge-base) diff

Run configuration
run benchmark filter_kernels
shards: 1

Results will be posted when all workers finish.


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Target: filter_kernels

Comparing rich-t-kid/filter-bits-gather-optimization (946cfbc) to 6b34163 (merge-base) diff

Run configuration
run benchmark filter_kernels
shards: 1
Details
group                                                                                base                                   changed
-----                                                                                ----                                   -------
filter context decimal128 (kept 1/2)                                                 1.00     20.5±0.07µs        ? ?/sec    1.01     20.7±0.06µs        ? ?/sec
filter context decimal128 high selectivity (kept 1023/1024)                          1.09     20.5±0.13µs        ? ?/sec    1.00     18.8±0.15µs        ? ?/sec
filter context decimal128 low selectivity (kept 1/1024)                              1.00    145.7±0.81ns        ? ?/sec    1.02    148.0±3.18ns        ? ?/sec
filter context f32 (kept 1/2)                                                        1.48     41.7±0.09µs        ? ?/sec    1.00     28.3±0.02µs        ? ?/sec
filter context f32 high selectivity (kept 1023/1024)                                 1.00      5.5±0.01µs        ? ?/sec    1.00      5.5±0.02µs        ? ?/sec
filter context f32 low selectivity (kept 1/1024)                                     1.00    322.0±1.47ns        ? ?/sec    1.04    334.7±4.96ns        ? ?/sec
filter context fsb with value length 20 (kept 1/2)                                   1.00     45.7±0.15µs        ? ?/sec    1.03     47.3±0.32µs        ? ?/sec
filter context fsb with value length 20 high selectivity (kept 1023/1024)            1.00     24.2±0.57µs        ? ?/sec    1.02     24.8±0.29µs        ? ?/sec
filter context fsb with value length 20 low selectivity (kept 1/1024)                1.01    280.1±0.54ns        ? ?/sec    1.00    278.0±0.66ns        ? ?/sec
filter context fsb with value length 20 w NULLs (kept 1/2)                           1.23     77.1±0.32µs        ? ?/sec    1.00     62.5±0.18µs        ? ?/sec
filter context fsb with value length 20 w NULLs high selectivity (kept 1023/1024)    1.03     26.5±0.30µs        ? ?/sec    1.00     25.8±0.19µs        ? ?/sec
filter context fsb with value length 20 w NULLs low selectivity (kept 1/1024)        1.00    477.5±1.73ns        ? ?/sec    1.00    478.3±2.59ns        ? ?/sec
filter context fsb with value length 5 (kept 1/2)                                    1.00     44.3±0.04µs        ? ?/sec    1.04     46.2±0.71µs        ? ?/sec
filter context fsb with value length 5 high selectivity (kept 1023/1024)             1.00      4.8±0.01µs        ? ?/sec    1.01      4.8±0.01µs        ? ?/sec
filter context fsb with value length 5 low selectivity (kept 1/1024)                 1.02    237.9±3.05ns        ? ?/sec    1.00    234.2±0.88ns        ? ?/sec
filter context fsb with value length 5 w NULLs (kept 1/2)                            1.23     75.8±0.06µs        ? ?/sec    1.00     61.7±0.82µs        ? ?/sec
filter context fsb with value length 5 w NULLs high selectivity (kept 1023/1024)     1.00      6.6±0.01µs        ? ?/sec    1.01      6.7±0.01µs        ? ?/sec
filter context fsb with value length 5 w NULLs low selectivity (kept 1/1024)         1.00    413.4±3.27ns        ? ?/sec    1.02    419.7±2.43ns        ? ?/sec
filter context fsb with value length 50 (kept 1/2)                                   1.00     85.7±0.57µs        ? ?/sec    1.01     86.6±0.52µs        ? ?/sec
filter context fsb with value length 50 high selectivity (kept 1023/1024)            1.00     71.8±1.11µs        ? ?/sec    1.00     71.7±1.45µs        ? ?/sec
filter context fsb with value length 50 low selectivity (kept 1/1024)                1.08    313.6±0.80ns        ? ?/sec    1.00    290.2±0.42ns        ? ?/sec
filter context fsb with value length 50 w NULLs (kept 1/2)                           1.12    117.1±1.10µs        ? ?/sec    1.00    104.3±0.53µs        ? ?/sec
filter context fsb with value length 50 w NULLs high selectivity (kept 1023/1024)    1.00     67.4±0.91µs        ? ?/sec    1.00     67.3±0.82µs        ? ?/sec
filter context fsb with value length 50 w NULLs low selectivity (kept 1/1024)        1.01    503.9±0.99ns        ? ?/sec    1.00    501.1±2.15ns        ? ?/sec
filter context i32 (kept 1/2)                                                        1.01     12.4±0.04µs        ? ?/sec    1.00     12.2±0.01µs        ? ?/sec
filter context i32 high selectivity (kept 1023/1024)                                 1.01      3.7±0.01µs        ? ?/sec    1.00      3.7±0.00µs        ? ?/sec
filter context i32 low selectivity (kept 1/1024)                                     1.05    139.2±1.14ns        ? ?/sec    1.00    132.9±1.07ns        ? ?/sec
filter context i32 w NULLs (kept 1/2)                                                1.48     43.4±0.36µs        ? ?/sec    1.00     29.3±0.49µs        ? ?/sec
filter context i32 w NULLs at end (kept 1/2)                                         1.48     43.2±0.12µs        ? ?/sec    1.00     29.3±0.04µs        ? ?/sec
filter context i32 w NULLs at end high selectivity (kept 1023/1024)                  1.00      5.5±0.01µs        ? ?/sec    1.00      5.4±0.01µs        ? ?/sec
filter context i32 w NULLs high selectivity (kept 1023/1024)                         1.01      5.5±0.01µs        ? ?/sec    1.00      5.5±0.02µs        ? ?/sec
filter context i32 w NULLs low selectivity (kept 1/1024)                             1.00    327.0±1.30ns        ? ?/sec    1.00    327.4±1.38ns        ? ?/sec
filter context i32 w NULLs, only valid (kept 1/4)                                    1.00     19.6±0.10µs        ? ?/sec    1.11     21.8±0.02µs        ? ?/sec
filter context i32 w NULLs, only valid high selectivity (kept 1023/2048)             1.44     42.6±0.06µs        ? ?/sec    1.00     29.6±0.04µs        ? ?/sec
filter context i32 w NULLs, only valid low selectivity (kept 1/2048)                 1.01    231.1±1.41ns        ? ?/sec    1.00    228.8±2.21ns        ? ?/sec
filter context mixed string view (kept 1/2)                                          1.39     50.9±0.50µs        ? ?/sec    1.00     36.8±0.10µs        ? ?/sec
filter context mixed string view high selectivity (kept 1023/1024)                   1.00     21.2±0.13µs        ? ?/sec    1.03     21.8±0.13µs        ? ?/sec
filter context mixed string view low selectivity (kept 1/1024)                       1.00    320.9±2.19ns        ? ?/sec    1.02    326.3±2.73ns        ? ?/sec
filter context short string view (kept 1/2)                                          1.33     50.1±0.15µs        ? ?/sec    1.00     37.6±0.37µs        ? ?/sec
filter context short string view high selectivity (kept 1023/1024)                   1.00     21.2±0.18µs        ? ?/sec    1.01     21.5±0.13µs        ? ?/sec
filter context short string view low selectivity (kept 1/1024)                       1.00    321.2±2.44ns        ? ?/sec    1.01    322.8±2.18ns        ? ?/sec
filter context string (kept 1/2)                                                     1.04    384.9±1.65µs        ? ?/sec    1.00    370.0±1.61µs        ? ?/sec
filter context string dictionary (kept 1/2)                                          1.15     14.2±0.02µs        ? ?/sec    1.00     12.3±0.01µs        ? ?/sec
filter context string dictionary high selectivity (kept 1023/1024)                   1.00      3.7±0.01µs        ? ?/sec    1.00      3.7±0.00µs        ? ?/sec
filter context string dictionary low selectivity (kept 1/1024)                       1.00    195.8±1.31ns        ? ?/sec    1.00    195.2±1.12ns        ? ?/sec
filter context string dictionary w NULLs (kept 1/2)                                  1.50     43.7±0.17µs        ? ?/sec    1.00     29.2±0.05µs        ? ?/sec
filter context string dictionary w NULLs high selectivity (kept 1023/1024)           1.01      5.6±0.02µs        ? ?/sec    1.00      5.6±0.01µs        ? ?/sec
filter context string dictionary w NULLs low selectivity (kept 1/1024)               1.02    394.0±3.06ns        ? ?/sec    1.00    386.7±4.47ns        ? ?/sec
filter context string high selectivity (kept 1023/1024)                              1.00    327.1±4.51µs        ? ?/sec    1.04    341.5±4.55µs        ? ?/sec
filter context string low selectivity (kept 1/1024)                                  1.02    762.4±1.69ns        ? ?/sec    1.00    749.0±2.86ns        ? ?/sec
filter context u8 (kept 1/2)                                                         1.00     12.1±0.01µs        ? ?/sec    1.11     13.5±0.88µs        ? ?/sec
filter context u8 high selectivity (kept 1023/1024)                                  1.00   1047.4±3.03ns        ? ?/sec    1.05   1104.8±2.40ns        ? ?/sec
filter context u8 low selectivity (kept 1/1024)                                      1.00    123.5±0.93ns        ? ?/sec    1.02    125.4±0.91ns        ? ?/sec
filter context u8 w NULLs (kept 1/2)                                                 1.45     42.5±0.07µs        ? ?/sec    1.00     29.4±0.07µs        ? ?/sec
filter context u8 w NULLs high selectivity (kept 1023/1024)                          1.00      2.8±0.01µs        ? ?/sec    1.04      2.9±0.02µs        ? ?/sec
filter context u8 w NULLs low selectivity (kept 1/1024)                              1.00    308.4±1.78ns        ? ?/sec    1.01    312.4±4.41ns        ? ?/sec
filter decimal128 (kept 1/2)                                                         1.05     32.2±0.07µs        ? ?/sec    1.00     30.7±0.64µs        ? ?/sec
filter decimal128 high selectivity (kept 1023/1024)                                  1.07     21.4±0.40µs        ? ?/sec    1.00     20.1±0.12µs        ? ?/sec
filter decimal128 low selectivity (kept 1/1024)                                      1.00  1244.6±11.71ns        ? ?/sec    1.02   1270.4±3.15ns        ? ?/sec
filter f32 (kept 1/2)                                                                1.24     59.9±0.35µs        ? ?/sec    1.00     48.2±0.03µs        ? ?/sec
filter fsb with value length 20 (kept 1/2)                                           1.07     62.2±0.14µs        ? ?/sec    1.00     58.0±0.11µs        ? ?/sec
filter fsb with value length 20 high selectivity (kept 1023/1024)                    1.00     23.2±0.13µs        ? ?/sec    1.03     23.9±0.27µs        ? ?/sec
filter fsb with value length 20 low selectivity (kept 1/1024)                        1.01   1254.9±0.84ns        ? ?/sec    1.00   1239.7±2.50ns        ? ?/sec
filter fsb with value length 20 w NULLs (kept 1/2)                                   1.25     93.5±0.17µs        ? ?/sec    1.00     75.0±0.13µs        ? ?/sec
filter fsb with value length 20 w NULLs high selectivity (kept 1023/1024)            1.00     25.9±0.16µs        ? ?/sec    1.02     26.5±0.23µs        ? ?/sec
filter fsb with value length 20 w NULLs low selectivity (kept 1/1024)                1.00   1758.5±2.26ns        ? ?/sec    1.03   1804.5±1.66ns        ? ?/sec
filter fsb with value length 5 (kept 1/2)                                            1.09     61.0±0.05µs        ? ?/sec    1.00     55.8±0.06µs        ? ?/sec
filter fsb with value length 5 high selectivity (kept 1023/1024)                     1.03      5.7±0.02µs        ? ?/sec    1.00      5.5±0.01µs        ? ?/sec
filter fsb with value length 5 low selectivity (kept 1/1024)                         1.01   1202.5±3.33ns        ? ?/sec    1.00   1192.0±2.90ns        ? ?/sec
filter fsb with value length 5 w NULLs (kept 1/2)                                    1.26     91.7±0.09µs        ? ?/sec    1.00     72.7±0.11µs        ? ?/sec
filter fsb with value length 5 w NULLs high selectivity (kept 1023/1024)             1.00      8.0±0.02µs        ? ?/sec    1.01      8.1±0.01µs        ? ?/sec
filter fsb with value length 5 w NULLs low selectivity (kept 1/1024)                 1.00   1686.3±8.89ns        ? ?/sec    1.03   1739.7±3.22ns        ? ?/sec
filter fsb with value length 50 (kept 1/2)                                           1.00     94.0±0.35µs        ? ?/sec    1.00     94.2±0.70µs        ? ?/sec
filter fsb with value length 50 high selectivity (kept 1023/1024)                    1.00     72.8±0.97µs        ? ?/sec    1.00     73.0±1.09µs        ? ?/sec
filter fsb with value length 50 low selectivity (kept 1/1024)                        1.02   1280.6±1.40ns        ? ?/sec    1.00   1251.7±2.41ns        ? ?/sec
filter fsb with value length 50 w NULLs (kept 1/2)                                   1.14    126.4±0.83µs        ? ?/sec    1.00    111.4±0.60µs        ? ?/sec
filter fsb with value length 50 w NULLs high selectivity (kept 1023/1024)            1.01     70.6±0.74µs        ? ?/sec    1.00     70.1±0.72µs        ? ?/sec
filter fsb with value length 50 w NULLs low selectivity (kept 1/1024)                1.00   1779.1±1.55ns        ? ?/sec    1.01   1804.0±1.84ns        ? ?/sec
filter i32 (kept 1/2)                                                                1.00     27.3±0.04µs        ? ?/sec    1.03     28.2±0.02µs        ? ?/sec
filter i32 high selectivity (kept 1023/1024)                                         1.01      4.4±0.01µs        ? ?/sec    1.00      4.4±0.01µs        ? ?/sec
filter i32 low selectivity (kept 1/1024)                                             1.00   1156.5±4.62ns        ? ?/sec    1.05   1214.6±2.84ns        ? ?/sec
filter optimize (kept 1/2)                                                           1.00     26.3±0.38µs        ? ?/sec    1.02     26.8±0.20µs        ? ?/sec
filter optimize high selectivity (kept 1023/1024)                                    1.00   1486.5±0.79ns        ? ?/sec    1.01   1499.1±0.97ns        ? ?/sec
filter optimize low selectivity (kept 1/1024)                                        1.00    805.3±2.96ns        ? ?/sec    1.00    804.1±0.43ns        ? ?/sec
filter run array (kept 1/2)                                                          1.01    285.2±1.64µs        ? ?/sec    1.00    283.5±2.02µs        ? ?/sec
filter run array high selectivity (kept 1023/1024)                                   1.00    285.2±3.58µs        ? ?/sec    1.00    284.8±3.47µs        ? ?/sec
filter run array low selectivity (kept 1/1024)                                       1.00    235.0±0.86µs        ? ?/sec    1.00    235.6±1.06µs        ? ?/sec
filter single record batch                                                           1.00     28.6±0.07µs        ? ?/sec    1.10     31.5±0.04µs        ? ?/sec
filter u8 (kept 1/2)                                                                 1.08     28.9±0.07µs        ? ?/sec    1.00     26.7±0.04µs        ? ?/sec
filter u8 high selectivity (kept 1023/1024)                                          1.02   1839.0±8.53ns        ? ?/sec    1.00   1811.0±9.81ns        ? ?/sec
filter u8 low selectivity (kept 1/1024)                                              1.01  1129.6±18.70ns        ? ?/sec    1.00  1123.1±13.92ns        ? ?/sec

Per-runner information
filter_kernels — shard 1/1

Node: gk3-benchmark-cluster-nap-k44dd5ur-8831919f-z9gk

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

uname:

Linux bench-c6039780110-3120-frz2r 6.12.94+ #1 SMP Fri Aug 21 08:00:16 UTC 2026 aarch64 GNU/Linux

BENCH_COMMAND:

cargo bench --features=arrow,async,test_common,experimental,object_store --bench filter_kernels
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Resource Usage

base (merge-base)

Metric Value
Wall time 890.2s
Peak memory 35.3 MiB
Avg memory 22.7 MiB
CPU user 887.8s
CPU sys 0.1s
Peak spill 0 B

branch

Metric Value
Wall time 880.2s
Peak memory 35.2 MiB
Avg memory 22.2 MiB
CPU user 874.0s
CPU sys 0.1s
Peak spill 0 B

File an issue against this benchmark runner

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Improve performance of filter kernel (when pext instruction is not available)

6 participants