Skip to content

perf(arrow-buffer): faster portable fallback for bit_util::compress - #11322

Open
mightsleep wants to merge 2 commits into
apache:mainfrom
mightsleep:compress-portable
Open

mightsleep wants to merge 2 commits into
apache:mainfrom
mightsleep:compress-portable

Conversation

@mightsleep

Copy link
Copy Markdown
Contributor

Which issue does this PR close?

Rationale for this change

Without pext (every aarch64 build, and x86-64 builds without BMI2, the default target) compress walks the kept bits one at a time. With masks from independent rows the number of kept bits changes from word to word, so the loop exit mispredicts on almost every word, and a dense word takes up to 64 steps.

What changes are included in this PR?

The fallback dispatches on the number of kept bits k:

  • k <= 2: the lowest two kept bits, no loop
  • k <= 16: the existing loop, still the cheapest when its branches are predictable (clustered or periodic rows)
  • k >= 62: the word with its at most two dropped bits removed, no loop
  • otherwise compress_bytes: constant time, no table. The parallel suffix network of HD 7-4 inside each byte (three rounds on all eight bytes at once), then the bytes joined at offsets from one multiply by 0x0101..01.

Builds with BMI2 enabled keep pext and are unchanged.

Are these changes tested?

test_compress_portable checks every path against a reference: every mask with at most two kept or two dropped bits, random masks at every popcount. The existing test only draws uniform masks (about 32 kept bits), which reach ne path of four. Filter tests pass on x86-64, x86-64-v2 and with BMI2.

Measurements

main / this PR, above 1 is faster. GitHub runners: Neoverse-N2 (ubuntu-24.04-arm) and AMD EPYC 7763 (ubuntu-latest), both default target, plus x86-64-v2 (which has POPCNT). Each variant built once, three interleaved rounds. Between rounds a row moves by 0.6 % (median), at most 5.5 %.

filter_bits batches from #11271 (512 batches of 8K rows):

mask Neoverse-N2 EPYC 7763 EPYC 7763, v2
random, kept 1/1024 0.99 1.01 0.95
random, kept 1/256 1.03 0.96 1.03
random, kept 1/64 1.21 1.20 1.27
random, kept 1/16 0.97 0.91 0.97
random, kept 1/4 1.21 1.12 1.19
random, kept 1/2 3.37 2.65 2.74
random, kept 3/4 4.68 3.54 3.64
random, kept 15/16 5.40 3.61 3.94
runs of 64, kept 1/8 2.11 1.77 1.87
runs of 64, kept 1/2 3.00 2.35 2.51
runs of 64, kept 7/8 4.75 3.46 3.78
runs of 512, kept 1/2 6.99 5.44 5.70

The other callers of compress:

benchmark Neoverse-N2 EPYC 7763
filter_kernels: filter context i32 w NULLs (kept 1/2) 1.80 1.69
filter_kernels: filter context u8 w NULLs (kept 1/2) 1.83 1.66
filter_kernels: filter context string dictionary w NULLs (kept 1/2) 1.77 1.63
filter_kernels: other w NULLs cases 0.96 to 1.78 0.96 to 1.64
arrow_reader: ListArray and struct (definition levels) 0.99 to 1.11 0.98 to 1.07

What gets slower, ideas welcome

Sparse masks. Up to 9 % on EPYC 7763 (random, kept 1/16), 7 % with x86-64-v2 (filter_bits indices, kept 1/10), 5 % on Neoverse-N2 (indices, kept 1/1024). On a word with one or two kept bits the old loop did almost nothing, and it did it very well. The new code first counts the bits and branches on the count, and on the default x86-64 target that count is a software popcount, so every mispredicted dispatch also waits for it.

What I tried, so nobody has to again:

  • Counting one word ahead, so the count is ready before its branch. perf stat had shown the same instructions and the same mispredictions as main, just about five cycles more per misprediction: the software popcount the branch waits on. The lookahead hid it, and Zen 5 hated it (EPYC 9V45, a large mask kept 1/256: 0.56 against main instead of 0.89).
  • Finding the k <= 2 and k >= 62 cases without a count, m & (m - 1) twice. Cheaper for those words, but every other word paid both tests before its count; words with exactly three kept bits went to 0.59 to 0.75.
  • The same tests plus the previous word's count as a hint for the rest: smaller gains everywhere, one slow case fixed and several new ones. On random rows the previous word is a poor hint.

What I have not tried, in case it itches someone:

  • Pipeline by blocks instead of words: popcount four to eight words at once (SWAR works even on the SSE2 baseline) and dispatch on the block. A sparse block could then walk its set bits across words in one loop, and pay for the loop exit once per block, not once per word.
  • Anything that gets the sparse case back to the naive loop and then past it.

The runners are shared VMs. The macos-15 runner is not in the tables: its rounds differed by up to 30 % even on cases that never reach compress.

Are there any user-facing changes?

No. compress keeps its signature and its results; only its speed without BMI2 changes.

@mightsleep

Copy link
Copy Markdown
Contributor Author

run benchmark filter_bits
env:
BENCH_FILTER: batches

@adriangbot

This comment has been minimized.

@alamb

alamb commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

run benchmark filter_bits
env:
BENCH_FILTER: batches

@adriangbot

This comment has been minimized.

@alamb

alamb commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

run benchmark filter_bits

env:
  BENCH_FILTER: batches

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark starting (GKE) | trigger

Target: filter_bits

Sharding factor: 1 (1 workers).

Comparing compress-portable (2e6f713) to 08651bc (merge-base) diff

Run configuration
run benchmark filter_bits
env:
  BENCH_FILTER: batches
shards: 1

Results will be posted when all workers finish.


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Target: filter_bits

Comparing compress-portable (2e6f713) to 08651bc (merge-base) diff

Run configuration
run benchmark filter_bits
env:
  BENCH_FILTER: batches
shards: 1
Details
group                                          base                                   changed
-----                                          ----                                   -------
filter_bits batches (random, kept 1/1024)      1.00     90.0±0.54µs        ? ?/sec    1.01     91.1±0.26µs        ? ?/sec
filter_bits batches (random, kept 1/16)        1.00    651.3±0.94µs        ? ?/sec    1.07    695.8±1.08µs        ? ?/sec
filter_bits batches (random, kept 1/2)         3.94      2.3±0.00ms        ? ?/sec    1.00    591.0±0.33µs        ? ?/sec
filter_bits batches (random, kept 1/256)       1.04    151.7±1.09µs        ? ?/sec    1.00    145.4±0.25µs        ? ?/sec
filter_bits batches (random, kept 1/4)         1.25   1428.3±2.20µs        ? ?/sec    1.00   1139.3±0.96µs        ? ?/sec
filter_bits batches (random, kept 1/64)        1.26    359.8±0.92µs        ? ?/sec    1.00    286.4±0.57µs        ? ?/sec
filter_bits batches (random, kept 15/16)       6.03      3.8±0.00ms        ? ?/sec    1.00    630.0±1.44µs        ? ?/sec
filter_bits batches (random, kept 3/4)         5.44      3.3±0.00ms        ? ?/sec    1.00    599.2±0.46µs        ? ?/sec
filter_bits batches (runs of 512, kept 1/2)    7.67   1886.1±0.79µs        ? ?/sec    1.00    245.8±0.71µs        ? ?/sec
filter_bits batches (runs of 512, kept 1/8)    4.81    538.4±0.65µs        ? ?/sec    1.00    111.8±0.43µs        ? ?/sec
filter_bits batches (runs of 512, kept 7/8)    9.63      3.3±0.00ms        ? ?/sec    1.00    338.8±0.39µs        ? ?/sec
filter_bits batches (runs of 64, kept 1/2)     3.17      2.3±0.00ms        ? ?/sec    1.00    713.1±1.39µs        ? ?/sec
filter_bits batches (runs of 64, kept 1/8)     2.26    657.7±0.77µs        ? ?/sec    1.00    290.6±0.37µs        ? ?/sec
filter_bits batches (runs of 64, kept 7/8)     5.17      3.5±0.00ms        ? ?/sec    1.00    679.4±2.81µs        ? ?/sec

Per-runner information
filter_bits — shard 1/1

Node: gk3-benchmark-cluster-nap-k44dd5ur-758ad03e-dbbg

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

uname:

Linux bench-c6039874955-3121-rmh6z 6.12.94+ #1 SMP Fri Aug 21 08:00:16 UTC 2026 aarch64 GNU/Linux

BENCH_COMMAND:

cargo bench --features=arrow,async,test_common,experimental,object_store --bench filter_bits
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Resource Usage

base (merge-base)

Metric Value
Wall time 155.0s
Peak memory 21.5 MiB
Avg memory 20.2 MiB
CPU user 151.0s
CPU sys 0.0s
Peak spill 0 B

branch

Metric Value
Wall time 155.0s
Peak memory 21.5 MiB
Avg memory 20.5 MiB
CPU user 152.2s
CPU sys 0.0s
Peak spill 0 B

File an issue against this benchmark runner

@mightsleep

Copy link
Copy Markdown
Contributor Author

ty for running. Could you also run filter_kernels? #11055 has it on the same runner, and that would put both compress versions side by side on the null-mask paths too

@alamb

alamb commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

run benchmark filter_kernels

1 similar comment
@alamb

alamb commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

run benchmark filter_kernels

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark starting (GKE) | trigger

Target: filter_kernels

Sharding factor: 1 (1 workers).

Comparing compress-portable (2e6f713) to 08651bc (merge-base) diff

Run configuration
run benchmark filter_kernels
shards: 1

Results will be posted when all workers finish.


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark starting (GKE) | trigger

Target: filter_kernels

Sharding factor: 1 (1 workers).

Comparing compress-portable (2e6f713) to 08651bc (merge-base) diff

Run configuration
run benchmark filter_kernels
shards: 1

Results will be posted when all workers finish.


File an issue against this benchmark runner

@alamb

alamb commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

FYI @devanbenz

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Target: filter_kernels

Comparing compress-portable (2e6f713) to 08651bc (merge-base) diff

Run configuration
run benchmark filter_kernels
shards: 1
Details
group                                                                                base                                   changed
-----                                                                                ----                                   -------
filter context decimal128 (kept 1/2)                                                 1.03     21.0±0.06µs        ? ?/sec    1.00     20.3±0.07µs        ? ?/sec
filter context decimal128 high selectivity (kept 1023/1024)                          1.11     20.6±0.08µs        ? ?/sec    1.00     18.6±0.17µs        ? ?/sec
filter context decimal128 low selectivity (kept 1/1024)                              1.02    148.3±1.81ns        ? ?/sec    1.00    145.6±0.83ns        ? ?/sec
filter context f32 (kept 1/2)                                                        1.93     43.0±0.05µs        ? ?/sec    1.00     22.3±0.28µs        ? ?/sec
filter context f32 high selectivity (kept 1023/1024)                                 1.05      5.8±0.01µs        ? ?/sec    1.00      5.5±0.01µs        ? ?/sec
filter context f32 low selectivity (kept 1/1024)                                     1.01    329.0±5.87ns        ? ?/sec    1.00    324.8±1.94ns        ? ?/sec
filter context fsb with value length 20 (kept 1/2)                                   1.01     46.1±0.20µs        ? ?/sec    1.00     45.6±0.14µs        ? ?/sec
filter context fsb with value length 20 high selectivity (kept 1023/1024)            1.00     24.2±0.52µs        ? ?/sec    1.02     24.8±0.30µs        ? ?/sec
filter context fsb with value length 20 low selectivity (kept 1/1024)                1.05    276.3±0.62ns        ? ?/sec    1.00    262.3±0.31ns        ? ?/sec
filter context fsb with value length 20 w NULLs (kept 1/2)                           1.38     77.1±0.21µs        ? ?/sec    1.00     55.7±0.12µs        ? ?/sec
filter context fsb with value length 20 w NULLs high selectivity (kept 1023/1024)    1.00     25.0±0.64µs        ? ?/sec    1.07     26.8±0.39µs        ? ?/sec
filter context fsb with value length 20 w NULLs low selectivity (kept 1/1024)        1.02    474.3±2.91ns        ? ?/sec    1.00    463.4±1.63ns        ? ?/sec
filter context fsb with value length 5 (kept 1/2)                                    1.00     44.2±0.02µs        ? ?/sec    1.00     44.3±0.03µs        ? ?/sec
filter context fsb with value length 5 high selectivity (kept 1023/1024)             1.01      4.8±0.00µs        ? ?/sec    1.00      4.7±0.01µs        ? ?/sec
filter context fsb with value length 5 low selectivity (kept 1/1024)                 1.00    228.2±2.81ns        ? ?/sec    1.02    233.2±0.87ns        ? ?/sec
filter context fsb with value length 5 w NULLs (kept 1/2)                            1.43     75.8±0.06µs        ? ?/sec    1.00     52.8±0.03µs        ? ?/sec
filter context fsb with value length 5 w NULLs high selectivity (kept 1023/1024)     1.01      6.6±0.01µs        ? ?/sec    1.00      6.6±0.01µs        ? ?/sec
filter context fsb with value length 5 w NULLs low selectivity (kept 1/1024)         1.00    413.8±3.41ns        ? ?/sec    1.01    417.7±1.62ns        ? ?/sec
filter context fsb with value length 50 (kept 1/2)                                   1.00     89.4±0.56µs        ? ?/sec    1.00     89.4±0.17µs        ? ?/sec
filter context fsb with value length 50 high selectivity (kept 1023/1024)            1.02     70.1±1.14µs        ? ?/sec    1.00     68.9±0.55µs        ? ?/sec
filter context fsb with value length 50 low selectivity (kept 1/1024)                1.08    315.4±0.64ns        ? ?/sec    1.00    291.7±0.73ns        ? ?/sec
filter context fsb with value length 50 w NULLs (kept 1/2)                           1.25    121.2±0.66µs        ? ?/sec    1.00     96.6±0.21µs        ? ?/sec
filter context fsb with value length 50 w NULLs high selectivity (kept 1023/1024)    1.01     66.9±0.62µs        ? ?/sec    1.00     66.0±0.38µs        ? ?/sec
filter context fsb with value length 50 w NULLs low selectivity (kept 1/1024)        1.01    496.1±2.31ns        ? ?/sec    1.00    492.9±1.32ns        ? ?/sec
filter context i32 (kept 1/2)                                                        1.00     12.3±0.01µs        ? ?/sec    1.00     12.3±0.01µs        ? ?/sec
filter context i32 high selectivity (kept 1023/1024)                                 1.00      3.7±0.00µs        ? ?/sec    1.00      3.7±0.00µs        ? ?/sec
filter context i32 low selectivity (kept 1/1024)                                     1.00    138.6±1.21ns        ? ?/sec    1.00    138.2±1.14ns        ? ?/sec
filter context i32 w NULLs (kept 1/2)                                                2.08     43.4±0.04µs        ? ?/sec    1.00     20.9±0.02µs        ? ?/sec
filter context i32 w NULLs at end (kept 1/2)                                         2.10     44.5±0.27µs        ? ?/sec    1.00     21.2±0.02µs        ? ?/sec
filter context i32 w NULLs at end high selectivity (kept 1023/1024)                  1.01      5.5±0.01µs        ? ?/sec    1.00      5.4±0.01µs        ? ?/sec
filter context i32 w NULLs high selectivity (kept 1023/1024)                         1.01      5.5±0.01µs        ? ?/sec    1.00      5.4±0.01µs        ? ?/sec
filter context i32 w NULLs low selectivity (kept 1/1024)                             1.01    327.9±2.34ns        ? ?/sec    1.00    326.0±1.38ns        ? ?/sec
filter context i32 w NULLs, only valid (kept 1/4)                                    1.22     19.7±0.15µs        ? ?/sec    1.00     16.2±0.14µs        ? ?/sec
filter context i32 w NULLs, only valid high selectivity (kept 1023/2048)             2.09     43.5±0.06µs        ? ?/sec    1.00     20.8±0.01µs        ? ?/sec
filter context i32 w NULLs, only valid low selectivity (kept 1/2048)                 1.01    232.5±2.20ns        ? ?/sec    1.00    230.2±1.67ns        ? ?/sec
filter context mixed string view (kept 1/2)                                          1.84     52.6±0.10µs        ? ?/sec    1.00     28.7±0.07µs        ? ?/sec
filter context mixed string view high selectivity (kept 1023/1024)                   1.01     21.2±0.18µs        ? ?/sec    1.00     21.0±0.17µs        ? ?/sec
filter context mixed string view low selectivity (kept 1/1024)                       1.00    319.9±2.36ns        ? ?/sec    1.01    324.1±2.45ns        ? ?/sec
filter context short string view (kept 1/2)                                          1.83     52.6±0.05µs        ? ?/sec    1.00     28.8±0.09µs        ? ?/sec
filter context short string view high selectivity (kept 1023/1024)                   1.04     21.7±0.28µs        ? ?/sec    1.00     20.9±0.12µs        ? ?/sec
filter context short string view low selectivity (kept 1/1024)                       1.00    321.7±2.88ns        ? ?/sec    1.00    322.2±2.83ns        ? ?/sec
filter context string (kept 1/2)                                                     1.09    386.4±2.18µs        ? ?/sec    1.00    355.8±0.78µs        ? ?/sec
filter context string dictionary (kept 1/2)                                          1.00     12.3±0.01µs        ? ?/sec    1.01     12.4±0.02µs        ? ?/sec
filter context string dictionary high selectivity (kept 1023/1024)                   1.00      3.7±0.00µs        ? ?/sec    1.00      3.7±0.01µs        ? ?/sec
filter context string dictionary low selectivity (kept 1/1024)                       1.00    194.5±1.26ns        ? ?/sec    1.00    194.2±1.87ns        ? ?/sec
filter context string dictionary w NULLs (kept 1/2)                                  2.16     45.5±0.03µs        ? ?/sec    1.00     21.1±0.01µs        ? ?/sec
filter context string dictionary w NULLs high selectivity (kept 1023/1024)           1.00      5.6±0.01µs        ? ?/sec    1.01      5.7±0.01µs        ? ?/sec
filter context string dictionary w NULLs low selectivity (kept 1/1024)               1.01    392.2±1.87ns        ? ?/sec    1.00    387.4±3.86ns        ? ?/sec
filter context string high selectivity (kept 1023/1024)                              1.01    316.5±1.14µs        ? ?/sec    1.00    314.3±2.14µs        ? ?/sec
filter context string low selectivity (kept 1/1024)                                  1.02    758.5±0.93ns        ? ?/sec    1.00    745.6±0.86ns        ? ?/sec
filter context u8 (kept 1/2)                                                         1.14     13.8±0.62µs        ? ?/sec    1.00     12.1±0.01µs        ? ?/sec
filter context u8 high selectivity (kept 1023/1024)                                  1.00   1050.7±2.73ns        ? ?/sec    1.05   1106.9±3.11ns        ? ?/sec
filter context u8 low selectivity (kept 1/1024)                                      1.00    123.6±0.76ns        ? ?/sec    1.01    125.3±0.80ns        ? ?/sec
filter context u8 w NULLs (kept 1/2)                                                 2.10     43.4±0.03µs        ? ?/sec    1.00     20.7±0.03µs        ? ?/sec
filter context u8 w NULLs high selectivity (kept 1023/1024)                          1.00      2.8±0.01µs        ? ?/sec    1.04      2.9±0.01µs        ? ?/sec
filter context u8 w NULLs low selectivity (kept 1/1024)                              1.02    315.4±6.17ns        ? ?/sec    1.00    308.7±1.55ns        ? ?/sec
filter decimal128 (kept 1/2)                                                         1.03     31.6±0.08µs        ? ?/sec    1.00     30.7±0.07µs        ? ?/sec
filter decimal128 high selectivity (kept 1023/1024)                                  1.07     21.4±0.09µs        ? ?/sec    1.00     19.9±0.14µs        ? ?/sec
filter decimal128 low selectivity (kept 1/1024)                                      1.01   1229.8±4.53ns        ? ?/sec    1.00   1219.7±3.71ns        ? ?/sec
filter f32 (kept 1/2)                                                                1.61     61.0±0.13µs        ? ?/sec    1.00     37.8±0.08µs        ? ?/sec
filter fsb with value length 20 (kept 1/2)                                           1.09     63.0±0.07µs        ? ?/sec    1.00     57.8±0.12µs        ? ?/sec
filter fsb with value length 20 high selectivity (kept 1023/1024)                    1.00     23.9±0.14µs        ? ?/sec    1.01     24.1±0.14µs        ? ?/sec
filter fsb with value length 20 low selectivity (kept 1/1024)                        1.06   1314.5±4.64ns        ? ?/sec    1.00   1244.4±2.08ns        ? ?/sec
filter fsb with value length 20 w NULLs (kept 1/2)                                   1.43     95.2±0.11µs        ? ?/sec    1.00     66.4±0.16µs        ? ?/sec
filter fsb with value length 20 w NULLs high selectivity (kept 1023/1024)            1.00     26.5±0.13µs        ? ?/sec    1.01     26.8±0.11µs        ? ?/sec
filter fsb with value length 20 w NULLs low selectivity (kept 1/1024)                1.04   1829.9±7.35ns        ? ?/sec    1.00   1755.9±3.21ns        ? ?/sec
filter fsb with value length 5 (kept 1/2)                                            1.09     60.9±0.02µs        ? ?/sec    1.00     55.8±0.05µs        ? ?/sec
filter fsb with value length 5 high selectivity (kept 1023/1024)                     1.02      5.5±0.01µs        ? ?/sec    1.00      5.4±0.01µs        ? ?/sec
filter fsb with value length 5 low selectivity (kept 1/1024)                         1.07   1275.3±8.23ns        ? ?/sec    1.00   1194.9±2.36ns        ? ?/sec
filter fsb with value length 5 w NULLs (kept 1/2)                                    1.44     93.3±0.24µs        ? ?/sec    1.00     64.8±0.07µs        ? ?/sec
filter fsb with value length 5 w NULLs high selectivity (kept 1023/1024)             1.02      8.1±0.01µs        ? ?/sec    1.00      8.0±0.01µs        ? ?/sec
filter fsb with value length 5 w NULLs low selectivity (kept 1/1024)                 1.03  1761.9±10.64ns        ? ?/sec    1.00   1704.3±8.25ns        ? ?/sec
filter fsb with value length 50 (kept 1/2)                                           1.01     97.4±0.26µs        ? ?/sec    1.00     96.1±0.34µs        ? ?/sec
filter fsb with value length 50 high selectivity (kept 1023/1024)                    1.02     71.4±0.51µs        ? ?/sec    1.00     70.1±0.38µs        ? ?/sec
filter fsb with value length 50 low selectivity (kept 1/1024)                        1.08   1348.7±9.16ns        ? ?/sec    1.00   1252.5±1.73ns        ? ?/sec
filter fsb with value length 50 w NULLs (kept 1/2)                                   1.24    128.8±0.31µs        ? ?/sec    1.00    103.9±0.35µs        ? ?/sec
filter fsb with value length 50 w NULLs high selectivity (kept 1023/1024)            1.01     69.5±0.29µs        ? ?/sec    1.00     69.1±0.51µs        ? ?/sec
filter fsb with value length 50 w NULLs low selectivity (kept 1/1024)                1.05   1850.6±7.12ns        ? ?/sec    1.00   1766.5±2.63ns        ? ?/sec
filter i32 (kept 1/2)                                                                1.00     27.3±0.07µs        ? ?/sec    1.01     27.7±0.24µs        ? ?/sec
filter i32 high selectivity (kept 1023/1024)                                         1.01      4.3±0.01µs        ? ?/sec    1.00      4.3±0.01µs        ? ?/sec
filter i32 low selectivity (kept 1/1024)                                             1.00   1152.9±2.97ns        ? ?/sec    1.05   1212.1±0.95ns        ? ?/sec
filter optimize (kept 1/2)                                                           1.01     27.8±0.05µs        ? ?/sec    1.00     27.6±0.04µs        ? ?/sec
filter optimize high selectivity (kept 1023/1024)                                    1.00   1487.4±0.93ns        ? ?/sec    1.01   1499.2±1.21ns        ? ?/sec
filter optimize low selectivity (kept 1/1024)                                        1.00    801.8±0.44ns        ? ?/sec    1.00    802.5±0.64ns        ? ?/sec
filter run array (kept 1/2)                                                          1.00    285.2±0.79µs        ? ?/sec    1.00    284.2±2.32µs        ? ?/sec
filter run array high selectivity (kept 1023/1024)                                   1.01    290.0±0.79µs        ? ?/sec    1.00    286.6±4.22µs        ? ?/sec
filter run array low selectivity (kept 1/1024)                                       1.00    235.1±0.89µs        ? ?/sec    1.00    235.7±0.90µs        ? ?/sec
filter single record batch                                                           1.00     29.0±0.05µs        ? ?/sec    1.01     29.3±0.05µs        ? ?/sec
filter u8 (kept 1/2)                                                                 1.00     27.4±0.03µs        ? ?/sec    1.00     27.3±0.06µs        ? ?/sec
filter u8 high selectivity (kept 1023/1024)                                          1.00   1817.0±3.89ns        ? ?/sec    1.02   1851.0±6.46ns        ? ?/sec
filter u8 low selectivity (kept 1/1024)                                              1.00  1160.8±13.67ns        ? ?/sec    1.00  1159.7±10.98ns        ? ?/sec

Per-runner information
filter_kernels — shard 1/1

Node: gk3-benchmark-cluster-nap-138ltj7s-756b5f7d-kl2n

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

uname:

Linux bench-c6058396235-3165-hv46c 6.12.94+ #1 SMP Fri Aug 21 08:00:16 UTC 2026 aarch64 GNU/Linux

BENCH_COMMAND:

cargo bench --features=arrow,async,test_common,experimental,object_store --bench filter_kernels
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Resource Usage

base (merge-base)

Metric Value
Wall time 890.2s
Peak memory 35.3 MiB
Avg memory 22.7 MiB
CPU user 887.8s
CPU sys 0.1s
Peak spill 0 B

branch

Metric Value
Wall time 890.2s
Peak memory 37.2 MiB
Avg memory 23.5 MiB
CPU user 887.0s
CPU sys 0.0s
Peak spill 0 B

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Target: filter_kernels

Comparing compress-portable (2e6f713) to 08651bc (merge-base) diff

Run configuration
run benchmark filter_kernels
shards: 1
Details
group                                                                                base                                   changed
-----                                                                                ----                                   -------
filter context decimal128 (kept 1/2)                                                 1.01     20.2±0.07µs        ? ?/sec    1.00     20.0±0.05µs        ? ?/sec
filter context decimal128 high selectivity (kept 1023/1024)                          1.04     19.5±0.06µs        ? ?/sec    1.00     18.7±0.16µs        ? ?/sec
filter context decimal128 low selectivity (kept 1/1024)                              1.00    145.7±0.65ns        ? ?/sec    1.00    145.0±0.59ns        ? ?/sec
filter context f32 (kept 1/2)                                                        2.19     45.5±0.25µs        ? ?/sec    1.00     20.8±1.14µs        ? ?/sec
filter context f32 high selectivity (kept 1023/1024)                                 1.01      5.6±0.01µs        ? ?/sec    1.00      5.5±0.01µs        ? ?/sec
filter context f32 low selectivity (kept 1/1024)                                     1.00    321.8±1.07ns        ? ?/sec    1.01    326.4±1.39ns        ? ?/sec
filter context fsb with value length 20 (kept 1/2)                                   1.01     45.7±0.15µs        ? ?/sec    1.00     45.4±0.16µs        ? ?/sec
filter context fsb with value length 20 high selectivity (kept 1023/1024)            1.00     23.5±0.53µs        ? ?/sec    1.00     23.5±0.17µs        ? ?/sec
filter context fsb with value length 20 low selectivity (kept 1/1024)                1.06    281.3±5.28ns        ? ?/sec    1.00    266.0±1.88ns        ? ?/sec
filter context fsb with value length 20 w NULLs (kept 1/2)                           1.46     79.2±0.84µs        ? ?/sec    1.00     54.5±0.77µs        ? ?/sec
filter context fsb with value length 20 w NULLs high selectivity (kept 1023/1024)    1.02     26.4±0.28µs        ? ?/sec    1.00     25.9±0.16µs        ? ?/sec
filter context fsb with value length 20 w NULLs low selectivity (kept 1/1024)        1.00    478.1±2.40ns        ? ?/sec    1.00    475.9±1.12ns        ? ?/sec
filter context fsb with value length 5 (kept 1/2)                                    1.00     44.4±0.04µs        ? ?/sec    1.03     45.9±0.14µs        ? ?/sec
filter context fsb with value length 5 high selectivity (kept 1023/1024)             1.02      4.8±0.00µs        ? ?/sec    1.00      4.7±0.00µs        ? ?/sec
filter context fsb with value length 5 low selectivity (kept 1/1024)                 1.00    231.4±6.35ns        ? ?/sec    1.01    234.1±1.50ns        ? ?/sec
filter context fsb with value length 5 w NULLs (kept 1/2)                            1.38     76.1±0.29µs        ? ?/sec    1.00     55.0±0.14µs        ? ?/sec
filter context fsb with value length 5 w NULLs high selectivity (kept 1023/1024)     1.02      6.6±0.01µs        ? ?/sec    1.00      6.5±0.01µs        ? ?/sec
filter context fsb with value length 5 w NULLs low selectivity (kept 1/1024)         1.00    414.9±5.23ns        ? ?/sec    1.02    421.8±2.52ns        ? ?/sec
filter context fsb with value length 50 (kept 1/2)                                   1.04     90.4±0.50µs        ? ?/sec    1.00     86.8±0.19µs        ? ?/sec
filter context fsb with value length 50 high selectivity (kept 1023/1024)            1.00     69.5±1.24µs        ? ?/sec    1.01     70.3±0.81µs        ? ?/sec
filter context fsb with value length 50 low selectivity (kept 1/1024)                1.08    316.3±0.66ns        ? ?/sec    1.00    294.2±1.31ns        ? ?/sec
filter context fsb with value length 50 w NULLs (kept 1/2)                           1.24    117.6±1.07µs        ? ?/sec    1.00     95.1±0.27µs        ? ?/sec
filter context fsb with value length 50 w NULLs high selectivity (kept 1023/1024)    1.01     66.2±0.97µs        ? ?/sec    1.00     65.8±0.88µs        ? ?/sec
filter context fsb with value length 50 w NULLs low selectivity (kept 1/1024)        1.00    496.3±1.34ns        ? ?/sec    1.01    500.3±1.52ns        ? ?/sec
filter context i32 (kept 1/2)                                                        1.00     12.3±0.01µs        ? ?/sec    1.00     12.3±0.01µs        ? ?/sec
filter context i32 high selectivity (kept 1023/1024)                                 1.00      3.7±0.00µs        ? ?/sec    1.00      3.7±0.01µs        ? ?/sec
filter context i32 low selectivity (kept 1/1024)                                     1.00    133.5±1.00ns        ? ?/sec    1.01    135.1±0.57ns        ? ?/sec
filter context i32 w NULLs (kept 1/2)                                                2.12     45.4±0.64µs        ? ?/sec    1.00     21.4±0.69µs        ? ?/sec
filter context i32 w NULLs at end (kept 1/2)                                         2.05     44.5±1.18µs        ? ?/sec    1.00     21.7±0.05µs        ? ?/sec
filter context i32 w NULLs at end high selectivity (kept 1023/1024)                  1.02      5.5±0.01µs        ? ?/sec    1.00      5.4±0.01µs        ? ?/sec
filter context i32 w NULLs high selectivity (kept 1023/1024)                         1.01      5.5±0.01µs        ? ?/sec    1.00      5.4±0.01µs        ? ?/sec
filter context i32 w NULLs low selectivity (kept 1/1024)                             1.00    327.7±1.59ns        ? ?/sec    1.00    328.5±1.41ns        ? ?/sec
filter context i32 w NULLs, only valid (kept 1/4)                                    1.20     19.7±0.05µs        ? ?/sec    1.00     16.3±0.18µs        ? ?/sec
filter context i32 w NULLs, only valid high selectivity (kept 1023/2048)             2.13     44.8±0.08µs        ? ?/sec    1.00     21.0±0.02µs        ? ?/sec
filter context i32 w NULLs, only valid low selectivity (kept 1/2048)                 1.00    227.3±0.57ns        ? ?/sec    1.01    229.3±0.99ns        ? ?/sec
filter context mixed string view (kept 1/2)                                          1.77     51.5±0.25µs        ? ?/sec    1.00     29.2±0.11µs        ? ?/sec
filter context mixed string view high selectivity (kept 1023/1024)                   1.00     20.5±0.19µs        ? ?/sec    1.07     21.9±0.11µs        ? ?/sec
filter context mixed string view low selectivity (kept 1/1024)                       1.00    323.7±3.74ns        ? ?/sec    1.00    322.6±1.80ns        ? ?/sec
filter context short string view (kept 1/2)                                          1.82     51.4±0.15µs        ? ?/sec    1.00     28.3±0.10µs        ? ?/sec
filter context short string view high selectivity (kept 1023/1024)                   1.00     19.7±0.11µs        ? ?/sec    1.04     20.4±0.11µs        ? ?/sec
filter context short string view low selectivity (kept 1/1024)                       1.00    322.7±2.36ns        ? ?/sec    1.00    323.4±2.21ns        ? ?/sec
filter context string (kept 1/2)                                                     1.08    381.6±1.01µs        ? ?/sec    1.00    354.2±0.89µs        ? ?/sec
filter context string dictionary (kept 1/2)                                          1.10     13.6±0.78µs        ? ?/sec    1.00     12.4±0.02µs        ? ?/sec
filter context string dictionary high selectivity (kept 1023/1024)                   1.00      3.7±0.01µs        ? ?/sec    1.00      3.7±0.00µs        ? ?/sec
filter context string dictionary low selectivity (kept 1/1024)                       1.00    195.3±1.83ns        ? ?/sec    1.00    194.8±1.65ns        ? ?/sec
filter context string dictionary w NULLs (kept 1/2)                                  1.98     43.5±0.13µs        ? ?/sec    1.00     22.0±0.37µs        ? ?/sec
filter context string dictionary w NULLs high selectivity (kept 1023/1024)           1.00      5.6±0.01µs        ? ?/sec    1.01      5.7±0.01µs        ? ?/sec
filter context string dictionary w NULLs low selectivity (kept 1/1024)               1.01    392.6±1.91ns        ? ?/sec    1.00    388.7±3.37ns        ? ?/sec
filter context string high selectivity (kept 1023/1024)                              1.00    309.8±0.96µs        ? ?/sec    1.00    310.9±0.89µs        ? ?/sec
filter context string low selectivity (kept 1/1024)                                  1.02    760.8±1.37ns        ? ?/sec    1.00    744.6±1.10ns        ? ?/sec
filter context u8 (kept 1/2)                                                         1.00     12.1±0.01µs        ? ?/sec    1.00     12.1±0.01µs        ? ?/sec
filter context u8 high selectivity (kept 1023/1024)                                  1.00   1049.9±2.42ns        ? ?/sec    1.06   1112.4±2.30ns        ? ?/sec
filter context u8 low selectivity (kept 1/1024)                                      1.00    124.9±0.68ns        ? ?/sec    1.01    125.8±0.80ns        ? ?/sec
filter context u8 w NULLs (kept 1/2)                                                 2.12     43.7±0.07µs        ? ?/sec    1.00     20.7±0.01µs        ? ?/sec
filter context u8 w NULLs high selectivity (kept 1023/1024)                          1.00      2.8±0.01µs        ? ?/sec    1.03      2.9±0.01µs        ? ?/sec
filter context u8 w NULLs low selectivity (kept 1/1024)                              1.01    310.3±1.65ns        ? ?/sec    1.00    308.2±0.88ns        ? ?/sec
filter decimal128 (kept 1/2)                                                         1.01     30.9±0.04µs        ? ?/sec    1.00     30.7±0.32µs        ? ?/sec
filter decimal128 high selectivity (kept 1023/1024)                                  1.01     20.2±0.19µs        ? ?/sec    1.00     20.1±0.10µs        ? ?/sec
filter decimal128 low selectivity (kept 1/1024)                                      1.01   1231.6±2.79ns        ? ?/sec    1.00   1215.4±3.30ns        ? ?/sec
filter f32 (kept 1/2)                                                                1.61     60.8±0.11µs        ? ?/sec    1.00     37.7±0.02µs        ? ?/sec
filter fsb with value length 20 (kept 1/2)                                           1.08     62.0±0.06µs        ? ?/sec    1.00     57.4±0.07µs        ? ?/sec
filter fsb with value length 20 high selectivity (kept 1023/1024)                    1.00     22.7±0.14µs        ? ?/sec    1.04     23.7±0.20µs        ? ?/sec
filter fsb with value length 20 low selectivity (kept 1/1024)                        1.05   1302.9±3.07ns        ? ?/sec    1.00   1243.3±3.41ns        ? ?/sec
filter fsb with value length 20 w NULLs (kept 1/2)                                   1.41     93.2±0.09µs        ? ?/sec    1.00     66.1±0.16µs        ? ?/sec
filter fsb with value length 20 w NULLs high selectivity (kept 1023/1024)            1.00     25.4±0.35µs        ? ?/sec    1.03     26.1±0.15µs        ? ?/sec
filter fsb with value length 20 w NULLs low selectivity (kept 1/1024)                1.03   1821.3±2.94ns        ? ?/sec    1.00   1766.1±2.60ns        ? ?/sec
filter fsb with value length 5 (kept 1/2)                                            1.09     60.9±0.05µs        ? ?/sec    1.00     55.8±0.05µs        ? ?/sec
filter fsb with value length 5 high selectivity (kept 1023/1024)                     1.05      5.7±0.01µs        ? ?/sec    1.00      5.4±0.01µs        ? ?/sec
filter fsb with value length 5 low selectivity (kept 1/1024)                         1.07   1273.5±8.07ns        ? ?/sec    1.00   1195.4±3.37ns        ? ?/sec
filter fsb with value length 5 w NULLs (kept 1/2)                                    1.45     93.1±1.13µs        ? ?/sec    1.00     64.4±0.09µs        ? ?/sec
filter fsb with value length 5 w NULLs high selectivity (kept 1023/1024)             1.01      8.1±0.01µs        ? ?/sec    1.00      8.0±0.01µs        ? ?/sec
filter fsb with value length 5 w NULLs low selectivity (kept 1/1024)                 1.03  1757.4±10.83ns        ? ?/sec    1.00   1704.7±8.68ns        ? ?/sec
filter fsb with value length 50 (kept 1/2)                                           1.04     98.0±0.26µs        ? ?/sec    1.00     94.1±0.35µs        ? ?/sec
filter fsb with value length 50 high selectivity (kept 1023/1024)                    1.01     71.7±0.61µs        ? ?/sec    1.00     70.9±0.66µs        ? ?/sec
filter fsb with value length 50 low selectivity (kept 1/1024)                        1.07   1346.4±9.05ns        ? ?/sec    1.00   1261.9±3.67ns        ? ?/sec
filter fsb with value length 50 w NULLs (kept 1/2)                                   1.23    125.9±0.36µs        ? ?/sec    1.00    102.4±0.42µs        ? ?/sec
filter fsb with value length 50 w NULLs high selectivity (kept 1023/1024)            1.01     68.6±0.32µs        ? ?/sec    1.00     68.2±0.56µs        ? ?/sec
filter fsb with value length 50 w NULLs low selectivity (kept 1/1024)                1.04   1840.1±6.91ns        ? ?/sec    1.00   1774.9±2.89ns        ? ?/sec
filter i32 (kept 1/2)                                                                1.01     27.6±0.23µs        ? ?/sec    1.00     27.3±0.03µs        ? ?/sec
filter i32 high selectivity (kept 1023/1024)                                         1.00      4.3±0.01µs        ? ?/sec    1.00      4.3±0.01µs        ? ?/sec
filter i32 low selectivity (kept 1/1024)                                             1.00   1154.3±3.30ns        ? ?/sec    1.05   1212.2±1.63ns        ? ?/sec
filter optimize (kept 1/2)                                                           1.00     27.8±0.03µs        ? ?/sec    1.00     27.9±0.04µs        ? ?/sec
filter optimize high selectivity (kept 1023/1024)                                    1.00   1487.5±0.77ns        ? ?/sec    1.01   1499.5±0.79ns        ? ?/sec
filter optimize low selectivity (kept 1/1024)                                        1.00    801.9±0.36ns        ? ?/sec    1.00    802.6±0.59ns        ? ?/sec
filter run array (kept 1/2)                                                          1.01    285.6±0.95µs        ? ?/sec    1.00    283.3±1.78µs        ? ?/sec
filter run array high selectivity (kept 1023/1024)                                   1.01    289.3±1.11µs        ? ?/sec    1.00    285.6±3.46µs        ? ?/sec
filter run array low selectivity (kept 1/1024)                                       1.00    235.1±0.96µs        ? ?/sec    1.00    235.3±0.89µs        ? ?/sec
filter single record batch                                                           1.00     29.0±0.11µs        ? ?/sec    1.01     29.2±0.03µs        ? ?/sec
filter u8 (kept 1/2)                                                                 1.00     27.4±0.02µs        ? ?/sec    1.01     27.7±0.03µs        ? ?/sec
filter u8 high selectivity (kept 1023/1024)                                          1.00   1813.8±6.78ns        ? ?/sec    1.00   1812.3±7.50ns        ? ?/sec
filter u8 low selectivity (kept 1/1024)                                              1.00  1160.6±14.27ns        ? ?/sec    1.00  1159.1±10.20ns        ? ?/sec

Per-runner information
filter_kernels — shard 1/1

Node: gk3-benchmark-cluster-nap-138ltj7s-6d311157-rlzv

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

uname:

Linux bench-c6058397541-3166-vbsqc 6.12.94+ #1 SMP Fri Aug 21 08:00:16 UTC 2026 aarch64 GNU/Linux

BENCH_COMMAND:

cargo bench --features=arrow,async,test_common,experimental,object_store --bench filter_kernels
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Resource Usage

base (merge-base)

Metric Value
Wall time 895.2s
Peak memory 35.3 MiB
Avg memory 22.7 MiB
CPU user 890.8s
CPU sys 0.1s
Peak spill 0 B

branch

Metric Value
Wall time 890.2s
Peak memory 44.1 MiB
Avg memory 23.8 MiB
CPU user 886.0s
CPU sys 0.1s
Peak spill 0 B

File an issue against this benchmark runner

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

arrow Changes to the arrow crate arrow-buffer performance

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Improve performance of filter kernel (when pext instruction is not available)

3 participants