Repository navigation
perf(filter): replace bit-at-a-time null bitmap filtering - #11055
Rich-T-kid wants to merge 2 commits into
Conversation
|
run benchmark filter_kernels |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing rich-t-kid/filter-bits-gather-optimization (75281be) to 2078680 (merge-base) diff Run configurationrun benchmark filter_kernels
env:
BENCH_FILTER: "NULL"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench filter_kernels File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing rich-t-kid/filter-bits-gather-optimization (75281be) to 2078680 (merge-base) diff Run configurationrun benchmark filter_kernels
env:
BENCH_FILTER: "NULL"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
run benchmark filter_kernels |
1 similar comment
|
run benchmark filter_kernels |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing rich-t-kid/filter-bits-gather-optimization (200598b) to 2078680 (merge-base) diff Run configurationrun benchmark filter_kernels
env:
BENCH_FILTER: "NULL"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench filter_kernels File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing rich-t-kid/filter-bits-gather-optimization (200598b) to 2078680 (merge-base) diff Run configurationrun benchmark filter_kernels
env:
BENCH_FILTER: "NULL"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench filter_kernels File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing rich-t-kid/filter-bits-gather-optimization (200598b) to 2078680 (merge-base) diff Run configurationrun benchmark filter_kernels
env:
BENCH_FILTER: "NULL"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing rich-t-kid/filter-bits-gather-optimization (200598b) to 2078680 (merge-base) diff Run configurationrun benchmark filter_kernels
env:
BENCH_FILTER: "NULL"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
c275451 to
4727e6a
Compare
|
run benchmark filter_kernels |
1 similar comment
|
run benchmark filter_kernels |
|
@sdf-jkl could you take a look 🚀 |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing rich-t-kid/filter-bits-gather-optimization (4727e6a) to 2078680 (merge-base) diff Run configurationrun benchmark filter_kernels
env:
BENCH_FILTER: "NULL"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench filter_kernels File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing rich-t-kid/filter-bits-gather-optimization (4727e6a) to 2078680 (merge-base) diff Run configurationrun benchmark filter_kernels
env:
BENCH_FILTER: "NULL"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench filter_kernels File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing rich-t-kid/filter-bits-gather-optimization (4727e6a) to 2078680 (merge-base) diff Run configurationrun benchmark filter_kernels
env:
BENCH_FILTER: "NULL"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing rich-t-kid/filter-bits-gather-optimization (4727e6a) to 2078680 (merge-base) diff Run configurationrun benchmark filter_kernels
env:
BENCH_FILTER: "NULL"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
I'll take a look this weekend |
|
|
||
| /// Collects the bits of `val` wherever `mask` is 1, packed into the low bits of the result. | ||
| #[inline(always)] | ||
| fn pext64(val: u64, mut mask: u64) -> u64 { |
There was a problem hiding this comment.
There already is an implementation of PEXT in the codebase --
arrow-rs/parquet/src/util/bit_util.rs
Lines 959 to 984 in 4727e6a
Yours seems to perform better though (on my machine 🤓 ) If anything we can drop drop the parquet one and make it import your implementation.
Rust has a nightly impl-- doc.rust-lang.org/std/primitive.u64.html#method.extract_bits
The issue here - rust-lang/rust#149069 - explains that the implementation is waiting on supporting the new LLVM intrinsics that support automatically using the PEXT/PDEP machine instructions if machine supports them. When the instructions are available the perf is pretty epic.
There was a problem hiding this comment.
yea I think moving this out so it can be re-used is a good idea. Would be nice to see how much of a speed up parquet gets from this as well
There was a problem hiding this comment.
There was a problem hiding this comment.
arrow-buffer/src/util/bit_util.rs is where @devanbenz put it
| } | ||
| }; | ||
|
|
||
| for (filter_word, src_word) in filter_chunks.iter().zip(src_chunks.iter()) { |
There was a problem hiding this comment.
Processing each word is independent from each other so this could be paralellizable.
I killed some time on it today. You can take a look here -- #11086
There was a problem hiding this comment.
a great follow on task perhaps
There was a problem hiding this comment.
I agree, can open up a follow on PR
There was a problem hiding this comment.
Yea sure, I'll keep a look out
There was a problem hiding this comment.
@sdf-jkl mhm I cant seem to get similar results to what you did on your branch. everything was pretty much within noise. maybe you can take a closer look
|
@sdf-jkl your PR seems to be faster than this one, should I close this in favor of that one? |
|
Mine was just pathfinding and only better when SIMD enabled, not on scalar. You can port smth / take inspiration from my PR here. |
alamb
left a comment
There was a problem hiding this comment.
This is very neat -- thank you @Rich-T-kid @sdf-jkl and @devanbenz
It will be pretty amazing to get 75% faster on some filter kernels 🤯
I'll keep an eye on this one
FYI @jhorstmann and @hhhizzz you may be interested too
| } | ||
| }; | ||
|
|
||
| for (filter_word, src_word) in filter_chunks.iter().zip(src_chunks.iter()) { |
There was a problem hiding this comment.
a great follow on task perhaps
|
|
||
| /// Collects the bits of `val` wherever `mask` is 1, packed into the low bits of the result. | ||
| #[inline(always)] | ||
| fn pext64(val: u64, mut mask: u64) -> u64 { |
There was a problem hiding this comment.
arrow-buffer/src/util/bit_util.rs is where @devanbenz put it
alamb
left a comment
There was a problem hiding this comment.
it might also make sense to ensure the benchmarks cover this case as well (clearly some do but maybe not all)
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing rich-t-kid/filter-bits-gather-optimization (b64eae3) to 1c2c390 (merge-base) diff Run configurationrun benchmark filter_kernels
env:
BENCH_FILTER: "NULL"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench filter_kernels File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing rich-t-kid/filter-bits-gather-optimization (b64eae3) to 1c2c390 (merge-base) diff Run configurationrun benchmark filter_kernels
env:
BENCH_FILTER: "NULL"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing rich-t-kid/filter-bits-gather-optimization (b64eae3) to 1c2c390 (merge-base) diff Run configurationrun benchmark filter_kernels
env:
BENCH_FILTER: "NULL"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
Yes hello,
The draft is w/o any LUT table for nibbles; it dispatches on the kept count k, with the bit loop for k ≤ 16 and a direct formula when k ≤ 2 or k ≥ 62. In between, each kept bit moves down by the number of dropped bits below it within its byte, done as shifts by 1, 2 and 4 on all eight bytes at once. Byte i then lands at the sum of the counts of bytes below it, and one multiply by The benchmarks in #11271 might be useful here too: the existing
Would it work to keep Runs: https://github.com/mightsleep/arrow-rs/actions/runs/36493097865 Is this ok? |
|
interesting, would be nice to have the benchmarks from #11271 |
|
run benchmark arrow-select |
|
🤖 Benchmark starting (GKE) | trigger Target: arrow-select Sharding factor: 1 (1 workers). Comparing rich-t-kid/filter-bits-gather-optimization (b64eae3) to 1c2c390 (merge-base) diff Run configurationrun benchmark arrow-select
env:
BENCH_FILTER: filter_bits batches
shards: 1
Results will be posted when all workers finish. File an issue against this benchmark runner |
|
🤖 Benchmark failed or incomplete (GKE) | trigger Target: arrow-select Comparing rich-t-kid/filter-bits-gather-optimization (b64eae3) to 1c2c390 (merge-base) diff Run configurationrun benchmark arrow-select
env:
BENCH_FILTER: filter_bits batches
shards: 1
Errors / missing shardsPer-runner informationarrow-select — shard 1/1Node: gk3-benchmark-cluster-nap-5e8o8x4q-b79dfcaa-hpwl Instance: c4a-highmem-16 (12 vCPU / 65 GiB) uname: BENCH_COMMAND: cargo bench --features=arrow,async,test_common,experimental,object_store --bench arrow-selectCPU Details (lscpu)Resource UsageNo completed measurement resource samples. File an issue against this benchmark runner |
|
@mightsleep do you remember the name of the benchmarks you added? |
|
It's The portable |
|
the bot doesn't take the command from me; could you trigger the same |
|
run benchmark filter_bits |
|
🤖 Benchmark starting (GKE) | trigger Target: filter_bits Sharding factor: 1 (1 workers). Comparing rich-t-kid/filter-bits-gather-optimization (b64eae3) to 1c2c390 (merge-base) diff Run configurationrun benchmark filter_bits
env:
BENCH_FILTER: batches
shards: 1
Results will be posted when all workers finish. File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Target: filter_bits Comparing rich-t-kid/filter-bits-gather-optimization (b64eae3) to 1c2c390 (merge-base) diff Run configurationrun benchmark filter_bits
env:
BENCH_FILTER: batches
shards: 1
DetailsPer-runner informationfilter_bits — shard 1/1Node: gk3-benchmark-cluster-nap-wik5zfr4-ca7d2def-nrtv Instance: c4a-highmem-16 (12 vCPU / 65 GiB) uname: BENCH_COMMAND: cargo bench --features=arrow,async,test_common,experimental,object_store --bench filter_bitsCPU Details (lscpu)Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
your branch is from before the batched benches landed, merge main for fix. |
|
run benchmark filter_bits |
1 similar comment
|
run benchmark filter_bits |
|
🤖 Benchmark starting (GKE) | trigger Target: filter_bits Sharding factor: 1 (1 workers). Comparing rich-t-kid/filter-bits-gather-optimization (946cfbc) to 6b34163 (merge-base) diff Run configurationrun benchmark filter_bits
env:
BENCH_FILTER: batches
shards: 1
Results will be posted when all workers finish. File an issue against this benchmark runner |
|
🤖 Benchmark starting (GKE) | trigger Target: filter_bits Sharding factor: 1 (1 workers). Comparing rich-t-kid/filter-bits-gather-optimization (946cfbc) to 6b34163 (merge-base) diff Run configurationrun benchmark filter_bits
env:
BENCH_FILTER: batches
shards: 1
Results will be posted when all workers finish. File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Target: filter_bits Comparing rich-t-kid/filter-bits-gather-optimization (946cfbc) to 6b34163 (merge-base) diff Run configurationrun benchmark filter_bits
env:
BENCH_FILTER: batches
shards: 1
DetailsPer-runner informationfilter_bits — shard 1/1Node: gk3-benchmark-cluster-nap-1w7yp11l-147db3f1-zq6k Instance: c4a-highmem-16 (12 vCPU / 65 GiB) uname: BENCH_COMMAND: cargo bench --features=arrow,async,test_common,experimental,object_store --bench filter_bitsCPU Details (lscpu)Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Target: filter_bits Comparing rich-t-kid/filter-bits-gather-optimization (946cfbc) to 6b34163 (merge-base) diff Run configurationrun benchmark filter_bits
env:
BENCH_FILTER: batches
shards: 1
DetailsPer-runner informationfilter_bits — shard 1/1Node: gk3-benchmark-cluster-nap-1w7yp11l-7d112169-cc8b Instance: c4a-highmem-16 (12 vCPU / 65 GiB) uname: BENCH_COMMAND: cargo bench --features=arrow,async,test_common,experimental,object_store --bench filter_bitsCPU Details (lscpu)Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
pretty nice results |
|
run benchmark filter_kernels |
|
🤖 Benchmark starting (GKE) | trigger Target: filter_kernels Sharding factor: 1 (1 workers). Comparing rich-t-kid/filter-bits-gather-optimization (946cfbc) to 6b34163 (merge-base) diff Run configurationrun benchmark filter_kernels
shards: 1
Results will be posted when all workers finish. File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Target: filter_kernels Comparing rich-t-kid/filter-bits-gather-optimization (946cfbc) to 6b34163 (merge-base) diff Run configurationrun benchmark filter_kernels
shards: 1
DetailsPer-runner informationfilter_kernels — shard 1/1Node: gk3-benchmark-cluster-nap-k44dd5ur-8831919f-z9gk Instance: c4a-highmem-16 (12 vCPU / 65 GiB) uname: BENCH_COMMAND: cargo bench --features=arrow,async,test_common,experimental,object_store --bench filter_kernelsCPU Details (lscpu)Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
Which issue does this PR close?
filterkernel (whenpextinstruction is not available) #11213.Rationale for this change
When filtering an array that has a null bitmap, Arrow walks the null bitmap
one bit at a time, one branch and one memory access per selected row. For a
65 536-row array at 50% filter density that is ~32 000 sequential bit reads.
On ARM (Neoverse-V2, Apple Silicon) where there is no hardware PEXT
instruction, the software fallback also serialises all 64 bit-positions in a
word into a loop-carried dependency chain, leaving gains on the table
What changes are included in this PR?
gather_bitsreplaces the bit-at-a-time loops infilter_bits_strategy.Instead of one
get_bit_rawcall per selected row, it zips 64-bit chunks ofthe filter and source bitmaps, calls
compress()once per word, and streamsthe packed result into the output. The
Indicesbranch retains the oldindex-lookup path only for filters with < 1/64 bits set, where scanning the
full bitmap costs more than jumping to the few set positions.
Sparse path (
compress_sparse, ≤ 8 bits set): keeps the samebit-by-bit loop but it now terminates in at most 8 iterations. For 1/1024
selectivity, most 64-bit mask words have 0–1 bits set, so this costs 0–1
iterations instead of 0–64.
Dense path (
compress_dense, > 8 bits set): usesNIBBLE_PEXT, a256-byte compile-time lookup table covering all 4-bit mask/value pairs. The
64-bit word is processed as 8 bytes × 2 nibbles = 16 independent table
lookups. Per-byte output offsets are precomputed so all 8 writes into
resultare data-independent (hopfully CPU canissue them in parallel)
Are these changes tested?
yes, existing test
Are there any user-facing changes?
no