Repository navigation
feat(perf): Improve filter performance with per word bit filtering (and BMI when supported) - #10136
Conversation
|
run benchmark arrow-select |
|
Hi @devanbenz, thanks for the request (#10136 (comment)). Only whitelisted users can trigger benchmarks. Allowed users: Dandandan, Fokko, Jefffrey, Omega359, adriangb, alamb, asubiotto, brunal, buraksenn, cetra3, codephage2020, coderfender, comphead, erenavsarogullari, etseidl, friendlymatthew, gabotechs, geoffreyclaude, grtlr, haohuaijin, jonathanc-n, kevinjqliu, klion26, kosiew, kumarUjjawal, kunalsinghdadhwal, liamzwbao, mbutrovich, mkleen, mzabaluev, neilconway, rluvaton, sdf-jkl, timsaucer, xudong963, zhuqi-lucas. File an issue against this benchmark runner |
|
run benchmark arrow-select |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing db/10098/bmi-null-bool-filter (76c7efc) to 1ba5d48 (merge-base) diff File an issue against this benchmark runner |
|
Benchmark for this request failed. Last 20 lines of output: Click to expandFile an issue against this benchmark runner |
|
run benchmark arrow_select |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing db/10098/bmi-null-bool-filter (76c7efc) to 1ba5d48 (merge-base) diff File an issue against this benchmark runner |
|
Benchmark for this request failed. Last 20 lines of output: Click to expandFile an issue against this benchmark runner |
@alamb my apologies, the bench is called |
|
run benchmark filter_bits |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing db/10098/bmi-null-bool-filter (76c7efc) to 1ba5d48 (merge-base) diff File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: CPU Details (lscpu)New benchmark — branch-only results (no baseline comparison) Details
Resource Usagebranch
File an issue against this benchmark runner |
|
@alamb Could you please run the benchmark again but setting |
This comment was marked as outdated.
This comment was marked as outdated.
This comment was marked as outdated.
This comment was marked as outdated.
|
run benchmark filter_bits env: |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing db/10098/bmi-null-bool-filter (8e326b4) to 9f37683 (merge-base) diff File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: CPU Details (lscpu)New benchmark — branch-only results (no baseline comparison) Details
Resource Usagebranch
File an issue against this benchmark runner |
|
@alamb Taking a look at the host machines. It doesn't look like the processor supports BMI2 intrinsics. So the benches are basically equal. |
|
Now that #9848 Is merged in, I'm going to merge this PR with it and move the |
# Which issue does this PR close? - follow up to review discussion on #10136 # Rationale for this change The "Test Release Mode" job in `arrow.yml` compiles with the default `x86_64` target, which only enables SSE2: ``` $ rustc | grep target_feature target_feature="fxsr" target_feature="sse" target_feature="sse2" ``` So any code gated on `cfg(target_feature = ...)`, such as the AVX / AVX2 paths in `arrow-arith/src/aggregate.rs` and `arrow-array/src/array/union_array.rs`, or the BMI2 `pext` path proposed in #10136, is never compiled or tested on CI. # What changes are included in this PR? Add `-C target-cpu=native` to `RUSTFLAGS` for the `linux-release-test` job, so we test with every instruction set extension of the runner CPU. # Are these changes tested? By CI: this job should now build and run the tests with the target features enabled. Example: https://github.com/apache/arrow-rs/actions/runs/35922205002/job/107388798752?pr=11192 <img width="1160" height="624" alt="Screenshot 2026-09-23 at 5 29 40 PM" src="https://github.com/user-attachments/assets/dac161bf-a918-4e86-88c3-e3584ecc0cf2" /> # Are there any user-facing changes? No --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
| /// bits of the result, in their original order: | ||
| /// | ||
| /// ```text | ||
| /// bit: 7 6 5 4 3 2 1 0 |
There was a problem hiding this comment.
Apparently there is no CI coverage for this yet (the existing release job doesn't use all the fancy architectural features). I made a PR to propose doing this
In my opinion, we should merge this PR and then improve the fallback code as a follow on issue / PR It seems like this PR is already faster even with the somewhat basic scalar fallback. I am sure we can all then geek out trying to improve the performance of the fallback with more crazy bithacks (this is a good thing) I looked over comments, and it looks to me like these are the only remaining outstanding comments about adding some additional asserts
For fun, I will re-run my benchmark run with the latest fixes |
Reading hackers delight as we speak |
…nbenz/arrow-rs into db/10098/bmi-null-bool-filter
I've gone ahead and added the suggested |
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
I merged up from main to resolve a merge conflict. I also re-ran this code on There is quite a bit of noise, but I think overall this looks like a win to me Details |
|
run benchmark filter_kernels |
|
run benchmark filter_kernels env: |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing db/10098/bmi-null-bool-filter (c7118c6) to c77f08d (merge-base) diff Run configurationrun benchmark filter_kernelsBENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench filter_kernels File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing db/10098/bmi-null-bool-filter (c7118c6) to c77f08d (merge-base) diff Run configurationrun benchmark filter_kernels
env:
BENCH_FILTER: "NULL"BENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench filter_kernels File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing db/10098/bmi-null-bool-filter (c7118c6) to c77f08d (merge-base) diff Run configurationrun benchmark filter_kernels
env:
BENCH_FILTER: "NULL"CPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing db/10098/bmi-null-bool-filter (c7118c6) to c77f08d (merge-base) diff Run configurationrun benchmark filter_kernelsCPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
run benchmark filter_kernels |
|
🤖 Arrow criterion benchmark running (GKE) | trigger CPU Details (lscpu)Comparing db/10098/bmi-null-bool-filter (c7118c6) to c77f08d (merge-base) diff Run configurationrun benchmark filter_kernelsBENCH_COMMAND=cargo bench --features=arrow,async,test_common,experimental,object_store --bench filter_kernels File an issue against this benchmark runner |
|
🤖 Arrow criterion benchmark completed (GKE) | trigger Instance: Comparing db/10098/bmi-null-bool-filter (c7118c6) to c77f08d (merge-base) diff Run configurationrun benchmark filter_kernelsCPU Details (lscpu)Details
Resource Usagebase (merge-base)
branch
File an issue against this benchmark runner |
|
I think this one has seen enough now; let's merge it in and improve things as a follow on PR |
|
Nice work @devanbenz @Rich-T-kid and @mbutrovich |
Applies review suggestion from apache#10136 (comment) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…nd `BMI` when supported) (apache#10136) - close apache#11060 This commit adds the ability for bit filtering to be done using the [`_pext_u64`](https://www.intel.com/content/www/us/en/docs/intrinsics-guide/index.html#text=_pext_u64) BMI with a scalar fallback. We cannot run the BMI2 feature on the benchmark host machine, here are the benchmarks from my machine: Specs: ``` Architecture: x86_64 CPU op-mode(s): 32-bit, 64-bit Address sizes: 46 bits physical, 48 bits virtual Byte Order: Little Endian Vendor ID: GenuineIntel Model name: 12th Gen Intel(R) Core(TM) i7-12700K ``` | Case | Build | main | apache#10136 | apache#10136 vs main | |---|---|---:|---:|---:| | filter context i32 w NULLs (kept 1/2) | scalar | 47.54 µs | 28.52 µs | -40.0% | | | bmi2 | 53.28 µs | 10.63 µs | -80.1% | | filter context u8 w NULLs (kept 1/2) | scalar | 39.32 µs | 26.41 µs | -32.9% | | | bmi2 | 37.90 µs | 9.10 µs | -76.0% | | filter context string dictionary w NULLs (kept 1/2) | scalar | 36.39 µs | 28.40 µs | -21.9% | | | bmi2 | 44.88 µs | 10.69 µs | -76.2% | | filter f32 (kept 1/2) | scalar | 61.87 µs | 39.97 µs | -35.4% | | | bmi2 | 61.96 µs | 22.74 µs | -63.3% | | filter context f32 (kept 1/2) | scalar | 44.43 µs | 28.40 µs | -36.1% | | | bmi2 | 36.93 µs | 10.80 µs | -70.8% | | filter context short string view (kept 1/2) | scalar | 51.30 µs | 44.48 µs | -13.3% | | | bmi2 | 50.56 µs | 29.11 µs | -42.4% | | filter context mixed string view (kept 1/2) | scalar | 54.59 µs | 43.77 µs | -19.8% | | | bmi2 | 56.13 µs | 26.35 µs | -53.1% | | boolean 65 536 bits, kept 1/2 (filter context fsb …, 9 rows) | scalar | 28.74–39.63 µs | 19.15–19.39 µs | -51…-33% | | | bmi2 | 32.59–41.92 µs | 1.71–1.77 µs | -96…-95% | With these changes we see the following improvements for filter kernels | Build | Avg improvement | Including boolean row | |---|---:|---:| | scalar | ~28.5% faster | ~30.2% faster | | bmi2 | ~66.0% faster | ~69.7% faster | | overall | ~47.2% faster | ~49.9% faster | --------- Co-authored-by: Andrew Lamb <andrew@nerdnetworks.org> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
# Which issue does this PR close? - Follow on to #10136 # Rationale for this change Applies the review suggestion in #10136 (comment) from @mbutrovich to add a doc example for the new public `BitChunks::chunk` API. # What changes are included in this PR? Adds a doctest to `BitChunks::chunk`. # Are these changes tested? Yes, by the new doctest. # Are there any user-facing changes? Docs only. Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
This commit adds the ability for bit filtering to be done using the
_pext_u64BMI with a scalar fallback.We cannot run the BMI2 feature on the benchmark host machine, here are the benchmarks from my machine:
Specs:
With these changes we see the following improvements for filter kernels