In the wheel variant design, it's often hard to motivate what complexities we need and what we can keep simple. In this issue, I want to collect use cases to reference when designing features and the scope the eventual PEP. Ideally, we can include a shortened version of this list in the motivation of the PEP and a longer one in the appendix.
Please feel free to edit this Post. When you add a use case, please write a comment with what you added so the change gets visibility.
Important Use Cases
These are cases that projects/users have been asking for, cases that are currently causing problems, where wheel tags are insufficient.
torch: GPU For CUDA, we require a minimum driver version, i.e. a greater-than-equal relation with driver version (See https://github.com/Slicer/light-the-torch/blob/33397cbe45d07b51ad8ee76b004571a4c236e37f/light_the_torch/_cb.py#L150-L213). This is different than https://pytorch.org/get-started/locally/, where the CUDA major.minor version is modeled: torch ships with CUDA, so it’s actually independent of the CUDA install. There are however cases where users need to select a specific CUDA version, for example if there are regressions in a later version (need for a manual override). For hosts without a GPU, we should install the CPU wheel. Open Question: Can there be more than one CUDA version in a Python process?
torch: compute capability Each torch wheel is compiled against a range of microarchitectures (sm). The host has a single sm (again assuming you're not mixing GPUs), and only wheels that contain the relevant sm can be installed. By having a list of microarchitectures that the wheel supports and having the microarchitecture of the host, we can ensure only a compatible wheel gets installed (not operator, just matching from a list). We can use this to split the large torch wheel into smaller wheels that contain less or only one microarchitecture.
x86_64 microarchitecture levels There are four microarchitecture levels of x86_64: x86_64_v1 (or x86_64 baseline), x86_64_v2, x86_64_v3 and x86_64_v4. For example, v4 includes support for AVX512 that is not yet widespread but can speed up some hashing functions significantly. Example:

Pillow benchmarks, the bars are baseline/SSE4/AVX2.
OpenMP OpenMP is a shared-memory multiprocessing API. There are multiple implementation such as the GNU OpenMP, Intel OpenMP, LLVM OpenMP and Microsoft OpenMP. Intel OpenMP has better performance on Intel chips than GNU OpenMP [TODO: Benchmarks for Python code]. GNU OpenMP is often used with GCC. OpenMP also can be used through a system component, e.g. /lib/x86_64-linux-gnu/libgomp.so.1 for GNU OpenMP on Ubuntu, can be vendored, or can be in a separate wheel. Currently, it is neither possible to have builds for different OpenMP implementations for the same platform, nor is it possible to take the installed OpenMP implementation into consideration for resolution nor installation.
BLAS Similar to OpenMP, BLAS is a specification for linear algebra operations such as matrix multiplication, which are used in many scientific and machine learning applications, including numpy and scipy. There are several implementations, such as OpenBLAS and Intel MKL. Different from OpenMP, there is no system component, but there is desire to have BLAS in a wheel. The desired functionality is in this case not running detection code, but declaring a dependency on a BLAS wheel depending on which wheel is installed, independent of the platform.
Variants
torch nvidia-cuda packages torch depends on a number of CUDA packages, but only for GPU backend and CUDA major version specific. These are packages such as nvidia-cublas-cu12 and nvidia-cuda-runtime-cu12. To have consistent metadata between wheels, we need a way, such as a marker, to select the nvidia-*-cu12 packages only if a CUDA 12 variant was selected.
Solved
This section largely describes the existing wheel tags and what they solve. Features that are currently in the wheel tag may have been part of variants if variants existed earlier, while we know about their solution and its rollout, so they are good reference examples.
libc On Linux, extension modules can be built against glibc Python or against musl Python. The by far more popular one, glibc, is forward compatible with newer versions and uses the manylinux_x_y tag, i.e. it works like a semver caret (or a >=, considering that glibc is perpetually 2.x).
Python ABI Python extension modules can use libpython symbols in two different ways: requiring a specific Python minor version (cp310-cp310) or requiring a Python major with a minimum minor (cp310-abi). Those are similar to a semver tilde and semver caret respectively.
Universal2 and multiple libc A wheel can be compatible with both x86_64 and aarch64 macs. A statically linked binary in a wheel can support both manylinux and musllinux. The first one can be expressed either by listing both compatible tags (macosx_10_12_x86_64.macosx_11_0_arm64) or by using the special universal2 architecture (macosx_10_12_universal2). For statically linked binaries, we can tag them with both the glibc and musl tags (x86_64.manylinux2010_x86_64.musllinux_1_1_x86_64.whl).
In the wheel variant design, it's often hard to motivate what complexities we need and what we can keep simple. In this issue, I want to collect use cases to reference when designing features and the scope the eventual PEP. Ideally, we can include a shortened version of this list in the motivation of the PEP and a longer one in the appendix.
Please feel free to edit this Post. When you add a use case, please write a comment with what you added so the change gets visibility.
Important Use Cases
These are cases that projects/users have been asking for, cases that are currently causing problems, where wheel tags are insufficient.
torch: GPU For CUDA, we require a minimum driver version, i.e. a greater-than-equal relation with driver version (See https://github.com/Slicer/light-the-torch/blob/33397cbe45d07b51ad8ee76b004571a4c236e37f/light_the_torch/_cb.py#L150-L213). This is different than https://pytorch.org/get-started/locally/, where the CUDA major.minor version is modeled: torch ships with CUDA, so it’s actually independent of the CUDA install. There are however cases where users need to select a specific CUDA version, for example if there are regressions in a later version (need for a manual override). For hosts without a GPU, we should install the CPU wheel. Open Question: Can there be more than one CUDA version in a Python process?
torch: compute capability Each torch wheel is compiled against a range of microarchitectures (
sm). The host has a singlesm(again assuming you're not mixing GPUs), and only wheels that contain the relevantsmcan be installed. By having a list of microarchitectures that the wheel supports and having the microarchitecture of the host, we can ensure only a compatible wheel gets installed (not operator, just matching from a list). We can use this to split the large torch wheel into smaller wheels that contain less or only one microarchitecture.x86_64 microarchitecture levels There are four microarchitecture levels of x86_64: x86_64_v1 (or x86_64 baseline), x86_64_v2, x86_64_v3 and x86_64_v4. For example, v4 includes support for AVX512 that is not yet widespread but can speed up some hashing functions significantly. Example:
Pillow benchmarks, the bars are baseline/SSE4/AVX2.
OpenMP OpenMP is a shared-memory multiprocessing API. There are multiple implementation such as the GNU OpenMP, Intel OpenMP, LLVM OpenMP and Microsoft OpenMP. Intel OpenMP has better performance on Intel chips than GNU OpenMP [TODO: Benchmarks for Python code]. GNU OpenMP is often used with GCC. OpenMP also can be used through a system component, e.g.
/lib/x86_64-linux-gnu/libgomp.so.1for GNU OpenMP on Ubuntu, can be vendored, or can be in a separate wheel. Currently, it is neither possible to have builds for different OpenMP implementations for the same platform, nor is it possible to take the installed OpenMP implementation into consideration for resolution nor installation.BLAS Similar to OpenMP, BLAS is a specification for linear algebra operations such as matrix multiplication, which are used in many scientific and machine learning applications, including numpy and scipy. There are several implementations, such as OpenBLAS and Intel MKL. Different from OpenMP, there is no system component, but there is desire to have BLAS in a wheel. The desired functionality is in this case not running detection code, but declaring a dependency on a BLAS wheel depending on which wheel is installed, independent of the platform.
Variants
torch nvidia-cuda packages torch depends on a number of CUDA packages, but only for GPU backend and CUDA major version specific. These are packages such as
nvidia-cublas-cu12andnvidia-cuda-runtime-cu12. To have consistent metadata between wheels, we need a way, such as a marker, to select thenvidia-*-cu12packages only if a CUDA 12 variant was selected.Solved
This section largely describes the existing wheel tags and what they solve. Features that are currently in the wheel tag may have been part of variants if variants existed earlier, while we know about their solution and its rollout, so they are good reference examples.
libc On Linux, extension modules can be built against glibc Python or against musl Python. The by far more popular one, glibc, is forward compatible with newer versions and uses the
manylinux_x_ytag, i.e. it works like a semver caret (or a >=, considering that glibc is perpetually 2.x).Python ABI Python extension modules can use libpython symbols in two different ways: requiring a specific Python minor version (
cp310-cp310) or requiring a Python major with a minimum minor (cp310-abi). Those are similar to a semver tilde and semver caret respectively.Universal2 and multiple libc A wheel can be compatible with both x86_64 and aarch64 macs. A statically linked binary in a wheel can support both manylinux and musllinux. The first one can be expressed either by listing both compatible tags (
macosx_10_12_x86_64.macosx_11_0_arm64) or by using the special universal2 architecture (macosx_10_12_universal2). For statically linked binaries, we can tag them with both the glibc and musl tags (x86_64.manylinux2010_x86_64.musllinux_1_1_x86_64.whl).