Repository navigation
Shuffle/swizzle operations #29
Description
Activity
Thanks for opening this issue!
We briefly spoke about this at the rendering office hours last week (#office hours > Renderer 2025-07-09 @ 💬).At a high level, this is definitely something that would be a very welcome addition 😄 !
As an additional reference, portable SIMD has a carefully designed swizzling API. The open question is what API enables LLVM to recognize the shuffle/swizzle so it can potentially optimize it.Would you be interested in drafting a proof-of-concept?
I can look into drafting one in the next few days.
I'm not sure how much special stuff needs to be done for LLVM to optimize shuffles. In my testing, it can optimize
_mm_permutevar_psinto a constant shuffle. Not sure about_mm_shuffle_epi8on non-AVX2 targets, or WASMi8x16_swizzle->i8x16_shuffle. As far as I know, NEON doesn't even have a constant swizzle instruction.There are currently two issues here:
-
As mentioned above, we are limited to dynamic shuffle intrinsics because different architectures use different types for the indices, and we can't convert between them without
generic_const_exprs. This shouldn't be a problem performance-wise since LLVM is good at optimizing them into their constant forms, but if there happens to be some shuffle operation that has no dynamic equivalent, we can't support it. I don't think there are any such operations. -
Most of AVX2's shuffle operations operate within 128-bit lanes--you can't shuffle an element from the lower half of a vector to the upper half, or vice versa. It has "permute" operations that operate at 32-bit granularity and above, but I don't think there's any way to implement a generalized 256-bit shuffle in AVX2 that operates on 16-bit or 8-bit elements.
I think there are two different operations we could expose here: a "block-wise" permute that operates within 128-bit blocks, and a full-width permute that operates across the entire vector but only at 32-bit granularity and above.
I think a block-wise permute is pretty straightforward, but the full-width permute seems tricky if the native vector width isn't big enough.
-
I'm still thinking about how to best expose shuffles. There's some kinda annoying stuff here:
-
LLVM has a
shufflevectorintrinsic, and Rust has asimd_shuffleintrinsic that maps to it. However,simd_shuffleis forever-unstable, and we can only access it through idiosyncratic architecture-specific intrinsics.- Those architecture-specific intrinsics cannot be abstracted over, because they all construct their indices in a different way. For example, x86 takes a single i32 shuffle mask that we'd have to construct using bitwise operations. That cannot be done without
generic_const_exprs.
- Those architecture-specific intrinsics cannot be abstracted over, because they all construct their indices in a different way. For example, x86 takes a single i32 shuffle mask that we'd have to construct using bitwise operations. That cannot be done without
-
Falling back to array indexing operations sometimes gets optimized to a single
shufflevectorintrinsic. Sometimes, it doesn't..- Interestingly, the version that uses
movddup+insertpsis faster, which would point to the autovectorizer doing a good job. But if that's the case, then why isn't the_mm_shuffle_psversion compiled down to that as well?
- Interestingly, the version that uses
-
Dynamic permute intrinsics on x86 are optimized to
shufflevectoroperations. This is not yet the case in WebAssembly, and seemingly not on AArch64, either. The fact that this optimization is architecture-specific is not a good sign.- See also this hack to prevent LLVM from generating a useless
memseton AArch64. I believe the popular perception of "LLVM's autovectorizer/optimizer is magic" comes from everyone using and compiling for x86_64, which has a ton of backend-specific optimization passes in LLVM. Other architectures, even major ones like AArch64, are not so lucky.
- See also this hack to prevent LLVM from generating a useless
-
All of the above limitations lead to real missed optimizations and slower code. End-to-end, worse code is generated as a direct result of us being unable to access a generic vector shuffle operation.
Here is a simple example. There are 3 different versions of a function that multiplies four floats by 2.5, then shuffles the lanes according to the indices
[1, 2, 1, 3].- The first version,
shuffle, is completely scalar and is autovectorized. The resulting LLVM IR involves twoloadoperations, a 2xf32shufflevector, aninsertelement, and finally a 4xf32shufflevector. This lowers to some pretty gnarly code. - The second version,
shuffle2, uses AArch64 NEON intrinsics and the best swizzle operation available in stable Rust,vqtbl1q_u8. Even if we didn't care about abstracting over architecture, this is the best we can do at all, becausecore::arch::aarch64does not expose any constant vector shuffle operation! This is unfortunate, because LLVM cannot replace thevqtbl1q_u8intrinsic with a constantvectorshuffleoperation, and it compiles down to atbl. - The third version,
shuffle3, uses the nightly-onlystd::simdAPI. Their "swizzle" operation does compile down to ashufflevector! Even better, it compiles down to the cleanest code so far, using twozipinstructions. Unfortunatelystd::simd, like most Rust RFCs, is stuck in development hell and won't ship anytime soon.
- The first version,
I think the least-bad course of action right now is to land LLVM optimization passes that convert dynamic shuffles to static
vectorshuffles whenever possible, as early in the optimization pipeline as possible, in lieu of Rust shippingstd::simd. Then we can just implement our shuffles on top of that.This does have the annoying property that we can't implement two-source shuffles, which LLVM's
vectorshuffleand Rust'ssimd_shuffleintrinsic allow. Unfortunately, I think that's basically unimplementable right now :(-
I've implemented the missing "dynamic shuffle ->
shufflevector" optimizations for WebAssembly and AArch64 on LLVM's side in llvm/llvm-project#169110 and llvm/llvm-project#169748; right now they're waiting on review.- added a commit that references this issue
on Feb 2, 2026 - linked a pull request that will close this issueAdd swizzle_dyn_precise, mirroring swizzle_dyn from `std::simd` #276
on Aug 3, 2026 - removed a link to a pull requestAdd swizzle_dyn_precise, mirroring swizzle_dyn from `std::simd` #276
on Aug 3, 2026 Cheap dynamic shuffles within blocks are implemented in #266 (merged)
Arbitrary dynamic shuffles are implemented in #276 (waiting on review)
Constant-index shuffles are indeed tricky. They can be implemented on stable with a proc macro, but it's questionable whether that's worth the complexity and the compile-time hit.
Reacted by Tom- added a commit that references this issue
on Aug 5, 2026 #354 adds
concat_swizzle_dyn(), and I think we can call this done once that's merged.The only missing piece is a dedicated API for swizzles with constant values. But I don't see a good reason to implement it:
std::simdonly has this dedicated API because it cannot properly supportswizzle_dyn, and LLVM will optimize dynamic swizzles with constant indices already.Larger-than-native widths might still present an optimization problem so it may be worth pursuing a proc macro like
swizzle!if that turns out to be true.
I'd like to replace my homegrown SIMD implementation with this crate. My use case requires support for a "shuffle" / "swizzle" / "permute" operation, which this crate currently doesn't appear to have.
One obstacle I encountered when implementing shuffles in my SIMD implementation is that different architectures' intrinsics take different types for shuffle indices (they differ in signedness and size), so we can't use const generics to specify them currently. However, LLVM appears to optimize constant permutes into shuffle instructions.
My thought is to give each SIMD type an associated
Indicestype (name not final), which is like itsMaskorBlocktypes, and is a[usize; N]. A shuffle function would take it as input.