Skip to content

Shuffle/swizzle operations #29

Description

@valadaptive

I'd like to replace my homegrown SIMD implementation with this crate. My use case requires support for a "shuffle" / "swizzle" / "permute" operation, which this crate currently doesn't appear to have.

One obstacle I encountered when implementing shuffles in my SIMD implementation is that different architectures' intrinsics take different types for shuffle indices (they differ in signedness and size), so we can't use const generics to specify them currently. However, LLVM appears to optimize constant permutes into shuffle instructions.

My thought is to give each SIMD type an associated Indices type (name not final), which is like its Mask or Block types, and is a [usize; N]. A shuffle function would take it as input.

Activity

  1. AndrewJakubowicz commented on Jul 12, 2025

    @AndrewJakubowicz

    Thanks for opening this issue!
    We briefly spoke about this at the rendering office hours last week (#office hours > Renderer 2025-07-09 @ 💬).

    At a high level, this is definitely something that would be a very welcome addition 😄 !
    As an additional reference, portable SIMD has a carefully designed swizzling API. The open question is what API enables LLVM to recognize the shuffle/swizzle so it can potentially optimize it.

    Would you be interested in drafting a proof-of-concept?

  2. valadaptive commented on Jul 12, 2025

    @valadaptive
    ContributorAuthor

    I can look into drafting one in the next few days.

    I'm not sure how much special stuff needs to be done for LLVM to optimize shuffles. In my testing, it can optimize _mm_permutevar_ps into a constant shuffle. Not sure about _mm_shuffle_epi8 on non-AVX2 targets, or WASM i8x16_swizzle -> i8x16_shuffle. As far as I know, NEON doesn't even have a constant swizzle instruction.

  3. valadaptive commented on Nov 12, 2025

    @valadaptive
    ContributorAuthor

    There are currently two issues here:

    • As mentioned above, we are limited to dynamic shuffle intrinsics because different architectures use different types for the indices, and we can't convert between them without generic_const_exprs. This shouldn't be a problem performance-wise since LLVM is good at optimizing them into their constant forms, but if there happens to be some shuffle operation that has no dynamic equivalent, we can't support it. I don't think there are any such operations.

    • Most of AVX2's shuffle operations operate within 128-bit lanes--you can't shuffle an element from the lower half of a vector to the upper half, or vice versa. It has "permute" operations that operate at 32-bit granularity and above, but I don't think there's any way to implement a generalized 256-bit shuffle in AVX2 that operates on 16-bit or 8-bit elements.

    I think there are two different operations we could expose here: a "block-wise" permute that operates within 128-bit blocks, and a full-width permute that operates across the entire vector but only at 32-bit granularity and above.

    I think a block-wise permute is pretty straightforward, but the full-width permute seems tricky if the native vector width isn't big enough.

  4. valadaptive commented on Nov 25, 2025

    @valadaptive
    ContributorAuthor

    I'm still thinking about how to best expose shuffles. There's some kinda annoying stuff here:

    • LLVM has a shufflevector intrinsic, and Rust has a simd_shuffle intrinsic that maps to it. However, simd_shuffle is forever-unstable, and we can only access it through idiosyncratic architecture-specific intrinsics.

      • Those architecture-specific intrinsics cannot be abstracted over, because they all construct their indices in a different way. For example, x86 takes a single i32 shuffle mask that we'd have to construct using bitwise operations. That cannot be done without generic_const_exprs.
    • Falling back to array indexing operations sometimes gets optimized to a single shufflevector intrinsic. Sometimes, it doesn't..

      • Interestingly, the version that uses movddup+insertps is faster, which would point to the autovectorizer doing a good job. But if that's the case, then why isn't the _mm_shuffle_ps version compiled down to that as well?
    • Dynamic permute intrinsics on x86 are optimized to shufflevector operations. This is not yet the case in WebAssembly, and seemingly not on AArch64, either. The fact that this optimization is architecture-specific is not a good sign.

      • See also this hack to prevent LLVM from generating a useless memset on AArch64. I believe the popular perception of "LLVM's autovectorizer/optimizer is magic" comes from everyone using and compiling for x86_64, which has a ton of backend-specific optimization passes in LLVM. Other architectures, even major ones like AArch64, are not so lucky.
    • All of the above limitations lead to real missed optimizations and slower code. End-to-end, worse code is generated as a direct result of us being unable to access a generic vector shuffle operation.

      Here is a simple example. There are 3 different versions of a function that multiplies four floats by 2.5, then shuffles the lanes according to the indices [1, 2, 1, 3].

      1. The first version, shuffle, is completely scalar and is autovectorized. The resulting LLVM IR involves two load operations, a 2xf32 shufflevector, an insertelement, and finally a 4xf32 shufflevector. This lowers to some pretty gnarly code.
      2. The second version, shuffle2, uses AArch64 NEON intrinsics and the best swizzle operation available in stable Rust, vqtbl1q_u8. Even if we didn't care about abstracting over architecture, this is the best we can do at all, because core::arch::aarch64 does not expose any constant vector shuffle operation! This is unfortunate, because LLVM cannot replace the vqtbl1q_u8 intrinsic with a constant vectorshuffle operation, and it compiles down to a tbl.
      3. The third version, shuffle3, uses the nightly-only std::simd API. Their "swizzle" operation does compile down to a shufflevector! Even better, it compiles down to the cleanest code so far, using two zip instructions. Unfortunately std::simd, like most Rust RFCs, is stuck in development hell and won't ship anytime soon.

    I think the least-bad course of action right now is to land LLVM optimization passes that convert dynamic shuffles to static vectorshuffles whenever possible, as early in the optimization pipeline as possible, in lieu of Rust shipping std::simd. Then we can just implement our shuffles on top of that.

    This does have the annoying property that we can't implement two-source shuffles, which LLVM's vectorshuffle and Rust's simd_shuffle intrinsic allow. Unfortunately, I think that's basically unimplementable right now :(

  5. valadaptive commented on Nov 27, 2025

    @valadaptive
    ContributorAuthor

    I've implemented the missing "dynamic shuffle -> shufflevector" optimizations for WebAssembly and AArch64 on LLVM's side in llvm/llvm-project#169110 and llvm/llvm-project#169748; right now they're waiting on review.

  6. Shnatsel commented on Aug 3, 2026

    @Shnatsel
    Contributor

    Cheap dynamic shuffles within blocks are implemented in #266 (merged)

    Arbitrary dynamic shuffles are implemented in #276 (waiting on review)

    Constant-index shuffles are indeed tricky. They can be implemented on stable with a proc macro, but it's questionable whether that's worth the complexity and the compile-time hit.

  7. added a commit that references this issue on Aug 5, 2026
  8. Shnatsel commented on Aug 29, 2026

    @Shnatsel
    Contributor

    #354 adds concat_swizzle_dyn(), and I think we can call this done once that's merged.

    The only missing piece is a dedicated API for swizzles with constant values. But I don't see a good reason to implement it: std::simd only has this dedicated API because it cannot properly support swizzle_dyn, and LLVM will optimize dynamic swizzles with constant indices already.

  9. Shnatsel commented on Aug 29, 2026

    @Shnatsel
    Contributor

    Larger-than-native widths might still present an optimization problem so it may be worth pursuing a proc macro like swizzle! if that turns out to be true.

  10. Shnatsel commented on Sep 29, 2026

    @Shnatsel
    Contributor

    concat_swizzle_dyn is merged in #354, convenient compile-time swizzles are proposed in #391

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions