Skip to content

Speed up A-law and mu-law encoding by counting leading zeros - #3

Merged
karip merged 2 commits into
karip:mainfrom
HEnquist:faster-g711-encoders
Aug 18, 2026
Merged

karip merged 2 commits into
karip:mainfrom
HEnquist:faster-g711-encoders

Conversation

@HEnquist

Copy link
Copy Markdown
Contributor

Replaces the range match in encode_alaw and encode_ulaw with a segment number computed from
leading_zeros(). Output is unchanged, the reference vector tests pass untouched.

  • the segment (chord) is the position of the highest set bit, so leading_zeros() gives it directly
  • SPRA163A does the same on page 18, with the EXP instruction, which "allows the extraction of the
    most significant bits without requiring a look-up table". The chord is derived that way for mu-law
    on page 20 and for A-law on page 23, so the comments now cite those pages alongside the encoding
    tables on pages 13 and 16.
  • cargo bench on an Apple M1, rustc 1.97: encode_alaw 143.1 → 92.5 µs (-34%), encode_ulaw
    155.8 → 95.0 µs (-39%)
  • cargo clippy and the internal-no-panic release build pass unchanged
  • decoding is untouched, the 256-entry tables are faster than arithmetic there
  • the downside is that the match arms are the application note's tables transcribed and can be
    checked by eye, while the arithmetic cannot. The G.191 reference vectors are what makes that safe.

@karip

karip commented Aug 17, 2026

Copy link
Copy Markdown
Owner

That looks like a good improvement. I benchmarked it on different CPUs:

MacBook Air M1 (ARM aarch64):

encode_alaw             time:   [91.427 µs 91.555 µs 91.737 µs]
                        change: [−34.764% −34.610% −34.467%] (p = 0.00 < 0.05)
                        Performance has improved.
Found 12 outliers among 100 measurements (12.00%)
  2 (2.00%) low mild
  10 (10.00%) high severe

encode_ulaw             time:   [96.779 µs 96.856 µs 96.961 µs]
                        change: [−37.609% −37.503% −37.404%] (p = 0.00 < 0.05)
                        Performance has improved.
Found 11 outliers among 100 measurements (11.00%)
  1 (1.00%) low severe
  2 (2.00%) low mild
  1 (1.00%) high mild
  7 (7.00%) high severe

MacBook 2017 (Intel Core i5 x86_64):

encode_alaw             time:   [226.79 µs 227.49 µs 228.68 µs]
                        change: [+5.6396% +6.3861% +7.2148%] (p = 0.00 < 0.05)
                        Performance has regressed.
Found 11 outliers among 100 measurements (11.00%)
  2 (2.00%) high mild
  9 (9.00%) high severe

encode_ulaw             time:   [235.45 µs 236.46 µs 238.12 µs]
                        change: [−29.665% −29.053% −28.467%] (p = 0.00 < 0.05)
                        Performance has improved.
Found 10 outliers among 100 measurements (10.00%)
  1 (1.00%) high mild
  9 (9.00%) high severe

On Intel x86_64, encode_alaw is now slightly slower. Looking at the code, I didn't find any reason why ulaw performs better but alaw doesn't. Do you have any ideas why this is happening?

It's not a significant regression, so maybe I'll just merge your changes..

@HEnquist

Copy link
Copy Markdown
Contributor Author

Thanks for the quick feedback, and for benchmarking on Intel too. I only had the M1 here and did not
think to check another architecture.

It is not the algorithm, it is how leading_zeros() lowers on x86-64 without LZCNT, which is what
the default target gives you.

In encode_ulaw the counted value is absval + 33, and absval comes from a u16, so LLVM knows it
is nonzero and at least 6 bits long. It can emit a bare bsr, and it can also drop the
saturating_sub(6) because it cannot underflow:

bsrl  %edx, %esi
xorl  $31, %esi
movb  $27, %cl
subb  %sil, %cl

In encode_alaw the value can be 0 (any input in -8..=7), so LLVM has to emit the zero-safe form,
and it cannot prove the saturating_sub(5) is safe either:

movl   $63, %ecx
bsrl   %eax, %ecx
xorl   $31, %ecx
xorl   %edx, %edx
movl   $27, %esi
subl   %ecx, %esi
cmovbl %edx, %esi

That is four extra instructions in the dependency chain, all serial. On aarch64 none of it costs
anything, since clz is a single instruction and is defined for 0, which is why the M1 sees no
difference.

Fix is to or in the bits below the first segment before counting. It never moves the highest set
bit, so the segment is unchanged, but the value becomes nonzero and at least 5 bits long:

let segment = (u32::BITS - (inputval | 0b11111).leading_zeros()) - 5;

LLVM then takes the bit index straight from bsr:

orl   $31, %ecx
bsrl  %ecx, %edx
addl  $-4, %edx

Pushed as a second commit. The reference vectors still pass and the M1 is unchanged, 95.0 µs against
143.1 µs on main. I have no Intel machine here, so could you re-run the 2017 MacBook?

Also worth knowing for both versions: the benchmark walks the inputs in order, so the branches in the
old match are nearly always predicted right. Input that jumps between segments would look worse
for it.

@karip

karip commented Aug 18, 2026

Copy link
Copy Markdown
Owner

Thanks for improving the alaw encoder. The results for Intel Mac compared to the version 0.8.0 look good now:

encode_alaw             time:   [209.86 µs 210.72 µs 211.90 µs]
                        change: [−4.1381% −3.2145% −2.3579%] (p = 0.00 < 0.05)
                        Performance has improved.
Found 16 outliers among 100 measurements (16.00%)
  2 (2.00%) low mild
  4 (4.00%) high mild
  10 (10.00%) high severe

encode_ulaw             time:   [234.74 µs 237.15 µs 240.11 µs]
                        change: [−31.770% −31.017% −30.261%] (p = 0.00 < 0.05)
                        Performance has improved.
Found 11 outliers among 100 measurements (11.00%)
  5 (5.00%) high mild
  6 (6.00%) high severe

I think it's enough to benchmark x86_64 and aarch64, because they are the most used tier 1 architectures. If someone notices that the performance has significantly regressed on other architectures (Risc-V, PowerPC, ..), they can open issues about them.

Also worth knowing for both versions: the benchmark walks the inputs in order, so the branches in the
old match are nearly always predicted right. Input that jumps between segments would look worse
for it.

That's a good point. I have tried to avoid optimizing the code because the benchmarking results are affected by so many variables, such as CPU architecture, inlining, lto, CPU throttling, etc.

@karip
karip merged commit 4721e3e into karip:main Aug 18, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants