Repository navigation
Speed up A-law and mu-law encoding by counting leading zeros - #3
Conversation
|
That looks like a good improvement. I benchmarked it on different CPUs: MacBook Air M1 (ARM aarch64): MacBook 2017 (Intel Core i5 x86_64): On Intel x86_64, It's not a significant regression, so maybe I'll just merge your changes.. |
|
Thanks for the quick feedback, and for benchmarking on Intel too. I only had the M1 here and did not It is not the algorithm, it is how In In That is four extra instructions in the dependency chain, all serial. On aarch64 none of it costs Fix is to or in the bits below the first segment before counting. It never moves the highest set LLVM then takes the bit index straight from Pushed as a second commit. The reference vectors still pass and the M1 is unchanged, 95.0 µs against Also worth knowing for both versions: the benchmark walks the inputs in order, so the branches in the |
|
Thanks for improving the alaw encoder. The results for Intel Mac compared to the version 0.8.0 look good now: I think it's enough to benchmark x86_64 and aarch64, because they are the most used tier 1 architectures. If someone notices that the performance has significantly regressed on other architectures (Risc-V, PowerPC, ..), they can open issues about them.
That's a good point. I have tried to avoid optimizing the code because the benchmarking results are affected by so many variables, such as CPU architecture, inlining, lto, CPU throttling, etc. |
Replaces the range
matchinencode_alawandencode_ulawwith a segment number computed fromleading_zeros(). Output is unchanged, the reference vector tests pass untouched.leading_zeros()gives it directlyEXPinstruction, which "allows the extraction of themost significant bits without requiring a look-up table". The chord is derived that way for mu-law
on page 20 and for A-law on page 23, so the comments now cite those pages alongside the encoding
tables on pages 13 and 16.
cargo benchon an Apple M1, rustc 1.97:encode_alaw143.1 → 92.5 µs (-34%),encode_ulaw155.8 → 95.0 µs (-39%)
cargo clippyand theinternal-no-panicrelease build pass unchangedmatcharms are the application note's tables transcribed and can bechecked by eye, while the arithmetic cannot. The G.191 reference vectors are what makes that safe.