Skip to content

fix(sonora): flush denormals during stream processing, like upstream - #38

Open
dignifiedquire wants to merge 7 commits into
sync/post-m145-fixesfrom
fix/issue-34-denormals
Open

dignifiedquire wants to merge 7 commits into
sync/post-m145-fixesfrom
fix/issue-34-denormals

Conversation

@dignifiedquire

@dignifiedquire dignifiedquire commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Summary

This PR fixes the AEC3 slowdown on x86 while the render reference is silent (Refs #34).

Cause: on silent input, several recursive filter states never decay to exact zero. Instead they settle in the subnormal float range. That includes:

  • the QMF all-pass states;
  • the decimator and high-pass biquads;
  • the reverb models;
  • some smoothers.

On x86 every operation on a subnormal takes a slow microcode assist. Upstream WebRTC avoids this with DenormalDisabler at every AudioProcessingImpl stream entry point. The port had no equivalent.

Fix: port that guard as a private scoped guard (crates/sonora/src/denormal_disabler.rs). Each of the four stream entry points creates it first, and every public process_* call and the C API reach those four.

Target Guard
x86/x86_64 with SSE MXCSR FTZ+DAZ (0x8040, upstream's mask)
aarch64 with NEON FPCR.FZ (bit 24)
32-bit ARM, arm64ec, wasm, i586, soft-float, Miri no-op; logs "Denormal disabler unsupported", as upstream

As upstream does, the guard sets the bits only if they aren't already set. It restores the caller's exact control word on drop, including during unwinding, so shared threads (for example tokio workers) are unaffected.

Stacked on #36. Merge after #36; the base is its branch, sync/post-m145-fixes. This PR adds only the two commits below.

Results on real x86_64

The aec3_render benchmark (added in the first commit) measures AEC3 only at 48 kHz mono. The job builds and runs it twice on the same GitHub runner, first with the four guard lines removed ("before") and then as committed ("after"). Each variant builds in its own target dir, and a check confirms that only the "after" binary contains ldmxcsr. Run 36598783119, attempt 2.

CPU Silent frame, before → after Silent/active, before → after Speed-up
Intel Xeon Platinum 8573C 734.6 → 71.3 µs 8.85 → 0.86 10.3×
Intel Xeon Platinum 8370C 737.3 → 92.4 µs 7.09 → 0.89 8.0×
AMD EPYC 7763 (Windows) 92.6 → 82.0 µs 0.97 → 0.86 1.13×
AMD EPYC 9V74 (Windows) 72.6 → 67.3 µs 0.94 → 0.88 1.08×
  • Intel: the before/after ratio reproduces the report, a silent frame costing about 9× an active one, and the guard removes it.
  • AMD: it penalizes subnormals far less.
  • Every runner: active frames were unchanged.

Verification

What Where Result
Workspace tests macOS aarch64 780 passed, 0 failed, 0 ignored
8 guard tests real x86 runners (4× Linux/Windows), Docker amd64 (AVX2+FMA backend; SSE2 backend via QEMU), Docker arm64, Windows ARM64 VM (x64 emulated + ARM64 native), MSRV 1.91.1 all 8/8
Regression test silent_render_leaves_no_subnormal_state same passes. It counts subnormal values in the processor state (no timing). A control run proves 12 named sources still go subnormal without the guard.
Mutations local, Docker, Windows VM Removing the guard from any entry point, dropping the restore, or moving the checkpoint makes the right tests fail
Disassembly x86_64 Linux/MSVC, i686, aarch64 darwin/linux/android; rlib, thin LTO, fat LTO The flag is set before any DSP work and restored on every return and unwind path (MSVC funclets included), with no FP instruction outside the window. The no-op targets emit no control-register access. An independent critic re-checked this.
Rust vs C++ comparisons local 27/27, no regression

The guard is effectively free. Measured locally it cost 36–61 ns per call, and the active-frame benchmarks didn't move.

Caveats (documented in code and in the new "Floating-point environment" section on AudioProcessing)

  • Formally undefined behavior. Rust treats changing the FP environment as UB (RFC 3514), the same as upstream, nih-plug and no_denormals. The guard's # Safety contract names the residual hazards. The disassembly shows no FP operation moved across the boundary.
  • Code that runs inside the window runs in flush-to-zero mode. That includes tracing subscribers, the global allocator and the panic hook. A comparison with a subnormal operand can change its result there. On Linux, a thread spawned inside the window inherits flush-to-zero (verified on arm64 Linux).
  • 32-bit ARM is a no-op. Upstream flushes there.
  • Windows with a static CRT (+crt-static) links a software fmaf that ignores FTZ/DAZ, so subnormals can survive in the biquads. perf: use C++-order arithmetic except on AArch64 and x86 with fma #37 fixes this by removing the fmaf calls; with both applied, the regression test passes under +crt-static.
  • Output differs by at most about 1e-38 on all-zero input. Speech output is bit-identical.

Commits

  • 1fc169e bench(sonora): add silent-render AEC3 benchmark
  • 7428065 fix(sonora): flush denormals during stream processing, like upstream
  • 00d55ec test(sonora): check the denormal guard restores a non-default word exactly
  • 48d497b docs(sonora): note the undefined-behavior status and clang-only x86 upstream
  • 20280ab docs(sonora): say that threads started during a call can keep flush-to-zero

The branch also merges #36's branch twice (e59dc1f, cc961e6), so the CHANGELOG stays conflict-free as #36 changes.

Updates after the review pass

  • New test guard_restores_non_default_word_exactly: it starts from a control word the guard must change (FTZ only on x86; FPCR.DN on aarch64) and checks that the word is restored bit for bit. It catches a Drop that clears the mask (x86) or writes a fixed default word (both architectures). The API-level test now also checks that DAZ doesn't leak. There are 9 guard tests now.
  • Public docs and CHANGELOG now say that changing the floating-point environment is formally undefined behavior in Rust, as upstream's DenormalDisabler, nih-plug and no_denormals also do. They also say that upstream sets the x86 bits only in clang builds, so callers replacing a GCC-built C++ library will see a mode change.
  • The containment sentence is corrected: flush-to-zero is set only on the calling thread and restored when the call returns, but threads started during the call can inherit it.

🤖 Generated with Claude Code

https://claude.ai/code/session_01H29DamSLugosGXJSz1e5Yx

dignifiedquire and others added 2 commits September 29, 2026 22:40
Add an `aec3_render` group with two cases, `48k_mono_active` and
`48k_mono_silent`. Each runs AEC3 only at 48 kHz mono: 1 s of active
render and echo, then 6 s (600 frames) of the measured condition, then
times one render plus one capture call per iteration. The silent case
feeds an exact-zero render reference and low-level capture noise (about
-60 dBFS peak), which is the case reported in issue #34.

The 6 s settle matters. Without flush-to-zero, the fast render states
(QMF, decimators, buffers) turn subnormal 0.1-0.4 s into silence, the
two reverb models about 2 s in, and the low-render detector's average
power about 4 s in. Timing starts about 1.9 s after the latest onset,
so a before-fix baseline measures the full slowdown.

The benchmark lands before the fix so that x86 hardware can record a
baseline and then compare the silent/active ratio after it. It is a
manual tool, not a CI gate.

Refs #34

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H29DamSLugosGXJSz1e5Yx
Root cause: several recursive states never decay to zero on silent input
under IEEE gradual underflow. The two-band QMF all-pass states, the AEC3
render and capture decimator biquads, the capture high-pass filter, the
reverb models and some smoothers settle a few ulps above zero, in the
subnormal range, and stay there. On x86 every operation on a subnormal
operand takes a microcode assist, which is why AEC3 ran 10-50x slower
while the render reference was silent. Upstream WebRTC hides this with
DenormalDisabler at every AudioProcessingImpl stream entry point; the
port had no equivalent.

Approach: port DenormalDisabler as a private scoped guard in
crates/sonora/src/denormal_disabler.rs, and create it first in
process_stream, process_reverse_stream, process_stream_i16 and
process_reverse_stream_i16. The four functions are #[inline(never)] so
caller code cannot be inlined into the flush-to-zero window.
- x86 and x86_64 with SSE (i686 included): MXCSR FTZ and DAZ (0x8040),
  upstream's mask.
- aarch64 with NEON: FPCR.FZ (bit 24).
- 32-bit ARM, arm64ec, wasm32, i586, soft-float aarch64 and Miri:
  no-op. AudioProcessing logs "Denormal disabler unsupported" at
  construction, as upstream does.
As upstream, the guard sets the bits only if they are not all set, and
restores the caller's exact control word only if it changed it, also
during unwinding. A caller that already runs with FTZ keeps it.

What runs under the guard: all work a process call does itself,
including re-initialization after a stream-format change and queued
runtime settings. Upstream does the same, except that its i16
ProcessStream re-initializes before it creates the guard
(audio_processing_impl.cc:1232-1235). Direct calls to initialize() and
apply_config(), and the public DSP building blocks, run unguarded, as
upstream.

Documentation: a "Floating-point environment" section on
AudioProcessing states which targets flush, that the caller's setting
is restored (also on panic), that subnormal input samples read as zero,
that code the call reaches (tracing subscribers, the allocator, the
panic hook) runs in flush-to-zero mode, that a comparison with a
subnormal operand can change its result, and that on Linux a thread
started during the call inherits flush-to-zero for its whole life. The
CHANGELOG records the change under Unreleased.

Tests (in the new module; none are skipped; the target-dependent ones
assert against is_supported(), so no-op targets check the no-op):
- FTZ and DAZ take effect, Inf and NaN are kept, nested guards keep the
  outer state, and the caller's state is restored on drop and on unwind.
- All four public process calls leave the caller's environment
  unchanged, both from a default caller and from one that set FTZ.
- A deterministic regression test runs the issue scenario (speech, then
  silent render with low-level capture noise of about -75 dBFS RMS,
  then all-zero) through the f32 and i16 APIs and requires zero
  subnormal values in the processor state. A control run with the guard
  bypassed (test builds only) must still find subnormals in 12 named
  sources, so the scenario cannot silently stop exercising the bug. It
  uses no timing, so it runs on real x86_64 in CI.

Measured on GitHub's Intel runners (Xeon Platinum 8573C and 8370C), the
silent-render benchmark drops from about 735 us to 71-92 us per call
(8-10x), back to the active-render cost. AMD EPYC runners, which
penalize subnormals far less, gain about 1.1x.

Caveat: changing the FP environment is formally undefined behavior in
Rust (RFC 3514), as it is for nih-plug and no_denormals. The guard's
# Safety contract also names the hazard that LLVM may move
register-only float operations across the guard boundary; disassembly
of the release builds shows no such motion today.

Refs #34

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H29DamSLugosGXJSz1e5Yx
@codecov-commenter

codecov-commenter commented Sep 29, 2026 •

Copy link
Copy Markdown

⚠️ Please install the 'codecov app svg image' to ensure uploads and comments are reliably processed by Codecov.

Codecov Report

❌ Patch coverage is 99.25373% with 2 lines in your changes missing coverage. Please review.
✅ Project coverage is 93.84%. Comparing base (72366c3) to head (20280ab).

Additional details and impacted files
@@                   Coverage Diff                    @@
##           sync/post-m145-fixes      #38      +/-   ##
========================================================
+ Coverage                 93.69%   93.84%   +0.14%     
========================================================
  Files                       146      147       +1     
  Lines                     31484    31752     +268     
  Branches                  31484    31752     +268     
========================================================
+ Hits                      29499    29797     +298     
+ Misses                     1762     1731      -31     
- Partials                    223      224       +1     
Flag Coverage Δ
rust 93.84% <99.25%> (+0.14%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

dignifiedquire and others added 5 commits October 1, 2026 13:22
…actly

The guard promises to restore the exact control word it read. Every
restore test started from the default word, where that word, the default
and the word with the flush bits cleared are all equal, so a Drop that
wrote `word & !MASK` or a fixed default word still passed.

guard_restores_non_default_word_exactly starts from a word the guard must
change and that such a restore loses: MXCSR with FTZ set and DAZ clear on
x86/x86_64, FPCR.DN set on aarch64. It fails when Drop writes a fixed
default word (both targets) or `word & !MASK` (x86_64). On aarch64 the
second is the same as a correct restore, because `new` saves only words
with FZ clear.

The API-level test now also checks that DAZ does not leak after each
process call, and the subnormal scan documents that it also counts f64
values, which can only cause a false failure.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H29DamSLugosGXJSz1e5Yx
…pstream

The public "Floating-point environment" section and the CHANGELOG now say
that changing the floating-point environment is formally undefined
behavior in Rust, that upstream's DenormalDisabler, nih-plug and
no_denormals also change it, and how sonora limits the change. Before,
only the private guard docs said so.

The CHANGELOG also notes that upstream sets the x86 bits only in clang
builds (rtc_base/denormal_disabler.cc checks __clang__), so a caller that
replaces a GCC-built C++ library gets FTZ/DAZ on x86 where it had IEEE
gradual underflow.

The guard's safety docs no longer say that only register-only float
operations can cross the guard boundary. Like the asm comments, they now
cover values on the stack and behind & or &mut.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H29DamSLugosGXJSz1e5Yx
…o-zero

48d497b added "sonora limits the change to the calling thread for the
length of each call" to the floating-point environment section on
AudioProcessing (audio_processing.rs:432-433) and to the CHANGELOG. In
the docs this contradicts the sentence just before it: a thread started
during the call, on platforms where new threads inherit the
floating-point environment, keeps flush-to-zero for its whole life
(denormal_disabler.rs:47-50 says the same). The CHANGELOG bullet left
out that caveat, so there the claim read as absolute.

Both now say that sonora sets flush-to-zero only on the calling thread
and restores that thread's setting after each call. The docs point to
the inheritance caveat above; the CHANGELOG states it. The rest of the
doc paragraph is rewrapped, with no other wording changes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H29DamSLugosGXJSz1e5Yx
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants