Search for MIME boundaries with memchr::memmem (#120, #204) - #213
Merged
Merged
Conversation
The codegen cliff #120 and #204 were circling has one cause, and it is a loop in a dependency. Sampling a metadata-mode parse of the 767 KiB fixture put 96.5% of the time in mailparse's find_from_u8, the byte-by-byte scan parse_mail runs for every MIME boundary. Its x86-64 instruction stream is byte-identical under rustc 1.97.1 and 1.98.0: 88 instructions, only label hashes differ. Yet two runners measured the metadata path at +96% under 1.98.0. The compiler had not made the loop slower; it had moved it. The crate hash changes with the rustc version -- and with the package version, which is #204 -- which changes symbol names, link order, and the loop's address. On the runners' Zen CPUs a scalar loop straddling the wrong 64-byte boundary falls out of the micro-op cache and runs at half speed. An Apple M4 measured the same two builds at +/-0.2%, which is why the effect never showed locally. The fix is a memchr::memmem search: vectorised, and laid out independently of this crate. It lives in vendor/mailparse -- upstream 0.16.1 with that one function changed and memchr added -- applied through [patch.crates-io] so the crate name and version stay what every other dependency expects. The same change is upstream as staktrace/mailparse#142; PATCH.md records the removal steps for when a release carries it. Interleaved A/B against master on the M4, controls within 1.2%: parse_email(mode="metadata") 0.365 ms -> 0.030 ms parse_email (full) 1.094 ms -> 0.743 ms parse_email_tree 1.124 ms -> 0.736 ms parse_many, 8 x 767 KiB 9.082 ms -> 6.189 ms Every mode faster, no output change: 803 tests pass and the vendored crate's own suite passes. memchr 2.8.3 is the only new crate in the lockfile. Two things the vendoring needed that a [patch] alone does not give. maturin follows [dependencies] path entries into the sdist but not [patch] ones, so without an explicit include the sdist would have built from the unpatched crates.io copy -- silently, and passing every test. The include is now explicit and the sdist job's install-from-source is what checks it. And the fuzz crate resolves mailparse on its own, so it carries the same patch, or it would be fuzzing a different parser than the one shipped. The lint job runs the vendored crate's tests, with a separate target dir: a target/ inside vendor/mailparse would be swept into the sdist by the include glob, which is also why .gitignore now names it. Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>
Every figure in these passages says "on the CI runner", so they change with the gate that measured this branch: an AMD EPYC 9V45, median of three interleaved rounds. The M4 numbers stay in the changelog as the local confirmation, with the reason they are smaller -- that machine never had the slow layout to lose. The hardware paragraph under the headline table used to cite the old ratios as its example of CPU dependence. It now cites them as the before, since they are no longer what any runner will measure. Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>
This was referenced Aug 27, 2026
kurok
added a commit
that referenced
this pull request
Aug 27, 2026
The pin was a holding action against a measured 15-30% slowdown under 1.98.0. The cause was never the compiler's codegen. Two byte-at-a-time loops in mailparse ran at half speed when the linker placed them across a 64-byte boundary on the runners' CPUs, and a new rustc moved them -- so did a version bump. Their x86-64 instruction streams were byte-identical under both compilers; only their addresses differed. With both loops replaced (vendor/mailparse, #213 and #214) the toolchain A/B on the CPU that showed the worst of it reads 1.98.0 within +/-0.5% of 1.97.1 on every parse path, and the same source under both compilers on an M4 is within 0.9%. So the pin moves. It stays a pin rather than floating to stable. The benchmark gate builds a revision and its base with the same toolchain, so a compiler change is the one regression it cannot see -- both sides move together. A pin makes the compiler a change like any other: proposed, measured with toolchain-ab.yml, reviewed. The toolchain file now describes that procedure instead of quoting a ~26% figure that turned out to be a measurement artifact and a 7.0x floor the gate no longer has. Every workflow that names the version follows, because components have to be installed for the toolchain cargo actually uses; the A/B workflow's baseline default follows the pin. Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>
This was referenced Aug 27, 2026
kurok
added a commit
that referenced
this pull request
Sep 17, 2026
…228) After #213/#214/#218 fixed the two byte-at-a-time loops, what was left of a full parse was the base64 decode proper: 53% of it in data_encoding's scalar loop, four table lookups per three bytes, at scalar peak on both CPUs tried. The 0.9.0 changelog said as much in one line. The only lever left on it is wider lanes. decode_base64 now decodes the whitespace-free buffer with base64_simd::STANDARD -- AVX2/SSE4.1 with runtime detection on x86-64, NEON on aarch64 -- and falls back to data_encoding::BASE64_MIME_PERMISSIVE on any rejection. The fallback is what makes this a speed change rather than a change to what this library accepts. STANDARD accepts a strict subset on a whitespace-free buffer: same alphabet, padding required by both, and STANDARD is additionally strict about '=' appearing mid-stream and about non-zero trailing bits -- the two lenient cases this library has always accepted. So the set "SIMD accepts, data-encoding rejects" is empty, anything both accept decodes to the same bytes, and every rejection still carries data_encoding's own DecodeError { position, kind } into MailParseError::Base64DecodeError and out to Python with identical text. Three things hold that up rather than one. simd_and_data_encoding_agree in body.rs sweeps a built corpus: every length to 200 padded and unpadded, every byte at every position of a one- and two-block body, mid-stream padding on and off boundaries, both trailing-bit shapes, every whitespace byte interleaved, and 0x0b/0x00 which are not whitespace and must survive the strip to be rejected. The new base64_agreement fuzz target asserts the same three-way agreement -- acceptance, bytes, and error text -- on arbitrary input through the public entry point; it ran 3,578,785 executions clean in 60 s, and its canary was checked to still fire. Three Python tests pin the lenient cases that now reach the fallback, so a fast path that swallowed them would not look green. On the shape of the call, which is not cosmetic. The out-of-place decode is what ships. The in-place form -- decode_inplace over the stripped buffer, then truncate -- reuses the strip's allocation instead of making a second one and reads strictly better. Built as the extension is built (lto = true, codegen-units = 1) it made the whole parse 9% SLOWER, twice, where this form makes it 51% faster; standalone the two decoders are within 10% of each other on the same payloads. That gap is the code-layout sensitivity of #204, and the function and PATCH.md both say so, because the tidier version is the one a future reader will reach for. Measured on an Apple M4 (10 vCPU), two independent interleaved A/Bs pooled to 8 rounds per side, pure-Python controls within 1.1%: parse_email 0.254 -> 0.168 ms -34% parse_email_tree 0.246 -> 0.159 ms -35% full_read 0.261 -> 0.172 ms -34% parse_many (8x) 1.989 -> 1.318 ms -34% mode="metadata" and untouched lazy parses never call this function and are flat, as they must be. base64-simd is MIT and pulls vsimd and outref, both MIT, all from crates.io; cargo deny's allowlist covers them. Its last release is 2022-12 and it is unsafe-heavy SIMD -- PATCH.md says so plainly rather than burying it, and notes that the fallback bounds a rejection bug to a slow path but bounds nothing about a wrong-bytes bug, which is what the differential test and the fuzz target are for. The PR fuzz budget stays at a minute: three targets at 20 s rather than two at 30 s. deep-fuzz.yml gains the target in its matrix. .gitignore gains the artifacts cargo fuzz leaves behind, since running the new target locally is now something a contributor will do. Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>
kurok
added a commit
that referenced
this pull request
Sep 17, 2026
…228) After #213/#214/#218 fixed the two byte-at-a-time loops, what was left of a full parse was the base64 decode proper: 53% of it in data_encoding's scalar loop, four table lookups per three bytes, at scalar peak on both CPUs tried. The 0.9.0 changelog said as much in one line. The only lever left on it is wider lanes. decode_base64 now decodes the whitespace-free buffer with base64_simd::STANDARD -- AVX2/SSE4.1 with runtime detection on x86-64, NEON on aarch64 -- and falls back to data_encoding::BASE64_MIME_PERMISSIVE on any rejection. The fallback is what makes this a speed change rather than a change to what this library accepts. STANDARD accepts a strict subset on a whitespace-free buffer: same alphabet, padding required by both, and STANDARD is additionally strict about '=' appearing mid-stream and about non-zero trailing bits -- the two lenient cases this library has always accepted. So the set "SIMD accepts, data-encoding rejects" is empty, anything both accept decodes to the same bytes, and every rejection still carries data_encoding's own DecodeError { position, kind } into MailParseError::Base64DecodeError and out to Python with identical text. Three things hold that up rather than one. simd_and_data_encoding_agree in body.rs sweeps a built corpus: every length to 200 padded and unpadded, every byte at every position of a one- and two-block body, mid-stream padding on and off boundaries, both trailing-bit shapes, every whitespace byte interleaved, and 0x0b/0x00 which are not whitespace and must survive the strip to be rejected. The new base64_agreement fuzz target asserts the same three-way agreement -- acceptance, bytes, and error text -- on arbitrary input through the public entry point; it ran 3,578,785 executions clean in 60 s, and its canary was checked to still fire. Three Python tests pin the lenient cases that now reach the fallback, so a fast path that swallowed them would not look green. On the shape of the call, which is not cosmetic. The out-of-place decode is what ships. The in-place form -- decode_inplace over the stripped buffer, then truncate -- reuses the strip's allocation instead of making a second one and reads strictly better. Built as the extension is built (lto = true, codegen-units = 1) it made the whole parse 9% SLOWER, twice, where this form makes it 51% faster; standalone the two decoders are within 10% of each other on the same payloads. That gap is the code-layout sensitivity of #204, and the function and PATCH.md both say so, because the tidier version is the one a future reader will reach for. Measured on an Apple M4 (10 vCPU), two independent interleaved A/Bs pooled to 8 rounds per side, pure-Python controls within 1.1%: parse_email 0.254 -> 0.168 ms -34% parse_email_tree 0.246 -> 0.159 ms -35% full_read 0.261 -> 0.172 ms -34% parse_many (8x) 1.989 -> 1.318 ms -34% mode="metadata" and untouched lazy parses never call this function and are flat, as they must be. base64-simd is MIT and pulls vsimd and outref, both MIT, all from crates.io; cargo deny's allowlist covers them. Its last release is 2022-12 and it is unsafe-heavy SIMD -- PATCH.md says so plainly rather than burying it, and notes that the fallback bounds a rejection bug to a slow path but bounds nothing about a wrong-bytes bug, which is what the differential test and the fuzz target are for. The PR fuzz budget stays at a minute: three targets at 20 s rather than two at 30 s. deep-fuzz.yml gains the target in its matrix. .gitignore gains the artifacts cargo fuzz leaves behind, since running the new target locally is now something a contributor will do. Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>
kurok
added a commit
that referenced
this pull request
Sep 17, 2026
…228) (#244) * Decode base64 bodies with SIMD, keeping data-encoding as the arbiter (#228) After #213/#214/#218 fixed the two byte-at-a-time loops, what was left of a full parse was the base64 decode proper: 53% of it in data_encoding's scalar loop, four table lookups per three bytes, at scalar peak on both CPUs tried. The 0.9.0 changelog said as much in one line. The only lever left on it is wider lanes. decode_base64 now decodes the whitespace-free buffer with base64_simd::STANDARD -- AVX2/SSE4.1 with runtime detection on x86-64, NEON on aarch64 -- and falls back to data_encoding::BASE64_MIME_PERMISSIVE on any rejection. The fallback is what makes this a speed change rather than a change to what this library accepts. STANDARD accepts a strict subset on a whitespace-free buffer: same alphabet, padding required by both, and STANDARD is additionally strict about '=' appearing mid-stream and about non-zero trailing bits -- the two lenient cases this library has always accepted. So the set "SIMD accepts, data-encoding rejects" is empty, anything both accept decodes to the same bytes, and every rejection still carries data_encoding's own DecodeError { position, kind } into MailParseError::Base64DecodeError and out to Python with identical text. Three things hold that up rather than one. simd_and_data_encoding_agree in body.rs sweeps a built corpus: every length to 200 padded and unpadded, every byte at every position of a one- and two-block body, mid-stream padding on and off boundaries, both trailing-bit shapes, every whitespace byte interleaved, and 0x0b/0x00 which are not whitespace and must survive the strip to be rejected. The new base64_agreement fuzz target asserts the same three-way agreement -- acceptance, bytes, and error text -- on arbitrary input through the public entry point; it ran 3,578,785 executions clean in 60 s, and its canary was checked to still fire. Three Python tests pin the lenient cases that now reach the fallback, so a fast path that swallowed them would not look green. On the shape of the call, which is not cosmetic. The out-of-place decode is what ships. The in-place form -- decode_inplace over the stripped buffer, then truncate -- reuses the strip's allocation instead of making a second one and reads strictly better. Built as the extension is built (lto = true, codegen-units = 1) it made the whole parse 9% SLOWER, twice, where this form makes it 51% faster; standalone the two decoders are within 10% of each other on the same payloads. That gap is the code-layout sensitivity of #204, and the function and PATCH.md both say so, because the tidier version is the one a future reader will reach for. Measured on an Apple M4 (10 vCPU), two independent interleaved A/Bs pooled to 8 rounds per side, pure-Python controls within 1.1%: parse_email 0.254 -> 0.168 ms -34% parse_email_tree 0.246 -> 0.159 ms -35% full_read 0.261 -> 0.172 ms -34% parse_many (8x) 1.989 -> 1.318 ms -34% mode="metadata" and untouched lazy parses never call this function and are flat, as they must be. base64-simd is MIT and pulls vsimd and outref, both MIT, all from crates.io; cargo deny's allowlist covers them. Its last release is 2022-12 and it is unsafe-heavy SIMD -- PATCH.md says so plainly rather than burying it, and notes that the fallback bounds a rejection bug to a slow path but bounds nothing about a wrong-bytes bug, which is what the differential test and the fuzz target are for. The PR fuzz budget stays at a minute: three targets at 20 s rather than two at 30 s. deep-fuzz.yml gains the target in its matrix. .gitignore gains the artifacts cargo fuzz leaves behind, since running the new target locally is now something a contributor will do. Signed-off-by: kurok <22548029+kurok@users.noreply.github.com> * Quote the x86 benchmark gate alongside the M4 numbers (#228) The changelog entry carried only the Apple M4 A/B. The gate's EPYC run on PR #244 measured the same change interleaved against the merge base -- parse_email 0.536 -> 0.316 ms (-41%), parse_many 4.382 -> 2.583 ms (-41%), controls within 1.1% -- which is the number a reader on x86 wants and the one the issue asks to be recorded here. It also settles the one thing the local run left open: parse_many_metadata read +2.8% on the M4, above that run's noise floor, and reads +0.9% on x86. Metadata mode never calls decode_base64, so that was the measurement. Signed-off-by: kurok <22548029+kurok@users.noreply.github.com> --------- Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
One loop, two issues
Sampling a metadata-mode parse of the 767 KiB fixture put 96.5% of the time in one function: mailparse's
find_from_u8, the byte-by-byte scanparse_mailruns for every MIME boundary.Its x86-64 instruction stream is byte-identical under rustc 1.97.1 and 1.98.0 — 88 instructions, only label hashes differ — yet two runners (EPYC 9V74, EPYC 7763) both measured the metadata path at +96% under 1.98.0 (run 1, run 2). The compiler didn't make the loop slower; it moved it. The crate hash changes with the rustc version (#120) and with the package version (#204) → symbol names → link order → the loop's address; on Zen a scalar loop straddling the wrong 64-byte boundary falls out of the µop cache and runs at half speed. An Apple M4 measured the same two builds at ±0.2%, which is why it never reproduced locally.
The fix
memchr::memmem::find— vectorised with runtime CPU dispatch, and laid out independently of this crate. It lives invendor/mailparse: upstream 0.16.1 with that one function changed andmemchradded, applied via[patch.crates-io]so the crate name and version stay what everything else expects. Upstream as staktrace/mailparse#142;vendor/mailparse/PATCH.mdhas the removal steps for when a release carries it.Local interleaved A/B vs master (Apple M4, 3 rounds, controls within 1.2%):
parse_email(mode="metadata")parse_email(mode="lazy"), untouchedparse_many(mode="metadata")parse_email(full)parse_email_treeparse_many, 8 × 767 KiBEvery mode faster; no output change — 803 tests pass, and the vendored crate's own 75 pass.
memchr2.8.3 (MIT OR Unlicense) is the only new crate in the lockfile. The x86 gain should be larger than the M4's, since it also removes the 2x slow state; the gate on this PR is the measurement of that, and the README figures will be updated from it before merge.What the vendoring needed beyond
[patch][patch.crates-io]into the sdist. Without an explicitinclude, the sdist would have built from the unpatched crates.io copy — silently, passing every test. It is explicit now, and thesdist installs from sourcejob is what checks it. (Verified locally: 13 vendor entries in the tarball, install from sdist compilesmemchr.)target/insidevendor/mailparsegets swept into the sdist by the include glob, which is also why.gitignorenames it.Follow-ups (separate PRs)
toolchain-ab.ymlon the merged master. If 1.98.0 no longer regresses, unpin — that closes Investigate the ~26% slowdown under rustc 1.98.0, then unpin the toolchain #120.