Vendored mailparse: dependency-free whitespace strip; memchr stays for the byte search - #218
Merged
Merged
Conversation
Upstream declined the memchr version of the two mailparse fixes as an added dependency. This replaces the two vendored functions with a dependency-free implementation in one new module, vendor/mailparse/src/bytescan.rs, and drops memchr from the lockfile. The same code is what the upstream proposal now carries, so if a release includes it the vendor directory can go; PATCH.md has the removal steps back alongside the sync procedure. The scans read a usize at a time in plain std with no unsafe: chunks_exact and from_le_bytes for the loads, the exact haszero / hasless word tests to decide whether a word needs a closer look, and trailing_zeros to go straight to a hit rather than rescan the word. The boundary search tests four words per branch in the no-hit case; the whitespace strip, whose hits come every line, one word, walking the mask's set bits and re-checking each with is_ascii_whitespace so 0x0B and the other control bytes are kept as before. Interleaved A/B against master's memchr version on the M4, both rustc 1.98.0: metadata paths within 0.4%, full parse 0.256 -> 0.240 ms, parse_many 2.13 -> 1.97 ms. The strip is faster because it is one pass where the memchr version made two searches per run. Three earlier variants were measured and rejected on the way: single-word steps cost the metadata path 28%, two-word steps that rescanned on a hit cost the full parse 7%, and splitting the no-hit test into two branches cost the metadata path 18%. 803 tests pass on the built wheel; the vendored crate's suite passes, including differential tests over every alignment and byte value. Whether the new loops are themselves placement-sensitive on x86 is the one question a local A/B cannot answer; one toolchain A/B on the merged master will. Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>
The gate on the four-word version measured the metadata paths 13-14% behind the memchr build on an EPYC 7763, where memchr's AVX2 path compares 32 bytes per instruction. Doubling the words per branch to eight -- one cache line -- halves the branch count again and gives the CPU eight independent loads and tests to overlap. On the M4 this variant is 4-5% ahead of memchr on the metadata paths and 12% ahead on the decoding paths; whether the same holds on x86 is what this push asks the gate. Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>
The gate measured the dependency-free byte search behind memchr on x86 twice: +13-14% on the metadata paths with four words per branch, +9-11% with eight. A word-at-a-time scan tops out below a 32-byte AVX2 compare, and the rule for this copy is no degradation, so find_from_u8 goes back to memchr::memmem. The strip stays dependency-free. It is one pass where the memchr version made two searches per run, and it measured faster on both CPUs tried: -9 to -12% on the decoding paths on the M4, -2.5 to -2.9% on the gate's EPYC 7763. So the vendored copy now carries the half of the upstream proposal that is strictly no slower here, and PATCH.md records what switching the other half would cost if a mailparse release ever includes it. The unused search functions and their tests leave bytescan.rs in this copy; the upstream branch keeps them. Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>
This was referenced Aug 30, 2026
kurok
added a commit
that referenced
this pull request
Sep 17, 2026
…228) After #213/#214/#218 fixed the two byte-at-a-time loops, what was left of a full parse was the base64 decode proper: 53% of it in data_encoding's scalar loop, four table lookups per three bytes, at scalar peak on both CPUs tried. The 0.9.0 changelog said as much in one line. The only lever left on it is wider lanes. decode_base64 now decodes the whitespace-free buffer with base64_simd::STANDARD -- AVX2/SSE4.1 with runtime detection on x86-64, NEON on aarch64 -- and falls back to data_encoding::BASE64_MIME_PERMISSIVE on any rejection. The fallback is what makes this a speed change rather than a change to what this library accepts. STANDARD accepts a strict subset on a whitespace-free buffer: same alphabet, padding required by both, and STANDARD is additionally strict about '=' appearing mid-stream and about non-zero trailing bits -- the two lenient cases this library has always accepted. So the set "SIMD accepts, data-encoding rejects" is empty, anything both accept decodes to the same bytes, and every rejection still carries data_encoding's own DecodeError { position, kind } into MailParseError::Base64DecodeError and out to Python with identical text. Three things hold that up rather than one. simd_and_data_encoding_agree in body.rs sweeps a built corpus: every length to 200 padded and unpadded, every byte at every position of a one- and two-block body, mid-stream padding on and off boundaries, both trailing-bit shapes, every whitespace byte interleaved, and 0x0b/0x00 which are not whitespace and must survive the strip to be rejected. The new base64_agreement fuzz target asserts the same three-way agreement -- acceptance, bytes, and error text -- on arbitrary input through the public entry point; it ran 3,578,785 executions clean in 60 s, and its canary was checked to still fire. Three Python tests pin the lenient cases that now reach the fallback, so a fast path that swallowed them would not look green. On the shape of the call, which is not cosmetic. The out-of-place decode is what ships. The in-place form -- decode_inplace over the stripped buffer, then truncate -- reuses the strip's allocation instead of making a second one and reads strictly better. Built as the extension is built (lto = true, codegen-units = 1) it made the whole parse 9% SLOWER, twice, where this form makes it 51% faster; standalone the two decoders are within 10% of each other on the same payloads. That gap is the code-layout sensitivity of #204, and the function and PATCH.md both say so, because the tidier version is the one a future reader will reach for. Measured on an Apple M4 (10 vCPU), two independent interleaved A/Bs pooled to 8 rounds per side, pure-Python controls within 1.1%: parse_email 0.254 -> 0.168 ms -34% parse_email_tree 0.246 -> 0.159 ms -35% full_read 0.261 -> 0.172 ms -34% parse_many (8x) 1.989 -> 1.318 ms -34% mode="metadata" and untouched lazy parses never call this function and are flat, as they must be. base64-simd is MIT and pulls vsimd and outref, both MIT, all from crates.io; cargo deny's allowlist covers them. Its last release is 2022-12 and it is unsafe-heavy SIMD -- PATCH.md says so plainly rather than burying it, and notes that the fallback bounds a rejection bug to a slow path but bounds nothing about a wrong-bytes bug, which is what the differential test and the fuzz target are for. The PR fuzz budget stays at a minute: three targets at 20 s rather than two at 30 s. deep-fuzz.yml gains the target in its matrix. .gitignore gains the artifacts cargo fuzz leaves behind, since running the new target locally is now something a contributor will do. Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>
kurok
added a commit
that referenced
this pull request
Sep 17, 2026
…228) After #213/#214/#218 fixed the two byte-at-a-time loops, what was left of a full parse was the base64 decode proper: 53% of it in data_encoding's scalar loop, four table lookups per three bytes, at scalar peak on both CPUs tried. The 0.9.0 changelog said as much in one line. The only lever left on it is wider lanes. decode_base64 now decodes the whitespace-free buffer with base64_simd::STANDARD -- AVX2/SSE4.1 with runtime detection on x86-64, NEON on aarch64 -- and falls back to data_encoding::BASE64_MIME_PERMISSIVE on any rejection. The fallback is what makes this a speed change rather than a change to what this library accepts. STANDARD accepts a strict subset on a whitespace-free buffer: same alphabet, padding required by both, and STANDARD is additionally strict about '=' appearing mid-stream and about non-zero trailing bits -- the two lenient cases this library has always accepted. So the set "SIMD accepts, data-encoding rejects" is empty, anything both accept decodes to the same bytes, and every rejection still carries data_encoding's own DecodeError { position, kind } into MailParseError::Base64DecodeError and out to Python with identical text. Three things hold that up rather than one. simd_and_data_encoding_agree in body.rs sweeps a built corpus: every length to 200 padded and unpadded, every byte at every position of a one- and two-block body, mid-stream padding on and off boundaries, both trailing-bit shapes, every whitespace byte interleaved, and 0x0b/0x00 which are not whitespace and must survive the strip to be rejected. The new base64_agreement fuzz target asserts the same three-way agreement -- acceptance, bytes, and error text -- on arbitrary input through the public entry point; it ran 3,578,785 executions clean in 60 s, and its canary was checked to still fire. Three Python tests pin the lenient cases that now reach the fallback, so a fast path that swallowed them would not look green. On the shape of the call, which is not cosmetic. The out-of-place decode is what ships. The in-place form -- decode_inplace over the stripped buffer, then truncate -- reuses the strip's allocation instead of making a second one and reads strictly better. Built as the extension is built (lto = true, codegen-units = 1) it made the whole parse 9% SLOWER, twice, where this form makes it 51% faster; standalone the two decoders are within 10% of each other on the same payloads. That gap is the code-layout sensitivity of #204, and the function and PATCH.md both say so, because the tidier version is the one a future reader will reach for. Measured on an Apple M4 (10 vCPU), two independent interleaved A/Bs pooled to 8 rounds per side, pure-Python controls within 1.1%: parse_email 0.254 -> 0.168 ms -34% parse_email_tree 0.246 -> 0.159 ms -35% full_read 0.261 -> 0.172 ms -34% parse_many (8x) 1.989 -> 1.318 ms -34% mode="metadata" and untouched lazy parses never call this function and are flat, as they must be. base64-simd is MIT and pulls vsimd and outref, both MIT, all from crates.io; cargo deny's allowlist covers them. Its last release is 2022-12 and it is unsafe-heavy SIMD -- PATCH.md says so plainly rather than burying it, and notes that the fallback bounds a rejection bug to a slow path but bounds nothing about a wrong-bytes bug, which is what the differential test and the fuzz target are for. The PR fuzz budget stays at a minute: three targets at 20 s rather than two at 30 s. deep-fuzz.yml gains the target in its matrix. .gitignore gains the artifacts cargo fuzz leaves behind, since running the new target locally is now something a contributor will do. Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>
kurok
added a commit
that referenced
this pull request
Sep 17, 2026
…228) (#244) * Decode base64 bodies with SIMD, keeping data-encoding as the arbiter (#228) After #213/#214/#218 fixed the two byte-at-a-time loops, what was left of a full parse was the base64 decode proper: 53% of it in data_encoding's scalar loop, four table lookups per three bytes, at scalar peak on both CPUs tried. The 0.9.0 changelog said as much in one line. The only lever left on it is wider lanes. decode_base64 now decodes the whitespace-free buffer with base64_simd::STANDARD -- AVX2/SSE4.1 with runtime detection on x86-64, NEON on aarch64 -- and falls back to data_encoding::BASE64_MIME_PERMISSIVE on any rejection. The fallback is what makes this a speed change rather than a change to what this library accepts. STANDARD accepts a strict subset on a whitespace-free buffer: same alphabet, padding required by both, and STANDARD is additionally strict about '=' appearing mid-stream and about non-zero trailing bits -- the two lenient cases this library has always accepted. So the set "SIMD accepts, data-encoding rejects" is empty, anything both accept decodes to the same bytes, and every rejection still carries data_encoding's own DecodeError { position, kind } into MailParseError::Base64DecodeError and out to Python with identical text. Three things hold that up rather than one. simd_and_data_encoding_agree in body.rs sweeps a built corpus: every length to 200 padded and unpadded, every byte at every position of a one- and two-block body, mid-stream padding on and off boundaries, both trailing-bit shapes, every whitespace byte interleaved, and 0x0b/0x00 which are not whitespace and must survive the strip to be rejected. The new base64_agreement fuzz target asserts the same three-way agreement -- acceptance, bytes, and error text -- on arbitrary input through the public entry point; it ran 3,578,785 executions clean in 60 s, and its canary was checked to still fire. Three Python tests pin the lenient cases that now reach the fallback, so a fast path that swallowed them would not look green. On the shape of the call, which is not cosmetic. The out-of-place decode is what ships. The in-place form -- decode_inplace over the stripped buffer, then truncate -- reuses the strip's allocation instead of making a second one and reads strictly better. Built as the extension is built (lto = true, codegen-units = 1) it made the whole parse 9% SLOWER, twice, where this form makes it 51% faster; standalone the two decoders are within 10% of each other on the same payloads. That gap is the code-layout sensitivity of #204, and the function and PATCH.md both say so, because the tidier version is the one a future reader will reach for. Measured on an Apple M4 (10 vCPU), two independent interleaved A/Bs pooled to 8 rounds per side, pure-Python controls within 1.1%: parse_email 0.254 -> 0.168 ms -34% parse_email_tree 0.246 -> 0.159 ms -35% full_read 0.261 -> 0.172 ms -34% parse_many (8x) 1.989 -> 1.318 ms -34% mode="metadata" and untouched lazy parses never call this function and are flat, as they must be. base64-simd is MIT and pulls vsimd and outref, both MIT, all from crates.io; cargo deny's allowlist covers them. Its last release is 2022-12 and it is unsafe-heavy SIMD -- PATCH.md says so plainly rather than burying it, and notes that the fallback bounds a rejection bug to a slow path but bounds nothing about a wrong-bytes bug, which is what the differential test and the fuzz target are for. The PR fuzz budget stays at a minute: three targets at 20 s rather than two at 30 s. deep-fuzz.yml gains the target in its matrix. .gitignore gains the artifacts cargo fuzz leaves behind, since running the new target locally is now something a contributor will do. Signed-off-by: kurok <22548029+kurok@users.noreply.github.com> * Quote the x86 benchmark gate alongside the M4 numbers (#228) The changelog entry carried only the Apple M4 A/B. The gate's EPYC run on PR #244 measured the same change interleaved against the merge base -- parse_email 0.536 -> 0.316 ms (-41%), parse_many 4.382 -> 2.583 ms (-41%), controls within 1.1% -- which is the number a reader on x86 wants and the one the issue asks to be recorded here. It also settles the one thing the local run left open: parse_many_metadata read +2.8% on the M4, above that run's noise floor, and reads +0.9% on x86. Metadata mode never calls decode_base64, so that was the measurement. Signed-off-by: kurok <22548029+kurok@users.noreply.github.com> --------- Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Upstream declined the
memchrversion of the mailparse fix as an added dependency, and on request the upstream proposal (now staktrace/mailparse#143) is a dependency-free version of both loops — word-at-a-time scans in plainstd, nounsafe, in one modulebytescan.rs.This PR brings into the vendored copy the part of that which is strictly no slower here:
decode_base64's whitespace strip →bytescan::strip_ascii_whitespace(dependency-free). A word with no byte below0x21cannot contain whitespace and is skipped whole; for one that might, the mask's set bits say which bytes to check, each re-checked withis_ascii_whitespaceso0x0Band other control bytes are kept exactly as before; runs are copied in one piece. One pass where thememchrversion made two searches per run — faster on both CPUs measured.find_from_u8stays onmemchr::memmem. Three dependency-free variants were pushed through the gate on the way here: four words per branch measured +13–14% on the metadata paths on an EPYC 7763, eight words (a cache line) +9–11%. A word-at-a-time scan tops out below a 32-byte AVX2 compare, and the rule for this copy is no degradation, somemchrkeeps the byte search. The x86 cost of switching is recorded inPATCH.mdfor whenever upstream ships the pure-stdversion.Local interleaved A/B vs master, Apple M4, both rustc 1.98.0, controls within 2.4%:
parse_email(full)parse_email_treeparse_many, 8 × 767 KiBparse_email(mode="metadata")parse_email(mode="lazy"), untouchedThe gate's earlier runs on this PR (the two rejected variants) already showed the strip −2.5 to −2.9% on the EPYC 7763 too. Lockfile unchanged. 803 tests pass; the vendored crate's suite (51 + 25, with the differential test over every alignment and byte value) passes;
fmtclean; patched code clippy-clean.