Skip to content

Vendored mailparse: dependency-free whitespace strip; memchr stays for the byte search - #218

Merged
kurok merged 3 commits into
masterfrom
vendored-mailparse-swar
Aug 28, 2026
Merged

kurok merged 3 commits into
masterfrom
vendored-mailparse-swar

Conversation

@kurok

@kurok kurok commented Aug 28, 2026 •

Copy link
Copy Markdown
Contributor

Upstream declined the memchr version of the mailparse fix as an added dependency, and on request the upstream proposal (now staktrace/mailparse#143) is a dependency-free version of both loops — word-at-a-time scans in plain std, no unsafe, in one module bytescan.rs.

This PR brings into the vendored copy the part of that which is strictly no slower here:

  • decode_base64's whitespace strip → bytescan::strip_ascii_whitespace (dependency-free). A word with no byte below 0x21 cannot contain whitespace and is skipped whole; for one that might, the mask's set bits say which bytes to check, each re-checked with is_ascii_whitespace so 0x0B and other control bytes are kept exactly as before; runs are copied in one piece. One pass where the memchr version made two searches per run — faster on both CPUs measured.
  • find_from_u8 stays on memchr::memmem. Three dependency-free variants were pushed through the gate on the way here: four words per branch measured +13–14% on the metadata paths on an EPYC 7763, eight words (a cache line) +9–11%. A word-at-a-time scan tops out below a 32-byte AVX2 compare, and the rule for this copy is no degradation, so memchr keeps the byte search. The x86 cost of switching is recorded in PATCH.md for whenever upstream ships the pure-std version.

Local interleaved A/B vs master, Apple M4, both rustc 1.98.0, controls within 2.4%:

benchmark master this PR
parse_email (full) 0.254 ms 0.228 ms −10.4%
parse_email_tree 0.245 ms 0.224 ms −8.9%
parse_many, 8 × 767 KiB 2.085 ms 1.834 ms −12.0%
parse_email(mode="metadata") 0.030 ms 0.030 ms +0.8%
parse_email(mode="lazy"), untouched 0.044 ms 0.044 ms −0.7%

The gate's earlier runs on this PR (the two rejected variants) already showed the strip −2.5 to −2.9% on the EPYC 7763 too. Lockfile unchanged. 803 tests pass; the vendored crate's suite (51 + 25, with the differential test over every alignment and byte value) passes; fmt clean; patched code clippy-clean.

kurok added 3 commits August 28, 2026 08:17
Upstream declined the memchr version of the two mailparse fixes as an added
dependency. This replaces the two vendored functions with a dependency-free
implementation in one new module, vendor/mailparse/src/bytescan.rs, and drops
memchr from the lockfile. The same code is what the upstream proposal now
carries, so if a release includes it the vendor directory can go; PATCH.md has
the removal steps back alongside the sync procedure.

The scans read a usize at a time in plain std with no unsafe: chunks_exact and
from_le_bytes for the loads, the exact haszero / hasless word tests to decide
whether a word needs a closer look, and trailing_zeros to go straight to a hit
rather than rescan the word. The boundary search tests four words per branch
in the no-hit case; the whitespace strip, whose hits come every line, one
word, walking the mask's set bits and re-checking each with
is_ascii_whitespace so 0x0B and the other control bytes are kept as before.

Interleaved A/B against master's memchr version on the M4, both rustc 1.98.0:
metadata paths within 0.4%, full parse 0.256 -> 0.240 ms, parse_many
2.13 -> 1.97 ms. The strip is faster because it is one pass where the memchr
version made two searches per run. Three earlier variants were measured and
rejected on the way: single-word steps cost the metadata path 28%, two-word
steps that rescanned on a hit cost the full parse 7%, and splitting the no-hit
test into two branches cost the metadata path 18%.

803 tests pass on the built wheel; the vendored crate's suite passes,
including differential tests over every alignment and byte value. Whether
the new loops are themselves placement-sensitive on x86 is the one question
a local A/B cannot answer; one toolchain A/B on the merged master will.

Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>
The gate on the four-word version measured the metadata paths 13-14% behind
the memchr build on an EPYC 7763, where memchr's AVX2 path compares 32 bytes
per instruction. Doubling the words per branch to eight -- one cache line --
halves the branch count again and gives the CPU eight independent loads and
tests to overlap. On the M4 this variant is 4-5% ahead of memchr on the
metadata paths and 12% ahead on the decoding paths; whether the same holds on
x86 is what this push asks the gate.

Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>
The gate measured the dependency-free byte search behind memchr on x86 twice:
+13-14% on the metadata paths with four words per branch, +9-11% with eight.
A word-at-a-time scan tops out below a 32-byte AVX2 compare, and the rule for
this copy is no degradation, so find_from_u8 goes back to memchr::memmem.

The strip stays dependency-free. It is one pass where the memchr version made
two searches per run, and it measured faster on both CPUs tried: -9 to -12% on
the decoding paths on the M4, -2.5 to -2.9% on the gate's EPYC 7763. So the
vendored copy now carries the half of the upstream proposal that is strictly
no slower here, and PATCH.md records what switching the other half would cost
if a mailparse release ever includes it.

The unused search functions and their tests leave bytescan.rs in this copy;
the upstream branch keeps them.

Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>
@kurok kurok changed the title Vendored mailparse: word-at-a-time scans in plain std, memchr dropped Vendored mailparse: dependency-free whitespace strip; memchr stays for the byte search Aug 28, 2026
@kurok
kurok merged commit 454bd07 into master Aug 28, 2026
15 checks passed
@kurok
kurok deleted the vendored-mailparse-swar branch August 28, 2026 07:38
kurok added a commit that referenced this pull request Sep 17, 2026
…228)

After #213/#214/#218 fixed the two byte-at-a-time loops, what was left of a
full parse was the base64 decode proper: 53% of it in data_encoding's scalar
loop, four table lookups per three bytes, at scalar peak on both CPUs tried.
The 0.9.0 changelog said as much in one line.

The only lever left on it is wider lanes. decode_base64 now decodes the
whitespace-free buffer with base64_simd::STANDARD -- AVX2/SSE4.1 with runtime
detection on x86-64, NEON on aarch64 -- and falls back to
data_encoding::BASE64_MIME_PERMISSIVE on any rejection.

The fallback is what makes this a speed change rather than a change to what
this library accepts. STANDARD accepts a strict subset on a whitespace-free
buffer: same alphabet, padding required by both, and STANDARD is additionally
strict about '=' appearing mid-stream and about non-zero trailing bits -- the
two lenient cases this library has always accepted. So the set "SIMD accepts,
data-encoding rejects" is empty, anything both accept decodes to the same
bytes, and every rejection still carries data_encoding's own
DecodeError { position, kind } into MailParseError::Base64DecodeError and out
to Python with identical text.

Three things hold that up rather than one. simd_and_data_encoding_agree in
body.rs sweeps a built corpus: every length to 200 padded and unpadded, every
byte at every position of a one- and two-block body, mid-stream padding on and
off boundaries, both trailing-bit shapes, every whitespace byte interleaved,
and 0x0b/0x00 which are not whitespace and must survive the strip to be
rejected. The new base64_agreement fuzz target asserts the same three-way
agreement -- acceptance, bytes, and error text -- on arbitrary input through
the public entry point; it ran 3,578,785 executions clean in 60 s, and its
canary was checked to still fire. Three Python tests pin the lenient cases
that now reach the fallback, so a fast path that swallowed them would not look
green.

On the shape of the call, which is not cosmetic. The out-of-place decode is
what ships. The in-place form -- decode_inplace over the stripped buffer, then
truncate -- reuses the strip's allocation instead of making a second one and
reads strictly better. Built as the extension is built (lto = true,
codegen-units = 1) it made the whole parse 9% SLOWER, twice, where this form
makes it 51% faster; standalone the two decoders are within 10% of each other
on the same payloads. That gap is the code-layout sensitivity of #204, and the
function and PATCH.md both say so, because the tidier version is the one a
future reader will reach for.

Measured on an Apple M4 (10 vCPU), two independent interleaved A/Bs pooled to
8 rounds per side, pure-Python controls within 1.1%:

  parse_email        0.254 -> 0.168 ms   -34%
  parse_email_tree   0.246 -> 0.159 ms   -35%
  full_read          0.261 -> 0.172 ms   -34%
  parse_many (8x)    1.989 -> 1.318 ms   -34%

mode="metadata" and untouched lazy parses never call this function and are
flat, as they must be.

base64-simd is MIT and pulls vsimd and outref, both MIT, all from crates.io;
cargo deny's allowlist covers them. Its last release is 2022-12 and it is
unsafe-heavy SIMD -- PATCH.md says so plainly rather than burying it, and
notes that the fallback bounds a rejection bug to a slow path but bounds
nothing about a wrong-bytes bug, which is what the differential test and the
fuzz target are for.

The PR fuzz budget stays at a minute: three targets at 20 s rather than two at
30 s. deep-fuzz.yml gains the target in its matrix. .gitignore gains the
artifacts cargo fuzz leaves behind, since running the new target locally is
now something a contributor will do.

Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>
kurok added a commit that referenced this pull request Sep 17, 2026
…228)

After #213/#214/#218 fixed the two byte-at-a-time loops, what was left of a
full parse was the base64 decode proper: 53% of it in data_encoding's scalar
loop, four table lookups per three bytes, at scalar peak on both CPUs tried.
The 0.9.0 changelog said as much in one line.

The only lever left on it is wider lanes. decode_base64 now decodes the
whitespace-free buffer with base64_simd::STANDARD -- AVX2/SSE4.1 with runtime
detection on x86-64, NEON on aarch64 -- and falls back to
data_encoding::BASE64_MIME_PERMISSIVE on any rejection.

The fallback is what makes this a speed change rather than a change to what
this library accepts. STANDARD accepts a strict subset on a whitespace-free
buffer: same alphabet, padding required by both, and STANDARD is additionally
strict about '=' appearing mid-stream and about non-zero trailing bits -- the
two lenient cases this library has always accepted. So the set "SIMD accepts,
data-encoding rejects" is empty, anything both accept decodes to the same
bytes, and every rejection still carries data_encoding's own
DecodeError { position, kind } into MailParseError::Base64DecodeError and out
to Python with identical text.

Three things hold that up rather than one. simd_and_data_encoding_agree in
body.rs sweeps a built corpus: every length to 200 padded and unpadded, every
byte at every position of a one- and two-block body, mid-stream padding on and
off boundaries, both trailing-bit shapes, every whitespace byte interleaved,
and 0x0b/0x00 which are not whitespace and must survive the strip to be
rejected. The new base64_agreement fuzz target asserts the same three-way
agreement -- acceptance, bytes, and error text -- on arbitrary input through
the public entry point; it ran 3,578,785 executions clean in 60 s, and its
canary was checked to still fire. Three Python tests pin the lenient cases
that now reach the fallback, so a fast path that swallowed them would not look
green.

On the shape of the call, which is not cosmetic. The out-of-place decode is
what ships. The in-place form -- decode_inplace over the stripped buffer, then
truncate -- reuses the strip's allocation instead of making a second one and
reads strictly better. Built as the extension is built (lto = true,
codegen-units = 1) it made the whole parse 9% SLOWER, twice, where this form
makes it 51% faster; standalone the two decoders are within 10% of each other
on the same payloads. That gap is the code-layout sensitivity of #204, and the
function and PATCH.md both say so, because the tidier version is the one a
future reader will reach for.

Measured on an Apple M4 (10 vCPU), two independent interleaved A/Bs pooled to
8 rounds per side, pure-Python controls within 1.1%:

  parse_email        0.254 -> 0.168 ms   -34%
  parse_email_tree   0.246 -> 0.159 ms   -35%
  full_read          0.261 -> 0.172 ms   -34%
  parse_many (8x)    1.989 -> 1.318 ms   -34%

mode="metadata" and untouched lazy parses never call this function and are
flat, as they must be.

base64-simd is MIT and pulls vsimd and outref, both MIT, all from crates.io;
cargo deny's allowlist covers them. Its last release is 2022-12 and it is
unsafe-heavy SIMD -- PATCH.md says so plainly rather than burying it, and
notes that the fallback bounds a rejection bug to a slow path but bounds
nothing about a wrong-bytes bug, which is what the differential test and the
fuzz target are for.

The PR fuzz budget stays at a minute: three targets at 20 s rather than two at
30 s. deep-fuzz.yml gains the target in its matrix. .gitignore gains the
artifacts cargo fuzz leaves behind, since running the new target locally is
now something a contributor will do.

Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>
kurok added a commit that referenced this pull request Sep 17, 2026
…228) (#244)

* Decode base64 bodies with SIMD, keeping data-encoding as the arbiter (#228)

After #213/#214/#218 fixed the two byte-at-a-time loops, what was left of a
full parse was the base64 decode proper: 53% of it in data_encoding's scalar
loop, four table lookups per three bytes, at scalar peak on both CPUs tried.
The 0.9.0 changelog said as much in one line.

The only lever left on it is wider lanes. decode_base64 now decodes the
whitespace-free buffer with base64_simd::STANDARD -- AVX2/SSE4.1 with runtime
detection on x86-64, NEON on aarch64 -- and falls back to
data_encoding::BASE64_MIME_PERMISSIVE on any rejection.

The fallback is what makes this a speed change rather than a change to what
this library accepts. STANDARD accepts a strict subset on a whitespace-free
buffer: same alphabet, padding required by both, and STANDARD is additionally
strict about '=' appearing mid-stream and about non-zero trailing bits -- the
two lenient cases this library has always accepted. So the set "SIMD accepts,
data-encoding rejects" is empty, anything both accept decodes to the same
bytes, and every rejection still carries data_encoding's own
DecodeError { position, kind } into MailParseError::Base64DecodeError and out
to Python with identical text.

Three things hold that up rather than one. simd_and_data_encoding_agree in
body.rs sweeps a built corpus: every length to 200 padded and unpadded, every
byte at every position of a one- and two-block body, mid-stream padding on and
off boundaries, both trailing-bit shapes, every whitespace byte interleaved,
and 0x0b/0x00 which are not whitespace and must survive the strip to be
rejected. The new base64_agreement fuzz target asserts the same three-way
agreement -- acceptance, bytes, and error text -- on arbitrary input through
the public entry point; it ran 3,578,785 executions clean in 60 s, and its
canary was checked to still fire. Three Python tests pin the lenient cases
that now reach the fallback, so a fast path that swallowed them would not look
green.

On the shape of the call, which is not cosmetic. The out-of-place decode is
what ships. The in-place form -- decode_inplace over the stripped buffer, then
truncate -- reuses the strip's allocation instead of making a second one and
reads strictly better. Built as the extension is built (lto = true,
codegen-units = 1) it made the whole parse 9% SLOWER, twice, where this form
makes it 51% faster; standalone the two decoders are within 10% of each other
on the same payloads. That gap is the code-layout sensitivity of #204, and the
function and PATCH.md both say so, because the tidier version is the one a
future reader will reach for.

Measured on an Apple M4 (10 vCPU), two independent interleaved A/Bs pooled to
8 rounds per side, pure-Python controls within 1.1%:

  parse_email        0.254 -> 0.168 ms   -34%
  parse_email_tree   0.246 -> 0.159 ms   -35%
  full_read          0.261 -> 0.172 ms   -34%
  parse_many (8x)    1.989 -> 1.318 ms   -34%

mode="metadata" and untouched lazy parses never call this function and are
flat, as they must be.

base64-simd is MIT and pulls vsimd and outref, both MIT, all from crates.io;
cargo deny's allowlist covers them. Its last release is 2022-12 and it is
unsafe-heavy SIMD -- PATCH.md says so plainly rather than burying it, and
notes that the fallback bounds a rejection bug to a slow path but bounds
nothing about a wrong-bytes bug, which is what the differential test and the
fuzz target are for.

The PR fuzz budget stays at a minute: three targets at 20 s rather than two at
30 s. deep-fuzz.yml gains the target in its matrix. .gitignore gains the
artifacts cargo fuzz leaves behind, since running the new target locally is
now something a contributor will do.

Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>

* Quote the x86 benchmark gate alongside the M4 numbers (#228)

The changelog entry carried only the Apple M4 A/B. The gate's EPYC run on
PR #244 measured the same change interleaved against the merge base --
parse_email 0.536 -> 0.316 ms (-41%), parse_many 4.382 -> 2.583 ms (-41%),
controls within 1.1% -- which is the number a reader on x86 wants and the
one the issue asks to be recorded here.

It also settles the one thing the local run left open: parse_many_metadata
read +2.8% on the M4, above that run's noise floor, and reads +0.9% on x86.
Metadata mode never calls decode_base64, so that was the measurement.

Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>

---------

Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant