Skip to content

Search for MIME boundaries with memchr::memmem (#120, #204) - #213

Merged
kurok merged 2 commits into
masterfrom
boundary-search-memmem
Aug 27, 2026
Merged

kurok merged 2 commits into
masterfrom
boundary-search-memmem

Conversation

@kurok

@kurok kurok commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

One loop, two issues

Sampling a metadata-mode parse of the 767 KiB fixture put 96.5% of the time in one function: mailparse's find_from_u8, the byte-by-byte scan parse_mail runs for every MIME boundary.

Its x86-64 instruction stream is byte-identical under rustc 1.97.1 and 1.98.0 — 88 instructions, only label hashes differ — yet two runners (EPYC 9V74, EPYC 7763) both measured the metadata path at +96% under 1.98.0 (run 1, run 2). The compiler didn't make the loop slower; it moved it. The crate hash changes with the rustc version (#120) and with the package version (#204) → symbol names → link order → the loop's address; on Zen a scalar loop straddling the wrong 64-byte boundary falls out of the µop cache and runs at half speed. An Apple M4 measured the same two builds at ±0.2%, which is why it never reproduced locally.

The fix

memchr::memmem::find — vectorised with runtime CPU dispatch, and laid out independently of this crate. It lives in vendor/mailparse: upstream 0.16.1 with that one function changed and memchr added, applied via [patch.crates-io] so the crate name and version stay what everything else expects. Upstream as staktrace/mailparse#142; vendor/mailparse/PATCH.md has the removal steps for when a release carries it.

Local interleaved A/B vs master (Apple M4, 3 rounds, controls within 1.2%):

benchmark master this PR
parse_email(mode="metadata") 0.365 ms 0.030 ms −92%
parse_email(mode="lazy"), untouched 0.384 ms 0.049 ms −87%
parse_many(mode="metadata") 3.019 ms 0.242 ms −92%
parse_email (full) 1.094 ms 0.743 ms −32%
parse_email_tree 1.124 ms 0.736 ms −35%
parse_many, 8 × 767 KiB 9.082 ms 6.189 ms −32%

Every mode faster; no output change — 803 tests pass, and the vendored crate's own 75 pass. memchr 2.8.3 (MIT OR Unlicense) is the only new crate in the lockfile. The x86 gain should be larger than the M4's, since it also removes the 2x slow state; the gate on this PR is the measurement of that, and the README figures will be updated from it before merge.

What the vendoring needed beyond [patch]

  • maturin does not follow [patch.crates-io] into the sdist. Without an explicit include, the sdist would have built from the unpatched crates.io copy — silently, passing every test. It is explicit now, and the sdist installs from source job is what checks it. (Verified locally: 13 vendor entries in the tarball, install from sdist compiles memchr.)
  • The fuzz crate resolves mailparse on its own, so it carries the same patch — or it would fuzz a different parser than the one shipped.
  • The lint job runs the vendored crate's tests with a separate target dir: a target/ inside vendor/mailparse gets swept into the sdist by the include glob, which is also why .gitignore names it.

Follow-ups (separate PRs)

kurok added 2 commits August 27, 2026 16:42
The codegen cliff #120 and #204 were circling has one cause, and it is a loop
in a dependency. Sampling a metadata-mode parse of the 767 KiB fixture put
96.5% of the time in mailparse's find_from_u8, the byte-by-byte scan
parse_mail runs for every MIME boundary.

Its x86-64 instruction stream is byte-identical under rustc 1.97.1 and 1.98.0:
88 instructions, only label hashes differ. Yet two runners measured the
metadata path at +96% under 1.98.0. The compiler had not made the loop slower;
it had moved it. The crate hash changes with the rustc version -- and with the
package version, which is #204 -- which changes symbol names, link order, and
the loop's address. On the runners' Zen CPUs a scalar loop straddling the
wrong 64-byte boundary falls out of the micro-op cache and runs at half speed.
An Apple M4 measured the same two builds at +/-0.2%, which is why the effect
never showed locally.

The fix is a memchr::memmem search: vectorised, and laid out independently of
this crate. It lives in vendor/mailparse -- upstream 0.16.1 with that one
function changed and memchr added -- applied through [patch.crates-io] so the
crate name and version stay what every other dependency expects. The same
change is upstream as staktrace/mailparse#142; PATCH.md records the removal
steps for when a release carries it.

Interleaved A/B against master on the M4, controls within 1.2%:

  parse_email(mode="metadata")   0.365 ms -> 0.030 ms
  parse_email (full)             1.094 ms -> 0.743 ms
  parse_email_tree               1.124 ms -> 0.736 ms
  parse_many, 8 x 767 KiB        9.082 ms -> 6.189 ms

Every mode faster, no output change: 803 tests pass and the vendored crate's
own suite passes. memchr 2.8.3 is the only new crate in the lockfile.

Two things the vendoring needed that a [patch] alone does not give. maturin
follows [dependencies] path entries into the sdist but not [patch] ones, so
without an explicit include the sdist would have built from the unpatched
crates.io copy -- silently, and passing every test. The include is now
explicit and the sdist job's install-from-source is what checks it. And the
fuzz crate resolves mailparse on its own, so it carries the same patch, or it
would be fuzzing a different parser than the one shipped.

The lint job runs the vendored crate's tests, with a separate target dir: a
target/ inside vendor/mailparse would be swept into the sdist by the include
glob, which is also why .gitignore now names it.

Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>
Every figure in these passages says "on the CI runner", so they change with the
gate that measured this branch: an AMD EPYC 9V45, median of three interleaved
rounds. The M4 numbers stay in the changelog as the local confirmation, with the
reason they are smaller -- that machine never had the slow layout to lose.

The hardware paragraph under the headline table used to cite the old ratios as
its example of CPU dependence. It now cites them as the before, since they are
no longer what any runner will measure.

Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>
@kurok
kurok merged commit 98280bc into master Aug 27, 2026
15 checks passed
@kurok
kurok deleted the boundary-search-memmem branch August 27, 2026 15:52
kurok added a commit that referenced this pull request Aug 27, 2026
The pin was a holding action against a measured 15-30% slowdown under 1.98.0.
The cause was never the compiler's codegen. Two byte-at-a-time loops in
mailparse ran at half speed when the linker placed them across a 64-byte
boundary on the runners' CPUs, and a new rustc moved them -- so did a version
bump. Their x86-64 instruction streams were byte-identical under both
compilers; only their addresses differed.

With both loops replaced (vendor/mailparse, #213 and #214) the toolchain A/B
on the CPU that showed the worst of it reads 1.98.0 within +/-0.5% of 1.97.1
on every parse path, and the same source under both compilers on an M4 is
within 0.9%. So the pin moves.

It stays a pin rather than floating to stable. The benchmark gate builds a
revision and its base with the same toolchain, so a compiler change is the one
regression it cannot see -- both sides move together. A pin makes the compiler
a change like any other: proposed, measured with toolchain-ab.yml, reviewed.
The toolchain file now describes that procedure instead of quoting a ~26%
figure that turned out to be a measurement artifact and a 7.0x floor the gate
no longer has.

Every workflow that names the version follows, because components have to be
installed for the toolchain cargo actually uses; the A/B workflow's baseline
default follows the pin.

Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>
kurok added a commit that referenced this pull request Sep 17, 2026
…228)

After #213/#214/#218 fixed the two byte-at-a-time loops, what was left of a
full parse was the base64 decode proper: 53% of it in data_encoding's scalar
loop, four table lookups per three bytes, at scalar peak on both CPUs tried.
The 0.9.0 changelog said as much in one line.

The only lever left on it is wider lanes. decode_base64 now decodes the
whitespace-free buffer with base64_simd::STANDARD -- AVX2/SSE4.1 with runtime
detection on x86-64, NEON on aarch64 -- and falls back to
data_encoding::BASE64_MIME_PERMISSIVE on any rejection.

The fallback is what makes this a speed change rather than a change to what
this library accepts. STANDARD accepts a strict subset on a whitespace-free
buffer: same alphabet, padding required by both, and STANDARD is additionally
strict about '=' appearing mid-stream and about non-zero trailing bits -- the
two lenient cases this library has always accepted. So the set "SIMD accepts,
data-encoding rejects" is empty, anything both accept decodes to the same
bytes, and every rejection still carries data_encoding's own
DecodeError { position, kind } into MailParseError::Base64DecodeError and out
to Python with identical text.

Three things hold that up rather than one. simd_and_data_encoding_agree in
body.rs sweeps a built corpus: every length to 200 padded and unpadded, every
byte at every position of a one- and two-block body, mid-stream padding on and
off boundaries, both trailing-bit shapes, every whitespace byte interleaved,
and 0x0b/0x00 which are not whitespace and must survive the strip to be
rejected. The new base64_agreement fuzz target asserts the same three-way
agreement -- acceptance, bytes, and error text -- on arbitrary input through
the public entry point; it ran 3,578,785 executions clean in 60 s, and its
canary was checked to still fire. Three Python tests pin the lenient cases
that now reach the fallback, so a fast path that swallowed them would not look
green.

On the shape of the call, which is not cosmetic. The out-of-place decode is
what ships. The in-place form -- decode_inplace over the stripped buffer, then
truncate -- reuses the strip's allocation instead of making a second one and
reads strictly better. Built as the extension is built (lto = true,
codegen-units = 1) it made the whole parse 9% SLOWER, twice, where this form
makes it 51% faster; standalone the two decoders are within 10% of each other
on the same payloads. That gap is the code-layout sensitivity of #204, and the
function and PATCH.md both say so, because the tidier version is the one a
future reader will reach for.

Measured on an Apple M4 (10 vCPU), two independent interleaved A/Bs pooled to
8 rounds per side, pure-Python controls within 1.1%:

  parse_email        0.254 -> 0.168 ms   -34%
  parse_email_tree   0.246 -> 0.159 ms   -35%
  full_read          0.261 -> 0.172 ms   -34%
  parse_many (8x)    1.989 -> 1.318 ms   -34%

mode="metadata" and untouched lazy parses never call this function and are
flat, as they must be.

base64-simd is MIT and pulls vsimd and outref, both MIT, all from crates.io;
cargo deny's allowlist covers them. Its last release is 2022-12 and it is
unsafe-heavy SIMD -- PATCH.md says so plainly rather than burying it, and
notes that the fallback bounds a rejection bug to a slow path but bounds
nothing about a wrong-bytes bug, which is what the differential test and the
fuzz target are for.

The PR fuzz budget stays at a minute: three targets at 20 s rather than two at
30 s. deep-fuzz.yml gains the target in its matrix. .gitignore gains the
artifacts cargo fuzz leaves behind, since running the new target locally is
now something a contributor will do.

Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>
kurok added a commit that referenced this pull request Sep 17, 2026
…228)

After #213/#214/#218 fixed the two byte-at-a-time loops, what was left of a
full parse was the base64 decode proper: 53% of it in data_encoding's scalar
loop, four table lookups per three bytes, at scalar peak on both CPUs tried.
The 0.9.0 changelog said as much in one line.

The only lever left on it is wider lanes. decode_base64 now decodes the
whitespace-free buffer with base64_simd::STANDARD -- AVX2/SSE4.1 with runtime
detection on x86-64, NEON on aarch64 -- and falls back to
data_encoding::BASE64_MIME_PERMISSIVE on any rejection.

The fallback is what makes this a speed change rather than a change to what
this library accepts. STANDARD accepts a strict subset on a whitespace-free
buffer: same alphabet, padding required by both, and STANDARD is additionally
strict about '=' appearing mid-stream and about non-zero trailing bits -- the
two lenient cases this library has always accepted. So the set "SIMD accepts,
data-encoding rejects" is empty, anything both accept decodes to the same
bytes, and every rejection still carries data_encoding's own
DecodeError { position, kind } into MailParseError::Base64DecodeError and out
to Python with identical text.

Three things hold that up rather than one. simd_and_data_encoding_agree in
body.rs sweeps a built corpus: every length to 200 padded and unpadded, every
byte at every position of a one- and two-block body, mid-stream padding on and
off boundaries, both trailing-bit shapes, every whitespace byte interleaved,
and 0x0b/0x00 which are not whitespace and must survive the strip to be
rejected. The new base64_agreement fuzz target asserts the same three-way
agreement -- acceptance, bytes, and error text -- on arbitrary input through
the public entry point; it ran 3,578,785 executions clean in 60 s, and its
canary was checked to still fire. Three Python tests pin the lenient cases
that now reach the fallback, so a fast path that swallowed them would not look
green.

On the shape of the call, which is not cosmetic. The out-of-place decode is
what ships. The in-place form -- decode_inplace over the stripped buffer, then
truncate -- reuses the strip's allocation instead of making a second one and
reads strictly better. Built as the extension is built (lto = true,
codegen-units = 1) it made the whole parse 9% SLOWER, twice, where this form
makes it 51% faster; standalone the two decoders are within 10% of each other
on the same payloads. That gap is the code-layout sensitivity of #204, and the
function and PATCH.md both say so, because the tidier version is the one a
future reader will reach for.

Measured on an Apple M4 (10 vCPU), two independent interleaved A/Bs pooled to
8 rounds per side, pure-Python controls within 1.1%:

  parse_email        0.254 -> 0.168 ms   -34%
  parse_email_tree   0.246 -> 0.159 ms   -35%
  full_read          0.261 -> 0.172 ms   -34%
  parse_many (8x)    1.989 -> 1.318 ms   -34%

mode="metadata" and untouched lazy parses never call this function and are
flat, as they must be.

base64-simd is MIT and pulls vsimd and outref, both MIT, all from crates.io;
cargo deny's allowlist covers them. Its last release is 2022-12 and it is
unsafe-heavy SIMD -- PATCH.md says so plainly rather than burying it, and
notes that the fallback bounds a rejection bug to a slow path but bounds
nothing about a wrong-bytes bug, which is what the differential test and the
fuzz target are for.

The PR fuzz budget stays at a minute: three targets at 20 s rather than two at
30 s. deep-fuzz.yml gains the target in its matrix. .gitignore gains the
artifacts cargo fuzz leaves behind, since running the new target locally is
now something a contributor will do.

Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>
kurok added a commit that referenced this pull request Sep 17, 2026
…228) (#244)

* Decode base64 bodies with SIMD, keeping data-encoding as the arbiter (#228)

After #213/#214/#218 fixed the two byte-at-a-time loops, what was left of a
full parse was the base64 decode proper: 53% of it in data_encoding's scalar
loop, four table lookups per three bytes, at scalar peak on both CPUs tried.
The 0.9.0 changelog said as much in one line.

The only lever left on it is wider lanes. decode_base64 now decodes the
whitespace-free buffer with base64_simd::STANDARD -- AVX2/SSE4.1 with runtime
detection on x86-64, NEON on aarch64 -- and falls back to
data_encoding::BASE64_MIME_PERMISSIVE on any rejection.

The fallback is what makes this a speed change rather than a change to what
this library accepts. STANDARD accepts a strict subset on a whitespace-free
buffer: same alphabet, padding required by both, and STANDARD is additionally
strict about '=' appearing mid-stream and about non-zero trailing bits -- the
two lenient cases this library has always accepted. So the set "SIMD accepts,
data-encoding rejects" is empty, anything both accept decodes to the same
bytes, and every rejection still carries data_encoding's own
DecodeError { position, kind } into MailParseError::Base64DecodeError and out
to Python with identical text.

Three things hold that up rather than one. simd_and_data_encoding_agree in
body.rs sweeps a built corpus: every length to 200 padded and unpadded, every
byte at every position of a one- and two-block body, mid-stream padding on and
off boundaries, both trailing-bit shapes, every whitespace byte interleaved,
and 0x0b/0x00 which are not whitespace and must survive the strip to be
rejected. The new base64_agreement fuzz target asserts the same three-way
agreement -- acceptance, bytes, and error text -- on arbitrary input through
the public entry point; it ran 3,578,785 executions clean in 60 s, and its
canary was checked to still fire. Three Python tests pin the lenient cases
that now reach the fallback, so a fast path that swallowed them would not look
green.

On the shape of the call, which is not cosmetic. The out-of-place decode is
what ships. The in-place form -- decode_inplace over the stripped buffer, then
truncate -- reuses the strip's allocation instead of making a second one and
reads strictly better. Built as the extension is built (lto = true,
codegen-units = 1) it made the whole parse 9% SLOWER, twice, where this form
makes it 51% faster; standalone the two decoders are within 10% of each other
on the same payloads. That gap is the code-layout sensitivity of #204, and the
function and PATCH.md both say so, because the tidier version is the one a
future reader will reach for.

Measured on an Apple M4 (10 vCPU), two independent interleaved A/Bs pooled to
8 rounds per side, pure-Python controls within 1.1%:

  parse_email        0.254 -> 0.168 ms   -34%
  parse_email_tree   0.246 -> 0.159 ms   -35%
  full_read          0.261 -> 0.172 ms   -34%
  parse_many (8x)    1.989 -> 1.318 ms   -34%

mode="metadata" and untouched lazy parses never call this function and are
flat, as they must be.

base64-simd is MIT and pulls vsimd and outref, both MIT, all from crates.io;
cargo deny's allowlist covers them. Its last release is 2022-12 and it is
unsafe-heavy SIMD -- PATCH.md says so plainly rather than burying it, and
notes that the fallback bounds a rejection bug to a slow path but bounds
nothing about a wrong-bytes bug, which is what the differential test and the
fuzz target are for.

The PR fuzz budget stays at a minute: three targets at 20 s rather than two at
30 s. deep-fuzz.yml gains the target in its matrix. .gitignore gains the
artifacts cargo fuzz leaves behind, since running the new target locally is
now something a contributor will do.

Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>

* Quote the x86 benchmark gate alongside the M4 numbers (#228)

The changelog entry carried only the Apple M4 A/B. The gate's EPYC run on
PR #244 measured the same change interleaved against the merge base --
parse_email 0.536 -> 0.316 ms (-41%), parse_many 4.382 -> 2.583 ms (-41%),
controls within 1.1% -- which is the number a reader on x86 wants and the
one the issue asks to be recorded here.

It also settles the one thing the local run left open: parse_many_metadata
read +2.8% on the M4, above that run's noise floor, and reads +0.9% on x86.
Metadata mode never calls decode_base64, so that was the measurement.

Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>

---------

Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Investigate the ~26% slowdown under rustc 1.98.0, then unpin the toolchain

1 participant