Skip to content

Measure the core directly: a Rust bench crate with allocation counting (#225) - #256

Merged
kurok merged 1 commit into
masterfrom
feat/225-bench-crate
Sep 17, 2026
Merged

kurok merged 1 commit into
masterfrom
feat/225-bench-crate

Conversation

@kurok

@kurok kurok commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

Closes #225.

Branched from master (47965f8). Chosen partly because it adds a new crate rather than touching src/mail_parser.rs — #253 and #255 are both in that file, and this way nothing conflicts.

What it is for

The pytest benchmarks are the oracle for what the wheel costs a user, and a poor instrument for asking where that cost is: every number carries the FFI crossing, the GIL and the construction of Python objects. bench/ measures the core itself.

Isolating the base64 path. strip and b64-* time the whitespace strip and the decode separately — together most of a full parse on the large fixture. b64-simd against b64-scalar is what #228 bought on its own rather than diluted through a whole parse:

mode min (M4)
b64-simd (what ships) 82.6 µs
b64-scalar (what it replaced) 136.8 µs
strip 55.7 µs
decode-part 101.7 µs
full 178.7 µs
metadata 33.7 µs

1.66× for #228, measured on its own for the first time.

Counting allocations. A global allocator records allocs, bytes and peak live bytes per mode, and the modes' documented memory claims become assertions. On the 785 KiB fixture:

mode allocs bytes peak
full 389 1,401,650 1,060,757
metadata 346 37,985 11,929
lazy 387 836,013 790,993
tree 426 1,393,164 1,061,316
tree-deferred 423 823,108 798,111

Metadata mode peaks 89× lower than a full parse. That was expected and never measured.

Three things the measuring taught me

All now encoded rather than assumed — and each was a failing assertion before it was a finding:

A lazy tree holds more than a decoded one on a small attachment-bearing message — 6,696 against 6,323 bytes. Lazy retains the encoded bytes, and base64 is 4/3 of what it decodes to, so the trade only pays once the bodies are large. That ordering is now asserted only where it is actually a promise.

Metadata mode's peak ties a full parse's on a fixture whose bodies are a few hundred bytes, because the peak there is the parse itself. Demanding a strict drop everywhere demands something the mode never promised; <= plus a hard assertion on the large fixture is the honest pair.

The counters are process-global, so the four separate tests I wrote first counted each other's allocations and failed three assertions very convincingly — it looked exactly like the modes breaking their promises. A mutex does not fix that: it serialises the measuring, not the other threads' allocating. One test does.

The lockfile check earned itself immediately

bench/Cargo.lock is committed and pinned to the versions the wheel ships, with check_bench_lockfile.py in the lint job. A freshly resolved lockfile took data-encoding 2.11.1 against the shipped 2.6.0 (and encoding_rs 0.8.41 against 0.8.35) — so the b64 comparison above read 1.86× until it was pinned, comparing against a decoder the library does not use. The number in this PR is the corrected one.

Also

A profiling cargo profile — release codegen plus line tables — so samply and perf attribute time to a line instead of a stripped symbol. Documented as a different binary from the release one, since line tables change section sizes and therefore layout, and this crate is measurably sensitive to that (#240 measured up to 7.5%). For finding where the time goes, never for measuring how much.

Checks run locally

cargo test --release in bench/ (passes under parallel cargo test, which is the point), cargo fmt --check and cargo clippy --all-targets -D warnings for the bench crate, check_bench_lockfile.py, plus the root surface: pytest tests --ignore=tests/benchmark (836 passed, 3 skipped), cargo fmt --check, cargo clippy -D warnings, ruff check ..

No new runtime dependency. Nothing under src/, vendor/ or fast_mail_parser/ changes — the wheel is byte-identical, and bench/ is excluded from the sdist alongside fuzz/.

#225)

The pytest benchmarks are the oracle for what the wheel costs a user, and a
poor instrument for asking where that cost is: every number carries the FFI
crossing, the GIL and the construction of Python objects. bench/ measures the
core itself.

Two things it can do that nothing here could:

Isolate the halves of the base64 path. `strip` and `b64-*` time the
whitespace strip and the decode separately -- together most of a full parse
on the large fixture. b64-simd against b64-scalar is what #228 bought on its
own rather than diluted through a whole parse: 1.66x on an M4, 82.6 us
against 136.8 us.

Count allocations. A global allocator records allocs, bytes and peak live
bytes per mode, and the modes' documented memory claims become assertions
instead of prose. On the 785 KiB fixture a full parse peaks at 1,060,757
bytes and metadata mode at 11,929 -- 89x lower, because it copies no part
bodies. That ratio was expected but never measured.

Three things the measuring taught me, all now encoded rather than assumed:

A lazy tree holds MORE than a decoded one on a small attachment-bearing
message -- 6,696 against 6,323 bytes -- because lazy retains the encoded
bytes and base64 is 4/3 of what it decodes to. The trade only pays once the
bodies are large, so that ordering is asserted only where it is a promise.

Metadata mode's peak ties a full parse's on a fixture whose bodies are a few
hundred bytes, because the peak there is the parse itself. Asserting a strict
drop everywhere demands something the mode never promised.

The counters are process-global, so the four separate tests I wrote first
counted each other's allocations and failed three assertions convincingly.
A mutex does not fix that -- it serialises the measuring, not the other
threads' allocating. One test does.

bench/Cargo.lock is committed and pinned to the versions the wheel ships,
with a lint check that keeps it so. That check earned itself immediately: a
freshly resolved lockfile took data-encoding 2.11.1 against the shipped
2.6.0, and the b64 comparison above read 1.86x until it was pinned.

Also a `profiling` cargo profile -- release codegen plus line tables -- so
samply and perf can attribute time to a line. It is a different binary from
the release one, since line tables change section sizes and therefore layout,
so it is documented as being for finding where the time goes and never for
measuring how much.

No new runtime dependency. Nothing under src/, vendor/ or fast_mail_parser/
changes; the wheel is byte-identical and bench/ is excluded from the sdist.

Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>
@kurok

kurok commented Sep 17, 2026

Copy link
Copy Markdown
Contributor Author

Rebased onto 097394c (post-#254 and #255).

CI was already 15/15 green; the problem was that it had stopped being mergeable, and the rebase alone would not have fixed it.

.github/workflows/test.yml conflicted — both sides add a lint step, so both are kept. But #255 moved src/mail_parser.rs to crates/fast_mail_parser_core/src/lib.rs, and this crate reached for it by path:

#[path = "../../src/mail_parser.rs"]   // in both main.rs and tests/allocs.rs

A path in a string literal is invisible to a merge. The rebase resolved cleanly and left two dead includes behind it — I only found them by building afterwards rather than trusting the resolution.

So the bench crate now links fast_mail_parser_core instead, exactly as #255 did for the fuzz harness. That is the better shape anyway, and it is the argument #236 was making: a crate can be depended on, and a dependency survives its source file moving.

Two consequences of #255 that followed:

  • bench had to join fuzz and vendor/mailparse in the root [workspace] exclude, or cargo adopts it as a member and its pinned lockfile is overridden by the workspace one — which would defeat check_bench_lockfile.py. The comment there says so.
  • Relinking re-resolved the lockfile, so data-encoding and encoding_rs needed re-pinning to the shipped versions. The check caught it again; 12 shared packages now agree.

Re-verified after the rebase: bench builds, its tests pass, cargo fmt --check and clippy -D warnings clean for the bench crate and --workspace, check_bench_lockfile.py green, 854 Python tests pass, and the timing modes still produce numbers (b64-simd 93.0 µs against b64-scalar 148.9 µs on this run).

@kurok
kurok force-pushed the feat/225-bench-crate branch from ff22d3d to 9af50b9 Compare September 17, 2026 14:17
@kurok
kurok merged commit 997c7df into master Sep 17, 2026
15 checks passed
@kurok
kurok deleted the feat/225-bench-crate branch September 17, 2026 14:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add a std-only Rust bench crate (timing + counting allocator) and document a symbolisable release build

1 participant