Repository navigation
Measure the core directly: a Rust bench crate with allocation counting (#225) - #256
Conversation
#225) The pytest benchmarks are the oracle for what the wheel costs a user, and a poor instrument for asking where that cost is: every number carries the FFI crossing, the GIL and the construction of Python objects. bench/ measures the core itself. Two things it can do that nothing here could: Isolate the halves of the base64 path. `strip` and `b64-*` time the whitespace strip and the decode separately -- together most of a full parse on the large fixture. b64-simd against b64-scalar is what #228 bought on its own rather than diluted through a whole parse: 1.66x on an M4, 82.6 us against 136.8 us. Count allocations. A global allocator records allocs, bytes and peak live bytes per mode, and the modes' documented memory claims become assertions instead of prose. On the 785 KiB fixture a full parse peaks at 1,060,757 bytes and metadata mode at 11,929 -- 89x lower, because it copies no part bodies. That ratio was expected but never measured. Three things the measuring taught me, all now encoded rather than assumed: A lazy tree holds MORE than a decoded one on a small attachment-bearing message -- 6,696 against 6,323 bytes -- because lazy retains the encoded bytes and base64 is 4/3 of what it decodes to. The trade only pays once the bodies are large, so that ordering is asserted only where it is a promise. Metadata mode's peak ties a full parse's on a fixture whose bodies are a few hundred bytes, because the peak there is the parse itself. Asserting a strict drop everywhere demands something the mode never promised. The counters are process-global, so the four separate tests I wrote first counted each other's allocations and failed three assertions convincingly. A mutex does not fix that -- it serialises the measuring, not the other threads' allocating. One test does. bench/Cargo.lock is committed and pinned to the versions the wheel ships, with a lint check that keeps it so. That check earned itself immediately: a freshly resolved lockfile took data-encoding 2.11.1 against the shipped 2.6.0, and the b64 comparison above read 1.86x until it was pinned. Also a `profiling` cargo profile -- release codegen plus line tables -- so samply and perf can attribute time to a line. It is a different binary from the release one, since line tables change section sizes and therefore layout, so it is documented as being for finding where the time goes and never for measuring how much. No new runtime dependency. Nothing under src/, vendor/ or fast_mail_parser/ changes; the wheel is byte-identical and bench/ is excluded from the sdist. Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>
|
Rebased onto CI was already 15/15 green; the problem was that it had stopped being mergeable, and the rebase alone would not have fixed it.
#[path = "../../src/mail_parser.rs"] // in both main.rs and tests/allocs.rsA path in a string literal is invisible to a merge. The rebase resolved cleanly and left two dead includes behind it — I only found them by building afterwards rather than trusting the resolution. So the bench crate now links Two consequences of #255 that followed:
Re-verified after the rebase: bench builds, its tests pass, |
ff22d3d to
9af50b9
Compare
Closes #225.
Branched from
master(47965f8). Chosen partly because it adds a new crate rather than touchingsrc/mail_parser.rs— #253 and #255 are both in that file, and this way nothing conflicts.What it is for
The pytest benchmarks are the oracle for what the wheel costs a user, and a poor instrument for asking where that cost is: every number carries the FFI crossing, the GIL and the construction of Python objects.
bench/measures the core itself.Isolating the base64 path.
stripandb64-*time the whitespace strip and the decode separately — together most of a full parse on the large fixture.b64-simdagainstb64-scalaris what #228 bought on its own rather than diluted through a whole parse:b64-simd(what ships)b64-scalar(what it replaced)stripdecode-partfullmetadata1.66× for #228, measured on its own for the first time.
Counting allocations. A global allocator records allocs, bytes and peak live bytes per mode, and the modes' documented memory claims become assertions. On the 785 KiB fixture:
Metadata mode peaks 89× lower than a full parse. That was expected and never measured.
Three things the measuring taught me
All now encoded rather than assumed — and each was a failing assertion before it was a finding:
A lazy tree holds more than a decoded one on a small attachment-bearing message — 6,696 against 6,323 bytes. Lazy retains the encoded bytes, and base64 is 4/3 of what it decodes to, so the trade only pays once the bodies are large. That ordering is now asserted only where it is actually a promise.
Metadata mode's peak ties a full parse's on a fixture whose bodies are a few hundred bytes, because the peak there is the parse itself. Demanding a strict drop everywhere demands something the mode never promised;
<=plus a hard assertion on the large fixture is the honest pair.The counters are process-global, so the four separate tests I wrote first counted each other's allocations and failed three assertions very convincingly — it looked exactly like the modes breaking their promises. A mutex does not fix that: it serialises the measuring, not the other threads' allocating. One test does.
The lockfile check earned itself immediately
bench/Cargo.lockis committed and pinned to the versions the wheel ships, withcheck_bench_lockfile.pyin the lint job. A freshly resolved lockfile tookdata-encoding 2.11.1against the shipped 2.6.0 (andencoding_rs 0.8.41against 0.8.35) — so the b64 comparison above read 1.86× until it was pinned, comparing against a decoder the library does not use. The number in this PR is the corrected one.Also
A
profilingcargo profile — release codegen plus line tables — sosamplyandperfattribute time to a line instead of a stripped symbol. Documented as a different binary from the release one, since line tables change section sizes and therefore layout, and this crate is measurably sensitive to that (#240 measured up to 7.5%). For finding where the time goes, never for measuring how much.Checks run locally
cargo test --releaseinbench/(passes under parallelcargo test, which is the point),cargo fmt --checkandcargo clippy --all-targets -D warningsfor the bench crate,check_bench_lockfile.py, plus the root surface:pytest tests --ignore=tests/benchmark(836 passed, 3 skipped),cargo fmt --check,cargo clippy -D warnings,ruff check ..No new runtime dependency. Nothing under
src/,vendor/orfast_mail_parser/changes — the wheel is byte-identical, andbench/is excluded from the sdist alongsidefuzz/.