Skip to content

Stop allocating header keys twice, and test for a header without decoding it (#238) - #248

Merged
kurok merged 1 commit into
masterfrom
feat/238-header-fast-path
Sep 17, 2026
Merged

kurok merged 1 commit into
masterfrom
feat/238-header-fast-path

Conversation

@kurok

@kurok kurok commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

Closes #238.

Branched from master (e7bbb05). Independent of #245/#246/#247, but it edits vendor/mailparse/PATCH.md and the root [patch.crates-io] comment, which #245 also touches — see the rebase note.

What changed

Every header value went through mailparse's two-stage tokenizer: a Vec per physical line, an outer Vec, a second Vec from the whitespace pass, and a result String grown with no capacity hint. About six allocations to hand back the line it was given.

That machinery is for folded values and RFC 2047 encoded words, and the common case is neither. On the 767 KiB fixture the root block has 24 headers, 8 continuation lines and exactly one encoded word; every part-level Content-Type, Content-Transfer-Encoding, Content-Disposition, Content-ID, Date and MIME-Version is one line with no =?. And the tokenizer is paid more than once per header — collect_headers reads them all, then mailparse re-tokenises Content-Type per part, Content-Transfer-Encoding on every get_body_encoded(), Content-Disposition on every get_content_disposition(), and the three flat parsers add Subject, Date and Content-ID. Roughly sixty get_value calls per full parse, ~55 of them for values the fast path answers with one to_owned().

Three changes, no new dependency, Cargo.lock untouched, no output change:

  1. normalize_header (vendored) returns chars.trim_start().to_owned() when the value has no \n, no \r and no =?; otherwise it delegates to the unchanged tokenizer, kept as normalize_header_tokens.
  2. collect_headers keys its position map by get_key_raw() instead of a second owned String, and sizes both containers from part.headers.len(). Latin-1 decoding is injective, so raw-byte equality is the String equality it had — case-sensitive grouping included.
  3. disposition_token tests presence with get_first_header instead of get_first_value, which was building and dropping a normalised String on the next line.

Why the fast path is exactly equivalent

Not "close enough" — the comment and the test both make the argument:

tokenize_header runs value.lines().map(str::trim_start), so with no \n that yields this string once (and nothing for "", where trim_start is also empty). A value produced by parse_header never ends in \r (only a non-CR non-LF byte advances ix_value_end), so the line is the whole value. tokenize_header_line with no =? pushes exactly one maybe_whitespace(line) token. normalize_header_whitespace passes a lone Text or Whitespace through unchanged.

\r is checked anyway — redundant for values parse_header produces, but it keeps the fast path correct for a MailHeader built any other way.

fast_path_matches_the_tokenizer asserts it regardless, against normalize_header_tokens as oracle, over a hand-written edge table (empty values, lone \r, "=?" split across the boundary, leading tabs, "= ?", "?=") and a generated corpus at every length to 48 over an alphabet built from the bytes the check looks for.

Numbers

Apple M4 (10 vCPU), rustc 1.98.0, CPython 3.12, 5 interleaved rounds. Pure-Python controls within 1.3%.

Benchmark master this branch
parse_many_small_threads1 (new) 4.056 ms 3.133 ms −23%
parse_many_metadata_small_threads1 (new) 3.620 ms 2.754 ms −24%
parse_metadata (767 KiB) 0.037 ms 0.027 ms see below
parse_message (767 KiB) 0.164 ms 0.154 ms −6.4%

The two threads=1 batches are the reading that matters, and they are new here (the issue asks for them). Serial, so no scheduler; small messages with no attachments, so no transfer decode — very nearly pure header work, which is exactly what this changes. ~23–24% off both.

On the 767 KiB rows, a caveat I want to be explicit about. parse_metadata reads 0.037 → 0.027 ms, which looks like 36%. It is not: master's 0.037 includes the M4 code-placement artifact documented on #244 — that benchmark was 0.030 ms before base64-simd was linked, and x86 has measured it flat on #244, #245, #246 and #247. Against the artifact-free 0.030 the real gain is ~10%. The x86 gate is the number to quote for that row.

Full A/B report

Measured on Apple M4, 10 vCPU.

Median of 5 interleaved rounds per side; each value is a benchmark's minimum. Positive delta = master is slower.

Benchmark headers-fast master Delta
test__fast_mail_parser___attachment_reread 0.000 ms 0.000 ms +0.0%
test__fast_mail_parser___full_read 0.161 ms 0.171 ms +6.0%
test__fast_mail_parser___parse_lazy_all_attachments 0.172 ms 0.181 ms +4.9%
test__fast_mail_parser___parse_lazy_untouched 0.040 ms 0.050 ms +25.1%
test__fast_mail_parser___parse_many 1.309 ms 1.379 ms +5.3%
test__fast_mail_parser___parse_many_metadata 0.215 ms 0.294 ms +37.0%
test__fast_mail_parser___parse_message 0.154 ms 0.164 ms +6.4%
test__fast_mail_parser___parse_message_strict 0.154 ms 0.163 ms +6.0%
test__fast_mail_parser___parse_metadata 0.027 ms 0.037 ms +36.5%
test__fast_mail_parser___parse_metadata_str 0.027 ms 0.037 ms +36.1%
test__fast_mail_parser___parse_tree 0.157 ms 0.168 ms +7.4%
test__fast_mail_parser___parse_tree_lazy_untouched 0.038 ms 0.049 ms +28.3%
test__fast_mail_parser___parse_tree_metadata 0.027 ms 0.038 ms +38.4%
test__mail_parser___parse_message 5.562 ms 5.492 ms -1.3% control
test__mailparser_lib___full_read 5.897 ms 5.940 ms +0.7% control
test__stdlib_email___full_read 7.536 ms 7.477 ms -0.8% control
test__threaded___parse_many 0.622 ms 0.751 ms +20.8% informational
test__threaded___parse_many_metadata_small 2.279 ms 2.540 ms +11.4% informational
test__threaded___parse_many_metadata_small_threads1 2.754 ms 3.620 ms +31.5% informational
test__threaded___parse_many_small 2.442 ms 2.785 ms +14.0% informational
test__threaded___parse_many_small_threads1 3.133 ms 4.056 ms +29.4% informational
test__threaded___threadpool_parse_email 0.620 ms 0.646 ms +4.2% informational
test__threaded___threadpool_parse_email_small 14.442 ms 13.758 ms -4.7% informational

Noise floor from the pure-Python controls: 1.3% (they cannot be affected by the build, so this is measurement error).

Checks run locally

pytest tests --ignore=tests/benchmark (811 passed, 3 skipped — including test_headers.py, test_multivalue_headers.py, test_stdlib_parity.py with no DIVERGENCES change, and the attachment/metadata/lazy/tree suites that exercise disposition_token), the vendored suite (54 + 25 passed) via --target-dir outside the tree, cargo fmt for both manifests, cargo clippy --all-targets -- -D warnings -W clippy::cast_possible_truncation, mypy --strict, ruff check . — all green. git diff --stat master -- Cargo.lock is empty.

The M4 is not a proxy for x86; the Benchmark quality gate is the verdict.

Rebase note

#245 also edits vendor/mailparse/PATCH.md (making it "three functions changed" for the quoted-printable decoder) and the root [patch.crates-io] comment. Whichever of the two lands second needs those merged by hand — the conflicts are additive, but the "three functions" counts collide and want reading rather than taking one side. I have done this merge twice already this series and will handle it.

@kurok
kurok force-pushed the feat/238-header-fast-path branch from 942989b to 3c4a698 Compare September 17, 2026 09:59
@kurok

kurok commented Sep 17, 2026

Copy link
Copy Markdown
Contributor Author

Rebased onto 2aee3df now that #245 has landed — the PATCH.md collision this PR's Rebase note predicted. Three files conflicted (PATCH.md, the root [patch.crates-io] comment, CHANGELOG.md); the vendored copy is now four changed functions, and the diff -r claim and sync checklist name normalize_header/normalize_header_tokens alongside src/qp.rs.

The new baseline settles the caveat in the PR body. I wrote that parse_metadata's 0.037 → 0.027 was not really 36%, because master's 0.037 carried the M4 placement artifact and the artifact-free figure was 0.030, so the real gain was ~10%. On 2aee3df master measures 0.030 and the gain measures +10.4% — which is the number that reasoning predicted, arrived at from a different base.

Re-measured, 5 interleaved rounds, pure-Python controls within 2.1%:

Benchmark master 2aee3df this branch
parse_many_metadata_small_threads1 3.671 ms 2.746 ms +33.7%
parse_many_small_threads1 4.056 ms 3.171 ms +27.9%
parse_qp_message_metadata 0.022 ms 0.019 ms +18.2%
parse_metadata 0.030 ms 0.027 ms +10.4%
parse_message 0.158 ms 0.154 ms +2.7%
parse_qp_message 0.134 ms 0.130 ms +2.8%

The two threads=1 batches remain the reading that matters: serial, small messages, no transfer decode — very nearly pure header work.

Worth noting the quoted-printable rows, which did not exist when this branch was first cut: parse_qp_message_metadata is header-dominated on a message with 29 headers, and it moves 18%. #229 and #238 are independent changes to different parts of the same parse and they compose.

Re-verified after the rebase: 811 passed / 3 skipped, the vendored suite 58 + 25 (both fast_path_matches_the_tokenizer and all four of #229's QP differential tests), fmt for both manifests, clippy, mypy --strict, ruff.

@kurok

kurok commented Sep 17, 2026

Copy link
Copy Markdown
Contributor Author

The gate failed, and it was right to. test__fast_mail_parser___parse_qp_message read +9.2% slower on x86 (0.246 → 0.268 ms), past the 7% threshold, while everything else on the same run improved:

Benchmark base this revision
parse_qp_message 0.246 ms 0.268 ms +9.2% slower
parse_tree_metadata 0.055 ms 0.049 ms −10.0%
parse_many_metadata_small 4.165 ms 3.739 ms −10.2%
parse_metadata 0.052 ms 0.048 ms −8.1%
parse_many_small 4.705 ms 4.298 ms −8.7%

Two things make me read that as placement rather than as this change:

  1. There is no mechanism. Nothing here touches quoted-printable decoding. The QP path's only contact with this diff is get_body_encoded() → get_first_value("Content-Transfer-Encoding"), which the fast path makes cheaper.
  2. The two CPUs disagree in direction. The M4 reads that same benchmark 3.2% faster, over 5 interleaved rounds, twice. Code layout can swing the hot path 0-96% depending on the runner's CPU #204 is exactly this.

So: normalize_header_tokens is now marked #[inline(never)]. That is not a workaround dressed up — after the fast path it genuinely is the cold half, running only for folded values and encoded words, and leaving it inlinable put its two Vecs and its loop inside a function most callers leave after one branch.

What I cannot tell you is whether it works. The M4 does not reproduce the regression, so I have no local instrument for this: the branch measures the same before and after the attribute here. This is the principled change and the documented first thing to try, but the gate on x86 is the only thing that can decide it. If it does not clear, the next step is #240's salted layout A/B rather than another guess.

Everything else re-verified: 811 passed / 3 skipped, vendored suite 58 + 25, fmt both manifests, clippy, mypy --strict, ruff.

…ding it (#238)

Two of the three changes #238 asks for. The third -- a single-line fast path
in the vendored normalize_header -- is not here; see below.

collect_headers keyed its position map by a decoded String, so every header
allocated its key twice: once for the map and once for the table. Latin-1
decoding is injective, so the raw key bytes are exactly the same equivalence
the String gave, case-sensitive grouping included, and both containers are
now sized from part.headers.len() instead of growing.

disposition_token called get_first_value("Content-Disposition") purely to
test presence, normalising a value it dropped on the next line.
get_first_header answers the same question with the same case-insensitive
first-match semantics and no tokenizer.

Neither touches vendor/, so PATCH.md and the vendored suite are unchanged.

Apple M4 (10 vCPU), 5 interleaved rounds, controls within 1.7%:

  parse_many_small_threads1           4.048 -> 3.839 ms   -5.2%
  parse_many_metadata_small_threads1  3.622 -> 3.442 ms   -5.0%
  parse_qp_message_metadata           0.022 -> 0.021 ms   -6.0%

The two threads=1 batches are new here and are the reading that matters:
serial, so no scheduler, and small messages with no attachments, so no
decode -- very nearly pure header work.

Why the fast path is not in this PR. It was, and it was worth far more --
the same benchmarks read -28% to -34% with it. But the x86 gate failed it
three times on parse_qp_message, a benchmark this branch has no mechanism to
affect: the same message in metadata mode, which runs the headers and skips
the decode, got 6.5% FASTER on every one of those runs. Against an unmoving
0.246 ms base the three binaries measured 0.268, 0.265 and 0.277 -- the
second after pinning the now-cold tokenizer path, the third after also
pinning qp::decode_robust, which made it worse. The M4 reads that benchmark
3% faster throughout and cannot see any of it.

That is the code-layout sensitivity of #120 and #204, and three binaries is
enough to establish that guessing at it does not converge. #240 exists to
build the instrument -- a salted layout A/B on x86 -- and the fast path
should go back through the gate behind that, not behind a fourth guess.
These two changes are independent of it, need no vendored edit, and stand on
their own.

Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>
@kurok
kurok force-pushed the feat/238-header-fast-path branch from 4a3c866 to 34be2f4 Compare September 17, 2026 10:23
@kurok kurok changed the title Take single-line header values without the tokenizer (#238) Stop allocating header keys twice, and test for a header without decoding it (#238) Sep 17, 2026
@kurok

kurok commented Sep 17, 2026

Copy link
Copy Markdown
Contributor Author

Third gate run made it worse (+12.9%), so I have stopped guessing and rescoped this PR instead.

Against an unmoving 0.246 ms base, parse_qp_message measured across three binaries:

attempt change parse_qp_message head delta
1 as authored 0.268 ms +9.2%
2 #[inline(never)] on the now-cold tokenizer path 0.265 ms +7.5%
3 also #[inline(never)] on qp::decode_robust 0.277 ms +12.9%

On all three runs parse_qp_message_metadata — the same message, same headers, same structure, decode skipped — got 6.5% faster, and every metadata and small-batch benchmark improved 6–10%. The M4 reads parse_qp_message ~3% faster throughout and cannot see any of this. Three binaries is enough to say that guessing at placement does not converge.

What this PR is now. The two changes that need no vendored edit:

  • collect_headers keys its position map by raw key bytes instead of a second owned String, and pre-sizes both containers.
  • disposition_token uses get_first_header instead of building and dropping a normalised value.
Benchmark master this branch
parse_many_small_threads1 (new) 4.048 ms 3.839 ms −5.2%
parse_many_metadata_small_threads1 (new) 3.622 ms 3.442 ms −5.0%
parse_qp_message_metadata 0.022 ms 0.021 ms −6.0%

Controls within 1.7%. vendor/ is untouched — git diff origin/master -- vendor/ is empty — so PATCH.md, the root patch comment and the vendored suite are all back to master's state.

What came out, and what should happen to it. The normalize_header fast path was worth far more: the same two benchmarks read −28% to −34% with it. It is a real ~25-point win sitting behind a layout problem that this repo already has an issue for. #240 is exactly that instrument — a salted layout A/B on x86 — and the fast path should go back through the gate behind it rather than behind a fourth guess. I have not filed anything; say the word and I will open an issue carrying the three measurements, the equivalence argument and the differential test, so none of that work is lost.

I would rather land 5% that is understood than 30% that fails the gate for a reason I cannot name.

@kurok
kurok merged commit c403a69 into master Sep 17, 2026
15 checks passed
@kurok
kurok deleted the feat/238-header-fast-path branch September 17, 2026 10:28
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Header values: single-line fast path in the vendored get_value, drop the double key allocation in collect_headers

1 participant