Skip to content

Size parse_many's workers by bytes, and let the caller be one (#232) - #247

Merged
kurok merged 1 commit into
masterfrom
feat/232-parse-many-workers
Sep 17, 2026
Merged

kurok merged 1 commit into
masterfrom
feat/232-parse-many-workers

Conversation

@kurok

@kurok kurok commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

Closes #232.

Branched from master (e7bbb05). Independent of #245 and #246.

The bug, measured before the fix

parse_many capped its workers on message count alone, never on how much work there was. Sixteen one-kilobyte messages — an IMAP fetch page, a queue poll — got min(cores, 16) threads, each of which parsed roughly one message and exited. Creating and joining an OS thread is 15–40 µs against a parse of ~1.1 µs/KB, so the scheduler cost one to two orders of magnitude more than the work it scheduled.

On this machine, on master:

min mean
parse_many(page) — default threads 63.5 µs 108.4 µs
parse_many(page, threads=1) 28.4 µs 31.2 µs

parse_many was 2.2x slower than asking it not to use threads — for the shape a mail pipeline produces most often, against a README that says it "costs nothing at any size". No benchmark covered that shape, which is why it went unnoticed.

What changed

Three std-only changes. No new dependency, the dynamic cursor kept.

(a) Workers are capped by bytes as well as by count. MIN_BYTES_PER_WORKER = 64 KiB; a batch with less total work than one worker's worth takes the existing inline path. threads= stays an upper bound — the gate may lower it, never raise it.

(b) The calling thread is a worker. The claim loop is now a closure that both the spawned workers and the caller run; workers - 1 get spawned. The caller already sits inside py.detach with the GIL released, so blocking it on joins was one spawn and one join per call for nothing — half the thread cost when workers == 2. A panic on the caller unwinds out of thread::scope, which joins the spawned workers first, and reaches catch_panics exactly as a resume_unwind from a joined handle did.

(c) The default parallelism is probed once per process, in the binding layer. std does not memoise available_parallelism; on Linux it reads /proc/self/cgroup, /proc/self/mountinfo and the cgroup cpu.max files on every call. The cache cannot live in mail_parser — that module's docs rule out statics and OnceLock so its free-threading audit stays valid — so it lives next to the OnceLocks this layer already has, and the core keeps its own fallback for the fuzz targets that include it by path. The trade-off is stated at the function: a cgroup quota changed after the first call is not observed; pass threads= to override.

Numbers

Apple M4 (10 vCPU), rustc 1.98.0, CPython 3.12, 5 interleaved rounds. Pure-Python controls within 0.5%.

Benchmark master this branch
parse_many_small_page (new) 0.070 ms 0.032 ms 2.2x
parse_many_small_page_threads1 (new control) 0.029 ms 0.028 ms +1.0% flat
parse_many (16 × 767 KiB) 1.326 ms 1.333 ms −0.5% flat

The control is the point. Same batch, same parsing, scheduler forced off — it does not move. That is what says the 2.2x is scheduling and not parsing, and it is why the benchmark was added in a pair.

The 16 × 767 KiB batch is unaffected, as intended: 12 MB is far above the gate.

What I am not claiming from this table. parse_many_small and parse_many_metadata_small read ~27% faster, threaded___parse_many ~8%, and every metadata/lazy-untouched benchmark +16–24%. Those last ones are the M4 code-placement artifact documented on #244 — master sits at 0.037 ms on parse_metadata where it was 0.030 before base64-simd was linked, and x86 measured them flat on #244, #245 and #246. The 2000-message rows are above the byte gate and spawn the same worker count either way, so I do not have a mechanism that accounts for 27%; the plausible one (one fewer spawn out of ten) is worth well under 1%. Treat those rows as unexplained until the gate reports. The claim this PR makes is the first row and its control.

Full A/B report

Measured on Apple M4, 10 vCPU.

Median of 5 interleaved rounds per side; each value is a benchmark's minimum. Positive delta = master is slower.

Benchmark workers master Delta
test__fast_mail_parser___attachment_reread 0.000 ms 0.000 ms +0.0%
test__fast_mail_parser___full_read 0.163 ms 0.171 ms +4.7%
test__fast_mail_parser___parse_lazy_all_attachments 0.174 ms 0.181 ms +3.8%
test__fast_mail_parser___parse_lazy_untouched 0.043 ms 0.050 ms +16.5%
test__fast_mail_parser___parse_many 1.333 ms 1.326 ms -0.5%
test__fast_mail_parser___parse_many_metadata 0.241 ms 0.294 ms +21.9%
test__fast_mail_parser___parse_message 0.157 ms 0.163 ms +4.0%
test__fast_mail_parser___parse_message_strict 0.157 ms 0.163 ms +4.0%
test__fast_mail_parser___parse_metadata 0.030 ms 0.037 ms +24.0%
test__fast_mail_parser___parse_metadata_str 0.030 ms 0.037 ms +23.5%
test__fast_mail_parser___parse_tree 0.157 ms 0.165 ms +5.0%
test__fast_mail_parser___parse_tree_lazy_untouched 0.041 ms 0.049 ms +17.0%
test__fast_mail_parser___parse_tree_metadata 0.031 ms 0.038 ms +22.9%
test__mail_parser___parse_message 5.476 ms 5.497 ms +0.4% control
test__mailparser_lib___full_read 5.790 ms 5.821 ms +0.5% control
test__stdlib_email___full_read 7.450 ms 7.429 ms -0.3% control
test__threaded___parse_many 0.677 ms 0.731 ms +8.0% informational
test__threaded___parse_many_metadata_small 1.975 ms 2.521 ms +27.7% informational
test__threaded___parse_many_small 2.187 ms 2.789 ms +27.5% informational
test__threaded___parse_many_small_page 0.032 ms 0.070 ms +123.2% informational
test__threaded___parse_many_small_page_threads1 0.028 ms 0.029 ms +1.0% informational
test__threaded___threadpool_parse_email 0.631 ms 0.647 ms +2.5% informational
test__threaded___threadpool_parse_email_small 14.095 ms 14.203 ms +0.8% informational

Noise floor from the pure-Python controls: 0.5% (they cannot be affected by the build, so this is measurement error).

Contract

Unchanged: one result per input in input order, per-slot Ok/Err, raise_on_error, strict, mode=, threads=0 → ValueError, negative → OverflowError, the GIL released for the whole batch, a parser panic surfacing as ParseError.

Worker count is not observable from Python, so the new tests in tests/test_parse_many.py pin the only thing that is — that results do not depend on how the scheduler sized itself:

  • test__a_small_batch_with_default_threads_matches_threads_one — 16 messages, asserted under the gate, all three modes.
  • test__a_large_batch_with_default_threads_still_matches_threads_one — 400 × 20 messages, asserted above the gate, so the genuinely parallel path stays covered rather than being quietly serialised by the change under test.
  • test__threads_is_an_upper_bound_not_a_target — threads=64 on a tiny batch still returns correct results.
  • test__a_single_message_batch_still_parses and test__empty_and_tiny_batches_do_not_divide_by_zero — workers bottoms out through two .min()s and a .max(1), and total_bytes / MIN_BYTES_PER_WORKER is integer division, so these are where an off-by-one or a zero would show.

Docs

Readme.md said parse_many "costs nothing at any size". That was false for small batches, so the section now says workers are sized by bytes, that a sub-64-KiB batch parses inline, and what it used to cost. The parse_many docstring in __init__.pyi says threads is an upper bound and that small batches use fewer.

Checks run locally

pytest tests --ignore=tests/benchmark (820 passed, 3 skipped), cargo fmt --all --check, cargo clippy --all-targets -- -D warnings -W clippy::cast_possible_truncation, mypy --strict, ruff check . — all green. No vendor/ change, so no PATCH.md and no vendored cargo test. src/mail_parser.rs stays PyO3-free (it is compiled by path into two fuzz targets) and gains no shared state — the one cache added is in the binding layer.

@kurok

kurok commented Sep 17, 2026

Copy link
Copy Markdown
Contributor Author

Gate status, and why I am not patching it here.

This PR was green at 15/15 before I rebased it onto 2aee3df. The only change since is the base — #229's quoted-printable decoder landed in it — plus a one-line CHANGELOG.md conflict resolution. The gate now fails on test__fast_mail_parser___parse_qp_message at +19.0%.

This PR is the clearest evidence in the series that the failure is not the PR. It touches src/fast_mail_parser.rs and src/mail_parser.rs only — the thread scheduler. There is no path from a worker-sizing change to a quoted-printable decoder it never calls. And the run says as much itself:

  • parse_qp_message_metadata — same message, same headers, decode skipped — −0.3%, flat
  • parse_many_small_page, what this PR is for — −55.6%, with its forced-serial control at +1.4%
  • parse_qp_message — +19.0%, per-round minima 0.279 / 0.280 / 0.278

Rock-stable across rounds, so not noise; and the only work separating the flat row from the failing one is the decode.

#246 failed the same way on the same rebase (+7.3%), and #248 took three #[inline(never)] attempts that moved its number 0.268 → 0.265 → 0.277 against an unmoving 0.246 base before I stopped guessing. #249 adds the instrument: same source, K builds differing only in a layout salt, per-benchmark spread across salts. That number is what a 19% verdict here should be read against.

Leaving it failing and explained rather than passing by guesswork. Ready to rebase once there is a measured spread to judge against.

@kurok
kurok force-pushed the feat/232-parse-many-workers branch from 52ed36b to 75e4ac1 Compare September 17, 2026 10:48
@kurok

kurok commented Sep 17, 2026

Copy link
Copy Markdown
Contributor Author

Rebased onto c403a69. The code did not change. The gate's verdict changed completely — and that is the answer.

Same source as the previous run (a CHANGELOG line and a hand-merged benchmark file aside), different base commit, therefore a different binary layout. Compare the two runs on x86:

Benchmark previous run this run
parse_qp_message +19.0% (failed the gate) −6.9%
parse_qp_message_metadata −0.3% −10.4%
parse_tree_lazy_untouched — +15.2% (fails the gate)
parse_many_metadata — +11.9%
parse_lazy_untouched — +11.1%
parse_tree — +9.3%
parse_metadata_str — −9.7%
controls 3.5% 0.5%

The benchmark that fails moved. The one that failed at +19% is now 6.9% faster. The same unchanged scheduler change now scatters ±15% in both directions against a 0.5% control floor.

And #246 — rebased in the same batch, also no code change — went from failing at +7.3% to passing 15/15.

Two independent flips in one batch. This is not a property of either PR; it is the spread of the instrument. On this revision, on these runners, code placement alone moves the parse benchmarks by at least ±15%, which is more than twice the gate's 7% threshold. A 7% verdict on these paths is currently a coin toss with extra steps.

That is exactly the number #249 exists to measure properly — four salted builds of one source, interleaved, reporting the per-benchmark spread — instead of us inferring it from failures. The inference is now strong enough to act on: this PR should be judged against the measured spread, not against 7%, and I am not going to touch its code to make a placement lottery land the right way.

Its own benchmark, meanwhile, is doing exactly what it was written to do: parse_many_small_page −44.2%, with the forced-serial control at −0.4%.

parse_many capped its workers on message count alone, never on how much
work there was. Sixteen one-kilobyte messages -- an IMAP fetch page, a queue
poll -- got min(cores, 16) threads, each of which parsed about one message
and exited. Creating and joining an OS thread is 15-40 us against a parse of
roughly 1.1 us per KB, so for that shape the scheduler cost one to two
orders of magnitude more than the work it scheduled.

Measured, on an Apple M4 before this change: a 16-message page with the
default thread count took 0.070 ms, and the same page with threads=1 took
0.029 ms. parse_many was 2.2x slower than asking it not to use threads --
for the shape a mail pipeline produces most often, and against a README that
says it "costs nothing at any size".

Three std-only changes, no new dependency, the dynamic cursor kept:

  (a) Workers are capped at one per MIN_BYTES_PER_WORKER (64 KiB) of input as
      well as one per message, so a batch with less work than one worker's
      worth takes the existing inline path. threads= stays an upper bound:
      the gate may lower it and never raises it.
  (b) The claim loop is a closure both the spawned workers and the calling
      thread run. The caller already sits inside py.detach with the GIL
      released, so blocking it on joins was one spawn and one join per call
      for nothing -- half the thread cost at workers == 2. A panic on the
      caller unwinds out of thread::scope, which joins the spawned workers
      first, and reaches catch_panics exactly as a resume_unwind did.
  (c) available_parallelism is probed once per process, in the binding layer.
      std does not memoise it; on Linux it reads /proc/self/cgroup,
      /proc/self/mountinfo and the cgroup cpu.max files on every call. The
      cache cannot live in mail_parser -- that module's docs rule out statics
      and OnceLock so its free-threading audit stays valid -- and the core
      keeps its own fallback for the fuzz targets that include it by path.
      The trade-off is stated at the function: a cgroup quota changed after
      the first call is not observed; pass threads= to override.

Results, input ordering, per-slot Ok/Err, raise_on_error, strict, mode= and
the GIL being released for the whole batch are all unchanged. Worker count
is not observable from Python, so the new tests pin the only thing that is:
default threads and threads=1 agree, subject for subject, in every mode --
for a batch under the gate and for one above it, so the genuinely parallel
path stays covered rather than being quietly serialised by the gate.

Apple M4 (10 vCPU), 5 interleaved rounds, pure-Python controls within 0.5%:

  parse_many_small_page            0.070 -> 0.032 ms   2.2x
  parse_many_small_page_threads1   0.029 -> 0.028 ms   flat (control)
  parse_many (16 x 767 KiB)        1.326 -> 1.333 ms   flat

The control is the point: same batch, same parsing, scheduler forced off. It
does not move, so the saving is scheduling.

Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>
@kurok
kurok force-pushed the feat/232-parse-many-workers branch from 75e4ac1 to d9ef69f Compare September 17, 2026 11:02
@kurok
kurok merged commit 5027ca5 into master Sep 17, 2026
15 checks passed
@kurok
kurok deleted the feat/232-parse-many-workers branch September 17, 2026 11:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

parse_many: run the claim loop on the calling thread, gate workers on batch bytes, and stop probing available_parallelism() per call

1 participant