From 0a9a10a843240dc0051327de1ce26802b0c0c6d5 Mon Sep 17 00:00:00 2001 From: leynos Date: Fri, 18 Sep 2026 20:35:13 +0200 Subject: [PATCH 01/17] Record the deferred split-build-dir harness trim (#693) The 156s figure that motivated trimming `harness_compiles_under_a_split_build_dir` came from the contended distribution, where the two isolated-Cargo tests each roughly halved the other. #687 moved the other one off the Windows lane, so this test got faster without being touched and the saving a trim could return fell with it. Measured across the three runs after #687, the test's exclusive tail is 62.3s, 87.2s and 85.5s, so trimming it is worth about 85s rather than 156s. The lane itself fell from a 1468s median to about 848s across #687, #690 and #691, which makes that 85s roughly ten percent of what remains. Record the decision not to trim it now, the ten-run revisit gate under which that decision is reconsidered, and the three alternatives already ruled out with their measured costs: `cargo check` (114s against 102s cold, and it writes nothing into the target directory so the uplift the regression exists to catch stops happening), warming the compiler cache (about three percent), and sharing a target directory (blocked by E0460 races with the `#[once]` fixture). Describe what a fixture-crate replacement would have to carry, so the design is on record while the trim is deferred: the fidelity argument naming the regression it still guards and the coverage it drops, and the Windows response-file pressure, which a one-dependency fixture would stop exercising unless it generates enough search paths or the contract moves to its own dedicated test. --- docs/developers-guide.md | 116 +++++++++++++++++++++++++++++++++++++-- 1 file changed, 112 insertions(+), 4 deletions(-) diff --git a/docs/developers-guide.md b/docs/developers-guide.md index 6b35e0fdf..f59b5156c 100644 --- a/docs/developers-guide.md +++ b/docs/developers-guide.md @@ -3012,9 +3012,76 @@ Table: Windows durations before and after the verification build moved. | `Test` step | 471s | 260s | 368s | The 420s budget is therefore sized against the older, contended distribution -and is deliberately conservative while the new shape has two samples. It is a -candidate for tightening, or for deletion, once ten runs have accumulated under -it. +and is deliberately conservative while the new shape has three samples. It is a +candidate for tightening, or for deletion, once the revisit gate below is met. + +#### Deferring the split-build-dir harness trim + +The repository has decided **not** to trim +`harness_compiles_under_a_split_build_dir` yet. Trimming it to near zero is +worth about 85s now rather than the 156s an earlier reading implied, and that +smaller number is the whole reason the decision was to wait rather than to +build. + +The 156s came from the contended distribution. Before #687 the two +isolated-Cargo tests ran concurrently, each with four compile jobs on a +four-vCPU runner, so each roughly halved the other, and 156s was the implied +value of trimming both of them. Once +`packaged_manifest_retains_build_script_sources` left the Windows lane, as the +table above records, this test got faster without being touched and what a trim +could return fell with it. + +What a trim returns is the test's **exclusive tail**, the period after every +other test has reported, rather than its total duration, because the remaining +tests fill the run either way. Measured on the three runs after #687 — the pair +the table above draws on, plus the third under the new shape: + +Table: the split-build harness's duration and exclusive tail after #687. + +| Run | Test duration | Exclusive tail | +| ----------- | ------------- | -------------- | +| 34075197897 | 125.3s | 62.3s | +| 34079222917 | 170.6s | 87.2s | +| 34080385050 | 170.7s | 85.5s | + +The figure to plan against is therefore about 85s, not 156s, and nobody should +start this work expecting the larger one. For scale, the whole Windows lane +fell from a 1468s median to about 848s across #687, #690 and #691, so 85s is +roughly ten percent of what remains. + +Three alternatives were measured before the decision to defer, and each was +rejected on its own evidence rather than on preference: + +- **`cargo check` instead of `cargo build`.** Timed cold at `-j 4` on a 32-core + host: 114s against 102s, twelve percent. It also writes nothing into the + target directory, so the uplift the regression exists to catch stops + happening and the test passes vacuously. +- **Warming the lane's compiler cache.** The cache already reaches the spawned + build, because `ci-windows.yml` sets `RUSTC_WRAPPER` at job scope and the + test adds to the child environment rather than clearing it. Warming it is + worth about three percent: 281.0s on the cold run against a 271.9s warm + median. +- **Sharing a target directory.** The test needs private roots to avoid racing + the `#[once]` fixture with `E0460`, so its build cannot reuse the lane's + artefacts or the other test's. + +**Revisit gate.** This defers the trim; it does not close it. Wait until ten +runs of the split Windows lane exist, so the harness test's share of the +`build-test-windows` job is known under the new shape rather than estimated +from three runs. If its tail has settled below the 85s measured here, or if +`Test` has stopped being the lane's critical path, the trim is not worth the +fidelity risk and the work closes without it. Any replacement built at that +point inherits the constraints in +[what a fixture-crate replacement would have to preserve](#what-a-fixture-crate-replacement-would-have-to-preserve). + +Four references sit behind the figures above: +[#673](https://github.com/leynos/netsuke/issues/673) measured the lane, +[#687](https://github.com/leynos/netsuke/pull/687) relocated the packaging +verification build and so changed this test's cost, +[#690](https://github.com/leynos/netsuke/pull/690) folded the native-recipe +smoke job into the gate job, and +[#691](https://github.com/leynos/netsuke/issues/691) split the Windows lints +from the tests. ### How this relates to the isolation utilities @@ -4551,7 +4618,48 @@ That private build is why this test is the most expensive one on the Windows gate: `test_support` depends on `netsuke-build`, so a private root means compiling that crate and roughly 350 dependencies from scratch. Its measured budget is recorded in -[Windows budget for the isolated-Cargo-build tests][windows-test-budget]. +[Windows budget for the isolated-Cargo-build tests][windows-test-budget], which +also records the +[decision to defer a trim](#deferring-the-split-build-dir-harness-trim) and the +gate at which that decision is revisited. + +#### What a fixture-crate replacement would have to preserve + +If the trim is taken up after that gate, the obvious shape is a minimal fixture +crate built under the split layout in place of `test_support`. It needs at +least one dependency, so that dependency rlibs land in the split build +directory while the fixture's own uplifted rlib lands in the target directory. +That is precisely the arrangement the regression exists to catch: a single +derived `-L dependency=` directory that missed the dependencies entirely. This +section records what such a replacement must carry; nothing here is built while +the trim is deferred. + +**The fidelity argument.** The current test is a regression test for a defect +that was found once, and its subject is the real `test_support` build. Swapping +that subject for a stand-in weakens the test unless the argument for the swap +is explicit, in a doc comment beside the test, about exactly which regression +it still guards and what it no longer covers. A one-dependency fixture does +exercise the split-directory derivation — dependency artefacts in the build +directory, uplifted artefacts in the target directory — but it no longer covers +that derivation against the real crate's roughly 350-dependency scale, nor +against the proc-macro and dynamic-library artefacts described above. Those are +what makes the directory enumeration non-trivial, and the comment must say so +rather than let the coverage drop silently. + +**The Windows response-file pressure.** `TestSupportRlib::compile` passes its +arguments through a `rustc` response file, and the reason is a Windows command +line limit rather than a style choice. Cargo 1.99 gives every crate its own +artefact directory, so the `-L dependency=` set holds one entry per dependency; +this test adds long temporary roots on top of that. Passed directly, the result +exceeds the Windows `CreateProcess` command-line limit and the spawn fails with +`Os { code: 206 }` before `rustc` runs at all. A fixture crate with one +dependency produces far fewer directories and would stop exercising that +pressure, which is a measurable loss of coverage however cheap the fixture +becomes. So a replacement must either generate enough search paths to keep the +`@file` path genuinely exercised, or move the response-file contract into its +own dedicated test. Either way the doc comment above the replacement must say +which of the two it does, because the failure it guards is Windows-specific and +cannot be reproduced on most local hosts. ### Manifest `env()` reader From 997233101eb8fd22c8a856286a0d6e0afacaea9e Mon Sep 17 00:00:00 2001 From: leynos Date: Fri, 18 Sep 2026 20:45:54 +0200 Subject: [PATCH 02/17] Point the harness comments at the deferred trim (#693) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Neither `.config/nextest.toml` nor the test's doc comment recorded that the trim measured at 156s is now worth about 85s, so both still read as though the larger figure were available. Add the three post-#687 runs (125.3s, 170.6s and 170.7s against the 274.7s contended median) and the exclusive tails that give the 85s figure, and point the "candidate for tightening or deletion" note at the ten-run revisit gate that "Deferring the split-build-dir harness trim" in docs/developers-guide.md defines. The test's doc comment gains the fidelity argument — that the subject is the real `test_support` build rather than a fixture crate, and that the developers' guide holds the criteria for revisiting that — and the response-file note, that the long `-L dependency=` set plus the long temporary roots is what keeps the Windows response-file path exercised, so any replacement must preserve that pressure or move it to a dedicated test. No executable logic and no timeout value changes. --- .config/nextest.toml | 23 ++++++++++++++++++----- tests/locale_stub_ui_tests.rs | 18 ++++++++++++++++++ 2 files changed, 36 insertions(+), 5 deletions(-) diff --git a/.config/nextest.toml b/.config/nextest.toml index 8975b3c85..7d1bb8115 100644 --- a/.config/nextest.toml +++ b/.config/nextest.toml @@ -69,11 +69,24 @@ success-output = "immediate" # # Removing that second Cargo build also unblocked this one. The two used to run # concurrently, each with four compile jobs on a four-vCPU runner, and each -# roughly halved the other; on those same two runs this test took 125.3s and -# 170.6s rather than the 274.7s median above. 420s is therefore sized against -# the older, contended distribution and is deliberately conservative while the -# new shape has two samples. It is a candidate for tightening, or for deletion, -# once ten runs have accumulated under it. +# roughly halved the other; on those same runs this test took 125.3s, 170.6s +# and 170.7s rather than the 274.7s median above. 420s is therefore sized +# against the older, contended distribution and is deliberately conservative +# while the new shape has three samples. +# +# Trimming the test is worth about 85s, not the 156s that figure implied +# before the second build left. What a trim returns is the test's exclusive +# tail, the period after every other test has reported, measured at 62.3s, +# 87.2s and 85.5s on the three runs after that change, where 156s was the +# implied value of trimming both tests while they contended. The lane fell +# from a 1468s median to about 848s across the three changes, so 85s is +# roughly ten percent of what remains. The repository has decided not to spend +# the fidelity risk on that number yet: the trim is a candidate for tightening +# or deletion at the ten-run revisit gate that "Deferring the split-build-dir +# harness trim" in docs/developers-guide.md defines. That section also records +# the alternatives already measured and rejected — `cargo check` for +# `cargo build`, warming the compiler cache, and sharing a target directory — +# and the constraints a fixture-crate replacement would have to preserve. # # Windows only. On Ubicloud both tests finish in a fraction of the budget, and # widening the timeout there would blunt the hang detection this file exists to diff --git a/tests/locale_stub_ui_tests.rs b/tests/locale_stub_ui_tests.rs index d933e2707..4e2ca08bd 100644 --- a/tests/locale_stub_ui_tests.rs +++ b/tests/locale_stub_ui_tests.rs @@ -84,6 +84,24 @@ fn stub_env_builders_compile_under_the_same_harness( /// the collected `-L dependency=` set has to span the split for the control /// fixture to compile. This pins the regression where a single derived /// directory missed the dependencies entirely. +/// +/// The subject is the real `test_support` build, not a fixture crate, and that +/// is deliberate: it is what carries both the dependency artefacts and the +/// uplifted one that the split-directory derivation has to tell apart. The +/// cost of building it here is recorded, with the decision to defer trimming +/// it, under "Deferring the split-build-dir harness trim" in +/// docs/developers-guide.md; that section also states the fidelity argument +/// any fixture-crate replacement would owe and the response-file pressure it +/// would have to keep. +/// +/// It is also what keeps the Windows response-file path exercised. The long +/// `-L dependency=` set this build produces, plus the long temporary roots the +/// test adds, is why `TestSupportRlib::compile` sends its arguments through a +/// `rustc` response file at all rather than a command line. A fixture crate +/// with one dependency would produce far fewer directories and stop +/// exercising that, so any replacement must either generate enough search +/// paths to keep the pressure or move the response-file contract into its own +/// dedicated test. #[rstest] fn harness_compiles_under_a_split_build_dir() -> io::Result<()> { let subscriber = tracing_subscriber::fmt().with_test_writer().finish(); From 151d0bef86a85913f19fc642ed4126607a055add Mon Sep 17 00:00:00 2001 From: leynos Date: Fri, 18 Sep 2026 22:00:21 +0200 Subject: [PATCH 03/17] Address CodeRabbit findings on the trim record (#693) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two findings, both on the same passage of the nextest comment, and one prose nit on the guide. The contention narrative was genuinely ambiguous: "on those same runs" tied the 125.3s, 170.6s and 170.7s values to the runs where the two isolated-Cargo tests contended, when they are the three runs *after* the second build left the Windows lane. Say which direction the change ran — contention had been increasing each test's elapsed time, and its removal produced the shorter durations — so the sentence cannot be read as claiming the values were measured under contention. The guide's cross-reference to the new fixture-crate subsection used an inline link too long for mdtablefix to wrap at 80 columns, which check-fmt rejects. Move it to the file's reference-link convention. The link label is `fixture-constraints` rather than the longer `fixture-crate-constraints`: mdtablefix treats `[text][label]` as one atomic token, and the longer label leaves no valid wrap point, so it wanted to strand the sentence's full stop on a line of its own. The visible text, the anchor target, and the convention are unchanged. Also stop quoting the guide's subsection title inline in the nextest comment; naming the section rather than reproducing its heading reads the same and does not invite drift. --- .config/nextest.toml | 17 +++++++++-------- docs/developers-guide.md | 3 ++- 2 files changed, 11 insertions(+), 9 deletions(-) diff --git a/.config/nextest.toml b/.config/nextest.toml index 7d1bb8115..27cf28d94 100644 --- a/.config/nextest.toml +++ b/.config/nextest.toml @@ -67,12 +67,13 @@ success-output = "immediate" # default budget with room to spare: 4.2s and 7.6s on runs 34075197897 and # 34079222917, against a 244.0s median before. # -# Removing that second Cargo build also unblocked this one. The two used to run -# concurrently, each with four compile jobs on a four-vCPU runner, and each -# roughly halved the other; on those same runs this test took 125.3s, 170.6s -# and 170.7s rather than the 274.7s median above. 420s is therefore sized -# against the older, contended distribution and is deliberately conservative -# while the new shape has three samples. +# Removing that second Cargo build also unblocked this one. Contention had been +# increasing each test's elapsed time: the two ran concurrently, each with four +# compile jobs on a four-vCPU runner, and each roughly halved the other. With +# the second build gone from the Windows lane, this test took 125.3s, 170.6s +# and 170.7s on the three runs that followed, rather than the 274.7s median +# above. 420s is therefore sized against the older, contended distribution and +# is deliberately conservative while the new shape has three samples. # # Trimming the test is worth about 85s, not the 156s that figure implied # before the second build left. What a trim returns is the test's exclusive @@ -82,8 +83,8 @@ success-output = "immediate" # from a 1468s median to about 848s across the three changes, so 85s is # roughly ten percent of what remains. The repository has decided not to spend # the fidelity risk on that number yet: the trim is a candidate for tightening -# or deletion at the ten-run revisit gate that "Deferring the split-build-dir -# harness trim" in docs/developers-guide.md defines. That section also records +# or deletion at the ten-run revisit gate that the developers' guide defines +# under "Deferring the split-build-dir harness trim". That section also records # the alternatives already measured and rejected — `cargo check` for # `cargo build`, warming the compiler cache, and sharing a target directory — # and the constraints a fixture-crate replacement would have to preserve. diff --git a/docs/developers-guide.md b/docs/developers-guide.md index f59b5156c..56b7f5bd9 100644 --- a/docs/developers-guide.md +++ b/docs/developers-guide.md @@ -2912,6 +2912,7 @@ Linux runner. Both platforms still assert the packaged file list. See [Windows budget for the isolated-Cargo-build tests][windows-test-budget]. [windows-test-budget]: #windows-budget-for-the-isolated-cargo-build-tests +[fixture-constraints]: #what-a-fixture-crate-replacement-would-have-to-preserve `tests/workflow_contracts/test_execution_coverage_test.py` holds all of this: the coverage inputs, the denied warnings, the doctest pass and its position, @@ -3072,7 +3073,7 @@ from three runs. If its tail has settled below the 85s measured here, or if `Test` has stopped being the lane's critical path, the trim is not worth the fidelity risk and the work closes without it. Any replacement built at that point inherits the constraints in -[what a fixture-crate replacement would have to preserve](#what-a-fixture-crate-replacement-would-have-to-preserve). +[what a fixture-crate replacement would have to preserve][fixture-constraints]. Four references sit behind the figures above: [#673](https://github.com/leynos/netsuke/issues/673) measured the lane, From ede2819b74e6e8eeae7474e6e7ef3ceb3240ddc2 Mon Sep 17 00:00:00 2001 From: leynos Date: Fri, 18 Sep 2026 23:31:24 +0200 Subject: [PATCH 04/17] Record the serialization group in the deferred-trim evidence (#693) Main's nested-cargo-builds test group serializes this test with the other build-capable child Cargo tests. It landed after the three runs the deferred trim was sized from, so those samples describe a shape that no longer exists. Under the group the test holds the single slot and everything behind it waits, so a trim returns its whole occupancy rather than its exclusive tail: 113s to 152s on the three Windows runs after the group landed, against the 85s the uncontended reading gave. The decision to defer still stands and the trim remains open at the ten-run gate, now against the serialized shape. Recorded in the developers' guide decision record, the nextest.toml budget comment, and the harness test's doc comment. --- .config/nextest.toml | 32 ++++++++----- docs/developers-guide.md | 86 +++++++++++++++++++++++++++++++---- tests/locale_stub_ui_tests.rs | 6 +++ 3 files changed, 102 insertions(+), 22 deletions(-) diff --git a/.config/nextest.toml b/.config/nextest.toml index 27cf28d94..b7e472cbf 100644 --- a/.config/nextest.toml +++ b/.config/nextest.toml @@ -76,18 +76,26 @@ success-output = "immediate" # is deliberately conservative while the new shape has three samples. # # Trimming the test is worth about 85s, not the 156s that figure implied -# before the second build left. What a trim returns is the test's exclusive -# tail, the period after every other test has reported, measured at 62.3s, -# 87.2s and 85.5s on the three runs after that change, where 156s was the -# implied value of trimming both tests while they contended. The lane fell -# from a 1468s median to about 848s across the three changes, so 85s is -# roughly ten percent of what remains. The repository has decided not to spend -# the fidelity risk on that number yet: the trim is a candidate for tightening -# or deletion at the ten-run revisit gate that the developers' guide defines -# under "Deferring the split-build-dir harness trim". That section also records -# the alternatives already measured and rejected — `cargo check` for -# `cargo build`, warming the compiler cache, and sharing a target directory — -# and the constraints a fixture-crate replacement would have to preserve. +# before the second build left, on the uncontended reading those three runs +# support: what a trim returns there is the test's exclusive tail, the period +# after every other test has reported, measured at 62.3s, 87.2s and 85.5s, +# where 156s was the implied value of trimming both tests while they +# contended. The lane fell from a 1468s median to about 848s across the three +# changes, so 85s is roughly ten percent of what remains. +# +# That figure has since moved. This test is a member of `nested-cargo-builds` +# above, which serializes it with the other build-capable child Cargo tests, so +# it holds the group's single slot and every member behind it waits. Trimming +# it therefore returns its whole group occupancy rather than its exclusive +# tail: 113s, 149s and 152s on the three Windows runs after the group landed, +# against the 85s above, which was measured before it. The repository has still +# decided not to spend the fidelity risk yet, and the trim remains a candidate +# for tightening or deletion at the ten-run revisit gate that the developers' +# guide defines under "Deferring the split-build-dir harness trim"; that +# section carries the serialized measurements, the alternatives already +# measured and rejected — `cargo check` for `cargo build`, warming the compiler +# cache, and sharing a target directory — and the constraints a fixture-crate +# replacement would have to preserve. # # Windows only. On Ubicloud both tests finish in a fraction of the budget, and # widening the timeout there would blunt the hang detection this file exists to diff --git a/docs/developers-guide.md b/docs/developers-guide.md index 56b7f5bd9..100b5c7c4 100644 --- a/docs/developers-guide.md +++ b/docs/developers-guide.md @@ -2913,6 +2913,7 @@ Linux runner. Both platforms still assert the packaged file list. See [windows-test-budget]: #windows-budget-for-the-isolated-cargo-build-tests [fixture-constraints]: #what-a-fixture-crate-replacement-would-have-to-preserve +[serialized-value]: #the-nested-cargo-build-serialization-and-what-it-does-to-the-figure `tests/workflow_contracts/test_execution_coverage_test.py` holds all of this: the coverage inputs, the denied warnings, the doctest pass and its position, @@ -3016,13 +3017,20 @@ The 420s budget is therefore sized against the older, contended distribution and is deliberately conservative while the new shape has three samples. It is a candidate for tightening, or for deletion, once the revisit gate below is met. +Those three samples predate the serialization group described below, which +lands the harness test in a group of one-at-a-time build-capable tests. Under +that group the test is a serial link rather than a slow finisher, and the trim +is worth more than the 85s this table supports. The +[serialization subsection][serialized-value] holds the later measurements. + #### Deferring the split-build-dir harness trim The repository has decided **not** to trim -`harness_compiles_under_a_split_build_dir` yet. Trimming it to near zero is -worth about 85s now rather than the 156s an earlier reading implied, and that -smaller number is the whole reason the decision was to wait rather than to -build. +`harness_compiles_under_a_split_build_dir` yet. When the decision was taken the +trim was worth about 85s rather than the 156s an earlier reading implied, and +that smaller number was the whole reason to wait rather than to build. A +serialization group added to `.config/nextest.toml` afterwards changed the +arithmetic; the [subsection below][serialized-value] records the new value. The 156s came from the contended distribution. Before #687 the two isolated-Cargo tests ran concurrently, each with four compile jobs on a @@ -3066,13 +3074,71 @@ rejected on its own evidence rather than on preference: the `#[once]` fixture with `E0460`, so its build cannot reuse the lane's artefacts or the other test's. +#### The nested-Cargo-build serialization, and what it does to the figure + +A later change to `.config/nextest.toml` — `nested-cargo-builds`, a +`[test-groups]` entry with `max-threads = 1` — puts this test in a group with +the other fourteen tests that spawn a build-capable child Cargo command, so +only one of them runs at a time. It landed for the coverage lane's benefit: +four nextest workers each starting a four-job child Cargo build on a four-vCPU +runner is what the group exists to prevent. The Windows lane runs `make test` +with no `NEXTEST_PROFILE`, so it selects `[profile.default]` and inherits the +same group. + +That changes the figure this section was written around, and not in the +direction the deferral assumed. The 85s was measured on three runs from +2026-09-07; the group landed on 2026-09-17, so those samples predate it. Under +the group the test is no longer merely a slow finisher: it is a **serial +link**. Every other member is blocked while it holds the single slot, and the +chain cannot finish until it releases it. Removing it therefore returns its +whole group occupancy, not just its exclusive tail. + +Measured on the three Windows runs after the group landed, from the same +`build-test-windows` job logs: + +Table: the harness test as a serialized group member, after the group landed. + +| Run | Test duration | Chain end without it | Trim returns | +| ----------- | ------------- | -------------------- | ------------ | +| 35266003414 | 152.0s | 162.4s | 152.0s | +| 35266979317 | 149.0s | 178.3s | 149.0s | +| 35272793454 | 124.7s | 134.9s | 113.4s | + +The last column is the whole-run saving: the run's own end, less where the run +would end with the harness's occupancy removed from the group chain. It is 113s +to 152s, against the 85s the uncontended reading gave. In two of the three runs +the harness is not the last test to finish — the group's cheap tail members are +— but trimming it still returns its occupancy, because the tail cannot start +until the slot frees. + +The decision recorded above still stands, and this is a change to the evidence +for it, not to the decision: the trim remains deferred, and it remains a +question about fidelity rather than about seconds. What changes is that the +number is again large enough to be worth arguing about, so the revisit gate +below is now the thing that settles it rather than a formality. Two cautions +belong with the table. Three Windows runs under the group is a small sample, +and the group's own scheduling — not the test alone — produces the chain ends, +so the figures above are readings of a serialized system rather than isolated +measurements of the test. + +Anything that removes this test from the lane also removes the group's heaviest +member, which shortens the chain for every other member behind it. A +replacement that is cheap but still build-capable would return most of the +figure; one that stops spawning a child Cargo build at all would be lighter +still, and would lose the [response-file pressure][fixture-constraints] that +makes the group entry necessary in the first place. + **Revisit gate.** This defers the trim; it does not close it. Wait until ten -runs of the split Windows lane exist, so the harness test's share of the -`build-test-windows` job is known under the new shape rather than estimated -from three runs. If its tail has settled below the 85s measured here, or if -`Test` has stopped being the lane's critical path, the trim is not worth the -fidelity risk and the work closes without it. Any replacement built at that -point inherits the constraints in +runs of the split Windows lane exist under the serialization group, so the +harness test's share of the `build-test-windows` job is known under the shape +that now exists rather than estimated from three runs. The gate was written +against the uncontended reading and the criterion has moved with the evidence: +with the group in place the test holds a serial slot, so the question is no +longer whether its exclusive tail has settled below 85s — it plainly has not — +but whether the run still ends when it ends. If the group chain stops being +what bounds the run, or if `Test` stops being the lane's critical path, the +trim is not worth the fidelity risk and the work closes without it. Any +replacement built at that point inherits the constraints in [what a fixture-crate replacement would have to preserve][fixture-constraints]. Four references sit behind the figures above: diff --git a/tests/locale_stub_ui_tests.rs b/tests/locale_stub_ui_tests.rs index 4e2ca08bd..6d65138f9 100644 --- a/tests/locale_stub_ui_tests.rs +++ b/tests/locale_stub_ui_tests.rs @@ -94,6 +94,12 @@ fn stub_env_builders_compile_under_the_same_harness( /// any fixture-crate replacement would owe and the response-file pressure it /// would have to keep. /// +/// This test is a member of the `nested-cargo-builds` nextest group, so on +/// Windows it holds that group's single slot: every other build-capable test +/// waits for it, and a trim would return its whole occupancy rather than only +/// the tail it finishes on. The group's measurements are in the same +/// developers' guide section. +/// /// It is also what keeps the Windows response-file path exercised. The long /// `-L dependency=` set this build produces, plus the long temporary roots the /// test adds, is why `TestSupportRlib::compile` sends its arguments through a From f695757a2f3423de8ea735b80f86641ddcf09c09 Mon Sep 17 00:00:00 2001 From: leynos Date: Fri, 18 Sep 2026 23:36:26 +0200 Subject: [PATCH 05/17] Count the group without a fixed number (#693) Main added a member to the nested-cargo-builds group while this branch was open, so the prose count of the other members was already stale by one. Name the membership instead of counting it, so a later addition to the group does not silently falsify the sentence. The group's contract test discovers members rather than listing only the pinned ones, which is why it stays correct through such an addition. --- docs/developers-guide.md | 12 ++++++------ 1 file changed, 6 insertions(+), 6 deletions(-) diff --git a/docs/developers-guide.md b/docs/developers-guide.md index 100b5c7c4..8434118f7 100644 --- a/docs/developers-guide.md +++ b/docs/developers-guide.md @@ -3078,12 +3078,12 @@ rejected on its own evidence rather than on preference: A later change to `.config/nextest.toml` — `nested-cargo-builds`, a `[test-groups]` entry with `max-threads = 1` — puts this test in a group with -the other fourteen tests that spawn a build-capable child Cargo command, so -only one of them runs at a time. It landed for the coverage lane's benefit: -four nextest workers each starting a four-job child Cargo build on a four-vCPU -runner is what the group exists to prevent. The Windows lane runs `make test` -with no `NEXTEST_PROFILE`, so it selects `[profile.default]` and inherits the -same group. +the other tests that spawn a build-capable child Cargo command, so only one of +them runs at a time. It landed for the coverage lane's benefit: four nextest +workers each starting a four-job child Cargo build on a four-vCPU runner is +what the group exists to prevent. The Windows lane runs `make test` with no +`NEXTEST_PROFILE`, so it selects `[profile.default]` and inherits the same +group. That changes the figure this section was written around, and not in the direction the deferral assumed. The 85s was measured on three runs from From 7219c504f453654daf2fd276175d8c9c06723069 Mon Sep 17 00:00:00 2001 From: leynos Date: Fri, 18 Sep 2026 23:57:36 +0200 Subject: [PATCH 06/17] Correct the serialized-trim figures against the run logs (#693) Re-deriving every figure from the three Windows job logs showed three claims in the record were not what the logs support. The harness is not the last test to finish in any of the three runs; the prose said "two of the three". The group's cheap tail members trail it by under seven seconds in each case. The saving is not the group occupancy in every run. It is the run's end less whichever of the trimmed group chain and the last non-group test finishes later: 152.0s, 149.0s and 113.4s, where only the first two equal the occupancy and the third is capped because non-group work becomes the binding constraint once the harness is gone. The table's middle column is the group chain's end with the trim applied, not the run's end, so it is now named that way. All nine figures in the table reproduce from the logs; the changes are to what the prose claimed about them. --- .config/nextest.toml | 8 +++++--- docs/developers-guide.md | 27 +++++++++++++++------------ 2 files changed, 20 insertions(+), 15 deletions(-) diff --git a/.config/nextest.toml b/.config/nextest.toml index b7e472cbf..cd863ed01 100644 --- a/.config/nextest.toml +++ b/.config/nextest.toml @@ -86,9 +86,11 @@ success-output = "immediate" # That figure has since moved. This test is a member of `nested-cargo-builds` # above, which serializes it with the other build-capable child Cargo tests, so # it holds the group's single slot and every member behind it waits. Trimming -# it therefore returns its whole group occupancy rather than its exclusive -# tail: 113s, 149s and 152s on the three Windows runs after the group landed, -# against the 85s above, which was measured before it. The repository has still +# it therefore returns its group occupancy rather than its exclusive tail: +# 152s, 149s and 113s on the three Windows runs after the group landed, against +# the 85s above, which was measured before it. The last of those is below the +# test's own 124.7s duration because a trim there leaves non-group work as the +# run's next binding constraint. The repository has still # decided not to spend the fidelity risk yet, and the trim remains a candidate # for tightening or deletion at the ten-run revisit gate that the developers' # guide defines under "Deferring the split-build-dir harness trim"; that diff --git a/docs/developers-guide.md b/docs/developers-guide.md index 8434118f7..fb25c6f4d 100644 --- a/docs/developers-guide.md +++ b/docs/developers-guide.md @@ -3098,18 +3098,21 @@ Measured on the three Windows runs after the group landed, from the same Table: the harness test as a serialized group member, after the group landed. -| Run | Test duration | Chain end without it | Trim returns | -| ----------- | ------------- | -------------------- | ------------ | -| 35266003414 | 152.0s | 162.4s | 152.0s | -| 35266979317 | 149.0s | 178.3s | 149.0s | -| 35272793454 | 124.7s | 134.9s | 113.4s | - -The last column is the whole-run saving: the run's own end, less where the run -would end with the harness's occupancy removed from the group chain. It is 113s -to 152s, against the 85s the uncontended reading gave. In two of the three runs -the harness is not the last test to finish — the group's cheap tail members are -— but trimming it still returns its occupancy, because the tail cannot start -until the slot frees. +| Run | Test duration | Group chain end, trim applied | Trim returns | +| ----------- | ------------- | ----------------------------- | ------------ | +| 35266003414 | 152.0s | 162.4s | 152.0s | +| 35266979317 | 149.0s | 178.3s | 149.0s | +| 35272793454 | 124.7s | 134.9s | 113.4s | + +The rightmost column is the whole-run saving: the run's own end, less whichever +of the trimmed group chain and the last non-group test finishes later. It is +113s to 152s, against the 85s the uncontended reading gave. In none of the +three runs is the harness the last test to finish — the group's cheap tail +members trail it by under seven seconds — but a trim returns its occupancy +rather than its exclusive tail, because the tail cannot start until the slot +frees: in the first two runs the figure is that occupancy exactly, and in the +third it is 113s rather than 125s only because unrelated non-group work becomes +the run's next binding constraint once the harness is gone. The decision recorded above still stands, and this is a change to the evidence for it, not to the decision: the trim remains deferred, and it remains a From 69d460a299a421d74fa4cdfac5556b68425b2668 Mon Sep 17 00:00:00 2001 From: leynos Date: Sat, 19 Sep 2026 00:46:54 +0200 Subject: [PATCH 07/17] Record the fourth post-group sample for the harness trim (#693) The guide stated that trimming `harness_compiles_under_a_split_build_dir` returned 113s to 152s, from three Windows runs under the `nested-cargo-builds` group. This pull request's own Windows lane supplies a fourth: run 35400200137, the harness held the group's slot for 115.7s and a trim there returns 95.0s, below the low end of the published range. Add the row, widen the range to 95s to 152s, and correct the per-run analysis, which claimed the figure was the occupancy exactly in one run and short of it in one. It is short in the last two: 113s against 125s, and 95s against 116s, both because unrelated non-group work becomes the run's next binding constraint once the harness is gone. The four samples are also not uniform, so say so: the first three are trunk pushes and this one is a pull-request lane. The deferral decision is unchanged; only the evidence for it moves. Both files keep their existing shape, and `.config/nextest.toml` remains a comment-only change. Co-Authored-By: Claude Code --- .config/nextest.toml | 8 ++++---- docs/developers-guide.md | 30 +++++++++++++++++------------- 2 files changed, 21 insertions(+), 17 deletions(-) diff --git a/.config/nextest.toml b/.config/nextest.toml index cd863ed01..d8a97b302 100644 --- a/.config/nextest.toml +++ b/.config/nextest.toml @@ -87,10 +87,10 @@ success-output = "immediate" # above, which serializes it with the other build-capable child Cargo tests, so # it holds the group's single slot and every member behind it waits. Trimming # it therefore returns its group occupancy rather than its exclusive tail: -# 152s, 149s and 113s on the three Windows runs after the group landed, against -# the 85s above, which was measured before it. The last of those is below the -# test's own 124.7s duration because a trim there leaves non-group work as the -# run's next binding constraint. The repository has still +# 152s, 149s, 113s and 95s on the four Windows runs after the group landed, +# against the 85s above, which was measured before it. The last two are below +# the test's own 124.7s and 115.7s durations because a trim there leaves +# non-group work as the run's next binding constraint. The repository has still # decided not to spend the fidelity risk yet, and the trim remains a candidate # for tightening or deletion at the ten-run revisit gate that the developers' # guide defines under "Deferring the split-build-dir harness trim"; that diff --git a/docs/developers-guide.md b/docs/developers-guide.md index fb25c6f4d..bd5e77397 100644 --- a/docs/developers-guide.md +++ b/docs/developers-guide.md @@ -3093,7 +3093,7 @@ link**. Every other member is blocked while it holds the single slot, and the chain cannot finish until it releases it. Removing it therefore returns its whole group occupancy, not just its exclusive tail. -Measured on the three Windows runs after the group landed, from the same +Measured on the four Windows runs after the group landed, from the same `build-test-windows` job logs: Table: the harness test as a serialized group member, after the group landed. @@ -3103,26 +3103,30 @@ Table: the harness test as a serialized group member, after the group landed. | 35266003414 | 152.0s | 162.4s | 152.0s | | 35266979317 | 149.0s | 178.3s | 149.0s | | 35272793454 | 124.7s | 134.9s | 113.4s | +| 35400200137 | 115.7s | 154.1s | 95.0s | The rightmost column is the whole-run saving: the run's own end, less whichever of the trimmed group chain and the last non-group test finishes later. It is -113s to 152s, against the 85s the uncontended reading gave. In none of the -three runs is the harness the last test to finish — the group's cheap tail -members trail it by under seven seconds — but a trim returns its occupancy -rather than its exclusive tail, because the tail cannot start until the slot -frees: in the first two runs the figure is that occupancy exactly, and in the -third it is 113s rather than 125s only because unrelated non-group work becomes -the run's next binding constraint once the harness is gone. +95s to 152s, against the 85s the uncontended reading gave. In none of the four +runs is the harness the last test to finish — the group's cheap tail members +trail it by under seven seconds — but a trim returns its occupancy rather than +its exclusive tail, because the tail cannot start until the slot frees: in the +first two runs the figure is that occupancy exactly, and in the last two it +falls short of it — 113s against 125s, and 95s against 116s — because unrelated +non-group work becomes the run's next binding constraint once the harness is +gone. The decision recorded above still stands, and this is a change to the evidence for it, not to the decision: the trim remains deferred, and it remains a question about fidelity rather than about seconds. What changes is that the number is again large enough to be worth arguing about, so the revisit gate below is now the thing that settles it rather than a formality. Two cautions -belong with the table. Three Windows runs under the group is a small sample, -and the group's own scheduling — not the test alone — produces the chain ends, -so the figures above are readings of a serialized system rather than isolated -measurements of the test. +belong with the table. Four Windows runs under the group is a small sample, and +it is not a uniform one: the first three are trunk pushes and the fourth is a +pull-request lane, which starts from a different tree state. The group's own +scheduling — not the test alone — produces the chain ends, so the figures above +are readings of a serialized system rather than isolated measurements of the +test. Anything that removes this test from the lane also removes the group's heaviest member, which shortens the chain for every other member behind it. A @@ -3134,7 +3138,7 @@ makes the group entry necessary in the first place. **Revisit gate.** This defers the trim; it does not close it. Wait until ten runs of the split Windows lane exist under the serialization group, so the harness test's share of the `build-test-windows` job is known under the shape -that now exists rather than estimated from three runs. The gate was written +that now exists rather than estimated from four runs. The gate was written against the uncontended reading and the criterion has moved with the evidence: with the group in place the test holds a serial slot, so the question is no longer whether its exclusive tail has settled below 85s — it plainly has not — From 8a989b30efcda36554ee5743073f1b7d8ecd994c Mon Sep 17 00:00:00 2001 From: leynos Date: Sat, 19 Sep 2026 01:16:11 +0200 Subject: [PATCH 08/17] Anchor the post-group sample instead of counting it (#693) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The post-group write-up counted its own samples: "the four runs", "the last two", "estimated from five". Every push to this branch starts another Windows run, so each count was stale by the time the next one landed, and the fifth run — 35403273264, occupancy 141.6s, trim returns 136.3s — left three of them wrong. Anchor the claims rather than re-counting them. The table gains the fifth run, the range stays 95s to 152s, and the prose now says across-the-sample and names the date the sample was taken. The "falls short of occupancy" case is stated as a pattern rather than as an ordinal claim about which runs did it, since it has now happened in three of the five. The caution paragraph notes the sample mixes trunk pushes with pull-request lanes and is a snapshot rather than a running total, so later runs belong to the revisit gate rather than to this table. The revisit gate now reads "the sample above" instead of a figure, and the nextest.toml range carries the same 2026-09-18 date. The deferral decision is unchanged. `.config/nextest.toml` remains a comment-only change. Co-Authored-By: Claude Code --- .config/nextest.toml | 22 +++++++++--------- docs/developers-guide.md | 50 +++++++++++++++++++++------------------- 2 files changed, 37 insertions(+), 35 deletions(-) diff --git a/.config/nextest.toml b/.config/nextest.toml index d8a97b302..e3fd60438 100644 --- a/.config/nextest.toml +++ b/.config/nextest.toml @@ -87,17 +87,17 @@ success-output = "immediate" # above, which serializes it with the other build-capable child Cargo tests, so # it holds the group's single slot and every member behind it waits. Trimming # it therefore returns its group occupancy rather than its exclusive tail: -# 152s, 149s, 113s and 95s on the four Windows runs after the group landed, -# against the 85s above, which was measured before it. The last two are below -# the test's own 124.7s and 115.7s durations because a trim there leaves -# non-group work as the run's next binding constraint. The repository has still -# decided not to spend the fidelity risk yet, and the trim remains a candidate -# for tightening or deletion at the ten-run revisit gate that the developers' -# guide defines under "Deferring the split-build-dir harness trim"; that -# section carries the serialized measurements, the alternatives already -# measured and rejected — `cargo check` for `cargo build`, warming the compiler -# cache, and sharing a target directory — and the constraints a fixture-crate -# replacement would have to preserve. +# 95s to 152s across the Windows runs sampled on 2026-09-18, against the 85s +# above, which was measured before it, and at the low end of that range only +# because a trim there leaves non-group work as the run's next binding +# constraint. The repository has still decided not to spend the fidelity risk +# yet, and the trim remains a candidate for tightening or deletion at the +# ten-run revisit gate that the developers' guide defines under "Deferring the +# split-build-dir harness trim"; that section carries the serialized +# measurements, the alternatives already measured and rejected — `cargo check` +# for `cargo build`, warming the compiler cache, and sharing a target +# directory — and the constraints a fixture-crate replacement would have to +# preserve. # # Windows only. On Ubicloud both tests finish in a fraction of the budget, and # widening the timeout there would blunt the hang detection this file exists to diff --git a/docs/developers-guide.md b/docs/developers-guide.md index bd5e77397..af4c8963a 100644 --- a/docs/developers-guide.md +++ b/docs/developers-guide.md @@ -3093,8 +3093,8 @@ link**. Every other member is blocked while it holds the single slot, and the chain cannot finish until it releases it. Removing it therefore returns its whole group occupancy, not just its exclusive tail. -Measured on the four Windows runs after the group landed, from the same -`build-test-windows` job logs: +Measured from the same `build-test-windows` job logs, over the Windows runs +available on 2026-09-18: Table: the harness test as a serialized group member, after the group landed. @@ -3104,29 +3104,31 @@ Table: the harness test as a serialized group member, after the group landed. | 35266979317 | 149.0s | 178.3s | 149.0s | | 35272793454 | 124.7s | 134.9s | 113.4s | | 35400200137 | 115.7s | 154.1s | 95.0s | +| 35403273264 | 141.6s | 154.4s | 136.3s | The rightmost column is the whole-run saving: the run's own end, less whichever -of the trimmed group chain and the last non-group test finishes later. It is -95s to 152s, against the 85s the uncontended reading gave. In none of the four -runs is the harness the last test to finish — the group's cheap tail members -trail it by under seven seconds — but a trim returns its occupancy rather than -its exclusive tail, because the tail cannot start until the slot frees: in the -first two runs the figure is that occupancy exactly, and in the last two it -falls short of it — 113s against 125s, and 95s against 116s — because unrelated -non-group work becomes the run's next binding constraint once the harness is -gone. +of the trimmed group chain and the last non-group test finishes later. Across +the sample it is 95s to 152s, against the 85s the uncontended reading gave. The +harness is never the last test to finish — the group's cheap tail members trail +it by under seven seconds in every run — but a trim returns its occupancy +rather than its exclusive tail, because the tail cannot start until the slot +frees. Where the figure matches that occupancy exactly, the shortened chain is +still what bounds the run. Where it falls short — 113s against 125s, 95s +against 116s, and 136s against 142s — unrelated non-group work becomes the +run's next binding constraint once the harness is gone. The decision recorded above still stands, and this is a change to the evidence for it, not to the decision: the trim remains deferred, and it remains a question about fidelity rather than about seconds. What changes is that the number is again large enough to be worth arguing about, so the revisit gate below is now the thing that settles it rather than a formality. Two cautions -belong with the table. Four Windows runs under the group is a small sample, and -it is not a uniform one: the first three are trunk pushes and the fourth is a -pull-request lane, which starts from a different tree state. The group's own -scheduling — not the test alone — produces the chain ends, so the figures above -are readings of a serialized system rather than isolated measurements of the -test. +belong with the table. The sample is small and it is not a uniform one: it +mixes trunk pushes with pull-request lanes, which start from different tree +states. It is also a snapshot rather than a running total — it names the runs +it was taken from and does not claim to be current, and later runs belong to +the revisit gate below. And the group's own scheduling, not the test alone, +produces the chain ends, so the figures above are readings of a serialized +system rather than isolated measurements of the test. Anything that removes this test from the lane also removes the group's heaviest member, which shortens the chain for every other member behind it. A @@ -3138,13 +3140,13 @@ makes the group entry necessary in the first place. **Revisit gate.** This defers the trim; it does not close it. Wait until ten runs of the split Windows lane exist under the serialization group, so the harness test's share of the `build-test-windows` job is known under the shape -that now exists rather than estimated from four runs. The gate was written -against the uncontended reading and the criterion has moved with the evidence: -with the group in place the test holds a serial slot, so the question is no -longer whether its exclusive tail has settled below 85s — it plainly has not — -but whether the run still ends when it ends. If the group chain stops being -what bounds the run, or if `Test` stops being the lane's critical path, the -trim is not worth the fidelity risk and the work closes without it. Any +that now exists rather than estimated from the sample above. The gate was +written against the uncontended reading and the criterion has moved with the +evidence: with the group in place the test holds a serial slot, so the question +is no longer whether its exclusive tail has settled below 85s — it plainly has +not — but whether the run still ends when it ends. If the group chain stops +being what bounds the run, or if `Test` stops being the lane's critical path, +the trim is not worth the fidelity risk and the work closes without it. Any replacement built at that point inherits the constraints in [what a fixture-crate replacement would have to preserve][fixture-constraints]. From 727d827a0594b63a74c9c8db493f85db57b8ff64 Mon Sep 17 00:00:00 2001 From: leynos Date: Sat, 19 Sep 2026 01:39:20 +0200 Subject: [PATCH 09/17] Wrap the serialized-value link definition (#693) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The definition line ran to 88 columns. markdownlint's MD013 does not flag it — it permits an over-length line when nothing follows the last space before the limit, and this line's only space sits at column 19 — but the repository already wraps long definitions this way, at `[github-actions-validation-test]` further up the same file. Put the fragment on an indented continuation line. The reference still resolves: every `][serialized-value]` use has a matching definition, and no definition is left unused. Co-Authored-By: Claude Code --- docs/developers-guide.md | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/docs/developers-guide.md b/docs/developers-guide.md index af4c8963a..6979c6fe3 100644 --- a/docs/developers-guide.md +++ b/docs/developers-guide.md @@ -2913,7 +2913,8 @@ Linux runner. Both platforms still assert the packaged file list. See [windows-test-budget]: #windows-budget-for-the-isolated-cargo-build-tests [fixture-constraints]: #what-a-fixture-crate-replacement-would-have-to-preserve -[serialized-value]: #the-nested-cargo-build-serialization-and-what-it-does-to-the-figure +[serialized-value]: + #the-nested-cargo-build-serialization-and-what-it-does-to-the-figure `tests/workflow_contracts/test_execution_coverage_test.py` holds all of this: the coverage inputs, the denied warnings, the doctest pass and its position, From f43183d809b7eecd4331c77c83c10b7cf0ca99f6 Mon Sep 17 00:00:00 2001 From: leynos Date: Sat, 19 Sep 2026 02:25:28 +0200 Subject: [PATCH 10/17] Anchor the trim figure to its mechanism, not a running range (#693) The seventh Windows sample returned 162.9s, above the 152s upper bound the guide had published. The bound was the wrong shape: it read like a settled range when it was only a running maximum over a sample that keeps growing, since every push spawns a fresh Windows run. Replace it with the mechanism. A trim can never return more than the test's own duration, because the duration *is* the occupancy it gives back; and it returns exactly that whenever the shortened group chain is still what bounds the run. It returns less only when unrelated non-group work binds first. The figure therefore tracks the test's own cost rather than converging on the 85s exclusive tail, which belonged to an uncontended lane that no longer exists. The claim "four of the seven runs above" required the `35405043577` row, which the table was missing; added, and the count verified programmatically. Both the guide and the nextest comment are dated to the 2026-09-18 sample so later runs accrue to the revisit gate rather than falsifying a bound. --- .config/nextest.toml | 18 +++++++++-------- docs/developers-guide.md | 43 ++++++++++++++++++++++++++-------------- 2 files changed, 38 insertions(+), 23 deletions(-) diff --git a/.config/nextest.toml b/.config/nextest.toml index e3fd60438..0ef66016a 100644 --- a/.config/nextest.toml +++ b/.config/nextest.toml @@ -86,14 +86,16 @@ success-output = "immediate" # That figure has since moved. This test is a member of `nested-cargo-builds` # above, which serializes it with the other build-capable child Cargo tests, so # it holds the group's single slot and every member behind it waits. Trimming -# it therefore returns its group occupancy rather than its exclusive tail: -# 95s to 152s across the Windows runs sampled on 2026-09-18, against the 85s -# above, which was measured before it, and at the low end of that range only -# because a trim there leaves non-group work as the run's next binding -# constraint. The repository has still decided not to spend the fidelity risk -# yet, and the trim remains a candidate for tightening or deletion at the -# ten-run revisit gate that the developers' guide defines under "Deferring the -# split-build-dir harness trim"; that section carries the serialized +# it therefore returns its group occupancy rather than its exclusive tail, and +# the test's own duration caps that figure, because the duration is the +# occupancy it gives back. The trim returns exactly the duration whenever the +# shortened group chain still bounds the run. The Windows runs sampled on +# 2026-09-18 put the figure between 95s and 163s, against the 85s above, which +# was the exclusive tail under an uncontended lane that no longer exists. The +# repository has still decided not to spend the fidelity risk yet, and the trim +# remains a candidate for tightening or deletion at the ten-run revisit gate +# that the developers' guide defines under "Deferring the split-build-dir +# harness trim"; that section carries the serialized # measurements, the alternatives already measured and rejected — `cargo check` # for `cargo build`, warming the compiler cache, and sharing a target # directory — and the constraints a fixture-crate replacement would have to diff --git a/docs/developers-guide.md b/docs/developers-guide.md index 6979c6fe3..dc16ad303 100644 --- a/docs/developers-guide.md +++ b/docs/developers-guide.md @@ -3095,7 +3095,9 @@ chain cannot finish until it releases it. Removing it therefore returns its whole group occupancy, not just its exclusive tail. Measured from the same `build-test-windows` job logs, over the Windows runs -available on 2026-09-18: +available on 2026-09-18. The rows are a snapshot of those runs, not a running +total; what generalizes past them is the mechanism below, which is why the +figure is stated as a rule with the sample as its evidence: Table: the harness test as a serialized group member, after the group landed. @@ -3106,17 +3108,28 @@ Table: the harness test as a serialized group member, after the group landed. | 35272793454 | 124.7s | 134.9s | 113.4s | | 35400200137 | 115.7s | 154.1s | 95.0s | | 35403273264 | 141.6s | 154.4s | 136.3s | +| 35405043577 | 141.9s | 167.1s | 141.9s | +| 35407132087 | 162.9s | 174.6s | 162.9s | The rightmost column is the whole-run saving: the run's own end, less whichever -of the trimmed group chain and the last non-group test finishes later. Across -the sample it is 95s to 152s, against the 85s the uncontended reading gave. The -harness is never the last test to finish — the group's cheap tail members trail -it by under seven seconds in every run — but a trim returns its occupancy -rather than its exclusive tail, because the tail cannot start until the slot -frees. Where the figure matches that occupancy exactly, the shortened chain is -still what bounds the run. Where it falls short — 113s against 125s, 95s -against 116s, and 136s against 142s — unrelated non-group work becomes the -run's next binding constraint once the harness is gone. +of the trimmed group chain and the last non-group test finishes later. Read it +against the mechanism rather than as a distribution, because the mechanism is +what generalizes past the sample. The harness is never the last test to finish +— the group's cheap tail members trail it by under seven seconds in every run — +but a trim returns its occupancy rather than its exclusive tail, because the +tail cannot start until the slot frees. + +That gives the figure its ceiling and its shape. **A trim can never return more +than the test's own duration**, because the duration *is* the occupancy it +gives back; and it returns exactly that whenever the shortened group chain is +still what bounds the run, which is the common case — four of the seven runs +above. It returns less only when unrelated non-group work becomes the run's +next binding constraint once the harness is gone: 113s against 125s, 95s +against 116s, and 136s against 142s. So the figure tracks the test's own cost +and moves with it, which is why the sample's durations span 115.7s to 162.9s +while its savings span 95s to 163s. The 85s reading is not a floor this settles +back to — it was the exclusive tail under an uncontended lane that no longer +exists. The decision recorded above still stands, and this is a change to the evidence for it, not to the decision: the trim remains deferred, and it remains a @@ -3125,11 +3138,11 @@ number is again large enough to be worth arguing about, so the revisit gate below is now the thing that settles it rather than a formality. Two cautions belong with the table. The sample is small and it is not a uniform one: it mixes trunk pushes with pull-request lanes, which start from different tree -states. It is also a snapshot rather than a running total — it names the runs -it was taken from and does not claim to be current, and later runs belong to -the revisit gate below. And the group's own scheduling, not the test alone, -produces the chain ends, so the figures above are readings of a serialized -system rather than isolated measurements of the test. +states. So read the rule rather than the spread — the spread is what the rule +predicts, and later runs belong to the revisit gate below rather than to this +table. And the group's own scheduling, not the test alone, produces the chain +ends, so the figures above are readings of a serialized system rather than +isolated measurements of the test. Anything that removes this test from the lane also removes the group's heaviest member, which shortens the chain for every other member behind it. A From fc4954ab6f38fa7c31c504cf0bbdb2deac694ab6 Mon Sep 17 00:00:00 2001 From: leynos Date: Sat, 19 Sep 2026 02:46:07 +0200 Subject: [PATCH 11/17] State the tail invariant structurally, not as a 0.1s margin (#693) The paragraph below the post-group table claimed the group's tail members trail the harness by "under seven seconds in every run". Re-derived across all seven rows, the worst gap is 6.9s against that 7.0s bound. The claim is true, but it is the same shape of hazard that already forced two rewrites here: a hard numeric bound over a sample that grows by one run on every push. The invariant underneath it does not decay, because it follows from the group's structure rather than from the measurement -- tail members cannot start until the harness frees the single slot, so they necessarily finish after it. State that, and drop the number. The figures the paragraph exists to support are unchanged; the seven-row table, its trim-returns column and the mechanism below it were re-verified against the raw job logs while checking this. --- docs/developers-guide.md | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/docs/developers-guide.md b/docs/developers-guide.md index dc16ad303..f3a60ea37 100644 --- a/docs/developers-guide.md +++ b/docs/developers-guide.md @@ -3114,10 +3114,10 @@ Table: the harness test as a serialized group member, after the group landed. The rightmost column is the whole-run saving: the run's own end, less whichever of the trimmed group chain and the last non-group test finishes later. Read it against the mechanism rather than as a distribution, because the mechanism is -what generalizes past the sample. The harness is never the last test to finish -— the group's cheap tail members trail it by under seven seconds in every run — -but a trim returns its occupancy rather than its exclusive tail, because the -tail cannot start until the slot frees. +what generalizes past the sample. The harness is never the last test to finish: +the group's cheap tail members cannot start until it frees the slot, so they +necessarily finish after it. But a trim returns its occupancy rather than its +exclusive tail, because that same tail cannot start until the slot frees. That gives the figure its ceiling and its shape. **A trim can never return more than the test's own duration**, because the duration *is* the occupancy it From 6702226f1aaf035179e3d7c8ffbc592b97448531 Mon Sep 17 00:00:00 2001 From: leynos Date: Sat, 19 Sep 2026 03:35:21 +0200 Subject: [PATCH 12/17] Separate the group's rationale from the coverage it owes (#693) CodeRabbit found the paragraph that follows the post-group table claiming the response-file pressure "makes the group entry necessary in the first place". That conflates two independent things, and the guide says the opposite two sections earlier: `nested-cargo-builds` exists because four nextest workers each starting a four-job child Cargo build on four vCPUs is what it prevents. Response-file pressure is a coverage requirement a replacement would owe, not a reason to serialize anything. A build-capable replacement does not change the contention rationale at all, so stating them as one requirement would misdirect exactly the fixture work this section exists to constrain. Split them, and keep the response-file clause attached to the coverage obligation where the constraints section already develops it. --- docs/developers-guide.md | 6 ++++-- 1 file changed, 4 insertions(+), 2 deletions(-) diff --git a/docs/developers-guide.md b/docs/developers-guide.md index f3a60ea37..0028af8ed 100644 --- a/docs/developers-guide.md +++ b/docs/developers-guide.md @@ -3148,8 +3148,10 @@ Anything that removes this test from the lane also removes the group's heaviest member, which shortens the chain for every other member behind it. A replacement that is cheap but still build-capable would return most of the figure; one that stops spawning a child Cargo build at all would be lighter -still, and would lose the [response-file pressure][fixture-constraints] that -makes the group entry necessary in the first place. +still, but it would owe the [response-file pressure][fixture-constraints] this +build supplies, which is a coverage requirement rather than the group's +rationale. The group exists because concurrent child Cargo builds contend for +the runner, and a build-capable replacement does not change that. **Revisit gate.** This defers the trim; it does not close it. Wait until ten runs of the split Windows lane exist under the serialization group, so the From 57b3a77cfdb8aeca8cae183c0ffb549abf773250 Mon Sep 17 00:00:00 2001 From: leynos Date: Sat, 19 Sep 2026 11:37:55 +0200 Subject: [PATCH 13/17] Qualify the doc comment's occupancy claim the way the guide does (#693) The comment on `harness_compiles_under_a_split_build_dir` said a trim "would return its whole occupancy rather than only the tail it finishes on", unconditionally. The guide conditions exactly this figure: the trim returns the whole occupancy only when the shortened group chain still bounds the run, and returns less when unrelated work becomes the run's next binding constraint once the slot frees. That qualification is not decoration. It is the mechanism the guide was rewritten around, and the comment stated the stronger, unqualified form of it next to the measurements that contradict it. The two now agree. Raised as an unreproduced CodeRabbit finding on an aborted review pass. It was not corroborated on re-review, but it is correct on its merits, so it is fixed rather than dismissed on provenance. --- tests/locale_stub_ui_tests.rs | 6 ++++-- 1 file changed, 4 insertions(+), 2 deletions(-) diff --git a/tests/locale_stub_ui_tests.rs b/tests/locale_stub_ui_tests.rs index 6d65138f9..6ea639592 100644 --- a/tests/locale_stub_ui_tests.rs +++ b/tests/locale_stub_ui_tests.rs @@ -96,8 +96,10 @@ fn stub_env_builders_compile_under_the_same_harness( /// /// This test is a member of the `nested-cargo-builds` nextest group, so on /// Windows it holds that group's single slot: every other build-capable test -/// waits for it, and a trim would return its whole occupancy rather than only -/// the tail it finishes on. The group's measurements are in the same +/// waits for it, so a trim returns its whole occupancy rather than only the +/// tail it finishes on, whenever the shortened group chain still bounds the +/// run. It returns less when unrelated work becomes the run's next binding +/// constraint once the slot frees. The group's measurements are in the same /// developers' guide section. /// /// It is also what keeps the Windows response-file path exercised. The long From dcbe77c18a9ac0143535358f934458682980c72b Mon Sep 17 00:00:00 2001 From: leynos Date: Sat, 19 Sep 2026 12:09:08 +0200 Subject: [PATCH 14/17] Move the trim decision out of the guide and into ADR-027 (#693) The decision record was living in `docs/developers-guide.md` as two subsections. That is the wrong home for it. The style guide reserves the developer's guide for current responsibilities and points design rationale and trade-offs at decision records, and it asks the guide to stay synchronized with them. A deferral with a revisit gate is a decision with a status, not a description of how the suite works today. `docs/adr-027-defer-split-build-dir-harness-trim.md` now carries it: context, the options already measured and rejected, the rule that a trim can never return more than the test's own duration, the serialized-lane sample, the ten-run revisit gate, and the consequences. The guide keeps a short summary of the decision and links to the ADR; it still owns the constraints a fixture-crate replacement would have to preserve, because those are implementation requirements rather than decision rationale. The serialized-measurement table moved rather than being duplicated. Two copies of a seven-row sample that each push can date further is the drift this branch already corrected twice, so the ADR is the single owner and the guide points at it. `contents.md` gains the ADR-027 entry, and the three inbound references that named the old section -- in the guide's cross-reference block, in the Windows override comment in `.config/nextest.toml`, and in the doc comment on `harness_compiles_under_a_split_build_dir` -- now name the ADR. The nextest.toml change remains comment-only: verified structurally identical to `origin/main` under `tomllib`, and no non-comment line appears in its diff. Every figure in the ADR was re-derived programmatically from its own table: seven rows, four returning the full duration, the 115.7s to 162.9s duration span, the 95s to 163s saving span, the three short rows, and zero ceiling violations. Nine gates green on this tree: check-fmt, markdownlint, lint, typecheck, test (3196 passed, 5 skipped), doc-coverage (98.80%), and nixie. --- .config/nextest.toml | 11 +- ...-027-defer-split-build-dir-harness-trim.md | 250 ++++++++++++++++++ docs/contents.md | 3 + docs/developers-guide.md | 186 +++---------- tests/locale_stub_ui_tests.rs | 10 +- 5 files changed, 300 insertions(+), 160 deletions(-) create mode 100644 docs/adr-027-defer-split-build-dir-harness-trim.md diff --git a/.config/nextest.toml b/.config/nextest.toml index 0ef66016a..e43a9e17a 100644 --- a/.config/nextest.toml +++ b/.config/nextest.toml @@ -94,12 +94,11 @@ success-output = "immediate" # was the exclusive tail under an uncontended lane that no longer exists. The # repository has still decided not to spend the fidelity risk yet, and the trim # remains a candidate for tightening or deletion at the ten-run revisit gate -# that the developers' guide defines under "Deferring the split-build-dir -# harness trim"; that section carries the serialized -# measurements, the alternatives already measured and rejected — `cargo check` -# for `cargo build`, warming the compiler cache, and sharing a target -# directory — and the constraints a fixture-crate replacement would have to -# preserve. +# that ADR-027 defines; that record carries the serialized measurements, the +# alternatives already measured and rejected — `cargo check` for `cargo +# build`, warming the compiler cache, and sharing a target directory — and the +# ten-run gate. The developers' guide carries the constraints a fixture-crate +# replacement would have to preserve. # # Windows only. On Ubicloud both tests finish in a fraction of the budget, and # widening the timeout there would blunt the hang detection this file exists to diff --git a/docs/adr-027-defer-split-build-dir-harness-trim.md b/docs/adr-027-defer-split-build-dir-harness-trim.md new file mode 100644 index 000000000..50ae2036a --- /dev/null +++ b/docs/adr-027-defer-split-build-dir-harness-trim.md @@ -0,0 +1,250 @@ +# Architectural decision record (ADR) 027: Defer replacing the split-build-dir harness test + +## Status + +Accepted. The trim of `harness_compiles_under_a_split_build_dir` is deferred, +not closed: the Windows lane keeps the real `test_support` build and its 420s +budget, and the question reopens at the ten-run revisit gate below. + +## Date + +2026-09-17, with the measurements behind it taken through 2026-09-18. + +## Context and problem statement + +`harness_compiles_under_a_split_build_dir` is the regression test for the +Windows `CreateProcessW` command-line limit. It forces a split layout with its +own private `CARGO_TARGET_DIR` and `CARGO_BUILD_BUILD_DIR` roots, confirms the +collected dependency directories span the split, and compiles a fixture against +them. Every `-L dependency=` pair it produces is required to avoid `E0463`, so +the list cannot be shortened; it moves off the command line entirely into a +`rustc` response file. That is the behaviour under test. + +Because its roots are private — sharing the ambient target directory would race +the `#[once]` `test_support_rlib` fixture and fail with version-skew errors +(`E0460`) — the test compiles `test_support` and roughly 350 dependencies from +scratch. On the four-vCPU GitHub-hosted `windows-latest` gate, that build is +98.8% of the test's wall time, which makes it the most expensive test in the +lane and an obvious target for trimming: build a minimal fixture crate under +the split layout instead of the real `test_support`. + +This record exists because that obvious target was measured, and the +measurement did not support spending the fidelity risk when the ticket that +proposed it assumed. + +The figure has moved twice, in both cases because the system around the test +changed rather than the test itself. + +Before [#687](https://github.com/leynos/netsuke/pull/687), the two +isolated-Cargo tests ran concurrently, each with four compile jobs on a +four-vCPU runner, so each roughly halved the other. In that contended shape +trimming **both** tests was worth 156s. Once +`packaged_manifest_retains_build_script_sources` left the Windows lane, this +test got faster without being touched, and the implied value of trimming it +fell to about **85s** — measured as its exclusive tail, the period after every +other test had reported. + +A serialization group then changed the arithmetic a second time. A later change +to `.config/nextest.toml` added `nested-cargo-builds`, a `[test-groups]` entry +with `max-threads = 1`, which puts this test in a group with the other tests +that spawn a build-capable child Cargo command. It landed for the coverage +lane's benefit: four nextest workers each starting a four-job child Cargo build +on a four-vCPU runner is what the group exists to prevent. The Windows lane runs +`make test` with no `NEXTEST_PROFILE`, so it selects `[profile.default]` and +inherits the same group. + +Under that group the test is no longer merely a slow finisher. It is a **serial +link**: every other member waits while it holds the single slot, and the chain +cannot finish until it releases it. Removing it therefore returns its whole +group occupancy rather than its exclusive tail. + +The question this record answers is therefore not whether the test is expensive +— it is — but whether the fidelity risk of replacing it is currently worth +paying, given that the number has never been stable enough to plan against. + +## Decision drivers + +- The test guards a Windows-specific failure + (`Os { code: 206, kind: InvalidFilename }`) that cannot be reproduced on most + local hosts, so weakening it silently is the most expensive possible outcome. +- The figure the decision rests on has moved twice, both times because the + lane changed rather than the test. A decision taken against an unstable + number reopens on its own. +- The repository has already spent effort making this lane cheaper by other + means. The lane fell from a 1468s median to about 848s across + [#687](https://github.com/leynos/netsuke/pull/687), [#690](https://github.com/leynos/netsuke/pull/690) + and [#691](https://github.com/leynos/netsuke/issues/691), so the remaining + saving is roughly ten percent of what is left. +- Any replacement owes coverage that is not the group's rationale to supply. + The response-file pressure this build generates is a separate requirement + from the serialization the group provides. + +## Options considered + +### Option A: Replace the test with a minimal fixture crate + +Build a one-dependency crate under the split layout in place of `test_support`, +keeping the private roots and the split-directory assertion but dropping the +roughly 350-dependency compile. + +This returns most of the figure, because the test is the group's heaviest +member: removing it shortens the chain for every member behind it. It also +drops coverage in two ways, neither of which is optional. A one-dependency +fixture still exercises the split-directory derivation, but no longer covers it +at the real crate's scale, nor against the proc-macro and dynamic-library +artefacts that make the directory enumeration non-trivial. And it produces far +fewer `-L dependency=` entries, so it stops exercising the response-file path +that the test exists to protect. + +### Option B: `cargo check` instead of `cargo build` + +Timed cold at `-j 4` on a 32-core host: 114s against 102s, twelve percent. It +also writes nothing into the target directory, so the uplift the regression +exists to catch stops happening and the test passes vacuously. + +### Option C: Warm the lane's compiler cache + +The cache already reaches the spawned build, because `ci-windows.yml` sets +`RUSTC_WRAPPER` at job scope and the test adds to the child environment rather +than clearing it. Warming it is worth about three percent: 281.0s on the cold +run against a 271.9s warm median. There is no reuse left to claim. + +### Option D: Share a target directory + +The test needs private roots to avoid racing the `#[once]` fixture with +`E0460`, so its build cannot reuse the lane's artefacts or the other test's. + +### Option E: Keep the test and defer the decision to a gate + +Leave the test, its subject and its budget as they are, and record the +conditions under which the question is worth reopening. + +| Topic | A: fixture crate | B: `cargo check` | C: cache | D: shared target | E: defer | +| ------------------------- | ---------------- | ---------------- | -------- | ---------------- | -------------- | +| Returns the figure | Most of it | ~12%, then none | ~3% | None | No | +| Keeps split-dir coverage | At reduced scale | Vacuously | Yes | Yes | Yes | +| Keeps response-file path | No | No | Yes | Yes | Yes | +| Reproduction on this host | Yes | Yes | Yes | Fails `E0460` | Not applicable | + +_Table 1: Comparison of the measured options._ + +## Decision outcome + +**Option E.** The repository does not trim +`harness_compiles_under_a_split_build_dir` yet. The test keeps its real +`test_support` subject and its 420s budget, and the trimming question reopens +at the revisit gate below rather than on a schedule. + +## Rationale + +What a trim returns is the test's whole group **occupancy**, because the +duration _is_ the occupancy it gives back. That gives the figure both its +ceiling and its shape: **a trim can never return more than the test's own +duration**, and it returns exactly that whenever the shortened group chain is +still what bounds the run. It returns less only when unrelated non-group work +becomes the run's next binding constraint once the harness is gone. + +The harness is never the last test to finish. The group's cheap tail members +cannot start until it frees the slot, so they necessarily finish after it. The +mechanism is what generalizes past the sample; the sample is its evidence. +Measured from the same `build-test-windows` job logs, over the Windows runs +available on 2026-09-18: + +| Run | Test duration | Group chain end, trim applied | Trim returns | +| ----------- | ------------- | ----------------------------- | ------------ | +| 35266003414 | 152.0s | 162.4s | 152.0s | +| 35266979317 | 149.0s | 178.3s | 149.0s | +| 35272793454 | 124.7s | 134.9s | 113.4s | +| 35400200137 | 115.7s | 154.1s | 95.0s | +| 35403273264 | 141.6s | 154.4s | 136.3s | +| 35405043577 | 141.9s | 167.1s | 141.9s | +| 35407132087 | 162.9s | 174.6s | 162.9s | + +_Table 2: The harness test as a serialized group member, after the group +landed._ + +The rightmost column is the whole-run saving: the run's own end, less whichever +of the trimmed group chain and the last non-group test finishes later. Four of +the seven runs return the test's full duration; the other three return less, at +113s against 125s, 95s against 116s, and 136s against 142s. So the figure +tracks the test's own cost and moves with it, which is why the sample's +durations span 115.7s to 162.9s while its savings span 95s to 163s. + +The 85s reading is not a floor this settles back to. It was the exclusive tail +under an uncontended lane that no longer exists, and the group's arrival is +what retired it. + +Two cautions belong with that table, and they are why the decision is to defer +rather than to proceed on a larger number. The sample is small and it is not a +uniform one: it mixes trunk pushes with pull-request lanes, which start from +different tree states, so it sets an order of magnitude rather than a value. +And the group's own scheduling, not the test alone, produces the chain ends, so +those figures are readings of a serialized system rather than isolated +measurements of the test. + +The strongest argument for deferring is not the size of the number. It is that +the number has moved twice for reasons outside the test, so a decision taken +against it would be a decision taken against the lane's current shape. The +fidelity risk, by contrast, is real and one-directional: the coverage a fixture +crate would drop is exactly the coverage that fails only on Windows, where it +is least likely to be noticed. + +## Revisit gate + +Wait until **ten runs** of the split Windows lane exist under the serialization +group. That is enough for the harness test's share of the `build-test-windows` +job to be known under the shape that now exists, rather than estimated from the +sample above. + +The criterion has moved with the evidence. With the group in place the test +holds a serial slot, so the question is no longer whether its exclusive tail +has settled below 85s — it plainly has not. The question is whether the run +still ends when the test ends. If the group chain stops being what bounds the +run, or if the `Test` step stops being the lane's critical path, then the trim +is not worth the fidelity risk and the work closes without it. + +## Consequences + +- The Windows lane keeps its most expensive test, and the group keeps its + heaviest member. Every other member of `nested-cargo-builds` waits behind + that member, so the serialization the group provides costs the lane more than + the test's own duration. +- A future trim is not blocked, only deferred. Option A remains available and + its requirements are recorded in the developer's guide, which owns them: the + fidelity argument a replacement owes in a doc comment beside the test, and the + `rustc` response-file pressure it must either keep generating or move into a + dedicated test. +- The 420s budget stays sized against the older, contended distribution. It is + conservative by roughly a third against the measured 312.9s worst case, and + it remains a candidate for tightening or deletion once the gate is met. +- Any future change that removes the test from the lane also removes the + group's heaviest member, which changes every other member's scheduling. That + is a lane-wide effect, not a local one, and belongs in the decision that + takes it. +- Because the estimate is recorded as a mechanism with a dated sample rather + than a range, later runs do not falsify it. They belong to the revisit gate. + +## Related decisions + +- [ADR-011: Serial `deps` ordering via Ninja dyndep][adr-011] is the other + decision in this repository that trades throughput for ordering guarantees. +- [ADR-025: Persistent coverage data owned by `main`][adr-025] governs the + coverage lane whose contention the `nested-cargo-builds` group was added to + prevent. + +## References + +- [#673](https://github.com/leynos/netsuke/issues/673) measured the Windows + lane. +- [#687](https://github.com/leynos/netsuke/pull/687) relocated the packaging + verification build and so changed this test's cost. +- [#690](https://github.com/leynos/netsuke/pull/690) folded the native-recipe + smoke job into the gate job. +- [#691](https://github.com/leynos/netsuke/issues/691) split the Windows lints + from the tests. +- [Developer guide: Windows budget for the isolated-Cargo-build tests and what + a fixture-crate replacement would have to preserve][dev-guide]. + +[adr-011]: adr-011-use-ninja-dyndep-for-serial-dependency-ordering.md +[adr-025]: adr-025-main-owned-coverage-publication.md +[dev-guide]: developers-guide.md#what-a-fixture-crate-replacement-would-have-to-preserve diff --git a/docs/contents.md b/docs/contents.md index 7d78a7988..170dbc9bd 100644 --- a/docs/contents.md +++ b/docs/contents.md @@ -170,6 +170,9 @@ operator, user, and contributor references are easier to find. - [ADR-026](adr-026-manifest-environment-access-policy.md): Exact-name manifest environment policy evaluated before the reader, with project allow entries quarantined below the operator ceiling. +- [ADR-027](adr-027-defer-split-build-dir-harness-trim.md): Deferred trim of + the split-build-dir harness test, with the serialized-lane measurements that + made the figure unstable and the ten-run gate that reopens the question. ## Proposals diff --git a/docs/developers-guide.md b/docs/developers-guide.md index 0028af8ed..28bf6aad6 100644 --- a/docs/developers-guide.md +++ b/docs/developers-guide.md @@ -2913,8 +2913,7 @@ Linux runner. Both platforms still assert the packaged file list. See [windows-test-budget]: #windows-budget-for-the-isolated-cargo-build-tests [fixture-constraints]: #what-a-fixture-crate-replacement-would-have-to-preserve -[serialized-value]: - #the-nested-cargo-build-serialization-and-what-it-does-to-the-figure +[adr-027-trim]: adr-027-defer-split-build-dir-harness-trim.md `tests/workflow_contracts/test_execution_coverage_test.py` holds all of this: the coverage inputs, the denied warnings, the doctest pass and its position, @@ -3016,164 +3015,50 @@ Table: Windows durations before and after the verification build moved. The 420s budget is therefore sized against the older, contended distribution and is deliberately conservative while the new shape has three samples. It is a -candidate for tightening, or for deletion, once the revisit gate below is met. +candidate for tightening, or for deletion, once the [ADR-027][adr-027-trim] +revisit gate is met. Those three samples predate the serialization group described below, which lands the harness test in a group of one-at-a-time build-capable tests. Under that group the test is a serial link rather than a slow finisher, and the trim -is worth more than the 85s this table supports. The -[serialization subsection][serialized-value] holds the later measurements. +is worth more than the 85s this table supports. #### Deferring the split-build-dir harness trim The repository has decided **not** to trim -`harness_compiles_under_a_split_build_dir` yet. When the decision was taken the -trim was worth about 85s rather than the 156s an earlier reading implied, and -that smaller number was the whole reason to wait rather than to build. A -serialization group added to `.config/nextest.toml` afterwards changed the -arithmetic; the [subsection below][serialized-value] records the new value. - -The 156s came from the contended distribution. Before #687 the two -isolated-Cargo tests ran concurrently, each with four compile jobs on a -four-vCPU runner, so each roughly halved the other, and 156s was the implied -value of trimming both of them. Once -`packaged_manifest_retains_build_script_sources` left the Windows lane, as the -table above records, this test got faster without being touched and what a trim -could return fell with it. - -What a trim returns is the test's **exclusive tail**, the period after every -other test has reported, rather than its total duration, because the remaining -tests fill the run either way. Measured on the three runs after #687 — the pair -the table above draws on, plus the third under the new shape: - -Table: the split-build harness's duration and exclusive tail after #687. - -| Run | Test duration | Exclusive tail | -| ----------- | ------------- | -------------- | -| 34075197897 | 125.3s | 62.3s | -| 34079222917 | 170.6s | 87.2s | -| 34080385050 | 170.7s | 85.5s | - -The figure to plan against is therefore about 85s, not 156s, and nobody should -start this work expecting the larger one. For scale, the whole Windows lane -fell from a 1468s median to about 848s across #687, #690 and #691, so 85s is -roughly ten percent of what remains. - -Three alternatives were measured before the decision to defer, and each was -rejected on its own evidence rather than on preference: - -- **`cargo check` instead of `cargo build`.** Timed cold at `-j 4` on a 32-core - host: 114s against 102s, twelve percent. It also writes nothing into the - target directory, so the uplift the regression exists to catch stops - happening and the test passes vacuously. -- **Warming the lane's compiler cache.** The cache already reaches the spawned - build, because `ci-windows.yml` sets `RUSTC_WRAPPER` at job scope and the - test adds to the child environment rather than clearing it. Warming it is - worth about three percent: 281.0s on the cold run against a 271.9s warm - median. -- **Sharing a target directory.** The test needs private roots to avoid racing - the `#[once]` fixture with `E0460`, so its build cannot reuse the lane's - artefacts or the other test's. - -#### The nested-Cargo-build serialization, and what it does to the figure +`harness_compiles_under_a_split_build_dir` yet. The test keeps its real +`test_support` subject and its 420s budget. A later change to `.config/nextest.toml` — `nested-cargo-builds`, a -`[test-groups]` entry with `max-threads = 1` — puts this test in a group with +`[test-groups]` entry with `max-threads = 1` — put this test in a group with the other tests that spawn a build-capable child Cargo command, so only one of them runs at a time. It landed for the coverage lane's benefit: four nextest workers each starting a four-job child Cargo build on a four-vCPU runner is what the group exists to prevent. The Windows lane runs `make test` with no `NEXTEST_PROFILE`, so it selects `[profile.default]` and inherits the same -group. - -That changes the figure this section was written around, and not in the -direction the deferral assumed. The 85s was measured on three runs from -2026-09-07; the group landed on 2026-09-17, so those samples predate it. Under -the group the test is no longer merely a slow finisher: it is a **serial -link**. Every other member is blocked while it holds the single slot, and the -chain cannot finish until it releases it. Removing it therefore returns its -whole group occupancy, not just its exclusive tail. - -Measured from the same `build-test-windows` job logs, over the Windows runs -available on 2026-09-18. The rows are a snapshot of those runs, not a running -total; what generalizes past them is the mechanism below, which is why the -figure is stated as a rule with the sample as its evidence: - -Table: the harness test as a serialized group member, after the group landed. - -| Run | Test duration | Group chain end, trim applied | Trim returns | -| ----------- | ------------- | ----------------------------- | ------------ | -| 35266003414 | 152.0s | 162.4s | 152.0s | -| 35266979317 | 149.0s | 178.3s | 149.0s | -| 35272793454 | 124.7s | 134.9s | 113.4s | -| 35400200137 | 115.7s | 154.1s | 95.0s | -| 35403273264 | 141.6s | 154.4s | 136.3s | -| 35405043577 | 141.9s | 167.1s | 141.9s | -| 35407132087 | 162.9s | 174.6s | 162.9s | - -The rightmost column is the whole-run saving: the run's own end, less whichever -of the trimmed group chain and the last non-group test finishes later. Read it -against the mechanism rather than as a distribution, because the mechanism is -what generalizes past the sample. The harness is never the last test to finish: -the group's cheap tail members cannot start until it frees the slot, so they -necessarily finish after it. But a trim returns its occupancy rather than its -exclusive tail, because that same tail cannot start until the slot frees. - -That gives the figure its ceiling and its shape. **A trim can never return more -than the test's own duration**, because the duration *is* the occupancy it -gives back; and it returns exactly that whenever the shortened group chain is -still what bounds the run, which is the common case — four of the seven runs -above. It returns less only when unrelated non-group work becomes the run's -next binding constraint once the harness is gone: 113s against 125s, 95s -against 116s, and 136s against 142s. So the figure tracks the test's own cost -and moves with it, which is why the sample's durations span 115.7s to 162.9s -while its savings span 95s to 163s. The 85s reading is not a floor this settles -back to — it was the exclusive tail under an uncontended lane that no longer -exists. - -The decision recorded above still stands, and this is a change to the evidence -for it, not to the decision: the trim remains deferred, and it remains a -question about fidelity rather than about seconds. What changes is that the -number is again large enough to be worth arguing about, so the revisit gate -below is now the thing that settles it rather than a formality. Two cautions -belong with the table. The sample is small and it is not a uniform one: it -mixes trunk pushes with pull-request lanes, which start from different tree -states. So read the rule rather than the spread — the spread is what the rule -predicts, and later runs belong to the revisit gate below rather than to this -table. And the group's own scheduling, not the test alone, produces the chain -ends, so the figures above are readings of a serialized system rather than -isolated measurements of the test. - -Anything that removes this test from the lane also removes the group's heaviest -member, which shortens the chain for every other member behind it. A -replacement that is cheap but still build-capable would return most of the -figure; one that stops spawning a child Cargo build at all would be lighter -still, but it would owe the [response-file pressure][fixture-constraints] this -build supplies, which is a coverage requirement rather than the group's -rationale. The group exists because concurrent child Cargo builds contend for -the runner, and a build-capable replacement does not change that. - -**Revisit gate.** This defers the trim; it does not close it. Wait until ten -runs of the split Windows lane exist under the serialization group, so the -harness test's share of the `build-test-windows` job is known under the shape -that now exists rather than estimated from the sample above. The gate was -written against the uncontended reading and the criterion has moved with the -evidence: with the group in place the test holds a serial slot, so the question -is no longer whether its exclusive tail has settled below 85s — it plainly has -not — but whether the run still ends when it ends. If the group chain stops -being what bounds the run, or if `Test` stops being the lane's critical path, -the trim is not worth the fidelity risk and the work closes without it. Any -replacement built at that point inherits the constraints in -[what a fixture-crate replacement would have to preserve][fixture-constraints]. - -Four references sit behind the figures above: -[#673](https://github.com/leynos/netsuke/issues/673) measured the lane, -[#687](https://github.com/leynos/netsuke/pull/687) relocated the packaging -verification build and so changed this test's cost, -[#690](https://github.com/leynos/netsuke/pull/690) folded the native-recipe -smoke job into the gate job, and -[#691](https://github.com/leynos/netsuke/issues/691) split the Windows lints -from the tests. +group. Under it the test is a serial link rather than a slow finisher, and a +trim returns its whole group occupancy rather than only the exclusive tail the +85s above measures. That raises the ceiling on the saving to the test's own +duration, and it makes the saving track the test's own cost rather than an +uncontended lane's tail. + +The decision does not change with the number, and it is not the number that +settles it. The figure has moved twice already, both times because the lane +changed rather than the test, so a decision taken against it would be a +decision taken against the lane's current shape. The fidelity risk, by +contrast, is one-directional: the coverage a fixture crate would drop is +exactly the coverage that fails only on Windows, where it is least likely to be +noticed. + +[ADR-027][adr-027-trim] holds the decision itself: the measurements, the rule +that a trim can never return more than the test's own duration, the +alternatives already measured and rejected (`cargo check` for `cargo build`, +warming the compiler cache, and sharing a target directory), and the ten-run +revisit gate. Read it before reopening the question or changing the Windows +shape of this lane. Any replacement built after that gate inherits the +constraints in +[what a fixture-crate replacement would have to preserve][fixture-constraints] +below. ### How this relates to the isolation utilities @@ -4710,13 +4595,16 @@ That private build is why this test is the most expensive one on the Windows gate: `test_support` depends on `netsuke-build`, so a private root means compiling that crate and roughly 350 dependencies from scratch. Its measured budget is recorded in -[Windows budget for the isolated-Cargo-build tests][windows-test-budget], which -also records the -[decision to defer a trim](#deferring-the-split-build-dir-harness-trim) and the -gate at which that decision is revisited. +[Windows budget for the isolated-Cargo-build tests][windows-test-budget]. The +[decision to defer a trim][adr-027-trim] and the gate at which it is revisited +are recorded in ADR-027. #### What a fixture-crate replacement would have to preserve +This section is the constraint list for a future attempt, not a plan. The +decision to defer, and the gate that reopens it, are in [ADR-027][adr-027-trim] +; nothing here is built while the trim is deferred. + If the trim is taken up after that gate, the obvious shape is a minimal fixture crate built under the split layout in place of `test_support`. It needs at least one dependency, so that dependency rlibs land in the split build diff --git a/tests/locale_stub_ui_tests.rs b/tests/locale_stub_ui_tests.rs index 6ea639592..eeffb9364 100644 --- a/tests/locale_stub_ui_tests.rs +++ b/tests/locale_stub_ui_tests.rs @@ -88,11 +88,11 @@ fn stub_env_builders_compile_under_the_same_harness( /// The subject is the real `test_support` build, not a fixture crate, and that /// is deliberate: it is what carries both the dependency artefacts and the /// uplifted one that the split-directory derivation has to tell apart. The -/// cost of building it here is recorded, with the decision to defer trimming -/// it, under "Deferring the split-build-dir harness trim" in -/// docs/developers-guide.md; that section also states the fidelity argument -/// any fixture-crate replacement would owe and the response-file pressure it -/// would have to keep. +/// cost of building it here is recorded in +/// docs/developers-guide.md, and the decision to defer trimming it, the gate +/// that reopens the question, and the fidelity argument any fixture-crate +/// replacement would owe are in ADR-027 +/// (docs/adr-027-defer-split-build-dir-harness-trim.md). /// /// This test is a member of the `nested-cargo-builds` nextest group, so on /// Windows it holds that group's single slot: every other build-capable test From 553d9c6c8746c38beb3b19e3c17e76e2e3cc9bec Mon Sep 17 00:00:00 2001 From: leynos Date: Sat, 19 Sep 2026 12:12:53 +0200 Subject: [PATCH 15/17] Correct the ADR header and renumber it to 028 (#693) Two fixes from the first review round on the new ADR. CodeRabbit asked for a bare `Accepted.` status and a date-only `Date`. That was right, and it caught a real error rather than a style preference: the style guide defines `Date` as the date the ADR was *created*, and this one was created today, not on 2026-09-17 when the decision was taken. Every other ADR in the repository uses the bare form, with the decision summary in its own section, which is where the "deferred, not closed" sentence has moved. The measurement dates stay in the context section, where they are read as evidence rather than as the record's identity. The second fix is a collision. Three open branches claimed ADR 027 -- PR #739's `adr-027-windows-reparse-point-same-handle-open.md`, PR #699's `adr-027-command-placeholder-contract.md`, and this one. Checking only `origin/main` for the ceiling was not enough, and the note that triggered the check recorded that #739 had already taken 027. Per the convention that the unmerged branch renumbers, this one moves to 028, which no branch or merged tree holds. The file moves with `git mv`; the heading, the `[adr-028-trim]` link definition, and all seven inbound references across the guide, `contents.md`, `nextest.toml` and the test doc comment move with it. `make check-fmt` and `make markdownlint` pass, and `.config/nextest.toml` remains structurally identical to `origin/main` under `tomllib`. --- .config/nextest.toml | 2 +- ...r-028-defer-split-build-dir-harness-trim.md} | 17 ++++++++++++----- docs/contents.md | 2 +- docs/developers-guide.md | 12 ++++++------ tests/locale_stub_ui_tests.rs | 4 ++-- 5 files changed, 22 insertions(+), 15 deletions(-) rename docs/{adr-027-defer-split-build-dir-harness-trim.md => adr-028-defer-split-build-dir-harness-trim.md} (96%) diff --git a/.config/nextest.toml b/.config/nextest.toml index e43a9e17a..dbcbd908f 100644 --- a/.config/nextest.toml +++ b/.config/nextest.toml @@ -94,7 +94,7 @@ success-output = "immediate" # was the exclusive tail under an uncontended lane that no longer exists. The # repository has still decided not to spend the fidelity risk yet, and the trim # remains a candidate for tightening or deletion at the ten-run revisit gate -# that ADR-027 defines; that record carries the serialized measurements, the +# that ADR-028 defines; that record carries the serialized measurements, the # alternatives already measured and rejected — `cargo check` for `cargo # build`, warming the compiler cache, and sharing a target directory — and the # ten-run gate. The developers' guide carries the constraints a fixture-crate diff --git a/docs/adr-027-defer-split-build-dir-harness-trim.md b/docs/adr-028-defer-split-build-dir-harness-trim.md similarity index 96% rename from docs/adr-027-defer-split-build-dir-harness-trim.md rename to docs/adr-028-defer-split-build-dir-harness-trim.md index 50ae2036a..ee8f40458 100644 --- a/docs/adr-027-defer-split-build-dir-harness-trim.md +++ b/docs/adr-028-defer-split-build-dir-harness-trim.md @@ -1,17 +1,21 @@ -# Architectural decision record (ADR) 027: Defer replacing the split-build-dir harness test +# Architectural decision record (ADR) 028: Defer replacing the split-build-dir harness test ## Status -Accepted. The trim of `harness_compiles_under_a_split_build_dir` is deferred, -not closed: the Windows lane keeps the real `test_support` build and its 420s -budget, and the question reopens at the ten-run revisit gate below. +Accepted. ## Date -2026-09-17, with the measurements behind it taken through 2026-09-18. +2026-09-19. ## Context and problem statement +The decision this record captures was taken on 2026-09-17, when the +`nested-cargo-builds` serialization group landed and changed what a trim could +return. The measurements it rests on were taken up to 2026-09-18, and the +sections below say where each figure came from rather than dating the record by +it. + `harness_compiles_under_a_split_build_dir` is the regression test for the Windows `CreateProcessW` command-line limit. It forces a split layout with its own private `CARGO_TARGET_DIR` and `CARGO_BUILD_BUILD_DIR` roots, confirms the @@ -135,6 +139,9 @@ _Table 1: Comparison of the measured options._ `test_support` subject and its 420s budget, and the trimming question reopens at the revisit gate below rather than on a schedule. +The trim is deferred, not closed. The Windows lane keeps the real +`test_support` build, and the question reopens when the gate is met. + ## Rationale What a trim returns is the test's whole group **occupancy**, because the diff --git a/docs/contents.md b/docs/contents.md index 170dbc9bd..7bfc208f2 100644 --- a/docs/contents.md +++ b/docs/contents.md @@ -170,7 +170,7 @@ operator, user, and contributor references are easier to find. - [ADR-026](adr-026-manifest-environment-access-policy.md): Exact-name manifest environment policy evaluated before the reader, with project allow entries quarantined below the operator ceiling. -- [ADR-027](adr-027-defer-split-build-dir-harness-trim.md): Deferred trim of +- [ADR-028](adr-028-defer-split-build-dir-harness-trim.md): Deferred trim of the split-build-dir harness test, with the serialized-lane measurements that made the figure unstable and the ten-run gate that reopens the question. diff --git a/docs/developers-guide.md b/docs/developers-guide.md index 28bf6aad6..800890f3e 100644 --- a/docs/developers-guide.md +++ b/docs/developers-guide.md @@ -2913,7 +2913,7 @@ Linux runner. Both platforms still assert the packaged file list. See [windows-test-budget]: #windows-budget-for-the-isolated-cargo-build-tests [fixture-constraints]: #what-a-fixture-crate-replacement-would-have-to-preserve -[adr-027-trim]: adr-027-defer-split-build-dir-harness-trim.md +[adr-028-trim]: adr-028-defer-split-build-dir-harness-trim.md `tests/workflow_contracts/test_execution_coverage_test.py` holds all of this: the coverage inputs, the denied warnings, the doctest pass and its position, @@ -3015,7 +3015,7 @@ Table: Windows durations before and after the verification build moved. The 420s budget is therefore sized against the older, contended distribution and is deliberately conservative while the new shape has three samples. It is a -candidate for tightening, or for deletion, once the [ADR-027][adr-027-trim] +candidate for tightening, or for deletion, once the [ADR-028][adr-028-trim] revisit gate is met. Those three samples predate the serialization group described below, which @@ -3050,7 +3050,7 @@ contrast, is one-directional: the coverage a fixture crate would drop is exactly the coverage that fails only on Windows, where it is least likely to be noticed. -[ADR-027][adr-027-trim] holds the decision itself: the measurements, the rule +[ADR-028][adr-028-trim] holds the decision itself: the measurements, the rule that a trim can never return more than the test's own duration, the alternatives already measured and rejected (`cargo check` for `cargo build`, warming the compiler cache, and sharing a target directory), and the ten-run @@ -4596,13 +4596,13 @@ gate: `test_support` depends on `netsuke-build`, so a private root means compiling that crate and roughly 350 dependencies from scratch. Its measured budget is recorded in [Windows budget for the isolated-Cargo-build tests][windows-test-budget]. The -[decision to defer a trim][adr-027-trim] and the gate at which it is revisited -are recorded in ADR-027. +[decision to defer a trim][adr-028-trim] and the gate at which it is revisited +are recorded in ADR-028. #### What a fixture-crate replacement would have to preserve This section is the constraint list for a future attempt, not a plan. The -decision to defer, and the gate that reopens it, are in [ADR-027][adr-027-trim] +decision to defer, and the gate that reopens it, are in [ADR-028][adr-028-trim] ; nothing here is built while the trim is deferred. If the trim is taken up after that gate, the obvious shape is a minimal fixture diff --git a/tests/locale_stub_ui_tests.rs b/tests/locale_stub_ui_tests.rs index eeffb9364..7cc377305 100644 --- a/tests/locale_stub_ui_tests.rs +++ b/tests/locale_stub_ui_tests.rs @@ -91,8 +91,8 @@ fn stub_env_builders_compile_under_the_same_harness( /// cost of building it here is recorded in /// docs/developers-guide.md, and the decision to defer trimming it, the gate /// that reopens the question, and the fidelity argument any fixture-crate -/// replacement would owe are in ADR-027 -/// (docs/adr-027-defer-split-build-dir-harness-trim.md). +/// replacement would owe are in ADR-028 +/// (docs/adr-028-defer-split-build-dir-harness-trim.md). /// /// This test is a member of the `nested-cargo-builds` nextest group, so on /// Windows it holds that group's single slot: every other build-capable test From 5e85e855831ec0df035d8dfedeb7408193043437 Mon Sep 17 00:00:00 2001 From: leynos Date: Sat, 19 Sep 2026 12:24:07 +0200 Subject: [PATCH 16/17] Drop the full stop from the ADR's Date value (#693) The style guide gives the Date field the format `YYYY-MM-DD`, and the trailing full stop on `2026-09-19.` was mine rather than the convention's. The three most recent ADRs before this one -- 023, 024 and 026 -- carry the bare date; 022 and 025 carry the stop, so the repository is mixed and this is genuinely minor. The spec is unambiguous, so the bare form wins. `make check-fmt` and `make markdownlint` pass (143 files, 0 errors), with mdtablefix leaving all 142 files unchanged. The ADR's presence in markdownlint's scope was confirmed three ways: a count reconciliation against the config's ignore set, a glob check against its ignores, and a scoped probe linting the file alone. The change is a single character on one line. --- docs/adr-028-defer-split-build-dir-harness-trim.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/adr-028-defer-split-build-dir-harness-trim.md b/docs/adr-028-defer-split-build-dir-harness-trim.md index ee8f40458..7606507e7 100644 --- a/docs/adr-028-defer-split-build-dir-harness-trim.md +++ b/docs/adr-028-defer-split-build-dir-harness-trim.md @@ -6,7 +6,7 @@ Accepted. ## Date -2026-09-19. +2026-09-19 ## Context and problem statement From ce4fea259355376dd789bde8ce8a3c8108ad6876 Mon Sep 17 00:00:00 2001 From: leynos Date: Sat, 19 Sep 2026 12:56:42 +0200 Subject: [PATCH 17/17] Name the third Windows run behind the 170.7s figure (#693) CodeRabbit flagged an inconsistency: the guide said "the new shape has three samples" while its table carried only two run columns. The premise was wrong -- the third sample is real -- but the finding was right that the set was not traceable, and the fault was mine. My first commit replaced main's "on those same two runs ... 170.6s" with "125.3s, 170.6s and 170.7s on the three runs that followed" and never named the third run. The exclusive tails 62.3s, 87.2s and 85.5s came across from the measurement work, so the prose was consistent with itself; only the identifier was missing. The third run is 34080385050, recovered from the branch that produced the other two, `measure-windows-isolated-cargo-tests`. It is third in sequence after 34075197897 and 34079222917 and it reports the harness at 170.741s, which is the 170.7s the comment cites. It is now a table column rather than a bare number, with its own values measured the same way as the other two: 6.8s for the packaging test, a 253.8s nextest phase, and a 358s `Test` step. The extraction was validated before it was trusted. Run against the two runs whose values were already published, it reproduces the per-test durations to the decimal and the `Test` step to the second. It also caught its own first error: a naive phase measurement (`cargo nextest` invocation to summary) disagreed with both published values, because the published figure is nextest's own summary line, which excludes build time. The corrected reading matches to 0.1s on both, which is what licenses using it on the third. `nextest.toml` gains the three identifiers for the same reason, so the comment's 170.7s is traceable without the guide. The guide's lead-in sentence also said "the first two under the new shape" while enumerating two runs; it now names all three, which the scrutineer flagged as a stale clause after the first pass. `make check-fmt` leaves all 142 files unchanged and `make markdownlint` reports 0 errors across 143, so the hand-written table padding is canonical. `.config/nextest.toml` still parses under `tomllib`. --- .config/nextest.toml | 7 ++++--- docs/developers-guide.md | 16 ++++++++-------- 2 files changed, 12 insertions(+), 11 deletions(-) diff --git a/.config/nextest.toml b/.config/nextest.toml index dbcbd908f..23277db44 100644 --- a/.config/nextest.toml +++ b/.config/nextest.toml @@ -71,9 +71,10 @@ success-output = "immediate" # increasing each test's elapsed time: the two ran concurrently, each with four # compile jobs on a four-vCPU runner, and each roughly halved the other. With # the second build gone from the Windows lane, this test took 125.3s, 170.6s -# and 170.7s on the three runs that followed, rather than the 274.7s median -# above. 420s is therefore sized against the older, contended distribution and -# is deliberately conservative while the new shape has three samples. +# and 170.7s on runs 34075197897, 34079222917 and 34080385050, rather than the +# 274.7s median above. 420s is therefore sized against the older, contended +# distribution and is deliberately conservative while the new shape has three +# samples. # # Trimming the test is worth about 85s, not the 156s that figure implied # before the second build left, on the uncontended reading those three runs diff --git a/docs/developers-guide.md b/docs/developers-guide.md index 800890f3e..e62cbd907 100644 --- a/docs/developers-guide.md +++ b/docs/developers-guide.md @@ -3001,17 +3001,17 @@ budget, which clears the measured 312.9s worst case by 34%. Removing the second Cargo build sped this one up as well. The two used to run concurrently, each with four compile jobs on a four-vCPU runner, so each -roughly halved the other. Measured on runs 34075197897 and 34079222917, the -first two under the new shape: +roughly halved the other. Measured on runs 34075197897, 34079222917 and +34080385050, the first three under the new shape: Table: Windows durations before and after the verification build moved. -| Measure | Before (median) | Run 34075197897 | Run 34079222917 | -| ------------------------------------------------ | --------------- | --------------- | --------------- | -| `harness_compiles_under_a_split_build_dir` | 274.7s | 125.3s | 170.6s | -| `packaged_manifest_retains_build_script_sources` | 244.0s | 4.2s | 7.6s | -| nextest run phase | 365s | 185.5s | 258.0s | -| `Test` step | 471s | 260s | 368s | +| Measure | Before (median) | Run 34075197897 | Run 34079222917 | Run 34080385050 | +| ------------------------------------------------ | --------------- | --------------- | --------------- | --------------- | +| `harness_compiles_under_a_split_build_dir` | 274.7s | 125.3s | 170.6s | 170.7s | +| `packaged_manifest_retains_build_script_sources` | 244.0s | 4.2s | 7.6s | 6.8s | +| nextest run phase | 365s | 185.5s | 258.0s | 253.8s | +| `Test` step | 471s | 260s | 368s | 358s | The 420s budget is therefore sized against the older, contended distribution and is deliberately conservative while the new shape has three samples. It is a