Skip to content

ci: two reds on main fail every PR — gap-suite shard 5 parity regression (test_gap_10430) and a stale public benchmark baseline #10707

Description

@proggeramlug

Two reds on main are failing every PR's CI, and neither has an owner

Found while triaging CI on #10669. Both reproduce on unrelated PRs, so they are main's, not any one change's.

1. gap-suite shard 5 — a real parity regression

REGRESSIONS — these were expected to pass:
  - test_gap_10430_stream_module_constructor: pass -> parity_fail

Confirmed identical on #10669 (a GC change) and #10646 (a codegen change), with the same second mismatch (test_gap_2514_settracesigint) in both. Shard 5 reports 140 pass / 2 parity-fail, 98.5%.

Because the snapshot records it as expected-to-pass, every PR touching anything in that shard's scope now fails gap-suite (5) and therefore pr-gate, regardless of content. That makes a real regression indistinguishable from an inherited one at a glance, which is how a genuine one gets waved through.

Possibly related: #10568 (node:stream/web ReadableStream.from() returning an empty object).

2. lint — the public benchmark baseline is stale

public baseline error: public artifact benchmark inputs changed;
regenerate it with ./benchmarks/run_public_baseline.sh

Identical on #10643, #10651 and #10669. Every PR that touches crates/ fails lint on this regardless of whether it has a changelog fragment — which also means the changeset gate's own signal is buried behind a failure nobody can act on from their own branch.

Why this is worth fixing rather than routing around

Both make pr-gate red on essentially every PR. The cost is not the red itself, it is that "CI is red, but it's red for everyone" becomes the default reading — and the next genuinely broken PR looks exactly the same. I nearly made that mistake on #10669 in the other direction: I initially assumed all of its reds were inherited, when the changelog fragment was in fact mine.

A third main-health item with a fix already written is #10655 (two ImportedClass test initializers missing constructor_has_synthetic_arguments, breaking cargo check --all-targets).

Activity

  1. proggeramlug commented on Sep 19, 2026

    @proggeramlug
    ContributorAuthor

    There is a third standing red with the same shape, and it has a fix up: e2e-scoped.

    It fails at "Compute e2e suite scope" after ~16 seconds on every PR:

    ci_e2e_scope: these crates/perry-codegen/tests/*.rs suites are in neither
    SOURCE_SUITE_MAP nor SUITE_EXCLUSIONS: error_subclass_field_init,
    typed_collection_receiver_guard
    

    Two suites arrived unclassified in 6925754a7d (#10443/#10446), which landed in merge train 218 (v0.5.1596) — the same train that introduced the test_gap_10430 regression in section 1 above. Like the other two, it is content-independent: #10721 (a Python script), #10722 (a .ts fixture) and #10719 (a Python script) all carry it. Fix in #10723; both suites are mapped rather than excluded, since both pass and SUITE_EXCLUSIONS is for a named failing test with an issue number.

    Status on the rest of this issue:

    Worth stating plainly because it is the actual thesis of this issue: train 218 is now the origin of two of the four reds. That is not a coincidence about that train so much as evidence that the per-PR tier cannot see either class — an unclassified test suite and a cross-shard parity regression are both invisible to a diff-scoped gate.

  2. proggeramlug commented on Sep 19, 2026

    @proggeramlug
    ContributorAuthor

    Section 1 is stale: test_gap_10430 is already fixed. The gate is still red — for a different test, in a different shard.

    This supersedes the "Section 1 (test_gap_10430) is being worked / the culprit is one of train 218's 30 commits" bullet in the previous comment. That was correct when written: the regression did land in train 218. It has since self-resolved in train 219, and I did not cause that — it was already fixed before I built anything.

    Both halves of section 1's framing are now wrong: the test is wrong and the shard is wrong.

    test_gap_10430 — six failures, then four passes

    All rows are ubuntu CI on main except the last:

    SHA landed (UTC) test_gap_10430
    68a5454396 train 217 09-18 17:54 green (all shards)
    60922041cd train 218 09-18 17:56 FAIL — sweep runs 35377260905 and 35410163573
    — PR #10646 / #10669 09-18 18:00 / 09-19 06:56 FAIL — the two runs this issue was opened from
    8df83f8c12 09-19 03:54 FAIL — sweep runs 35419885881 and 35423189647
    b0af11e1ce train 219 09-19 08:25 pass — fast 3-shard 2 (green) and full 8-shard 7 report
    4715bc2fa1 train 220 09-19 09:58 pass — sweep 3-shard 2 green
    4715bc2fa1 local — pass — 1/1, 100%, node_fail: 0, macOS

    Six consecutive failures then four consecutive passes is a fix, not flake.

    The cleanest single datum is the train-219 full-tier shard-7 report, because at b0af11e1ce both tests happen to land in 8-shard 7. One build, one runner, one report:

    test_gap_10430_stream_module_constructor                          pass
    test_gap_array_side_mask_covers_a_pointer_stored_at_a_late_index  parity_fail
    

    The handoff is visible inside a single shard, so it does not rest on comparing across runs.

    Likely fix: 57506478c0 (#10606), "root event/listener dispatch copies across moving GC" — flagged explicitly as UNPROVEN. Its changelog names crates/perry-runtime/src/node_stream_event_emitter.rs as the primary site, "the path every class X extends EventEmitter subclass and every Node stream class (Readable, Writable, Duplex, Transform) actually dispatches through", which is exactly what the 10430 fixture exercises. That is good reasoning about a hypothesis I have not built, and it stays a hypothesis. I have not attributed the train-218 cause at all, and would rather leave that blank than guess.

    The current blocker is #10727, in shard 4

    Train 219 is also where test_gap_array_side_mask_covers_a_pointer_stored_at_a_late_index went pass -> parity_fail. It is absent from gap_snapshot.json, so it blocks. Two independent observations at b0af11e1ce, in two different harness modes with different shard splits (fast 3-shard 3; auto-optimize 8-shard 7), and it passed at 8df83f8c12 — so it is neither flake nor a sharding artifact. Filed as #10727; no issue tracked it until now.

    Notably, no GC, array, or slot-enumeration file changed in train 219 at all, so the cause is not a direct code path and will need an A/B rather than more reading. Details in #10727.

    Recomputing the shard, since it moves

    Shard membership is round-robin over the post-filter list and shifts whenever the fixture count changes, so section 1's "shard 5" will keep going stale. To redo it:

    1. find test-files -maxdepth 1 -type f \( -name '*.ts' -o -name '*.cts' -o -name '*.mts' \) | sort
    2. keep entries whose basename (extension stripped) contains test_gap_; call the 0-based position k
    3. shard i of M runs exactly the tests where k % M == i - 1 (run_parity_tests.sh, the SHARD_COUNTER loop)

    At 4715bc2fa1 (870 test_gap_ fixtures), locale-independent:

    test index PR tier (6) sweep (3) full (8)
    test_gap_10430_stream_module_constructor 22 5 2 7
    test_gap_array_side_mask_...late_index 351 4 1 8

    So the blocking shard has moved 5 → 4. A PR rebased onto train 219 or later that still shows gap-suite (5) red is showing something else, and should not be attributed to this issue.

    test_gap_2514_settracesigint

    Confirmed, matching the previous comment: a standing gap_snapshot.json entry, not a regression, and it shares no cause with 10430 — it is util.setTraceSigInt, untouched since 2026-06-01. Settled from the snapshot and git history without needing a run. Closing the question.

    Suggested disposition

    Section 1 can be struck and replaced with a pointer to #10727. Sections 2 (public baseline) and the e2e-scoped item above are unaffected by any of this.

  3. proggeramlug commented on Sep 19, 2026

    @proggeramlug
    ContributorAuthor

    Diagnosis for the stale-public-baseline half of this issue. It is a real regression with a specific first-bad commit, not a gate that is red by construction.

    A plausible theory was tested and is false. The hypothesis was that Cargo.toml sits in SOURCE_PATHS, so every merge train's version bump re-invalidates the artifact — which would make the gate permanently red and any regeneration pointless within a day. The source fingerprint is byte-identical across train 225's release commit and its parent, so version bumps do not touch it.

    The actual first-bad commit, bisected over the fingerprint: it matches the migration target exactly at 57d3de75d3 (#7958, where the migration was recorded, 2026-08-12) and diverges at e3bd92bf65 — "fix(bench): reject zero-time benchmark false greens", 2026-09-01. That commit edited benchmarks/suite/*.ts and benchmarks/polyglot/bench.*, which are measured sources. So the published numbers were genuinely taken against older benchmark code, and the gate is reporting correctly.

    What that means for disposition:

    • It is a real ~2 h regeneration on a quiet machine, not a structural impossibility.
    • It has been standing for 18 days.
    • Afterwards it stays green until the next benchmark-source edit — not until the next train. That is a much better proposition than "red forever", and it makes clearing it worthwhile rather than futile.

    One concrete follow-up worth doing regardless of when the regeneration happens: the error message points at the wrong files. It reports that benchmark inputs changed, which directs a reader to HARNESS_PATHS, the declarative config. But the harness fingerprint matches its migration target exactly — it is SOURCE_PATHS that drifted. Anyone debugging from the message alone inspects the wrong two files and concludes nothing changed. Having the message distinguish "inputs (harness config)" from "measured sources" would have saved this bisect.

    Note the regeneration itself is deliberately not being done on any agent's own initiative: it rewrites the published perry-vs-node comparison, which is a claim about the project rather than a cleanup, so it is the maintainer's call.

  4. proggeramlug commented on Sep 19, 2026

    @proggeramlug
    ContributorAuthor

    Update: one of these is fixed, and they do not block merging

    e2e-scoped is fixed on main. Both suites — error_subclass_field_init and typed_collection_receiver_guard — are now present in _CODEGEN_SUITES in scripts/ci_e2e_scope.py on origin/main. They arrived unclassified with 6925754a7 and have since been classified. PRs still showing this failure are branched from before that fix; rebasing clears it.

    They do not gate merging. PR #10751 was merged today with e2e-scoped and lint both failing. So the practical cost of these reds is not that they stop work landing — it is the one I described above: every PR shows red for reasons unrelated to itself, so "red, but red for everyone" becomes the default reading, and a genuinely broken PR looks the same. I nearly made that error in the other direction on #10669, where a missing changelog fragment really was mine.

    Still open:

    • lint — the stale public benchmark baseline. Confirmed still failing on a PR merged today. The error names its own fix (./benchmarks/run_public_baseline.sh), and committed logs under benchmarks/json_performance/results/ carry the same message dated 2026-09-10, so this has been red for at least nine days.
    • gap-suite shard 5 — the test_gap_10430_stream_module_constructor parity regression.

    Neither is mine to fix blind: regenerating the public baseline commits measurement artifacts I have not reviewed the provenance of, and the gap regression wants whoever owns node:stream/web (possibly related to #10568).

  5. proggeramlug commented on Sep 20, 2026

    @proggeramlug
    ContributorAuthor

    Both halves of this issue are now resolved or relocated, so closing it.

    Item 1 — gap-suite shard-5 parity regression (test_gap_10430): fixed. Compiled and ran test_gap_10430_stream_module_constructor directly against v0.5.1617: byte-identical to node, exit 0. The co-named test_gap_2514_settracesigint still diverges, but it is a recorded entry in test-parity/gap_snapshot.json, so it cannot redden the shard.

    Item 2 — stale public benchmark baseline: tracked as #10799. That is the lint step "Public benchmark evidence freshness", still failing on today's runs, and it needs Node v22.23.1 and Bun 1.3.14 rather than a quiet machine — which is the misdiagnosis that kept it open for weeks.

    A note for anyone who lands here from an old link: the shard number in this issue's title went stale on its own. Shard assignment is round-robin over fixture count, so it moves whenever the fixture count changes. Record the recipe, not the index — and when picking up a long-red gate, re-read which test is failing from the current run rather than inheriting the name from the issue. This gate stayed red across a fix-and-break boundary and the continuity hid the handoff.

    Main's overall CI health is tracked on #10030; security-audit's single advisory is #10791 (unblocked by the soak window on 2026-09-21).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions