Skip to content

test(daemon): wait on readiness in load-sensitive lib tests - #2674

Merged
ScriptedAlchemy merged 5 commits into
masterfrom
fleet/load-flake-sweep
Sep 29, 2026
Merged

ScriptedAlchemy merged 5 commits into
masterfrom
fleet/load-flake-sweep

Conversation

@ScriptedAlchemy

Copy link
Copy Markdown
Owner

What was wrong

These tracedecay lib tests failed only under host load (load 57–233 on 96 cores) and passed alone. Each one either raced a signal it never armed or measured host scheduling with a wall clock that has no product deadline. One more was a deterministic break left by #2651.

Test Cause Change
mcp::server::host_admission_tests::owned_project_replay_worker_continues_past_one_bounded_batch Test helper race. ProjectHostAdmissionReplayWorker::wait_idle (test-only) read busy/dirty/pending, awaited, and only then created idle.notified(). The worker signals with notify_waiters, which stores no permit, so an idle transition in that window was lost and the wait ran out its invented 5 s. Arm the idle and cancel Notifieds before reading state; drop the 5 s bound. Both callers still assert pending_count() == 0.
daemon::tests::replay::client_identity_startup_replays_retained_profile_receipts Test sequencing. The in-process "restart" only dropped the first administration, whose detached session-runtime workers still held the profile database writer (busy_timeout is 0 by design), so the restarted schema install hit database is locked. Run the production store shutdown (prepare_memory_graph_reconciliation_shutdown → shutdown → close_retained_graph_runtimes_for_shutdown) before the restart, as a daemon restart does.
daemon::tests::runtime_identity::concurrent_same_identity_worktrees_keep_exact_server_and_scheduler_bindings Deterministic break since #2651 moved tool problems into structuredContent. Read structuredContent.problem.
daemon::tests::lifecycle::one_shot_tool_call_allows_long_response_while_daemon_stays_live, …_preserves_response_split_across_liveness_poll, …_receives_a_matching_saturation_response Fixture contradicted the scenario. The fake daemon dropped its listener right after answering, so the client's 10 ms liveness probe could hit a closed socket after the response had already arrived. The fake daemon keeps accepting probes until the client returns. The product side of this race (a daemon that exits right after answering) is filed as #2654.
daemon::broker_stream_transport_tests::rmcp_peer_disconnect_mid_delivery_settles_dropped_rather_than_unknown The in-process "client" was dropped while another test thread was forking. The child briefly inherits every fd, so the peer socket stayed alive and the daemon's write succeeded, which counts as delivered. A client in another process has no such window. Shut the client socket down (shutdown(Both) acts on the socket, not the fd), here and in the full-close test.
daemon::http_application_tests::remote_tls::remote_tls_listener_expires_saturated_non_reading_responses 128 concurrent TLS handshakes on one thread raced the product's 5 s request-read and write-idle deadlines, guarded by invented 2 s windows. Connect peers one at a time (each finishes its request right after its accept), join the handler barrier on readiness, and pause time at its release so only the test's sleeps age write-idle.
daemon::invocation_state::shutdown_tests::cancel_admissions_then_empty_shutdown_is_prompt 500 ms wall-clock check. start_paused, measured with tokio::time::Instant: the same bound now counts only timer time the shutdown waited out.
mcp::server::cancel_candidate_journey::cancelled_large_candidate_search_stops_before_the_next_batch The pausing fixture gave the cancel 5 s, then let the scan continue. Under load the canceller ran later and the scan passed one more checkpoint (5 vs 4). The resume channel was never sent on. Hold the checkpoint until the observed CancellationSignal is cancelled; await cancelled() on the test side; delete the dead channel.
mcp::server::routing::tests::many_slow_initialize_roots_share_one_discovery_budget elapsed < 3s also counted un-deadlined registry lookups. Rely on the typed outcome: per-root budgets would resolve all three 1.5 s probes and answer NotFound; only a shared budget ends in deadline exceeded. In-flight interruption stays pinned by a_probe_over_its_budget_publishes_for_the_next_resolution.
daemon::production_harness::generation_retention_test::mounted_code_generation_retention_continues_capped_segment_reclamation Invented 20 s window on background reindexing. Wait on the published generation id.

Fail before / pass after

Whole lib suite, compiled test binary run directly with cargo's env, default parallelism, on the shared fleet host. Baseline is master 35fe79e; post-fix is this branch before the final two commits.

Test Baseline failed runs Post-fix failed runs
concurrent_same_identity_worktrees_… 19/20 0/16
owned_project_replay_worker_… 5/20 0/16
client_identity_startup_replays_… 5/20 0/16
one_shot_tool_call_allows_long_response_… 1/20 0/16
remote_tls_listener_expires_… 1/20 0/16
many_slow_initialize_roots_… 1/20 0/16
cancelled_large_candidate_search_… 1/5 (second loop) 0/16
rmcp_peer_disconnect_mid_delivery_… 1/5 (second loop) 0/16
cancel_admissions_then_empty_shutdown_is_prompt 1/1 (third loop) 0/16
generation_retention… 1/16 fixed in the last commit
…_receives_a_matching_saturation_response 1/1 (focused run of 62 touched tests) fixed in the last commit

Baseline loads 57–233; post-fix loads 72–161. Post-fix failures in this loop are all owned elsewhere: daemon_http_shutdown_releases_loopback_listener (fixed on master by #2664), the corrupt-graph restart (#2633, fixed on master by #2671), the advisory reopen (#2668), and schema_convergence_is_reported_and_completes_after_admission / maintenance_reclaims_a_removed_linked_worktree_…, which are tracked for follow-up in #2650.

After merging master: all 88 touched lib tests pass (host_admission_tests replay:: runtime_identity:: lifecycle:: broker_stream_transport_tests remote_tls:: invocation_state::shutdown_tests cancel_candidate_journey routing::tests generation_retention_test project_host_admission).

Checks

Part of #2650.

Replace test-invented wall-clock bounds and races in daemon and MCP
server lib tests with the readiness signal each one is waiting for:
armed idle notifications, the production store shutdown before an
in-process restart, a fake daemon that stays live, explicit socket
shutdown, paused clocks, and the cancellation signal itself. Also read
the routed refusal from structuredContent after #2651.
Connect the 128 non-reading peers one at a time so none waits out the
5 s request-read deadlines behind the others' handshakes, join the
handler barrier on its own readiness, and pause time at the release so
only the test's sleeps age the write-idle deadlines.
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

@changeset-bot

changeset-bot Bot commented Sep 29, 2026

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: 1f8089d

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

@ScriptedAlchemy
ScriptedAlchemy merged commit 6aace0f into master Sep 29, 2026
6 of 7 checks passed
@ScriptedAlchemy
ScriptedAlchemy deleted the fleet/load-flake-sweep branch September 29, 2026 19:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant