test(daemon): wait on readiness in load-sensitive lib tests - #2674
Merged
Merged
Conversation
Replace test-invented wall-clock bounds and races in daemon and MCP server lib tests with the readiness signal each one is waiting for: armed idle notifications, the production store shutdown before an in-process restart, a fake daemon that stays live, explicit socket shutdown, paused clocks, and the cancellation signal itself. Also read the routed refusal from structuredContent after #2651.
Connect the 128 non-reading peers one at a time so none waits out the 5 s request-read deadlines behind the others' handshakes, join the handler barrier on its own readiness, and pause time at the release so only the test's sleeps age the write-idle deadlines.
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What was wrong
These
tracedecaylib tests failed only under host load (load 57–233 on 96 cores) and passed alone. Each one either raced a signal it never armed or measured host scheduling with a wall clock that has no product deadline. One more was a deterministic break left by #2651.mcp::server::host_admission_tests::owned_project_replay_worker_continues_past_one_bounded_batchProjectHostAdmissionReplayWorker::wait_idle(test-only) readbusy/dirty/pending, awaited, and only then createdidle.notified(). The worker signals withnotify_waiters, which stores no permit, so an idle transition in that window was lost and the wait ran out its invented 5 s.Notifieds before reading state; drop the 5 s bound. Both callers still assertpending_count() == 0.daemon::tests::replay::client_identity_startup_replays_retained_profile_receiptsbusy_timeoutis 0 by design), so the restarted schema install hitdatabase is locked.prepare_memory_graph_reconciliation_shutdown→shutdown→close_retained_graph_runtimes_for_shutdown) before the restart, as a daemon restart does.daemon::tests::runtime_identity::concurrent_same_identity_worktrees_keep_exact_server_and_scheduler_bindingsstructuredContent.structuredContent.problem.daemon::tests::lifecycle::one_shot_tool_call_allows_long_response_while_daemon_stays_live,…_preserves_response_split_across_liveness_poll,…_receives_a_matching_saturation_responsedaemon::broker_stream_transport_tests::rmcp_peer_disconnect_mid_delivery_settles_dropped_rather_than_unknownshutdown(Both)acts on the socket, not the fd), here and in the full-close test.daemon::http_application_tests::remote_tls::remote_tls_listener_expires_saturated_non_reading_responsesdaemon::invocation_state::shutdown_tests::cancel_admissions_then_empty_shutdown_is_promptstart_paused, measured withtokio::time::Instant: the same bound now counts only timer time the shutdown waited out.mcp::server::cancel_candidate_journey::cancelled_large_candidate_search_stops_before_the_next_batchresumechannel was never sent on.CancellationSignalis cancelled; awaitcancelled()on the test side; delete the dead channel.mcp::server::routing::tests::many_slow_initialize_roots_share_one_discovery_budgetelapsed < 3salso counted un-deadlined registry lookups.deadline exceeded. In-flight interruption stays pinned bya_probe_over_its_budget_publishes_for_the_next_resolution.daemon::production_harness::generation_retention_test::mounted_code_generation_retention_continues_capped_segment_reclamationFail before / pass after
Whole lib suite, compiled test binary run directly with cargo's env, default parallelism, on the shared fleet host. Baseline is master 35fe79e; post-fix is this branch before the final two commits.
concurrent_same_identity_worktrees_…owned_project_replay_worker_…client_identity_startup_replays_…one_shot_tool_call_allows_long_response_…remote_tls_listener_expires_…many_slow_initialize_roots_…cancelled_large_candidate_search_…rmcp_peer_disconnect_mid_delivery_…cancel_admissions_then_empty_shutdown_is_promptgeneration_retention……_receives_a_matching_saturation_responseBaseline loads 57–233; post-fix loads 72–161. Post-fix failures in this loop are all owned elsewhere:
daemon_http_shutdown_releases_loopback_listener(fixed on master by #2664), the corrupt-graph restart (#2633, fixed on master by #2671), the advisory reopen (#2668), andschema_convergence_is_reported_and_completes_after_admission/maintenance_reclaims_a_removed_linked_worktree_…, which are tracked for follow-up in #2650.After merging master: all 88 touched lib tests pass (
host_admission_tests replay:: runtime_identity:: lifecycle:: broker_stream_transport_tests remote_tls:: invocation_state::shutdown_tests cancel_candidate_journey routing::tests generation_retention_test project_host_admission).Checks
cargo clippy -p tracedecay -p tracedecay-mcp --all-targets --features tracedecay/test-helpers,tracedecay/test-transport -- -D warnings: clean.cargo fmt --all -- --check: clean.wait_idleis gated totest/test-transport), so there is no shipped-CLI journey to run.cargo check -p tracedecay -p tracedecay-mcp --all-targets --target x86_64-pc-windows-gnu …failed only on pre-existingdaemon/tests/socket.rsunix-only helpers (fixed on master since by fix(test): compile Windows check past unix-only helpers #2670); nothing in these files. Tracked with fix(private-fs): NTFS ChangeTime does not witness mtime-restored rewrite #2081.Part of #2650.