Skip to content

fix(doctor): open the project a lone doctor asks about - #2743

Merged
ScriptedAlchemy merged 1 commit into
masterfrom
fleet/doctor-cold-mount
Sep 30, 2026
Merged

ScriptedAlchemy merged 1 commit into
masterfrom
fleet/doctor-cold-mount

Conversation

@ScriptedAlchemy

Copy link
Copy Markdown
Owner

Cause

Only the project owner publishes the canonical Doctor report. On both brokers (the Unix engine and the portable/Windows path), the Doctor probe only looked for the project server in the open-server cache (cached_project_server / portable_cached_project_server). On a fresh daemon nothing had opened the project, so the probe answered "not ready". The core then served doctor_report_owner_warming and never requested an open, so doctor printed "mounting" until some other request mounted the project (#2713). A single doctor run also returned on the first warming reply, so it could not have seen a report even if an open had started.

Change

  • Both broker probes now schedule the same admitted, single-flight background open that an MCP initialize schedules when no server is cached: schedule_project_server_warmup or schedule_portable_project_server_warmup, whose initialize_request is now optional. Admission still goes through ensure_registered_project_route, so an unenrolled directory is refused as before. The probe's typed not-enrolled, warming, discovery-deferred and reset-required responses are unchanged.
  • The doctor client polls a warming owner within its existing 15 s warm-up window, the same way it already polled missing telemetry, and reports "mounting" only if that window runs out. No limit was changed.
  • The acceptance helper that re-ran doctor until it stopped saying "mounting" (mounted_doctor_report) is deleted in favour of a single run.
  • A stored index-path read that cannot run (index.exclude.v1 / index.include.v1 over a reset-required or still-mounting store) is now a warning, not an issue. fix(config): report uncompilable stored index patterns #2728 counted it as an issue, which turned a pending reset into 2 issue(s). The store state is already reported, and doctor's own rule is that a diagnostic that could not run is not an issue.

Tests

  • New: typed_terminal_restart_acceptance::stale_sessions_store_reset::doctor_alone_on_a_fresh_daemon_opens_the_project_and_reports_findings. It inits a project, restarts the daemon, then runs one doctor --json. On master it fails with left: String("mounting") right: "observed", and it passes here (3/3 runs).
  • The four stale-store tests in the same module (stale_session_stores_refuse_sessions_only_until_their_scoped_reset, project_session_store_at_another_{lcm_schema,git_correlation}_version…, …workflow_schema_identity…) fail on current master at line 508 with 2 issue(s) from the index-path reads, and pass here.

Counts:

  • typed_terminal_restart_acceptance::: 11 passed, 0 failed. The CLI was built with test-transport, so the commit-barrier tests run too.
  • tracedecay lib: 775 passed, 0 failed.
  • storage_suite multi_connection_test:: (runs doctor): 2 passed.

Built-CLI journey (isolated HOME, one scoped daemon)

The journey used the debug tracedecay binary, a fresh HOME and a git project, with tracedecay daemon run under systemd-run --user --scope -p MemoryMax=6G -p MemorySwapMax=1G. After tracedecay init, I stopped the daemon, started a new one, and ran a single doctor with no other request:

--- lone doctor on the fresh daemon
exit 0 daemon_findings.state observed outcome healthy issues 0
elapsed 1.522862397s

Checks

  • cargo clippy -p tracedecay --features test-helpers,test-transport --all-targets -- -D warnings: clean.
  • cargo fmt --all -- --check: clean.
  • Windows check (--workspace --all-targets --target x86_64-pc-windows-gnu …): 0 errors. The warnings in the touched files are pre-existing unused imports on the import lines.
  • ripwire --edit-check: both schedulers have 0 incompatible callers. --quality-delta=<merge-base>..HEAD has one gating row: run_doctor_json in this acceptance suite duplicates doctor_json in tracedecay-cli's core_cli_suite. Both are single-shot doctor --json test helpers in different packages' test targets, which cannot share a helper.

Fixes #2713

On a fresh daemon the Doctor probe only looked the project server up in
the open-server cache. With nothing else opening the project, the core
answered doctor_report_owner_warming forever and doctor printed
"mounting" until another request mounted it.

The probe on both brokers (Unix engine and portable) now schedules the
same admitted background open an MCP initialize schedules when no server
is cached, and the doctor client polls a warming owner within its
existing warm-up window like missing telemetry. The acceptance helper
that re-ran doctor until it stopped reporting mounting is deleted.

Stored index-path reads that cannot run (a reset-required or mounting
store) are now warnings rather than issues: the store state is already
reported, and an unrunnable diagnostic is not an issue.

Fixes #2713
@changeset-bot

changeset-bot Bot commented Sep 30, 2026

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: 0163a61

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

@ScriptedAlchemy
ScriptedAlchemy merged commit 4624623 into master Sep 30, 2026
5 of 7 checks passed
@ScriptedAlchemy
ScriptedAlchemy deleted the fleet/doctor-cold-mount branch September 30, 2026 06:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

doctor on a cold daemon reports mounting until another request opens the project

1 participant