Skip to content

fix(feedback): answer the first advisory cycle after a reopen - #2228

Merged
ScriptedAlchemy merged 1 commit into
masterfrom
fleet/advisory-cycle-first-call
Sep 26, 2026
Merged

ScriptedAlchemy merged 1 commit into
masterfrom
fleet/advisory-cycle-first-call

Conversation

@ScriptedAlchemy

Copy link
Copy Markdown
Owner

Root cause

I first tried the reported scenario literally: a daemon left idle past the 10-minute resident-memory window. It did not reproduce. The decode and graph engine were released (resident_owner_released, graph_engine_released reason="hibernate"), and the next advisory-cycle call still answered evidence: latest_feedback_generation_for_scope keeps serving the retained text generation.

The failing first call comes from the project being (re)opened over a ready sealed generation, which is what a daemon restart or a project reopen does. The advisory cycle mounts after project open publishes, so a request arriving during the open met one of two retryable pre-mount answers:

  1. Before the proximity-only placeholder publishes: runtimes.advisory_cycle is None, which answers The advisory feedback cycle authority is unavailable.
  2. While the deferred mount upgrades that placeholder to the full cycle (generation selection, then GitHub discovery): The advisory feedback cycle mounts once the first code-index generation is sealed.

Both answers are retryable: true, retry_after_millis: 250, so the second call succeeded. Timed restart run on master (the daemon had already sealed the generation in an earlier session):

1 14:06:19.396 -> 14:06:20.189 "The advisory feedback cycle authority is unavailable"
2 14:06:20.202 -> 14:06:20.325 "The advisory feedback cycle authority is unavailable"
3 14:06:20.336 -> 14:06:20.454 "The advisory feedback cycle authority is unavailable"
4 14:06:20.466 -> 14:06:20.580 "The advisory feedback cycle mounts once the first code-index generation is sealed"
… (the pre-mount answer until the deferred mount published)

What changed

  • Registry signal. ProjectRuntimeRegistryV1 has a published_changed watch. It is bumped when a component is published, when publish_advisory_atomically swaps in the full cycle, and when a project's publication stage finishes. advisory_cycle_view returns the owner and the publication stage together with a receiver subscribed before the read, so no publication between the read and the wait is missed.
  • Mount readiness on the port. DaemonAdvisoryCycleInvocationPort::mount() reports Answers by default. The pre-mount owner reports Mounting exactly when a ready sealed generation exists for its scope. It still answers at once when the code index is disabled or there is no indexable source (those states are terminal).
  • Bounded wait in dispatch. answering_advisory_cycle_owner waits, within the request's own deadline, while the owner is Mounting, or while the owner is absent and the project's publication stage is still Warming. Then it invokes the owner that answers. A finished publication without an owner returns at once, and so does a mounted owner.
  • Known ceiling, marked with a ponytail: comment: if the deferred mount fails terminally, a waiting request runs out its own deadline before it gets the warming answer.

Fail before / pass after

daemon::production_harness::advisory_cycle_language_journey_test::first_advisory_cycle_after_a_reopen_answers_without_a_retry opens a Rust checkout, settles its advisory cycle, and shuts the composition down. It then reopens the same profile and makes exactly one tracedecay_feedback_advisory_cycle call.

Before (this branch with the dispatch wait reverted):

the first call after a reopen was refused: {… "problem":{"code":"feedback.advisory-cycle.unavailable", … "message":"The advisory feedback cycle mounts once the first code-index generation is sealed", "retry":"after_delay","retry_after_millis":250 …}}
test result: FAILED. 0 passed; 1 failed

After: ok, the outcome is evidence.

Runtime journey

Debug tracedecay-cli --no-default-features --features production, one daemon under systemd-run --user --scope -p MemoryMax=6G -p MemorySwapMax=1G. The profile is the isolated one from the #2221 journey: rust-lang/log at PR #741's head, with a sealed generation from an earlier daemon. The daemon was restarted, and then one call was made:

$ tracedecay tool feedback_advisory_cycle --document-uri file://…/log/src/lib.rs   # first call after the restart, 14:40:31.805 -> 14:40:33.624
"outcome":{"outcome":"evidence"
daemon: 14:40:32.509797Z event="feedback_proximity_mount" outcome="ready"
daemon: 14:40:33.220089Z event="github_pull_request_discovery" outcome="found"
daemon: 14:40:33.227232Z event="feedback_advisory_mount" outcome="mounted"

The call waited about 1.8 s for the mount and answered on the first try.

Checks

  • cargo test -p tracedecay --lib -- production_harness::advisory_cycle_language_journey_test (after rebase): 3 passed.
  • cargo test -p tracedecay --lib -- daemon::tests::invocation_ownership daemon::tests::feedback_impact daemon::project_open_owners production_harness: 40 passed.
  • cargo test -p tracedecay-daemon-service --lib: 319 passed, 1 failed. The failure, adoption_observation::tests::census_counts_each_composed_family_and_omits_uncomposed_families, is pre-existing: the retrieval catalog now has 41 operations and the test pins 34. This change does not touch the catalog or adoption census.
  • cargo clippy -p tracedecay-daemon-service -p tracedecay --all-targets --features tracedecay/test-helpers -- -D warnings: clean.
  • cargo fmt --all -- --check: clean.

@changeset-bot

changeset-bot Bot commented Sep 26, 2026

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: b5b1894

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

@ScriptedAlchemy
ScriptedAlchemy merged commit dfb4ef4 into master Sep 26, 2026
1 check passed
@ScriptedAlchemy
ScriptedAlchemy deleted the fleet/advisory-cycle-first-call branch September 26, 2026 15:05
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 26, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-26T15:08:21.689815Z b5b1894 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: b5b189498c

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +520 to +524
self.answering_advisory_cycle_owner(
registered_project_root.as_deref(),
&deadline,
)
.await

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Check cancellation before waiting for advisory publication

When an advisory-cycle payload arrives with an already-cancelled CancellationContext while the owner is mounting or publication is warming, this newly added wait runs before execute_feedback_advisory_cycle performs its cancellation check. The request can therefore remain in flight until owner publication or its deadline instead of immediately returning cancelled_before_admission; check the payload cancellation before entering this wait or make the wait cancellation-aware.

AGENTS.md reference: AGENTS.md:L201-L203

Useful? React with 👍 / 👎.

Comment on lines +2100 to +2106
match self
.code_index_schedulers
.latest_feedback_generation_for_scope(&self.project_root, &self.scope)
.await
{
Some(_) => DaemonAdvisoryCycleMountV1::Mounting,
None => DaemonAdvisoryCycleMountV1::Answers,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Stop reporting Mounting after the deferred task terminates

When deferred::try_mount returns Attempt::Terminal after an LSP grant, owner registration, or advisory composition failure, the sealed generation remains present but no task remains capable of publishing the full owner. This implementation consequently reports Mounting forever, so every subsequent advisory-cycle call waits out its deadline and returns a timeout rather than promptly surfacing the terminal unavailable state; track the deferred attempt's terminal outcome rather than inferring active mounting solely from generation presence.

AGENTS.md reference: AGENTS.md:L189-L190

Useful? React with 👍 / 👎.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant