Releases: serve staged releases before any release reaches 100% - #78
Merged
Merged
Conversation
A prefix with no release at 100% made the default lookup throw a 500 before eligibility was checked, so no device got the staged build. That is every device for a new prefix until its first release is fully rolled out, and every JetKVM device the day no app or system row sits at 100%. The default lookup now returns null when nothing is at 100%. A device inside the rollout bucket gets the staged release as before; a device outside it gets a 404 saying no release is rolled out yet for its SKU, instead of a 500. Found by Bugbot on the release PR (#77).
Contributor
Author
|
bugbot run |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 394c019. Configure here.
adamshiervani
marked this pull request as ready for review
September 18, 2026 20:45
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
adamshiervani
added a commit
that referenced
this pull request
Sep 18, 2026
…78) (#79) A prefix with no release at 100% made the default lookup throw a 500 before eligibility was checked, so no device got the staged build. That is every device for a new prefix until its first release is fully rolled out, and every JetKVM device the day no app or system row sits at 100%. The default lookup now returns null when nothing is at 100%. A device inside the rollout bucket gets the staged release as before; a device outside it gets a 404 saying no release is rolled out yet for its SKU, instead of a 500. Found by Bugbot on the release PR (#77).
adamshiervani
added a commit
that referenced
this pull request
Sep 19, 2026
* fix(releases): serve staged releases before any release reaches 100% (#78) A prefix with no release at 100% made the default lookup throw a 500 before eligibility was checked, so no device got the staged build. That is every device for a new prefix until its first release is fully rolled out, and every JetKVM device the day no app or system row sits at 100%. The default lookup now returns null when nothing is at 100%. A device inside the rollout bucket gets the staged release as before; a device outside it gets a 404 saying no release is rolled out yet for its SKU, instead of a 500. Found by Bugbot on the release PR (#77). * Releases: sync new R2 releases from the API every 30 minutes (#81) * feat(releases): sync new R2 releases from the API every 30 minutes The release sync only ran when an operator invoked scripts/sync-releases.ts by hand, so a firmware upload sat in R2 until someone remembered to run it. The API process now runs the same sync on start and every 30 minutes, registering each new stable version at the default 10% rollout. The non-interactive core (bucket listing, artifact collection, DB insert) moves from the script into src/release-sync.ts so the API can import it; the script keeps the GPG check and the confirmation prompt and passes them in as the release decider. Both entry points share one R2 client via src/s3.ts. Scheduled runs skip a tick while the previous run is still going, log and survive a failed run, and treat a unique-constraint hit on insert as "already synced" so several API instances can race without an error. The existence check now precedes the R2 artifact walk, so a tick over an already-synced bucket costs one list call per prefix plus one DB lookup per version. * fix(release-sync): defer versions whose upload is still settling A scheduled tick can land while a version is still being uploaded and register it with a partial SKU set or a hash that changes afterwards. Sync never rewrites a row, so that snapshot would be permanent. Before collecting artifacts, list every object under the version folder and skip the version when the newest object changed in the last 10 minutes. The window is shorter than the schedule interval, so a real upload costs at most one extra tick. * refactor(release-sync): share R2 config, paginate via SDK, probe SKUs concurrently * src/s3.ts now exports bucketName and baseUrl next to the client, so the request handlers, the scheduler and the script read R2 config from one place instead of three. * The upload settle window moves into SyncConfig.uploadSettleMs. The scheduler sets it; the one-shot operator script leaves it unset because an operator sees the artifact list before confirming and has no next run. * newestUploadTime uses the SDK's paginateListObjectsV2 instead of a hand-rolled continuation loop. * collectReleaseArtifacts probes SKUs concurrently and folds the results in SKU order, so artifact order and the primary artifact are unchanged. * Tests drop empty per-prefix listing stubs that the file-level default already covers. * fix(release-sync): take the settle listing after the artifact scan Listing the version folder before the scan left a gap: an upload that started between the listing and the scan passed the check and was captured partially. Listing after the scan closes it, because any upload active during the scan leaves an object newer than the window and the snapshot is discarded instead of registered. * fix(release-sync): paginate the version listing ListObjectsV2 returns at most 1,000 common prefixes per page. Once a release prefix grows past that, versions on later pages were never seen on any tick. listStableVersions now walks every page with the SDK paginator, as newestUploadTime already does, and the mock in the sync tests can serve a truncated listing. * Release sync: a unique violation is only a race when the row exists (#82) * fix(release-sync): only treat a unique violation as a race when the row exists Sync caught every P2002 from the release insert as "created concurrently elsewhere". On staging the id sequences were behind the rows after a data import, so each insert failed on the primary key, was logged as a race, and left nothing in the table. After a unique violation, createRelease now looks the (version, type) row up. Present means another instance registered it first; absent means the insert really failed, and the error is rethrown with the type and version in its message so the scheduled run log names the release. * fix(release-sync): name the release in every per-version failure Wrapping only the insert error left the row lookup, and the S3 scan before it, free to escape without the type and version. syncReleases now wraps whatever createRelease throws for a version. * feat(releases): POST /releases/sync for the upload script (#83) The scheduled tick found a new version at most 30 minutes plus the settle window after upload. The upload script knows when its last object is written, so it can trigger the sync itself. One runner now owns the in-progress flag and serves both the timer and the endpoint. The endpoint is guarded by RELEASE_SYNC_TOKEN as a bearer token, compared in constant time, and is not registered without it. It runs with the settle window off and answers with the per-outcome counts, or 409 while a run is in progress. * fix(auth): accept the bearer scheme in any case (#85) Scheme names are case-insensitive (RFC 9110). The token is still compared exactly. * fix(releases): scope the sync endpoint's settle bypass to one version (#84) The endpoint turned the settle window off for every version in the bucket. A call made while a different upload was still running registered that upload half done, and sync never revisits a row. The body now names the version the caller finished, and only that version skips the window. Without a body the call is a normal tick.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Bugbot finding on #77, High: with no release at 100% for a prefix,
/releasesreturned 500 for every device, including devices inside the rollout bucket that should have received the staged build. First hit for the Mini, whose prefix starts empty; also JetKVM's behaviour on any day no app or system row is at 100%.Tests: the case that asserted the 500 now asserts the staged build for an in-bucket device and the 404 for an out-of-bucket one; a Mini case covers the first release at 10%.
+44 / −11.
Note
Medium Risk
Changes stable
/releasesrollout/fallback behavior and HTTP status codes for edge cases with no fully rolled-out release; well covered by tests but affects device OTA eligibility.Overview
Fixes
/releasesbackground checks when a prefix has no release at 100% rollout yet (e.g. Mini’s first staged build).getDefaultReleasenow returnsnullinstead of raising 500;Retrievestill serves the latest row to devices in the rollout bucket, and returns 404 (No <prefix> release is rolled out yet for SKU …) when the device is out of bucket and there is no 100% fallback.Tests replace the old “expect 500” case with in-bucket staged delivery and out-of-bucket 404, plus a Mini scenario at 10% rollout.
Reviewed by Cursor Bugbot for commit 394c019. Bugbot is set up for automated code reviews on this repo. Configure here.