Skip to content

Release sync: a unique violation is only a race when the row exists - #82

Merged
adamshiervani merged 2 commits into
devfrom
fix/release-sync-unique-target
Sep 19, 2026
Merged

adamshiervani merged 2 commits into
devfrom
fix/release-sync-unique-target

Conversation

@adamshiervani

@adamshiervani adamshiervani commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

The scheduled sync caught every unique violation from the release insert as "created concurrently elsewhere". On staging the id sequences were behind the rows after a data import, so each insert failed on the primary key, was logged as a race, and wrote nothing. The log claimed mini 0.5.1 and 0.5.2 existed while the table had no mini rows at all.

What changed

 src/release-sync.ts
   createRelease
     catch (error)
-      P2002                         → "created concurrently elsewhere"
+      P2002 && row (version, type) exists → "created concurrently elsewhere"
+      otherwise                     → throw Error("<type> <version>: create failed", { cause })
+  releaseExists(prisma, type, version)   # findUnique on version_type
 test/sync-releases.test.ts
   race test: stub gains findUnique (delegates to the real DB)
+  new: "fails on a unique violation that left no release row behind"

Why look the row up instead of reading the constraint name

Prisma reports the failing columns in meta.target, but that field is untyped, derived from the Postgres error text, and changes shape across adapters and major versions. Asking the database whether the (version, type) row exists answers the question the branch actually cares about and does not depend on any of that.

Failure path

createRelease
  release.create  ─── P2002 ───▶ releaseExists?
                                   yes → log "created concurrently elsewhere", already-synced
                                   no  → throw "[sync-releases] mini 0.5.2: create failed" (cause: P2002)
                                           └─ scheduleReleaseSync logs "scheduled run failed" + error
                                              remaining versions in this run are not processed;
                                              the next tick retries

The rethrown error carries the type and version, so the scheduled run log names the release that failed. Before this change the only line naming the version was the misleading one.

Staging

The staging sequences were advanced by hand to the current max ids; both mini releases registered on the next start. Production sequences already match their tables.


Note

Medium Risk
Changes how release rows are created during scheduled sync; the fix reduces silent data loss but alters failure behavior on DB constraint errors.

Overview
Fixes release sync treating every Prisma P2002 on insert as a benign multi-instance race. After a DB restore, stale primary-key sequences can fail inserts with the same error code while leaving no (version, type) row—those cases were logged as “created concurrently elsewhere” and skipped.

createRelease now calls new releaseExists after a unique violation and only returns already-synced when that row is present; otherwise the error propagates.

syncReleases wraps per-version createRelease failures in an error that names type and version ([sync-releases] …: sync failed, with cause), so scheduled run logs point at the failing release without tracing the whole run.

Tests: the concurrent-create case stubs findUnique; a new case asserts a P2002 with no row rejects with the wrapped message.

Reviewed by Cursor Bugbot for commit 2f4a51d. Bugbot is set up for automated code reviews on this repo. Configure here.

…ow exists

Sync caught every P2002 from the release insert as "created concurrently
elsewhere". On staging the id sequences were behind the rows after a data
import, so each insert failed on the primary key, was logged as a race, and
left nothing in the table.

After a unique violation, createRelease now looks the (version, type) row
up. Present means another instance registered it first; absent means the
insert really failed, and the error is rethrown with the type and version
in its message so the scheduled run log names the release.
@adamshiervani
adamshiervani marked this pull request as ready for review September 19, 2026 12:23
@adamshiervani
adamshiervani requested a balanced review from Copilot September 19, 2026 12:23
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 19, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-19T12:25:41.304277Z a212e44 Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

A failed conflict lookup bypasses the contextual error handling and loses the release identity.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 1 Medium severity

Open (1)
What changed in this PR

Improves release synchronization by distinguishing genuine concurrent inserts from unrelated unique-constraint failures.

Changes:

  • Verifies that the expected release exists after a P2002.
  • Adds coverage for stale primary-key sequence failures.
File Description
src/​release-sync.ts Adds conflict verification and contextual errors.
test/​sync-releases.test.ts Tests race and stale-sequence scenarios.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread src/release-sync.ts
// this insert, in which case the row it wrote is the one we wanted. Any
// other unique violation (a stale id sequence after a data import, say)
// leaves no row and is a real failure.
if (isUniqueViolation(error) && (await releaseExists(clients.prisma, type, version))) {

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 2f4a51d. The contextual wrap moved up to the per-version loop in syncReleases, so any failure for a version carries its type and version: the insert, the row lookup after a unique violation, and the S3 artifact scan before either. createRelease now rethrows the original error unchanged.

Wrapping only the insert error left the row lookup, and the S3 scan
before it, free to escape without the type and version. syncReleases now
wraps whatever createRelease throws for a version.
@adamshiervani
adamshiervani merged commit 68c4dfc into dev Sep 19, 2026
3 checks passed
@adamshiervani
adamshiervani deleted the fix/release-sync-unique-target branch September 19, 2026 12:30
adamshiervani added a commit that referenced this pull request Sep 19, 2026
* fix(releases): serve staged releases before any release reaches 100% (#78)

A prefix with no release at 100% made the default lookup throw a 500
before eligibility was checked, so no device got the staged build. That
is every device for a new prefix until its first release is fully
rolled out, and every JetKVM device the day no app or system row sits
at 100%.

The default lookup now returns null when nothing is at 100%. A device
inside the rollout bucket gets the staged release as before; a device
outside it gets a 404 saying no release is rolled out yet for its SKU,
instead of a 500.

Found by Bugbot on the release PR (#77).

* Releases: sync new R2 releases from the API every 30 minutes (#81)

* feat(releases): sync new R2 releases from the API every 30 minutes

The release sync only ran when an operator invoked scripts/sync-releases.ts
by hand, so a firmware upload sat in R2 until someone remembered to run it.
The API process now runs the same sync on start and every 30 minutes,
registering each new stable version at the default 10% rollout.

The non-interactive core (bucket listing, artifact collection, DB insert)
moves from the script into src/release-sync.ts so the API can import it;
the script keeps the GPG check and the confirmation prompt and passes them
in as the release decider. Both entry points share one R2 client via src/s3.ts.

Scheduled runs skip a tick while the previous run is still going, log and
survive a failed run, and treat a unique-constraint hit on insert as
"already synced" so several API instances can race without an error.
The existence check now precedes the R2 artifact walk, so a tick over an
already-synced bucket costs one list call per prefix plus one DB lookup per
version.

* fix(release-sync): defer versions whose upload is still settling

A scheduled tick can land while a version is still being uploaded and
register it with a partial SKU set or a hash that changes afterwards. Sync
never rewrites a row, so that snapshot would be permanent.

Before collecting artifacts, list every object under the version folder
and skip the version when the newest object changed in the last 10
minutes. The window is shorter than the schedule interval, so a real
upload costs at most one extra tick.

* refactor(release-sync): share R2 config, paginate via SDK, probe SKUs concurrently

* src/s3.ts now exports bucketName and baseUrl next to the client, so the
  request handlers, the scheduler and the script read R2 config from one
  place instead of three.
* The upload settle window moves into SyncConfig.uploadSettleMs. The
  scheduler sets it; the one-shot operator script leaves it unset because an
  operator sees the artifact list before confirming and has no next run.
* newestUploadTime uses the SDK's paginateListObjectsV2 instead of a
  hand-rolled continuation loop.
* collectReleaseArtifacts probes SKUs concurrently and folds the results in
  SKU order, so artifact order and the primary artifact are unchanged.
* Tests drop empty per-prefix listing stubs that the file-level default
  already covers.

* fix(release-sync): take the settle listing after the artifact scan

Listing the version folder before the scan left a gap: an upload that
started between the listing and the scan passed the check and was captured
partially. Listing after the scan closes it, because any upload active
during the scan leaves an object newer than the window and the snapshot
is discarded instead of registered.

* fix(release-sync): paginate the version listing

ListObjectsV2 returns at most 1,000 common prefixes per page. Once a
release prefix grows past that, versions on later pages were never seen
on any tick. listStableVersions now walks every page with the SDK
paginator, as newestUploadTime already does, and the mock in the sync
tests can serve a truncated listing.

* Release sync: a unique violation is only a race when the row exists (#82)

* fix(release-sync): only treat a unique violation as a race when the row exists

Sync caught every P2002 from the release insert as "created concurrently
elsewhere". On staging the id sequences were behind the rows after a data
import, so each insert failed on the primary key, was logged as a race, and
left nothing in the table.

After a unique violation, createRelease now looks the (version, type) row
up. Present means another instance registered it first; absent means the
insert really failed, and the error is rethrown with the type and version
in its message so the scheduled run log names the release.

* fix(release-sync): name the release in every per-version failure

Wrapping only the insert error left the row lookup, and the S3 scan
before it, free to escape without the type and version. syncReleases now
wraps whatever createRelease throws for a version.

* feat(releases): POST /releases/sync for the upload script (#83)

The scheduled tick found a new version at most 30 minutes plus the settle
window after upload. The upload script knows when its last object is
written, so it can trigger the sync itself.

One runner now owns the in-progress flag and serves both the timer and
the endpoint. The endpoint is guarded by RELEASE_SYNC_TOKEN as a bearer
token, compared in constant time, and is not registered without it. It
runs with the settle window off and answers with the per-outcome counts,
or 409 while a run is in progress.

* fix(auth): accept the bearer scheme in any case (#85)

Scheme names are case-insensitive (RFC 9110). The token is still
compared exactly.

* fix(releases): scope the sync endpoint's settle bypass to one version (#84)

The endpoint turned the settle window off for every version in the
bucket. A call made while a different upload was still running
registered that upload half done, and sync never revisits a row.

The body now names the version the caller finished, and only that
version skips the window. Without a body the call is a normal tick.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants