Skip to content

fork_repo / create_repo orphan on-disk repos (and a spent iCaptcha proof / uploaded object) on late DB failure #205

Description

@beardthelion

Surfaced during review of #196 (pre-existing ordering; not in that PR's diff).

fork_repo ordering (api/repos.rs:1555 -> 1567 -> 1585 -> 1605): spend proof -> git clone --mirror to disk -> release_after_write (Tigris upload) -> db.create_repo. If create_repo fails, the on-disk mirror and the already-uploaded Tigris object are left with no DB row, and the single-use iCaptcha proof is already consumed.

create_repo (api/repos.rs:208 -> 210-219 -> 236): proof.consume then repo_store.init (blocking git init) run before db.create_repo; a DB failure orphans the freshly-initialized directory. No cleanup-on-failure exists on either path.

Fix direction: make disk materialization + proof spend adjacent to a guaranteed-successful DB write; on db.create_repo failure best-effort remove_dir_all(disk_path), and prefer deferring the Tigris upload until after the row commits. Verified by tracing the ordering during the #196 review.

Activity

  1. added
    kind:bugDefect fix — wrong or unsafe behavior
    sev:mediumDegraded but workaround exists
    subsystem:apiNode REST API request/response surface
    subsystem:attestationCertificates, anchoring, per-ref attestation
    subsystem:storageBlob/object store, Arweave, IPFS, archives
    on Jul 15, 2026
  2. beardthelion commented on Jul 21, 2026

    @beardthelion
    CollaboratorAuthor

    Scope update: #196 (commit 8d2fec1) resolved the fork_repo half of this.

    On a late create_repo failure, the fork tail now removes the on-disk mirror (repos.rs:1814) and deletes the already-uploaded Tigris archive (repos.rs:1820) before returning, so no row-less disk dir or orphaned object is left. The clone-that-exits-non-zero arm also cleans the mirror now (it previously didn't, which wedged the fork name on retry). Both are covered by RED-verified tests: fork_row_insert_failure_rolls_back_mirror_and_archive and fork_nonzero_clone_exit_cleans_dest_and_allows_retry.

    Two things still open here:

    1. The non-fork create_repo path is unchanged (repos.rs:272, state.db.create_repo(&record).await? — still a bare ? with no remove_dir_all on the error arm). A transient insert failure after init_bare still orphans the freshly-initialized dir with no row; a retry then 500s at init_bare ("repository already exists") while get_repo reports the name free, wedging it until manual cleanup. This is the remaining work for fork_repo / create_repo orphan on-disk repos (and a spent iCaptcha proof / uploaded object) on late DB failure #205 — mirror the fork remediation onto this path.

    2. The spent iCaptcha proof on a failed fork is now a deliberate ordering, not an oversight. The proof is consumed before the durable upload, so a failed fork does spend it and a retry needs a fresh proof. I kept that order on purpose: deferring the consume until after the upload would let two requests carrying the same verified proof both pay for a git clone --mirror before one loses the consume race. If you'd still rather make the recovery idempotent (so a retry reuses the proof), that's a real option — flagging it as a design call rather than a straight bug.

    Suggest narrowing this issue to item 1 (the non-fork create_repo disk orphan).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    crate:nodegitlawb-node — the serving node and REST APIkind:bugDefect fix — wrong or unsafe behaviorsev:mediumDegraded but workaround existssubsystem:apiNode REST API request/response surfacesubsystem:attestationCertificates, anchoring, per-ref attestationsubsystem:storageBlob/object store, Arweave, IPFS, archives

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions