Skip to content

[Audit Finding] Make restore journal and component renames power-loss durable #110

Description

@Fooftilly

Status: Fixed by #133.

Finding ID: EF-017
Last verified: master@9e75411d05180252691751b2fdeebdfd1518644f
Source: repository crash-consistency / backup-restore audit, 2026-09-21

Security: This is a restore durability / crash-consistency finding, not a suspected vulnerability or sensitive security disclosure.

Finding

PRKS has a thoughtfully journaled restore transaction that can recover from process-level crashes injected around component renames. However, the persistence boundary is weaker than the recovery protocol assumes: journal updates and live/rollback component renames use os.replace() without fsyncing the containing directory/directories.

_atomic_write_json() fsyncs the temporary journal file and then replaces the final journal pathname, but does not fsync the journal parent directory. _move_old_component(), _install_new_component(), and _restore_rollback_component() similarly rename canonical/rollback/staged paths without syncing the affected parent directories.

On POSIX filesystems, fsync() of the file does not make the directory-entry rename durable. A sudden power loss or kernel/filesystem crash can therefore leave the on-disk journal and the component directory entries representing different transaction phases after reboot. The existing RestoreCrash tests model process interruption after system calls return; they do not model loss/reordering of directory metadata that was never durably flushed.

Evidence

backend/backup_restore.py::_atomic_write_json() currently does:

with open(tmp, "w", encoding="utf-8") as handle:
    json.dump(payload, handle, separators=(",", ":"), sort_keys=True)
    handle.write("\n")
    handle.flush()
    os.fsync(handle.fileno())
os.replace(tmp, path)

The file contents are flushed, but the parent directory is not fsynced after os.replace().

The restore state machine then relies on persisted journal flags around rename boundaries:

  • _move_old_component() persists old_move_started, performs os.replace(live, rollback), then persists old_moved.
  • _install_new_component() persists new_install_started, performs os.replace(staged, live), then persists new_installed.
  • _restore_rollback_component() moves rollback state back into the live location with os.replace().

Those operations can affect two different directories, for example live storage and .prks-maintenance/rollback/..., yet neither side is explicitly directory-fsynced.

PRKS already contains the correct pattern for a related canonical-file operation: backend/services/work_pdf_replace.py::fsync_managed_pdf_parent() exists specifically as a best-effort POSIX/Linux directory fsync after rename. Restore does not reuse or equivalently implement that durability boundary.

tests/test_backup_restore.py includes simulated crash coverage such as new_renamed:pdfs, which is valuable for logical recovery, but a Python exception after a completed rename cannot reproduce a machine crash that loses an unsynced directory entry or the final journal rename.

Why it matters

Backup restore is one of the highest-value crash-consistency boundaries in PRKS because it replaces the canonical database and managed file trees together.

The journal is intended to answer, after restart, which old components were moved and which new components were installed. That answer is trustworthy only if the journal transition and the filesystem rename it describes have explicit durability ordering.

Without directory fsyncs, an abrupt machine/power failure can theoretically produce states the logical crash tests cannot reach, for example:

  • a journal transition survives while the corresponding component rename does not;
  • a component rename survives while the subsequent journal transition does not;
  • rollback/live directory entries have different persistence outcomes across their two parent directories;
  • the recovery code makes a decision from a journal that is individually valid JSON but not durably ordered with the filesystem state it describes.

This is substantially rarer than an ordinary process crash, but restore is exactly the operation where rare power-loss inconsistency is worth handling deliberately.

Recommended direction

Define and implement an explicit durable-rename primitive for restore rather than scattering ad-hoc fsync() calls.

On POSIX/Linux, after a rename/replacement:

  • fsync the destination parent directory;
  • when moving between different parent directories, fsync both source and destination parent directories as required for the intended durability guarantee;
  • after atomically replacing the restore journal, fsync its parent directory before treating that journal phase as durable.

Use the ordering implied by the restore state machine: a journal phase should only be considered durably persisted after its file replacement and parent-directory fsync have completed, and component rename durability should be established before persisting the state that claims the rename completed.

Keep Windows behavior portable. Directory fsync is platform-specific and may be unsupported; use a small best-effort/platform-aware helper similar in spirit to fsync_managed_pdf_parent() rather than making restore unusable on Windows.

Do not replace the existing journal/state-machine design; it is useful and already covers process-level interruption. This finding is about making its persistence assumptions true for machine/power-loss crashes.

Add focused tests around the durability helper and ordering. Unit tests can spy on rename/fsync ordering; true power-loss behavior should not be represented by flaky destructive CI tests.

Also review _atomic_write_json() users such as staging metadata, but prioritize the restore journal and canonical component moves because they participate directly in crash recovery.

Existing roadmap / issue overlap

Open and closed issues plus active/recent pull requests were searched for restore journal durability, directory fsync, power-loss restore recovery, crash consistency, backup restore rename durability, and atomic journal writes.

No existing issue or PR was found that owns this concrete persistence-ordering gap.

Roadmap #57 covers broader Library Safety, History & Maintenance and backup health, but does not specify machine-crash durability of the restore transaction.

Existing restore crash tests cover logical/process interruption rather than filesystem metadata durability.

EF-011 / #91 concerns recoverability of post-delete cleanup and is a different transaction boundary.

Assessment

  • Priority: P2
  • Impact: high if triggered; low-frequency machine/power-loss restore inconsistency
  • Effort: low-medium
  • Change risk: medium because ordering and cross-platform filesystem semantics matter
  • Type: reliability / backup-restore / crash consistency / storage

This issue records a finding for review. It does not authorize implementation.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    acceptedarea:correctnessaudit-findingEvery EF-xxx issue gets thispriority:P2Important defect or architectural/reliability problem that should be addressed

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions