Status: Fixed by #107.
Finding ID: EF-008
Last verified: master@038f9fa6a6a9947cace0f77160a22c38276ed79b
Source: repository/Qodo follow-up audit, 2026-09-20
Finding
tests/e2e/policy.py:save_last_failed() writes JSON directly to the final last-failed file with Path.write_text().
An interruption, process termination, disk error, or partial write can leave the file truncated/corrupt. load_last_failed() catches JSON/read failures and returns None, so corruption is then indistinguishable from there being no saved failure state.
PRKS already uses the safer temporary-file + atomic replacement pattern for E2E timing persistence, so the two related persistence files currently have inconsistent durability semantics.
Why it matters
--last-failed is part of the recommended fast agent/developer loop. Losing its state after an interrupted write is not catastrophic, but it weakens a workflow specifically intended to avoid expensive full-suite reruns and can produce a misleading "no last-failed state" result.
Recommended direction
Write the serialized state to a sibling temporary file, flush/close it, and atomically replace the destination with os.replace() (or reuse a small existing atomic-JSON helper if one already fits).
Add tests covering a normal round trip and preservation of the previous valid file when replacement/write fails before commit.
Assessment
- Priority: P3
- Impact: low-medium reliability / developer experience
- Effort: low
- Change risk: very low
- Type: testing / reliability / developer workflow
This issue records a finding for review. It does not authorize implementation.
Finding
tests/e2e/policy.py:save_last_failed()writes JSON directly to the final last-failed file withPath.write_text().An interruption, process termination, disk error, or partial write can leave the file truncated/corrupt.
load_last_failed()catches JSON/read failures and returnsNone, so corruption is then indistinguishable from there being no saved failure state.PRKS already uses the safer temporary-file + atomic replacement pattern for E2E timing persistence, so the two related persistence files currently have inconsistent durability semantics.
Why it matters
--last-failedis part of the recommended fast agent/developer loop. Losing its state after an interrupted write is not catastrophic, but it weakens a workflow specifically intended to avoid expensive full-suite reruns and can produce a misleading "no last-failed state" result.Recommended direction
Write the serialized state to a sibling temporary file, flush/close it, and atomically replace the destination with
os.replace()(or reuse a small existing atomic-JSON helper if one already fits).Add tests covering a normal round trip and preservation of the previous valid file when replacement/write fails before commit.
Assessment
This issue records a finding for review. It does not authorize implementation.