Skip to content

v10 slice 2: Android queue store, executor, and events - #42

Open
dmurphy5 wants to merge 3 commits into
dylan/v10-1-js-layerfrom
dylan/v10-2-android
Open

dmurphy5 wants to merge 3 commits into
dylan/v10-1-js-layerfrom
dylan/v10-2-android

Conversation

@dmurphy5

@dmurphy5 dmurphy5 commented Sep 24, 2026 •

Copy link
Copy Markdown

Summary

This PR makes the Android side real. After it, a request that the app queues with mutate() is sent by Android, retried by Android, and reported back by Android, whether or not the app is open.

Before (v9): Android was a courier. JavaScript handed it one upload and listened for the result. Android did not know what a request was for, could not retry an HTTP failure, and could not answer "what is in the queue" after a restart.

After (v10): Android owns the queue. Each request is a small folder on disk with the request, its body, and its status. A background worker takes requests from that folder, sends them, and writes the result to a journal. JavaScript reads the journal at startup and gets every result it missed.

How a request moves through Android

flowchart LR
    A[mutate from JS] --> B[Write the request to disk]
    B --> C[Start a background worker]
    C --> D[Send the HTTP request]
    D --> E{Response?}
    E -- 2xx --> F[Write result to journal]
    E -- network error or 5xx --> G[Wait, then retry]
    G --> D
    E -- 401 or 403 --> H[Park until the app gives a new token]
    H --> D
    E -- other 4xx or deadline passed --> F
    F --> I[Tell JS]
    I --> J[JS acknowledges]
    J --> K[Delete the request and its files]
Loading

What each step guarantees:

  1. Write first. The request is on disk before the worker starts. A crash cannot lose it.
  2. Retry with backoff. Network errors and server errors retry with a growing wait, up to the request's deadline (14 days by default). The app does nothing.
  3. Auth parking. A 401 or 403 stops the request without retrying. When the app has a new token, it calls updateHeaders(), and every parked request continues.
  4. Journal before telling JS. The result is on disk before JavaScript hears about it. If the app is closed, the result waits.
  5. Exactly one result. A crash between "write result" and "delete request" cannot send the request twice. The worker checks the journal before it sends.

Status of a request

stateDiagram-v2
    [*] --> queued
    queued --> running
    running --> queued: retry after a wait
    running --> awaiting_auth: 401 or 403
    awaiting_auth --> queued: updateHeaders
    queued --> paused: pause()
    paused --> queued: resume()
    running --> completed
    running --> error: 4xx, missing file, or deadline
    running --> cancelled: cancel()
    completed --> [*]: after JS acknowledges
    cancelled --> [*]: after JS acknowledges
    error --> [*]: cancel() or a new mutate() with the same id
Loading

The app sees these states through getRequests() and the state event feed, so a screen can show a list of uploads with progress.

Things to know

  1. Upgrade from v9. At the first launch after the upgrade, results that v9 left in its journal become read-only rows so the app can see them. Unfinished v9 uploads are cancelled; the app re-sends them.
  2. Large files. A chunked upload sends a big file as parts, and parts that the server accepted are kept across restarts. The v9 chunked engine is reused.
  3. Ten-minute limit. When Android starts a worker in the background, the worker gets about ten minutes. A single large body that does not finish in that time restarts later. Large bodies should use parts. This is documented in the README.
  4. Backups. The queue holds auth headers. The host app must set allowBackup="false". Diana does.

What to look at

  1. The retry table in RetryClassifier.kt: does each HTTP status go where you expect?
  2. WorkerOps.kt, function settle: journal first, then tell JS.
  3. QueueStoreTest.kt: the crash-in-the-middle tests.

Test Plan

What's required for testing (prerequisites)?

JDK 17. For a device run, the example app in example/RNBGUExample; its README has a step-by-step script.

What are the steps to reproduce (after prerequisites)?

cd example/RNBGUExample/android
./gradlew :react-native-background-upload:testDebugUnitTest   # 307 tests, 0 failures

Nothing has run on a device yet. That is the next step.

Compatibility

OS Implemented
iOS ❌ see the iOS PR
Android ✅

Checklist

  • I have tested this on a device and a simulator
  • I added the documentation in README.md
  • I updated the typed files (TS)
  • I've added Detox End-to-End Test(s)
  • I've created a snack to demonstrate the changes

🤖 Generated with Claude Code

@dmurphy5
dmurphy5 added this pull request to stack #44 September 24, 2026 20:26
dmurphy5 and others added 2 commits September 24, 2026 16:41
Android implements the v10 contract pinned in the codegen spec. The v9
chunked manifest store generalizes into QueueStore: one durable entry per
mutate() (id, key, vars, descriptor, staged body, state, attempts, bytes,
expiresAt, header generation, deliveries) written tmp+fsync+rename before
any attempt. BodyStaging writes JSON and multipart bodies to library files,
copies single-file bodies, and moves chunked sources as v9 did.

WorkManager runs one worker per entry. WorkerOps applies the retry table
from the plan: accept rules with bodyIncludes, transient network/5xx/408/
429 with jittered backoff and nextAttemptAt, 401/403 parking with a header
generation that updateHeaders bumps, per-request exempt lists, expiry at
expiresAt with bytes kept. The transfer semaphore (4) and chunked window
(3) stay. Pause gates the queue; cancel settles a live entry and forgets a
settled one; ack is void and idempotent and forgets only a matching
generation. The journal writes every outcome before onSettled emits it;
onState carries a full RequestRow, onProgress is throttled 1 s / 10 min,
onAttempt is live-only. RequestIndex backs the synchronous getRequests.
First launch imports v9 journal entries as legacy rows and cancels v9 work.

Logic lives in JVM-testable classes (QueueController, WorkerOps,
EnqueueRules, EntryTransitions); the module and workers are thin shells.
217 unit tests, including crash-mid-write cases for the store.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Reliability: begin() applies an own-generation journal record instead of
re-running, so a crash between the journal write and the store update
cannot double-send. A failed journal write on settle holds the record in
memory and retries; nothing is acked that was not written. cancel() stops
the work in a finally block, journals before it saves, and applies an
existing record of its generation rather than adding a second outcome.
Listening state and the deliveries count are decided under one journal
lock, and a live settle goes to the listener that drained, not the newest
module instance.

Contract: vars and data arrive as JSON text; "null" is a body; GET with a
body rejects E_INVALID. Attempt events are completed or error only; pause,
cancel, and supersede emit none. A settled entry's cancel forgets its
unacked outcomes. attempts reset only on reopen. Under pause only an
accepted response settles. A same-id enqueue over a legacy row adopts its
v9 manifest. The response body cap applies while streaming. A prune never
deletes an eventId a row or a new record names. Rows carry live bytesSent.

Platform: a headless run stopped at JobScheduler's 10-minute limit takes
the transient path with backoff. Header validation messages carry the
name and offset, never the value. README documents the headless limit and
the allowBackup requirement.

Simplification: one EntryRun attempt path shared by simple and chunked
transfers; classifyFailure is the only failure rule; dead code removed.
Tests: 289 (was 217), with a TransferHost seam for the run loop.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@dmurphy5
dmurphy5 marked this pull request as ready for review September 25, 2026 15:56
The test cleared the listener with stopListening(listener), but from the
second iteration on the listener was the previous iteration's owner, so
nothing was cleared. When the settle won the race, the journal stamped one
live delivery and the drain counted a second one. CI failed about one run
in five. Now each iteration clears whatever listener is set and asserts
that none is set before the race starts.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants