Repository navigation
fix: bound activity retries instead of retrying forever - #216
Merged
Merged
Conversation
Every activity Flows schedules now uses a shared bounded retry policy (internal/retry: 2s initial interval, x2 backoff, 200s cap, 15 attempts), the one payment-initiation activities already used. - InfiniteRetryContext is renamed LedgerRetryContext and bounded; it keeps its VALIDATION/CONFLICT/NO_SCRIPT/COMPILATION_FAILED/INSUFFICIENT_FUND non-retryable codes. PaymentInitiationRetryContext shares the same base. - Trigger activities (ListTriggers, EvalTriggerVariables, InsertTriggerOccurrence, SendEventForTriggerTermination) had no RetryPolicy, i.e. unlimited attempts; production saw InsertTriggerOccurrence reach attempt 20,916. - Instance/stage bookkeeping activities in Initiate, Run and Config.run had no RetryPolicy either. Only activities scheduled after deploy are affected: already-scheduled activities keep the policy recorded in their ActivityTaskScheduled event, and activity options are not part of replay command matching.
Replace the per-package bookkeeping/trigger helpers and the hand-built stage contexts with retry.ActivityContext / retry.ShortActivityContext, unexport the policy constants and keep the timing rationale in one place.
NumaryBot
reviewed
Sep 29, 2026
NumaryBot
approved these changes
Sep 29, 2026
NumaryBot
left a comment
Contributor
There was a problem hiding this comment.
The required automated review completed with no remaining findings.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Most activities Flows schedules can retry forever.
InfiniteRetryContext(send/update stages:CreateTransaction,DebitWallet, …) sets noMaximumAttempts.ListTriggers,EvalTriggerVariables,InsertTriggerOccurrence,SendEventForTriggerTermination) and the instance/stage bookkeeping activities (InsertNewInstance,UpdateInstance,InsertNewStage, …) set noRetryPolicyat all, i.e. Temporal's unlimited default.Any deterministic failure missing from a
NonRetryableErrorTypeslist therefore wedges the workflow for good, with nothing surfaced to the caller. In production we have seenInsertTriggerOccurrencereach attempt 20,916 andEvalTriggerVariablesspin indefinitely.Fix
One bounded policy for every activity, matching the one
PaymentInitiationRetryContextalready used: 2s initial interval, ×2 backoff, 200s cap, 15 attempts (~30–43 min worst case before giving up).internal/retrypackage:retry.ActivityContext(ctx, timeout, nonRetryableCodes...)andretry.ShortActivityContext(ctx)(10s per attempt).InfiniteRetryContext→LedgerRetryContext, bounded, same non-retryable codes (VALIDATION, CONFLICT, NO_SCRIPT, COMPILATION_FAILED, INSUFFICIENT_FUND).PaymentInitiationRetryContext— same numbers, now built on the shared helper.retry.ShortActivityContext, keeping their 10s timeout.No activity is left unbounded. Signal waits and delays schedule no activities; child workflow options are unchanged.
Tests
stages/internal/context_test.go— both stage contexts pinned against literal values; policies don't share their non-retryable slices.stages/send/run_test.go—CreateTransactionalways failing retryably is called exactly 15 times, then the workflow fails with that error.internal/triggers/workflow_trigger_retry_test.go,internal/workflow/run_retry_test.go— trigger and bookkeeping activities give up after 15 attempts and fail their workflow.go build ./...,go vet -tags it ./internal/...andgo test -race -tags it ./...all pass.Deploying this
Changing activity options is replay-safe (they are not part of command matching). But a retry policy is captured when the activity is scheduled, so this only applies to activities scheduled after the deploy. Activities already retrying forever keep their old policy and must be terminated or reset by hand.
Behavioural changes worth a second opinion
RunID-ActivityID), so they can't double-post. A manual re-run (new instance or Temporal reset) gets new keys — check ledger/wallet state first if an attempt might have committed before failing.UpdateInstancenever succeeded). This is the main trade-off; I judged it better than wedging forever.Follow-up
The bound is still applied per call site. A
WorkflowOutboundInterceptordefaultingRetryPolicy/MaximumAttemptsfor everyExecuteActivitywould enforce it for future activities too.Companion PR: #215 (non-retryable expression errors in triggers, independent).