Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
56 changes: 36 additions & 20 deletions docs/best-practices/error-handling.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -20,26 +20,20 @@ For background on how Temporal represents and propagates failures, see

## Categorize failures {/* #categorize-failures */}

When an operation fails, the appropriate response depends on the nature of the failure. Failures fall into three
categories based on whether retrying can resolve them.
When an operation doesn't succeed, the appropriate response depends on whether retrying can resolve it.
[Durable Execution](/durable-execution) sorts these outcomes into three categories: transient failures, permanent
failures, and negative results.

### Transient failures

A transient failure is a one-off event that resolves on its own without intervention. For example, a Worker happens to
make a network request at the exact moment an administrator replaces a network cable. The cause is unlikely to affect
future requests.
A transient failure is one where retrying the operation can eventually succeed. Some transient failures are one-time
events. For example, a Worker happens to make a network request at the exact moment an administrator replaces a network
cable, and the next request succeeds. Others are intermittent. For example, a service that uses rate limiting rejects
requests once the threshold is reached, but accepts requests again after the rate limiter resets.

Transient failures are resolved by retrying the operation shortly after the failure. Temporal's default
[Retry Policy](/encyclopedia/retry-policies) handles transient failures automatically.

### Intermittent failures

An intermittent failure is one that recurs but resolves over time. For example, a service that uses rate limiting will
reject requests once the threshold is reached, but will accept requests again after the rate limiter resets.

Intermittent failures require retries spaced out over a longer period. Configure your
[Retry Policy](/encyclopedia/retry-policies) with an appropriate `backoffCoefficient` and `maximumInterval` to avoid
overwhelming the failing service.
Temporal's default [Retry Policy](/encyclopedia/retry-policies) handles transient failures automatically. For
intermittent failures, configure the Retry Policy with an appropriate `backoffCoefficient` and `maximumInterval` to
space retries out over a longer period and avoid overwhelming the failing service.

### Permanent failures

Expand All @@ -48,8 +42,30 @@ to an invalid email address will continue to fail no matter how many times the o
to correct the email address.

Permanent failures cannot be resolved through retries. They require different input data, a code fix, or some external
intervention. Mark these errors as [non-retryable](#non-retryable-errors) to fail fast instead of consuming resources on
retries that will not succeed.
intervention. Mark these errors as [non-retryable](#non-retryable-errors) so the Activity stops retrying instead of
consuming resources on retries that will not succeed.

Failing the Activity fast doesn't have to fail the Workflow. Durable Execution keeps the Workflow's state and progress,
so the Workflow can catch the failure, wait for corrected data through a Signal or Update, and then run the Activity
again. See the [Resumable Activity pattern](/design-patterns/resumable-activity) and
[Pause a Workflow on failure and resume it after a fix](/guides/recover-without-restart). Bugs in Workflow code work
differently: they fail the Workflow Task, not the Workflow Execution, so the Workflow Execution stays open until you
deploy a fix.

### Negative results

Some outcomes aren't failures. They are your business logic reaching an expected, but negative, result: a customer
outside the service area, an order exceeding a credit limit, an expired promotion code, or a card with insufficient
funds.

Durable Execution doesn't decide whether a negative result is a dead end or worth waiting on. That depends on your
domain, so the decision belongs in your code. You can:

- Return the result as a value and branch on it in the Workflow.
- Raise a [non-retryable](#non-retryable-errors) Application Failure and handle it in the Workflow, for example with
[compensation](#saga-pattern).
- Treat it as transient: wait, for example for a Signal that the customer updated their payment details, and then try
again.

## Mark permanent errors as non-retryable {/* #non-retryable-errors */}

Expand All @@ -60,8 +76,8 @@ background on what Application Failures are and how the `non_retryable` flag wor
Use non-retryable errors for situations like:

- **Invalid input data**: A malformed email address, a negative payment amount, or a missing required field.
- **Business rule violations**: A customer outside the service area, an order exceeding credit limits, or an expired
promotion code.
- **Negative results you handle as failures**: A customer outside the service area, an order exceeding credit limits,
or an expired promotion code. See [Negative results](#negative-results).
- **Authorization failures**: The caller does not have permission to perform the operation.
- **Data validation errors**: A referenced record does not exist, or data fails integrity checks.

Expand Down
16 changes: 9 additions & 7 deletions docs/develop/python/best-practices/error-handling.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -18,11 +18,12 @@ This page shows you how to build on these capabilities to create robust error ha
**Key concepts:**

Not all failures should be handled the same way.
**Transient failures** (like brief network hiccups) resolve on their own and should be retried immediately.
**Intermittent failures** (like rate limiting) need increasing delays between retries.
**Permanent failures** (like invalid input) won't resolve through retries and need different data or code changes.
**Transient failures** resolve if you retry. Some are one-time events (like brief network hiccups) that a quick retry resolves; others are intermittent (like rate limiting) and need increasing delays between retries.
**Permanent failures** (like invalid input or a bug) won't resolve through retries and need different data or code changes.
**Negative results** (like insufficient funds) aren't failures. Your code decides whether to return them, fail on them, or wait and try again.
See [Durable Execution](/durable-execution) for how Temporal handles each category.

Temporal distinguishes between **Workflow Task failures** (bugs that can be fixed with redeployment) and **Workflow Execution failures** (business logic failures that should stop the Workflow).
Temporal distinguishes between **Workflow Task failures** (bugs that can be fixed with redeployment) and **Workflow Execution failures** (failures your code raises on purpose to stop the Workflow).
Task failures retry automatically so you can fix and redeploy without losing state.
Execution failures require you to explicitly raise an `ApplicationError`.

Expand Down Expand Up @@ -165,7 +166,8 @@ This way, you can decide whether an error should not be retried automatically by
This can be useful for deliberately failing a Workflow due to bad input data, rather than waiting for a timeout to elapse. You can alternately specify a list of errors that are non-retryable in your Activity [Retry Policy](/develop/python/activities/timeouts#activity-retries).

This puts the Workflow Execution in "Failed" state with no automatic retries.
Use this for permanent failures where retrying won't help—like the customer being too far away.
Use this when retrying won't help and the Workflow should stop, like the customer being too far away.
That's a [negative result](/durable-execution#negative-results), and failing the Workflow is one way to handle it.

### Trigger a Workflow Task retry

Expand Down Expand Up @@ -364,7 +366,7 @@ An `ApplicationError` with `non_retryable=True` will never retry, regardless of

Use non-retryable errors for:
- Invalid input data that prevents the Activity from proceeding
- Business rule violations
- Negative results you choose to handle as failures, like business rule violations
- Authorization failures

**Use this sparingly.**
Expand Down Expand Up @@ -544,7 +546,7 @@ These retry automatically, letting you fix bugs and redeploy without losing Work
**Workflow Execution failures** occur when Workflow code raises a Temporal exception like `ApplicationError`.
These put the Workflow in "Failed" state with no automatic retries.

Example of a permanent failure that should fail the Workflow:
Example of a negative result that this Workflow handles by failing:

```python
if distance.kilometers > MAX_DELIVERY_DISTANCE:
Expand Down
15 changes: 12 additions & 3 deletions docs/encyclopedia/application-failures.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,9 @@ Understanding which failures Temporal handles and which ones your application mu
## Platform failures vs application failures {/* #platform-vs-application */}

Failures fall into two categories based on where they are detected and mitigated: platform failures and application failures.
This is a different question from whether retrying helps, which [Durable Execution](/durable-execution) answers with three categories: transient failures, permanent failures, and negative results.
Platform failures are transient.
Application failures can fall into any of the three.

### Platform failures

Expand All @@ -29,10 +32,16 @@ Platform failures are resolved through **forward recovery**: the system retries
### Application failures

Application failures are generated by your code.
They indicate an issue with your application logic, such as invalid input data, a business rule violation, or a failed call to an external service.
They indicate something your application logic has to account for, such as a failed call to an external service, invalid input data, a bug, or a business rule violation.

Application failures do not resolve on their own through retries alone.
Recovering from an application failure may require fixing a bug, passing different input data, or performing some external mitigation.
Some application failures are transient.
When an Activity's call to an external service fails because the service is temporarily down, the Retry Policy retries the Activity until the service recovers.

Others are permanent and do not resolve through retries alone.
Recovering from them may require fixing a bug, passing different input data, or performing some external mitigation.
Durable Execution keeps the Workflow's progress while that happens, so the Workflow can continue after the fix.

Still others are [negative results](/durable-execution#negative-results), such as a business rule violation, that your code chooses to represent as failures.

Application failures often involve **backward recovery**: the system undoes some of the work that has already been performed to return to a previous state.
For example, if a payment step fails after inventory has already been reserved, the application may need to release that inventory.
Expand Down
8 changes: 4 additions & 4 deletions docs/encyclopedia/detecting-workflow-failures.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -109,10 +109,10 @@ It's worth mentioning that although you can extend the timeout up to the maximum
## Detecting Workflow Task Failures

Use the `TemporalReportedProblems` Search Attribute to detect Workflows with failed Workflow Tasks.
A failed Workflow Task does not cause the Workflow to fail. Some Tasks within a Workflow may be intended to fail.
For example, a Workflow Task may check a remote data source for new messages. If there aren't any, the Task will fail as intended.
If your Task has code to handle the failure, the Workflow will proceed.
However, if your Workflow has a Task that fails and the failure is not handled, the Workflow will continue to run, but will not complete.
A failed Workflow Task does not cause the Workflow to fail.
A Workflow Task fails when the Worker can't complete it, for example because of a bug, such as an unhandled exception or panic in Workflow code, or a non-determinism error.
The Temporal Service retries the Workflow Task, and the Workflow Execution stays open with its progress intact, but it will not complete until you deploy a fix.
For causes and fixes, see the [Workflow Task errors reference](/references/workflow-task-errors).
Detecting Workflows in this state is a common troubleshooting issue.

To identify Workflows with Task failures, you can use the Temporal Web UI. See [Task Failures View](/web-ui/#task-failures-view) for more details.
Expand Down
15 changes: 9 additions & 6 deletions docs/encyclopedia/retry-policies.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -221,10 +221,12 @@ occur, the execution is not retried.

#### Non-Retryable Errors for Activities

When writing software applications, you will encounter three types of failures: transient, intermittent, and permanent.
While transient and intermittent failures may resolve themselves upon retrying without further intervention, permanent
failures will not. Permanent failures, by definition, require you to make some change to your logic or your input.
Therefore, it is better to surface them than to retry them.
When writing software applications, you will encounter two types of failures: transient and permanent. Transient
failures, including intermittent ones such as rate limiting, may resolve themselves upon retrying without further
intervention. [Permanent failures](/durable-execution#permanent-failures) will not. Permanent failures, by definition,
require you to make some change to your logic or your input. Therefore, it is better to surface them than to retry them.
Some business outcomes, such as insufficient funds, aren't failures at all, but you can model these
[negative results](/durable-execution#negative-results) as non-retryable errors too.

Non-Retryable Errors are errors that will not be retried, regardless of a Retry Policy.

Expand Down Expand Up @@ -363,8 +365,9 @@ set to `true` will always be non-retryable.
</Tabs>

For example, checking for bad input data is a reasonable time to use a non-retryable error. If the Activity cannot
proceed with the input it has, that error should be surfaced immediately so that the input can be corrected on the next
attempt.
proceed with the input it has, that error should be surfaced immediately so that the input can be corrected. The
Workflow keeps its progress, so it can wait for the corrected input and then run the Activity again. See the
[Resumable Activity pattern](/design-patterns/resumable-activity).

If responsibility for your application is distributed across multiple maintainers, or if you are developing a library to
integrate into somebody else's application, you can think of the decision to hardcode non-retryable errors as following
Expand Down
6 changes: 4 additions & 2 deletions docs/encyclopedia/workers/tasks.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -116,8 +116,10 @@ The table summarizes the differences:
**Workflow Task failure example:** A new deployment introduces non-determinism, existing Workflows fail Workflow Tasks,
and the executions stay Open and retry. After deploying a fix, the Workflows automatically continue.

**Workflow Execution failure example:** A payment Activity fails due to a declined card, the failure propagates
uncaught, and the Workflow closes as Failed. The customer updates payment details and restarts the order.
**Workflow Execution failure example:** A payment Activity raises a non-retryable failure for a declined card, the
failure propagates uncaught, and the Workflow closes as Failed. The customer updates payment details and restarts the
order. A declined card is a [negative result](/durable-execution#negative-results), so failing the Workflow is a choice.
The Workflow could instead catch the failure and wait for a Signal with new payment details, keeping its progress.

## What is an Activity Task? {/* #activity-task */}

Expand Down
Loading