diff --git a/docs/best-practices/error-handling.mdx b/docs/best-practices/error-handling.mdx index 467d18d179..1887d8108d 100644 --- a/docs/best-practices/error-handling.mdx +++ b/docs/best-practices/error-handling.mdx @@ -12,7 +12,7 @@ tags: --- Temporal automatically retries failed Activities and recovers from infrastructure failures through -[Durable Execution](/evaluate/why-temporal). But not all failures should be retried. This page covers how to categorize +[Durable Execution](/durable-execution). But not all failures should be retried. This page covers how to categorize failures, when to mark errors as non-retryable, and how to handle failures that retries cannot resolve. For background on how Temporal represents and propagates failures, see diff --git a/docs/develop/go/integrations/google-adk.mdx b/docs/develop/go/integrations/google-adk.mdx index 292309bef7..d14ecf2a3e 100644 --- a/docs/develop/go/integrations/google-adk.mdx +++ b/docs/develop/go/integrations/google-adk.mdx @@ -11,7 +11,7 @@ description: Run Google ADK agents with durable execution using the Temporal Go --- Temporal's integration with [Google ADK](https://google.github.io/adk-docs/) (`adk-go`) gives your agents -[Durable Execution](/temporal#durable-execution): the agent's orchestration loop runs inside a Temporal Workflow, each +[Durable Execution](/durable-execution): the agent's orchestration loop runs inside a Temporal Workflow, each LLM call becomes a durable Temporal Activity, and any tool that does I/O runs as an Activity too — so every step is retried, timed out, recorded in Workflow history, and replayable after crashes or restarts. diff --git a/docs/develop/python/integrations/google-adk.mdx b/docs/develop/python/integrations/google-adk.mdx index d799593373..0fafb8c862 100644 --- a/docs/develop/python/integrations/google-adk.mdx +++ b/docs/develop/python/integrations/google-adk.mdx @@ -14,7 +14,7 @@ description: Temporal's Google ADK integration lets you run [Google ADK](https://google.github.io/adk-docs/) agents inside Temporal Workflows, so an agent keeps its place across Worker restarts, deploys, and transient failures. -Temporal gives your agent code [Durable Execution](/temporal#durable-execution). The Google ADK +Temporal gives your agent code [Durable Execution](/durable-execution). The Google ADK gives you the agent itself: model calls, tools, multi-agent handoffs, and MCP. The integration connects the two so that you write an ordinary ADK agent and run it as a Workflow, without managing sessions, a database, or your own retry logic. diff --git a/docs/develop/python/integrations/google-genai.mdx b/docs/develop/python/integrations/google-genai.mdx index 79d5a9c577..9a15e66036 100644 --- a/docs/develop/python/integrations/google-genai.mdx +++ b/docs/develop/python/integrations/google-genai.mdx @@ -14,7 +14,7 @@ description: Temporal's Google GenAI integration lets you call [Gemini](https://ai.google.dev/gemini-api/docs) models from inside Temporal Workflows, so a sequence of model calls keeps its place across Worker restarts, deploys, and transient failures. -Temporal gives your code [Durable Execution](/temporal#durable-execution). The +Temporal gives your code [Durable Execution](/durable-execution). The [Google Gen AI SDK](https://googleapis.github.io/python-genai/) gives you the model API: content generation, automatic function calling, chat sessions, structured output, files, and MCP. The integration connects the two so that you write ordinary Gemini SDK code and run it as a Workflow, without writing your own retry loop or checkpointing. diff --git a/docs/develop/python/integrations/langsmith.mdx b/docs/develop/python/integrations/langsmith.mdx index 38d3123eff..f6178c8874 100644 --- a/docs/develop/python/integrations/langsmith.mdx +++ b/docs/develop/python/integrations/langsmith.mdx @@ -16,7 +16,7 @@ import { ReleaseNoteHeader } from '@site/src/components'; Temporal's LangSmith integration lets you trace AI agent Workflows in [LangSmith](https://smith.langchain.com/) alongside every LLM call, tool execution, and Temporal operation. -Temporal gives your agent code [durable execution](/temporal#durable-execution). +Temporal gives your agent code [durable execution](/durable-execution). LangSmith adds the observability side, so you can inspect LLM inputs and outputs, follow a request from the Client through to the model, and compare runs over time. diff --git a/docs/develop/python/integrations/strands-agents.mdx b/docs/develop/python/integrations/strands-agents.mdx index c1c5c4fa06..4834296445 100644 --- a/docs/develop/python/integrations/strands-agents.mdx +++ b/docs/develop/python/integrations/strands-agents.mdx @@ -13,7 +13,7 @@ description: Run Strands Agents AI workflows with durable execution using the Te import { ReleaseNoteHeader } from '@site/src/components'; Temporal's integration with [Strands Agents](https://strandsagents.com/) is an [SDK Plugin](/develop/plugins-guide) that -gives your Strands agents [Durable Execution](/temporal#durable-execution) via the Temporal platform. The plugin routes +gives your Strands agents [Durable Execution](/durable-execution) via the Temporal platform. The plugin routes model invocations, tool calls, MCP tool calls, and hooks through Temporal Activities, so every step your agent takes is recorded in Workflow history and can survive crashes, restarts, and infrastructure failures. diff --git a/docs/develop/typescript/integrations/langsmith.mdx b/docs/develop/typescript/integrations/langsmith.mdx index 276682906f..aab951d2ca 100644 --- a/docs/develop/typescript/integrations/langsmith.mdx +++ b/docs/develop/typescript/integrations/langsmith.mdx @@ -17,7 +17,7 @@ import { ReleaseNoteHeader } from '@site/src/components'; Temporal's LangSmith integration lets you trace AI agent Workflows in [LangSmith](https://smith.langchain.com/) alongside every LLM call, tool execution, and Temporal operation. -Temporal gives your agent code [durable execution](/temporal#durable-execution). LangSmith adds +Temporal gives your agent code [durable execution](/durable-execution). LangSmith adds the observability side, so you can inspect LLM inputs and outputs, follow a request from the Client through to the model, and compare runs over time. diff --git a/docs/develop/typescript/integrations/strands-agents.mdx b/docs/develop/typescript/integrations/strands-agents.mdx index df1fb43168..061d809624 100644 --- a/docs/develop/typescript/integrations/strands-agents.mdx +++ b/docs/develop/typescript/integrations/strands-agents.mdx @@ -11,7 +11,7 @@ description: Run Strands Agents AI Workflows with Durable Execution using the Te --- Temporal's integration with [Strands Agents](https://strandsagents.com/) is an [SDK Plugin](/develop/plugins-guide) that -gives your Strands agents [Durable Execution](/temporal#durable-execution) via the Temporal platform. The plugin routes +gives your Strands agents [Durable Execution](/durable-execution) via the Temporal platform. The plugin routes model invocations, tool calls, MCP tool calls, and hooks through Temporal Activities, so every step your agent takes is recorded in Workflow history and can survive crashes, restarts, and infrastructure failures. diff --git a/docs/encyclopedia/durable-execution.mdx b/docs/encyclopedia/durable-execution.mdx new file mode 100644 index 0000000000..fede5fed66 --- /dev/null +++ b/docs/encyclopedia/durable-execution.mdx @@ -0,0 +1,145 @@ +--- +id: durable-execution +title: What is Durable Execution? +sidebar_label: Durable Execution +description: Durable Execution preserves state and progress so code resumes after transient failures, waits out permanent ones, and leaves negative results to you. +slug: /durable-execution +toc_max_heading_level: 4 +tags: + - Durable Execution + - Concepts + - Failures +--- + +Durable Execution lets you write your code as though failures don't exist. +It preserves the state and progress of an operation as it runs. +When a failure occurs, execution resumes where it left off and continues until the operation completes. +You write the business logic, and the platform makes sure that logic runs to completion. + +In Temporal, the operation is a [Workflow Execution](/workflow-execution), and the steps it takes are +[Activities](/activities), Timers, and messages. + +What Durable Execution can do depends on the kind of failure: + +- **[Transient failures](#transient-failures)** go away if you try again. Durable Execution keeps trying until the + operation succeeds. +- **[Permanent failures](#permanent-failures)** never go away on their own. Durable Execution can't fix them, but it + keeps the operation's state so that the operation resumes once someone fixes the cause. +- **[Negative results](#negative-results)** aren't failures. Your code decides what to do with them. + +## How Durable Execution preserves progress {/* #how-it-works */} + +Temporal records each step of a Workflow Execution in its [Event History](/workflow-execution/event#event-history): +Activities scheduled and their results, Timers started and fired, and messages received. +If the Worker running a Workflow Execution stops, another Worker replays the Event History to rebuild the Workflow's +state, including local variables, and continues from the point where execution stopped. +Steps that already completed don't run again. Replay uses their recorded results instead. + +For a step-by-step walkthrough, see [How Temporal works](/encyclopedia/architecture/how-temporal-works) and +[Event History](/encyclopedia/event-history). + +## Transient failures {/* #transient-failures */} + +A transient failure is one where trying the operation again can eventually succeed. +Some transient failures are one-time events, such as a network request sent at the moment a cable is replaced. +Others are intermittent, such as a rate-limited API that rejects requests until its limit resets. +Durable Execution treats both the same way. + +Examples of transient failures: + +- Process crashes and Worker restarts +- Hardware failures +- Network outages +- Timeouts +- A temporary outage in a dependency, such as an external service that is down + +Durable Execution keeps trying until a transient failure clears, so you can write most of your code as though these +failures don't happen: + +- **Activities retry under a [Retry Policy](/encyclopedia/retry-policies).** The default Retry Policy retries with + exponential backoff and no limit on attempts. A Start-To-Close or Heartbeat Timeout counts as a failed attempt and is + retried too. +- **Work moves off a crashed Worker.** The Temporal Service hands the crashed Worker's Workflow Tasks to another Worker, + which replays the Event History and continues. Activities that were running on the crashed Worker time out and retry + on another Worker. + +You can customize how often Temporal tries again. +Set the Retry Policy's initial interval, backoff coefficient, and maximum interval to space out attempts. +This matters most for timeouts and rate limits, where retrying too often adds load to a dependency that is already +struggling. +If you cap retries with a maximum number of attempts or a Schedule-To-Close Timeout, a failure that outlasts the cap +reaches your Workflow code as an Activity Failure. + +## Permanent failures {/* #permanent-failures */} + +A permanent failure is one where trying again never succeeds. +Something has to change first: the input data, the code, or the dependency. +That intervention might be manual or automated, but without it the failure keeps happening. + +Examples of permanent failures: + +- Bad or invalid data +- A bug in the code +- An application or dependency that goes permanently offline + +In most systems, a permanent failure loses the work that was in progress. +With Durable Execution, the state and progress of the operation are stored durably. +Once an intervention clears the failure, the operation resumes where it left off and runs to completion. + +How that works in Temporal depends on where the failure happens: + +- **A bug in Workflow code.** An unhandled exception, a panic, or a non-determinism error fails the Workflow Task, not + the Workflow Execution. The Temporal Service retries the Workflow Task, and the Workflow Execution stays open with its + state intact. Deploy a fix, and the next attempt replays the Event History and + continues. See [Workflow Task failures vs Workflow Execution failures](/encyclopedia/application-failures#task-vs-execution). +- **A bug in Activity code.** Errors thrown from an Activity are retryable unless you mark them non-retryable, so the + Activity keeps retrying under its Retry Policy. Deploy a fix, and the next attempt runs the fixed code. +- **Bad input data.** Raise a non-retryable Application Failure from the Activity so it stops retrying. The Workflow can + catch the failure, wait for corrected data through a [Signal or Update](/encyclopedia/workflow-message-passing), and + run the Activity again. See the [Resumable Activity pattern](/design-patterns/resumable-activity) and + [Pause a Workflow on failure and resume it after a fix](/guides/recover-without-restart). +- **A dependency that is down for good.** [Pause the Activity](/activity-operations/pause) (Public Preview) to stop + retries while you repair or replace the dependency, then unpause it. +- **Steps that already ran with a bug.** [Reset](/workflow-execution/event#reset) the Workflow Execution to a point + before the bug affected it. The new run continues from that point with the fixed code. + +Durable Execution preserves progress only while the Workflow Execution stays open. +If Workflow code throws an Application Failure, or a Workflow timeout expires, the Workflow Execution closes, and +recovering means resetting it or starting a new one. +Decide which failures should end the Workflow Execution and which should wait for a fix. + +For specific errors and how to fix them, see +[Troubleshoot Workflow and Activity execution failures](/troubleshooting/execution-failures). + +## Negative results {/* #negative-results */} + +Some results aren't failures. +They are your business logic reaching an expected, but negative, outcome. + +Examples of negative results: + +- Inventory out of stock +- A credit card with insufficient funds +- No ride-share driver available + +Durable Execution doesn't decide whether a negative result is a dead end or worth waiting on. +The right response depends on your domain, so the decision belongs in your code. +Common choices are: + +- Return the result as a value, and branch on it in the Workflow. For example, notify the customer that an item is out + of stock. +- Raise a non-retryable Application Failure, and handle it in the Workflow. For example, run compensating Activities + with the [Saga pattern](/design-patterns/saga-pattern). +- Treat the result as transient. Retry, or wait for a Timer or a Signal, such as a restock notification, and then try + again. Durable Execution keeps the operation going until it completes. + +For guidance on modeling these choices, see [Error handling](/best-practices/error-handling). + +## Summary {/* #summary */} + +| | Transient failure | Permanent failure | Negative result | +| :-------------------- | :------------------------------------------------- | :-------------------------------------------------------------- | :------------------------------------ | +| **Does retrying help?** | Yes, eventually | No, not until something changes | Your code decides | +| **Who resolves it** | Temporal, by trying again | An intervention: a fix, new data, or a replacement dependency | Your Workflow code | +| **What Temporal does** | Retries Activities and moves work off failed Workers | Preserves state and progress so the operation resumes after the intervention | Records the result for your code to act on | +| **Examples** | Crashes, network outages, timeouts | Bad data, bugs, a dependency that is gone for good | Out of stock, insufficient funds | diff --git a/docs/encyclopedia/index.mdx b/docs/encyclopedia/index.mdx index 81e9d7af00..e6af4c168c 100644 --- a/docs/encyclopedia/index.mdx +++ b/docs/encyclopedia/index.mdx @@ -10,6 +10,7 @@ sidebar_label: Encyclopedia The following Encyclopedia pages describe the concepts, components, and features of Temporal in detail: - [Temporal](/temporal) +- [Durable Execution](/durable-execution) - [Temporal architecture](/encyclopedia/architecture/temporal-architecture) - [Temporal SDKs](/encyclopedia/architecture/temporal-sdks) - [Temporal Client](/encyclopedia/temporal-client) diff --git a/docs/encyclopedia/temporal.mdx b/docs/encyclopedia/temporal.mdx index d15e4c8735..b39209911f 100644 --- a/docs/encyclopedia/temporal.mdx +++ b/docs/encyclopedia/temporal.mdx @@ -28,9 +28,12 @@ The Temporal Platform handles these types of problems, allowing you to focus on ## Durable Execution {/* #durable-execution */} -Durable Execution in the context of Temporal refers to the ability of a Workflow Execution to maintain its state and progress even in the face of failures, crashes, or server outages. -This is achieved through Temporal's use of an [Event History](/workflow-execution/event#event-history), which records the state of a Workflow Execution at each step. -If a failure occurs, the Workflow Execution can resume from the last recorded event, ensuring that progress isn't lost. +Durable Execution preserves the state and progress of a Workflow Execution as it runs, so that after a failure the Workflow Execution resumes where it left off. +Temporal records each step in an [Event History](/workflow-execution/event#event-history) and replays it to restore state. +Transient failures, such as crashes and network outages, are retried until they succeed. +Permanent failures, such as bugs and bad data, need an intervention, but the Workflow Execution keeps its progress so it can resume after the fix. + +For details, see [What is Durable Execution?](/durable-execution) ## What is the Temporal Platform? {/* #temporal-platform */} diff --git a/docs/glossary.md b/docs/glossary.md index 11316d0f01..2f4ce0291a 100644 --- a/docs/glossary.md +++ b/docs/glossary.md @@ -188,10 +188,10 @@ in your Temporal Service to facilitate migrating your Visibility data from one d -#### [Durable Execution](/temporal#durable-execution) +#### [Durable Execution](/durable-execution) -Durable Execution in the context of Temporal refers to the ability of a Workflow Execution to maintain its state and -progress even in the face of failures, crashes, or server outages. +Durable Execution lets you write code as though failures don't exist. It preserves the state and progress of a Workflow +Execution so that, after a failure, execution resumes where it left off and continues until it completes. diff --git a/sidebars.js b/sidebars.js index bcb90c2f2f..3626b37cd1 100644 --- a/sidebars.js +++ b/sidebars.js @@ -1936,6 +1936,7 @@ module.exports = { }, items: [ 'encyclopedia/temporal', + 'encyclopedia/durable-execution', { type: 'category', label: 'Architecture', diff --git a/vercel.json b/vercel.json index dfc19b63c3..2c16d6d725 100644 --- a/vercel.json +++ b/vercel.json @@ -2515,12 +2515,7 @@ }, { "source": "/encyclopedia/durable-execution", - "destination": "/temporal#durable-execution", - "permanent": true - }, - { - "source": "/durable-execution", - "destination": "/temporal#durable-execution", + "destination": "/durable-execution", "permanent": true }, {