diff --git a/README.md b/README.md index 1d8fe8a..cd98e52 100644 --- a/README.md +++ b/README.md @@ -2,11 +2,14 @@ Deploy and operate [Temporal](https://temporal.io/) Workers on serverless compute with help from a coding agent. The skill guides an agent through the complete AWS Lambda lifecycle: scoping, access checks, Worker implementation, packaging, deployment, Temporal registration, verification, troubleshooting, updates, and rollback. +It also supports GCP Cloud Run Worker Pools through a separate pool-based path, without changing the Lambda workflow. + > [!WARNING] -> This skill is in Public Preview and will continue to evolve. Pin the Temporal SDK, serverless Worker package, and CLI versions for long-lived projects. +> This skill is in Public Preview and will continue to evolve. Pin the Temporal SDK and CLI versions for long-lived projects, plus any provider-specific package you use, such as the AWS Lambda serverless Worker package or a Cloud Run OpenTelemetry helper. > [!NOTE] -> Temporal Serverless Workers on AWS Lambda are in Public Preview and are available to all Temporal Cloud customers without an access request. AWS Lambda is currently the only compute provider supported by this skill. +> Temporal Serverless Workers on AWS Lambda are in Public Preview and are available to all Temporal Cloud customers without an access request. +> GCP Cloud Run is also in Public Preview and available without an access request. ## What the skill can do @@ -19,26 +22,39 @@ Deploy and operate [Temporal](https://temporal.io/) Workers on serverless comput - Publish immutable Lambda versions, update deployments, and roll back safely. - Add OpenTelemetry observability with the AWS Distro for OpenTelemetry. - Configure self-hosted Temporal deployments that meet the serverless prerequisites. +- Build ordinary long-lived Workers for GCP Cloud Run Worker Pools without a provider-specific package or handler. +- Containerize and deploy one immutable Worker Pool per build ID. +- Configure separate Cloud Run runner and invoker service accounts. +- Diagnose WCI-driven pool resizing, scale-in interruption, quotas, and pool annotations. ## Support | Area | Supported | |---|---| -| Compute | AWS Lambda — Public Preview | +| Compute | AWS Lambda and GCP Cloud Run — Public Preview | | Temporal | Temporal Cloud and self-hosted Temporal Service | | SDKs | Go, Python, TypeScript, Java, .NET | -| Other compute providers | Not currently supported | +| Other compute providers | Not currently supported by this skill | For Temporal Cloud, the Namespace must be hosted on AWS. The Namespace and Lambda function may be in different AWS regions. +For Cloud Run, the Namespace must be hosted on GCP. The Namespace and Worker Pool may be in different GCP regions. + ## Before you start -Before starting, make sure you can sign in to: +Before starting, make sure you can sign in to the accounts for the compute provider you plan to use. + +### AWS Lambda - An AWS account with permission to inspect and create the required Lambda, IAM, CloudFormation, and logging resources. - A Temporal Cloud Namespace hosted on AWS, or a compatible self-hosted Temporal Service. -You do not need to install or configure the AWS CLI, `tcld`, or the Temporal CLI before you begin. The skill checks what is already available and can help set up the tools and supported login flows needed for the task. If you prefer not to install a CLI, or a login method is unavailable, it can guide you through the corresponding Temporal Cloud UI or AWS console steps instead. It never asks you to paste credentials or secrets into the conversation. +### GCP Cloud Run + +- A GCP project with permission to inspect and create Worker Pools, Artifact Registry images, IAM bindings, Secret Manager secrets, and logs. +- A Temporal Cloud Namespace hosted on GCP, or a compatible self-hosted Temporal Service. + +You do not need to install or configure the AWS CLI, `gcloud`, `tcld`, or the Temporal CLI before you begin. The skill checks what is already available and can help set up the tools and supported login flows needed for the task. If you prefer not to install a CLI, or a login method is unavailable, it can guide you through the corresponding Temporal Cloud UI, AWS console, or Google Cloud console steps instead. It never asks you to paste credentials or secrets into the conversation. ## Installation @@ -98,7 +114,13 @@ Package this Java Worker as a shaded jar and deploy it to Lambda. Deploy this .NET Worker to Lambda with a runtime-specific publish. ``` -For a new deployment, the skill follows five stages: +```text +Deploy a Go Temporal Worker to a GCP Cloud Run Worker Pool. +``` + +For a new deployment, the skill follows five stages. + +### AWS Lambda 1. **Scope** — confirm the SDK, compute provider, Namespace, region, and resource-naming prefix. 2. **Access** — verify AWS and Temporal identities and permissions, then present the exact billable resources for approval. @@ -106,10 +128,20 @@ For a new deployment, the skill follows five stages: 4. **Connect** — configure Temporal's invocation role, register the Worker Deployment Version, validate the Task Queue binding, and set the version current. 5. **Verify and hand back** — run a Workflow, confirm two independent health signals, inventory every created resource, and offer teardown. +### GCP Cloud Run + +1. **Scope** — confirm the SDK, GCP-hosted Namespace, project, region, and resource-naming prefix. +2. **Access** — verify GCP and Temporal identities and permissions, then present the exact billable resources for approval. +3. **Build** — author an ordinary long-lived Worker, containerize it, and deploy a dedicated Worker Pool for the build ID. +4. **Connect** — configure the invoker identity, register the Worker Deployment Version, verify Task Queue bindings, and set the version current. +5. **Verify and hand back** — run a Workflow, confirm its history and pool logs, inventory every created resource, and offer teardown. + Nothing is created before you approve the resource list. Troubleshooting and inspection requests skip the deployment walkthrough and begin with read-only diagnostics. ## Important operating constraints +### AWS Lambda + - Serverless Workers and their APIs are Public Preview, not generally available. - Every Workflow must use a Worker Versioning behavior: `Pinned` or `AutoUpgrade`. - The deployment name and build ID in Worker code must exactly match the registered Worker Deployment Version. @@ -118,12 +150,20 @@ Nothing is created before you approve the resource list. Troubleshooting and ins - Secrets belong in a secret store for shared or production deployments, not plaintext environment variables. - Temporal creates and manages the Worker Controller Instance (WCI); this skill never creates or manages it directly. +### GCP Cloud Run + +- Cloud Run uses ordinary long-lived Worker APIs; there is no per-invocation handler or serverless Worker package. +- Every Workflow must use `Pinned` or `AutoUpgrade`, and each build ID maps to a dedicated immutable Worker Pool. +- Cloud Run has no invocation deadline, but scale-in can interrupt Activities; use graceful shutdown and Heartbeats for resumable work. +- Do not share the Task Queue with an independently managed long-lived fleet. +- A minimum instance count of zero permits scaling to zero; a nonzero minimum intentionally keeps capacity running. + ## Repository guide | Path | Contents | |---|---| | [`SKILL.md`](SKILL.md) | Core workflow, safety gates, provider rules, and reference routing | -| [`references/concepts.md`](references/concepts.md) | Architecture, invocation flow, autoscaling, lifecycle, constraints, and use cases | +| [`references/concepts.md`](references/concepts.md) | Shared serverless concepts (release status, compute providers, Worker Versioning, WCI) and the AWS Lambda invocation model: invocation flow, autoscaling, lifecycle, constraints, and use cases | | [`references/wci.md`](references/wci.md) | Shared WCI lifecycle, inputs, Workflow ID pattern, inspection commands, and health interpretation | | [`references/aws-lambda/sdk-go.md`](references/aws-lambda/sdk-go.md) | Go package, API, handler, build, packaging, Lambda deployment values, tuned defaults, connection configuration, and OpenTelemetry integration | | [`references/aws-lambda/sdk-python.md`](references/aws-lambda/sdk-python.md) | Python package, API, handler, build, packaging, Lambda deployment values, tuned defaults, connection configuration, OpenTelemetry integration, and diagnostics | @@ -136,6 +176,18 @@ Nothing is created before you approve the resource list. Troubleshooting and ins | [`references/aws-lambda/versioning.md`](references/aws-lambda/versioning.md) | Immutable releases, updates, and rollback | | [`references/aws-lambda/observability.md`](references/aws-lambda/observability.md) | Shared ADOT Collector configuration, X-Ray enablement, and IAM permissions | | [`references/aws-lambda/self-hosted.md`](references/aws-lambda/self-hosted.md) | Self-hosted Temporal prerequisites and configuration | +| [`references/gcp-cloud-run/sdk-go.md`](references/gcp-cloud-run/sdk-go.md) | Go Worker construction, versioning behavior, connection configuration, image packaging, scale-in safety, and observability | +| [`references/gcp-cloud-run/sdk-python.md`](references/gcp-cloud-run/sdk-python.md) | Python Worker construction, versioning behavior, connection configuration, image packaging, scale-in safety, and observability | +| [`references/gcp-cloud-run/sdk-typescript.md`](references/gcp-cloud-run/sdk-typescript.md) | TypeScript Worker construction, versioning behavior, connection configuration, image packaging, scale-in safety, and observability | +| [`references/gcp-cloud-run/sdk-java.md`](references/gcp-cloud-run/sdk-java.md) | Java Worker construction, versioning behavior, connection configuration, image packaging, scale-in safety, and observability | +| [`references/gcp-cloud-run/sdk-dotnet.md`](references/gcp-cloud-run/sdk-dotnet.md) | .NET Worker construction, versioning behavior, connection configuration, image packaging, scale-in safety, and observability | +| [`references/gcp-cloud-run/setup.md`](references/gcp-cloud-run/setup.md) | End-to-end Worker Pool deployment, registration, verification, and teardown | +| [`references/gcp-cloud-run/iam.md`](references/gcp-cloud-run/iam.md) | Operator permissions, runner and invoker service accounts, and Terraform IAM setup | +| [`references/gcp-cloud-run/constraints.md`](references/gcp-cloud-run/constraints.md) | Pool lifecycle, autoscaling, scale-in interruption, and mixed-fleet constraints | +| [`references/gcp-cloud-run/diagnostics.md`](references/gcp-cloud-run/diagnostics.md) | Worker Pool scaling, WCI Activity failures, annotations, quotas, and Worker logs | +| [`references/gcp-cloud-run/versioning.md`](references/gcp-cloud-run/versioning.md) | One Worker Pool per build ID, immutable releases, and rollback | +| [`references/gcp-cloud-run/observability.md`](references/gcp-cloud-run/observability.md) | Cloud Run OpenTelemetry helpers, Google-built Collector sidecar, Cloud Logging, and provider-specific scaling signals | +| [`references/gcp-cloud-run/self-hosted.md`](references/gcp-cloud-run/self-hosted.md) | Self-hosted Temporal Service prerequisites and GCP identity configuration | | [`assets/`](assets/) | CloudFormation templates for Temporal invocation roles | ## Feedback diff --git a/SKILL.md b/SKILL.md index 08ca47c..ea589b4 100644 --- a/SKILL.md +++ b/SKILL.md @@ -1,6 +1,6 @@ --- name: temporal-serverless -description: 'Deploy and operate Temporal Workers on serverless compute (AWS Lambda) driven by the Worker Controller Instance (WCI). Use when the user mentions: "serverless worker", "Temporal serverless", "Worker Controller Instance", "WCI", "deploy Temporal worker on Lambda", "Lambda packaging", "Lambda timeout", "WCI inspection", "CloudFormation Temporal".' +description: 'Deploy and operate Temporal Workers on serverless compute (AWS Lambda, GCP Cloud Run) driven by the Worker Controller Instance (WCI). Use when the user mentions: "serverless worker", "Temporal serverless", "Worker Controller Instance", "WCI", "deploy Temporal worker on Lambda", "Lambda packaging", "Lambda timeout", "WCI inspection", "CloudFormation Temporal", "Cloud Run worker", "deploy Temporal worker on Cloud Run".' disable-model-invocation: true --- @@ -8,16 +8,18 @@ disable-model-invocation: true ## Overview -This skill helps users deploy and operate Temporal Workers on serverless compute. Instead of a long-lived process, Temporal invokes the Worker on demand through the Worker Controller Instance (WCI); the Worker processes available Tasks and shuts down, scaling to zero when idle. The skill produces Worker code, deployment configuration, connection configs, and packaging steps for the chosen SDK, and walks users through troubleshooting when serverless Workers aren't picking up Tasks. +This skill helps users deploy and operate Temporal Workers on serverless compute. On AWS Lambda, Temporal invokes the Worker on demand through the Worker Controller Instance (WCI); the Worker processes available Tasks and shuts down, scaling to zero when idle. On GCP Cloud Run, the WCI instead resizes a Worker Pool whose instances run ordinary long-lived Workers. The skill produces Worker code, deployment configuration, connection configs, and packaging steps for the chosen SDK, and walks users through troubleshooting when serverless Workers aren't picking up Tasks. ## Supported compute providers | Cloud provider | Compute service | Support | Reference directory | |---|---|---|---| | AWS | Lambda | Supported — Public Preview, open to all Temporal Cloud customers | `references/aws-lambda/` | -| GCP | Cloud Run | Not supported | — | +| GCP | Cloud Run | Supported — Public Preview, open to all Temporal Cloud customers | `references/gcp-cloud-run/` | -Only a provider marked Supported is covered. If a request names another, say it is not supported and stop; do not adapt a supported provider's material to it. **Never let the provider be an unstated assumption:** when the request does not name one, it is confirmed in the step 1 questions, not silently defaulted. +Only a provider marked Supported is covered. If a request names another, say it is not supported and stop; do not adapt a supported provider's material to it. **Never let the provider be an unstated assumption:** when the request does not name one, it is settled in step 1, derived from the Namespace or asked, and stated to the user, not silently defaulted. + +**Select the provider before loading lifecycle guidance.** For AWS Lambda, read `references/concepts.md` (shared concepts and the AWS Lambda invocation model) plus the selected AWS SDK reference. For GCP Cloud Run, read `references/gcp-cloud-run/constraints.md` plus the selected Cloud Run SDK reference. Every supported provider's directory carries the same shared layout — `setup.md`, `iam.md`, `versioning.md`, `diagnostics.md`, `observability.md`, `self-hosted.md` — plus one `sdk-.md` file for each supported SDK. Paths below are written `references//…`; substitute the directory from the table. Provider-specific commands, templates, permissions, SDK APIs, and defaults live there — this file stays at the workflow level. When a step needs concrete commands or SDK details, go to the reference file named at the end of that step. @@ -29,13 +31,21 @@ Every supported provider's directory carries the same shared layout — `setup.m | Java | `references/aws-lambda/sdk-java.md` | | .NET | `references/aws-lambda/sdk-dotnet.md` | +| SDK language | GCP Cloud Run reference | +|---|---| +| Go | `references/gcp-cloud-run/sdk-go.md` | +| Python | `references/gcp-cloud-run/sdk-python.md` | +| TypeScript | `references/gcp-cloud-run/sdk-typescript.md` | +| Java | `references/gcp-cloud-run/sdk-java.md` | +| .NET | `references/gcp-cloud-run/sdk-dotnet.md` | + **Public Preview is not GA.** The APIs are still evolving and may change: pin SDK and CLI versions for anything long-lived, and read the installed package's actual API surface rather than writing from memory. ## Deployment workflow Follow these steps in order. Each step is provider-neutral; the concrete commands, templates, and options live in the reference file named at the end of the step. -**Open a new deployment with a plain-language summary of the run.** Before the step 1 questions, tell the user in a few sentences what is about to happen: that this creates real resources in their cloud account which cost money for as long as they exist; that you will ask about a handful of things, then show an exact list of what you are about to create and wait for approval, and that nothing is created before that approval; that the middle of the run is unattended; and that it ends with a Workflow they can watch execute, an inventory of everything created, and an offer to remove it all. Name the five stages below in ordinary words. Do not explain Temporal or serverless compute; keep it short enough to read at a glance. +**Open a new deployment with a plain-language summary of the run.** Before the step 1 questions, tell the user in a few sentences what is about to happen: that this creates real resources in their cloud account which cost money for as long as they exist; that you will ask about a handful of things, then show an exact list of what you are about to create and wait for approval, and that nothing is created before that approval; that the middle of the run is mostly unattended, though some providers need one short step in the user's own terminal, such as entering an API key, and you will say exactly when; and that it ends with a Workflow they can watch execute, an inventory of everything created, and an offer to remove it all. Name the five stages below in ordinary words. Do not explain Temporal or serverless compute; keep it short enough to read at a glance. **Lay it out as bullets, with the five stages as sub-bullets under "How it goes" — one stage per line, never chained into a single run-on bullet.** Follow this shape: @@ -48,10 +58,10 @@ Follow these steps in order. Each step is provider-neutral; the concrete command > - **Build** — write, package, deploy the Worker. > - **Connect** — bind the Task Queue, set the version current. > - **Verify and hand back.** -> - Nothing gets created before you approve that list. After approval the middle stretch runs unattended. +> - Nothing gets created before you approve that list. After approval the middle stretch runs mostly unattended; if I need you to run a short step in your own terminal, such as entering an API key, I'll say exactly when. > - At the end you get a Workflow you can watch execute, a full inventory of everything created, and an offer to remove it all. -**Write the summary provider-neutral, because at that point you do not know the provider.** It is one of the things step 1 asks. Say "your cloud account", never the name of a provider you have not been told. The same applies to the account, Namespace, and region: if a cheap read-only call has already told you (see step 1), name what you actually found; otherwise leave it out rather than filling it in with a plausible guess. +**Write the summary provider-neutral, because at that point you do not know the provider.** It is settled in step 1. Say "your cloud account", never the name of a provider you have not been told. The same applies to the account, Namespace, and region: if a cheap read-only call has already told you (see step 1), name what you actually found; otherwise leave it out rather than filling it in with a plausible guess. Skip the summary for troubleshooting, inspection, and configuration-change tasks. Someone whose Worker is not being invoked does not need an overview of a deployment they have already done. @@ -64,18 +74,18 @@ Where the harness has a todo list, use it *in addition to* the printed checklist **Word each item as plain language about what happens, not as a compressed step title,** and name both sides concretely — the confirmed compute provider and Temporal, never "both sides." Follow this shape: > **Scope** -> ✅ Confirm SDK (Go), compute provider (AWS Lambda), Namespace (``), and naming prefix (``) +> ✅ Confirm SDK (Go), compute provider (``), Namespace (``), and naming prefix (``) > > **Access** -> ⏳ Check credentials and permissions on AWS and on Temporal, then show the exact list of resources to be created and wait for your approval +> ⏳ Check credentials and permissions for `` and Temporal, then show the exact list of resources to be created and wait for your approval > > **Build** > ⬜ Write the Worker against the installed package's real API -> ⬜ Cross-compile, package, deploy the compute unit, wait for it to report ready +> ⬜ Build for the target platform, package, deploy the compute unit, wait for it to report ready > > **Connect** -> ⬜ Create the role Temporal assumes to invoke the Worker -> ⬜ Register the Worker Deployment Version, confirm the validation invocation bound the Task Queue, set it current +> ⬜ Grant Temporal permission to inspect and start or resize the compute unit +> ⬜ Register the Worker Deployment Version, confirm registration bound the Task Queue, set it current > > **Verify and hand back** > ⬜ Start a Workflow and confirm it executes, from both the Temporal side and the provider's logs @@ -85,26 +95,32 @@ Where the harness has a todo list, use it *in addition to* the printed checklist |---|---|---| | Scope | 1 | SDK, compute provider, Namespace, and naming prefix are all confirmed by the user. | | Access | 2 | Compute provider and Temporal both authenticated, permissions confirmed, and the list of resources to create approved. | -| Build | 3–4 | The compute unit is deployed and reports ready, built for the architecture it runs on. | +| Build | 3–4 | The provider-specific package or image is published, deployed, and reports ready for its target platform. | | Connect | 5–6 | The Task Queue is bound and the version is current. | | Verify and hand back | 7–8 | A Workflow completed, two independent signals agree, the inventory is delivered, and teardown has been offered. | **A step is complete when its verification passed — not when its command exited zero.** Several commands in this workflow exit clean having done nothing: the traffic-shifting and key-revocation commands no-op when their confirmation prompt goes unanswered, and providers return from create and update calls while the resource is still settling. Check an item off against state you read back, not against an exit code. When a step's verification fails, say which step you are on and what it is blocked on rather than moving down the list. -1. **Scope the task.** Identify the SDK language (Go, Python, TypeScript, Java, or .NET), the deployment target (Temporal Cloud or self-hosted — self-hosted has its own server prerequisites), the compute provider, and whether this is a new setup, a configuration change, or troubleshooting. Confirm the deployment target is compatible with the chosen provider — see "A Namespace on the target cloud provider is required" under Provider-neutral principles. Ensure a Temporal client/CLI is available and authenticated to the target. Each changes the specifics. → `references/concepts.md` for what the user is building; `references//setup.md` for the compatibility and client-setup details. +1. **Scope the task.** Identify the SDK language (Go, Python, TypeScript, Java, or .NET), the deployment target (Temporal Cloud or self-hosted — self-hosted has its own server prerequisites), the compute provider, and whether this is a new setup, a configuration change, or troubleshooting. Confirm the deployment target is compatible with the chosen provider — see "A Namespace on the target cloud provider is required" under Provider-neutral principles. Ensure a Temporal client/CLI is available. For Lambda, it must also be authenticated to the target at this point. For Cloud Run, check only that the `temporal` CLI is installed: its authenticated profile is created during the API-key hand-off after approval (`references/gcp-cloud-run/setup.md`). Each changes the specifics. For Lambda, read `references/concepts.md`; for Cloud Run, read `references/gcp-cloud-run/constraints.md`. Use `references//setup.md` for compatibility and client-setup details. + + **Derive the compute provider from the Namespace's cloud provider.** An AWS-hosted Namespace uses AWS Lambda; a GCP-hosted Namespace uses GCP Cloud Run. State the derived provider with the Namespace choice, and if the request names a provider, check that it matches. When `tcld` cannot be used, the Namespace's region settles it: `aws-us-east-1` is AWS, `gcp-us-central1` is GCP. Ask the provider as a structured question only for a self-hosted Temporal Service, where either is possible; carry each option's status from the support table in its description, and note that Activity duration can decide it, because Lambda caps an invocation at 15 minutes and Cloud Run does not. + + **Any `tcld` call can start a browser sign-in when its session has expired**, including a read-only one. Tell the user before the first `tcld` call, so a login is never a silent side effect of discovery. If you do not know whether the session is valid, ask the user to run `tcld login` in their own terminal first. If a `tcld` call stalls without output, stop it and ask the user to log in rather than waiting. + + **Let the user pick the Namespace from a list; never make them retype one.** Namespace names are long and error-prone — a generated suffix on an account ID, `-.`. Where control-plane access is available: - **Put the compute provider in that batch of questions as a confirmable default, not a free choice.** Pre-select the supported provider from the table above and carry its support status in the option's description. The user confirms rather than chooses, so it costs no extra turn, but the provider is never something they were assumed into. Skip the question only when the request already names a provider. Do not restate any of this in a paragraph before the questions; the option description is where it belongs. + - List names with `tcld namespace list`, following `nextPageToken` with `--page-token` when it is set. + - Run `tcld namespace get -n ` for each candidate. Read its provider from `.spec.regionId.provider` (for example `CloudProviderGcp`), its region from `.spec.regionId.name`, and its authentication method from `.spec.authMethod`. + - For Cloud Run, which authenticates with an API key, offer only Namespaces whose authentication method accepts API keys. The provider decides eligibility; the region does not. - **Let the user pick the Namespace from a list; never make them retype one.** Namespace names are long and error-prone — a generated suffix on an account ID, `-.`. Where control-plane access is available, `tcld namespace list` returns the full Namespace objects, so one call gives every name with its region — and a region ID is provider-prefixed (`aws-…`, `gcp-…`), so the same response tells you each Namespace's provider. Only the prefix carries meaning; the region itself imposes no constraint. + If the user names a Namespace that is not in the list, confirm it with `tcld namespace get` before continuing: `NotFound` means it does not exist in this account. Present it like this: - **Offer the eligible Namespaces as the options**, each labelled with its region. - - **Summarize the ineligible ones in a single line** — "you also have 2 Namespaces on \, which this skill does not support" — rather than listing them individually or hiding them. A user who knows they have a Namespace and cannot find it in the list concludes the tool is broken; one line keeps them informed and explains the constraint. - - **Name the account you are listing from and confirm it is the intended one** before showing anything. A stale credential lists a real account that is not the one the user means to deploy into, and every option under it looks authoritative. - - **If more Namespaces are eligible than the question format can hold, print the labelled list and ask the user to name one.** Do not silently show only the first few. - - This also settles the compute-provider answer, since a Namespace can only be served by compute on its own cloud provider — so a mismatch is caught here rather than at connection time, several steps later. + - **Summarize the ineligible ones in a single line, with the reason** — for example "you also have 2 Namespaces on \, which this skill does not support", "1 GCP Namespace accepts only mTLS, which the Cloud Run API-key path cannot use", or "3 AWS Namespaces, which do not match the Cloud Run deployment you asked for" — rather than listing them individually or hiding them. A user who knows they have a Namespace and cannot find it in the list concludes the tool is broken; one line keeps them informed and explains the constraint. + - **Name the account you are listing from and confirm it is the intended one** before showing anything. A stale credential lists a real account that is not the one the user means to deploy into, and every option under it looks authoritative. `tcld account get` prints no account ID; in practice it is the suffix after the last `.` of the account's Namespace names, so when you name the account that way, say it was inferred. Namespaces from other accounts can appear in the same list. + - **If all eligible Namespaces fit in a structured question, offer all of them as options. Otherwise, print the complete labelled list, numbered, and ask the user to reply with the number (or name) of the one to use.** Never offer only a subset as options. **Degrade gracefully if `tcld` is not authenticated.** Ask the user for the Namespace name rather than stopping to fix the login — they can copy it from the Cloud UI, where it appears on the Namespace page and in the URL. Ask for its region in the same batch of questions: the name alone does not tell you the provider, and a mismatch missed here surfaces at connection time instead. @@ -116,6 +132,8 @@ Where the harness has a todo list, use it *in addition to* the printed checklist Offer exactly two options plus the free-text escape: the identifying prefix, and "no prefix" — some users genuinely own the account. Do not offer a second prefix string; the consequential choice is prefix versus none, and anything else goes in free text. When you offer "no prefix," say what it risks in the same breath: unprefixed names can collide with or shadow an existing deployment, and that surfaces as another team's Worker behaving oddly rather than as an error you will see. + **For Cloud Run, ask the GCP project and the pool's region as their own structured questions.** Run `gcloud projects list`, `gcloud config get-value project`, and `gcloud config get-value run/region` first. When more than one project is visible, ask for the project with the current one as the first option, labelled by its source ("your current gcloud config"), plus "Other"; ask for the region the same way. Only when a single project is visible may you state it instead of asking. **Never treat a non-specific reply such as "go ahead" or "defaults OK" as confirming a project, region, or Namespace:** ask again with a structured question. Confirm the Namespace, project, region, and prefix explicitly, using more than one round of questions when needed and never omitting options. Do not proceed on an unconfirmed value without naming it in step 2's approval list, and name the project and region there in any case rather than relying on ambient `gcloud` configuration. + 2. **Confirm you can make the required changes — before making any.** Determine which credentials are available (for the compute provider and for Temporal) and confirm the active identity actually has permission to make the changes the task needs — creating or updating compute resources, creating roles, registering deployment versions. Verify *both* sides: the compute provider AND Temporal access. Do not run account-mutating commands and let them fail partway. **If access is missing or unconfirmed, stop and ask the user how they want to proceed** — extend their identity's permissions, have an administrator make the change and hand back the result, or generate the commands for the user to run under a privileged identity. Changing a user's cloud account is consequential; confirm authorization and the preferred method first. → `references//iam.md` (exact permissions, compute-provider preflight) and `references//setup.md` (Temporal connection preflight). **Classify an authentication failure before acting on it — "not signed in" and "not permitted" have different fixes.** A failed preflight does not automatically mean the bottom row of the table below. An absent or expired credential is usually recoverable in this session, in under a minute. A caller that resolves but is denied a specific action is a real permissions problem. Never collect credentials in the conversation: no interactive credential-configuration wizards, and never ask the user to paste access keys, API keys, or session tokens. → `references//iam.md` (credential recovery). @@ -138,7 +156,17 @@ Where the harness has a todo list, use it *in addition to* the printed checklist **Do not self-select a row.** Drop to a lower one only after the choice above has been put to the user and the browser path chosen, or the login attempted and failed. When you hand off a runbook, say the offer stands — if the user authenticates and comes back, take the work over rather than leaving them to run the steps by hand. - **Before the first account-mutating command, list what you are about to create — with final names — and get approval.** Name the target account and region, then every resource: compute unit, execution role, infrastructure stack, log group, deployment name, and Task Queue. Say plainly that they are live and billable. This is the mirror of the inventory in step 8, and it is worth more here than there: it makes the naming prefix concrete while changing it is still free, and the deployment name, build ID, and Task Queue become expensive to change once step 3 compiles them into the Worker. Skip it only when nothing will be created — a troubleshooting or inspection task. + **Before the first account-mutating command, list what you are about to create — with final names — and get approval.** Say plainly that they are live and billable. This is the mirror of the inventory in step 8, and it is worth more here than there: it makes the naming prefix concrete while changing it is still free, and the deployment name, build ID, and Task Queue become expensive to change once step 3 compiles them into the Worker. Skip it only when nothing will be created — a troubleshooting or inspection task. + + #### AWS Lambda + + Name the target account and region, then every resource: compute unit, execution role, infrastructure stack, log group, deployment name, and Task Queue. + + #### GCP Cloud Run + + List the project and region; any APIs that still need enabling, and the Cloud Build submission; the Artifact Registry repository, if new, and the image; the Worker Pool; the runner service account and the dedicated invoker, with its Terraform state directory; the secret; the local CLI profile; the deployment name, build ID, and Task Queue; and logs. Say that the user creates the Temporal API key and enters it during the hand-off, and that they may also need to run `terraform apply` or the secret IAM grant in their own terminal if the agent's environment blocks them. Note that the Temporal deployment name is checked right after the hand-off, because the CLI profile does not exist before approval. + +### AWS Lambda: steps 3–8 3. **Author the Worker.** *Install the SDK's serverless Worker package before writing any code* — it is usually shipped separately from the main SDK — sometimes on its own version line, sometimes in lockstep with it, and in one SDK not separately at all — so having the base SDK installed does not mean it is importable. Then read the installed package's actual API surface and write against that; these are Public Preview APIs that drift between versions, and generating code from memory costs a build cycle. Entry-point names are not consistent between SDKs, so inspect first rather than pattern-matching from another language. Every Workflow must declare a versioning behavior (`Pinned` or `AutoUpgrade`), per-Workflow or as a Worker-level default — code without it fails at runtime. → `references//sdk-.md` (package, install, API inspection, entry point, handler shape, versioning behavior, tuned defaults). @@ -154,6 +182,22 @@ Where the harness has a todo list, use it *in addition to* the printed checklist **Do not write a teardown script before the user asks for one.** Generating it unprompted buries the inventory under a file they did not request, and the inventory is what they need in order to decide. End with a single line — *"Let me know if you want a teardown script to remove these resources"* — and stop there. Write the script, or run the teardown, when they take you up on it. → `references//setup.md` (Teardown). +### GCP Cloud Run: steps 3–8 + +Steps 3 and 5 need no API key: write and build the image, and apply Step 5's Terraform, while the user runs the key hand-off. Step 4 waits for the hand-off, so the checklist may mark Step 5 done first. Deliver the hand-off as `references/gcp-cloud-run/setup.md` describes: end the turn with the user's action first, and repeat it in full while the run is blocked on it. + +3. **Author the Worker.** Write an ordinary long-lived Worker that starts polling when the container starts. There is no Cloud Run serverless Worker package or per-invocation handler. Set the required Worker Versioning behavior and make the deployment name and build ID match the version to be registered. → the selected `references/gcp-cloud-run/sdk-.md`. + +4. **Package and deploy the compute unit.** Build a container for the target architecture, publish an immutable image, deploy a dedicated Worker Pool for this build ID, and wait for the pool to report ready. Do not redeploy a new build into a pool used by a live Worker Deployment Version. → `references/gcp-cloud-run/setup.md`, `references/gcp-cloud-run/versioning.md`. + +5. **Grant Temporal permission to scale the pool.** Keep the runner service account used by pool instances separate from the invoker service account Temporal impersonates. Create a dedicated invoker named with the agreed prefix; reuse an existing one only when the user names it and confirms they own its Terraform state. → `references/gcp-cloud-run/iam.md`. + +6. **Register the Worker Deployment Version, verify the registration bootstrap, then set it current.** Point the version at the pool and invoker service account, provide the complete scaler group supported by the installed CLI or omit the group, and confirm the expected Task Queue types are bound before shifting traffic. → `references/gcp-cloud-run/setup.md`. + +7. **Verify.** Start a Workflow, confirm its history progresses, and confirm the Worker Pool logs show startup, polling, and Task execution. If it does not progress, start with the version's expected Task Queue bindings, then follow the WDV/WCI/provider decision table. → `references/gcp-cloud-run/diagnostics.md`. + +8. **Hand back the inventory first; offer teardown as the closing note.** Include the project and region; the Artifact Registry repository, image tag, and digest; the Worker Pool; the runner and invoker service accounts, and whether this run created the invoker; the Terraform state directory; the secret and its versions; the deployment name, build ID, and Task Queue; the local CLI profile; and who created the Temporal API key. Follow the same inventory-before-teardown and approval rules as the Lambda path. → `references/gcp-cloud-run/setup.md`. + ## Working practices How to move through the workflow above. @@ -171,7 +215,7 @@ How to move through the workflow above. For anything that creates, updates, or deletes, name the resource and the target account or Namespace explicitly — an approval prompt should arrive with its justification already on screen, not after it. - **Read the current state instead of recalling it.** Check the installed package's API, the CLI's own `--help` for the flags you are about to pass, the compute unit's reported state, and the CLI version. Each of these has drifted in practice: a Public Preview SDK whose fields moved, a CLI too old to have the serverless subcommand at all, a resource that reports success while still settling. - **Do not chain `cd` with commands that create or modify files.** A compound `cd && ` triggers a manual approval prompt no matter how the user's permissions are configured, so scaffolding a project this way asks for approval on every run. Use absolute paths, or the tool's own directory flag (`go -C …`), and rely on the shell's working directory persisting between calls — the `cd` buys nothing and costs a prompt. Keep the command count down for the same reason: one `go get` covering both packages beats two. -- **Verify each step before building the next on top of it.** Compile the Worker before packaging it, confirm the package's target architecture before uploading, wait for the compute unit to be ready before publishing a build, and confirm the Task Queue is bound before shifting traffic. Deployment failures here surface far from their cause — an architecture or dependency mismatch appears only at first invocation, and a first-invocation failure appears as "the Worker is never invoked", several steps later. +- **Verify each step before building the next on top of it.** Compile the Worker before packaging it, confirm the package or image targets the platform it will run on, publish the immutable build before deploying compute, wait for the compute unit to be ready, and confirm the Task Queue is bound before shifting traffic. Deployment failures here surface far from their cause — an architecture or dependency mismatch may appear only when compute first starts, and a runtime startup failure later appears as "the Worker never polls." - **When something fails, read the actual error before changing anything.** Fetch the failure reason from the provider (deployment events, logs, status fields) and fix that. Do not retry the same command with variations, and do not start editing permissions or trust policies on the theory that the problem might be access — most first-invocation failures are not permission problems, and some failures are on Temporal's side and will reproduce no matter what you change. - **Treat the user's account as shared and pre-existing.** Assume other deployments, roles, and stacks are already there. Look before creating, extend rather than duplicate, and never delete or repurpose something you did not create without asking. When you do work around existing infrastructure — a different name, a reused role — say so explicitly in your summary rather than leaving it as a silent deviation. - **Confirm the end state from two independent signals.** A Workflow that completes in the Temporal UI *and* the Worker's own logs showing startup, Task Queue registration, and Task execution. One signal alone can mislead: a system Workflow that exists and is running proves nothing about invocation health, and a command that exits zero may have done nothing at all if it was waiting on a confirmation prompt. @@ -179,7 +223,9 @@ How to move through the workflow above. ## Never create or manage the WCI -Temporal creates and manages the WCI automatically once a Worker Deployment Version has a compute provider. Never create, start, or manage it yourself. Read `references/wci.md` for its lifecycle, inputs, inspection commands, and health interpretation, then use `references//diagnostics.md` for provider-specific failures. +Temporal creates and manages the WCI automatically once a Worker Deployment Version has a compute provider. Never create, start, or manage it yourself. Read `references/wci.md` for its lifecycle, inputs, inspection commands, and health interpretation. + +Then use the selected provider's diagnostics: `references/aws-lambda/diagnostics.md` for Lambda invocation failures, or `references/gcp-cloud-run/diagnostics.md` for Cloud Run pool-resizing failures. ## Provider-neutral principles @@ -188,24 +234,42 @@ Surface these early — they apply regardless of compute provider: - **A Namespace on the target cloud provider is required.** A Serverless Worker runs only on the cloud provider that hosts its Temporal Cloud Namespace — there is no cross-cloud pairing. Confirm the user has a Namespace on the provider they intend to run compute on *before* building anything; without one, the work stops there and they need either a Namespace on that provider or a different provider. A mismatch is not caught at deploy time — it fails later, at connection time. **Regions do not have to match:** a Namespace in one region can drive a compute unit in another, so never tell a user to move or re-create a Namespace to line up regions. - **Use `tcld` for every Temporal Cloud control-plane operation** — accounts, Namespaces, API keys, users, service accounts. Do not use the unified CLI's `temporal cloud …` subcommands for them. Worker Deployments and Workflows are *not* control-plane operations: they live on the Namespace frontend, have no `tcld` equivalent, and use `temporal worker deployment …`. → `references//setup.md`. - **Versioning behavior is mandatory.** Every Workflow needs `Pinned` or `AutoUpgrade`, or the Worker sets a default. -- **Deployment name and build ID must match exactly** between the Worker code and the Worker Deployment Version. A mismatch causes an invocation loop (Temporal invokes → Worker polls with the wrong version → Task not processed → invoke again). Signature: rapid repeated invocations with no Workflow progress. -- **Set the invocation deadline high enough.** Providers often default to a very short timeout. If the first invocation times out before the Worker registers the Task Queue, the binding is never created and the Worker is never invoked again. → `references//setup.md` for the exact default. +- **Deployment name and build ID must match exactly** between the Worker code and the Worker Deployment Version. A mismatched Worker polls under another version and never creates the intended Task Queue binding. - **Use an immutable, versioned build per Build ID in production.** Pointing the provider at a mutable "latest" target lets code change under in-flight Workflows and cause non-determinism errors, even for Pinned Workflows. Keep a 1-to-1 mapping between each Build ID and one immutable build. → `references//versioning.md`. +- **Secrets belong in a secret store**, not plaintext environment variables. Provider docs and quickstarts commonly pass the API key or TLS key as a plaintext environment variable; that is acceptable in a throwaway development walkthrough *only if you say so explicitly at the time*. Anything the user describes as production, shared, or long-lived gets the secret store, loaded at cold start. Either way, keep key material out of shell history and command echoes. +- **Both CLIs prompt for confirmation before mutating state, and their flags differ.** Setting the current or ramping version, and revoking an API key, all ask interactively; run non-interactively without the flag, the command exits having done nothing, which reads as success. `temporal worker deployment …` takes `--yes`; `tcld` takes the global `--auto_confirm`. Pass the right one in scripts, CI, and agent shells, and confirm the resulting state rather than trusting the exit code. → `references//setup.md`. + +## Provider-specific principles + +### AWS Lambda + +- **A Lambda identity mismatch causes an invocation loop.** Temporal invokes, the Worker polls with the wrong version, the Task remains unprocessed, and Temporal invokes again. The signature is rapid repeated invocations with no Workflow progress. +- **Set the invocation deadline high enough.** Providers often default to a very short timeout. If the first invocation times out before the Worker registers the Task Queue, the binding is never created and the Worker is never invoked again. → `references//setup.md` for the exact default. - **Tune the timeout triple together for long-running Activities:** (1) worker stop timeout > longest Activity runtime, (2) shutdown deadline buffer > worker stop timeout + shutdown hook time, (3) invocation deadline > longest Activity runtime + shutdown deadline buffer. Raising one alone does not help. If the longest Activity exceeds half the maximum invocation deadline, recommend Activity Heartbeats. → `references/concepts.md`, `references//sdk-.md`. - **Eager Activities are always disabled** — serverless invocations don't maintain persistent connections. Don't suggest them as an optimization. - **Activities are bounded by the invocation limit** (minus the shutdown deadline buffer); Workflow duration is unbounded and can span many invocations. Flag Activities that approach the provider's limit early. → `references/concepts.md`. - **Mixed serverless + long-lived Workers on one Task Queue:** do not enable dynamic scaling on the long-lived Workers — the two groups can't coordinate scaling and will cause unnecessary invocations. -- **Secrets belong in a secret store**, not plaintext environment variables. Provider docs and quickstarts commonly pass the API key or TLS key as a plaintext environment variable; that is acceptable in a throwaway development walkthrough *only if you say so explicitly at the time*. Anything the user describes as production, shared, or long-lived gets the secret store, loaded at cold start. Either way, keep key material out of shell history and command echoes. -- **Both CLIs prompt for confirmation before mutating state, and their flags differ.** Setting the current or ramping version, and revoking an API key, all ask interactively; run non-interactively without the flag, the command exits having done nothing, which reads as success. `temporal worker deployment …` takes `--yes`; `tcld` takes the global `--auto_confirm`. Pass the right one in scripts, CI, and agent shells, and confirm the resulting state rather than trusting the exit code. → `references//setup.md`. + +### GCP Cloud Run + +Cloud Run Workers use ordinary long-lived Worker APIs and have no invocation deadline. Handle scale-in with graceful shutdown and Heartbeats, and give the pool a Task Queue separate from independently managed Workers. → `references/gcp-cloud-run/constraints.md`. ## Troubleshooting +### AWS Lambda + Start by determining whether the Worker is being invoked at all. Then, in priority order: (1) **Validate Connection** in the Temporal UI (Workers > Deployments > select > Actions > Validate Connection) — checks credentials, role assumption, and reachability in one step; (2) check whether the version's **Task Queue is bound** — if it is, invocation and Worker startup provably work and the fault is downstream, which rules out most of the surface in one command; (3) confirm the version is **current** (CLI-created versions are not automatic, and a confirmation-prompted command may have silently done nothing); (4) check the compute provider's logs for connection, auth, or TLS errors; (5) if rapid repeated invocations show no progress, check the deployment name/build ID match. Distinguish a Temporal-side failure (reproduces no matter what you change on the provider side) from a genuine user-permission problem before editing anything. → `references//diagnostics.md`, `references/concepts.md`. +### GCP Cloud Run + +Start with the intended version's expected Task Queue bindings. If they are absent, correlate the WCI registration result, requested pool count, `lastModifier`, image digest, and Worker startup identity log using the decision table. If they are present, continue with current-version routing and Task execution. Treat Validate Connection as a read-only pool lookup, not proof that Temporal can resize the pool. → `references/gcp-cloud-run/diagnostics.md`. + ## Common Pitfalls High-impact mistakes — warn the user proactively. Each is a symptom → cause → fix. +### AWS Lambda + 1. **Deployment name / build ID mismatch → invocation loop.** *Symptom:* rapid, repeated invocations with no Workflow progress. *Cause:* the name or build ID in the Worker code doesn't match the Worker Deployment Version, so the Worker polls with the wrong version, the Task isn't processed, and Temporal invokes again. *Fix:* make the values in code exactly match the version configuration. 2. **Version not set as current.** A version created through the CLI is not automatically current; without it, Tasks don't route to the version and the Worker is never invoked. *Fix:* set it current as a separate step (the UI does this automatically). 3. **Failed first invocation.** When a version is created, the WCI invokes the Worker once to validate. If that invocation fails — missing env vars, bad TLS/auth config, missing dependencies, or an invocation deadline too short for the Worker to start and register the Task Queue — the Worker never connects, never polls, the binding is never created, and the Worker is never automatically invoked again. *Fix:* diagnose by manually invoking the compute unit, and confirm the invocation deadline is set high. @@ -215,10 +279,21 @@ High-impact mistakes — warn the user proactively. Each is a symptom → cause 7. **Re-creating shared permission infrastructure that already exists.** *Symptom:* the infrastructure deployment fails outright and rolls back, or it succeeds and leaves a second, redundant grant behind. *Cause:* the permission grant Temporal assumes is account-wide with a fixed default name, so a previous serverless deployment already owns it. *Fix:* check whether it exists and what owns it *before* creating; extend the existing one to cover the new Worker, and fall back to a distinctly named parallel one only when the existing infrastructure is not yours to change — saying why when you do. A failed-and-rolled-back deployment must be deleted before the name can be reused; a successful one is live infrastructure and must not be. → `references//iam.md`. 8. **Invoke permission scoped to a single build.** *Symptom:* the deployment works, then the *next* release cannot be invoked, with an error that looks like a connection or configuration problem rather than a permissions one. *Cause:* the grant named one immutable build, and the new release is a different resource. *Fix:* scope the grant to cover the base resource and all its published builds. → `references//iam.md`. +### GCP Cloud Run + +1. **Deployment name / build ID mismatch.** One steady instance may start and announce another build while the intended version never binds its Task Queue. That instance polls and processes Tasks only for the version it announces; the immediate consequence is incorrect extra capacity requested by the intended version's WCI, not cross-version execution. Make both values match the registered version. +2. **Runner and invoker service accounts confused.** The runner is attached to instances; Temporal impersonates the invoker to read and resize the pool. → `references/gcp-cloud-run/iam.md`. +3. **Mutable live pool.** Redeploying a new image into a pool used by a live Worker Deployment Version changes code underneath that version. Use one pool per build ID. → `references/gcp-cloud-run/versioning.md`. +4. **Scale-in interrupts an Activity.** Graceful shutdown cannot guarantee completion. Heartbeat resumable progress and keep the shutdown timeout below Cloud Run's termination window. → `references/gcp-cloud-run/constraints.md`. +5. **Task Queue shared with an independently managed fleet.** The rate-based scaler sees the full queue workload and provisions duplicate capacity. Use a separate Task Queue. → `references/gcp-cloud-run/constraints.md`. +6. **Validate Connection passes but the pool never resizes.** For Cloud Run it proves impersonation and `run.workerPools.get`, but does not exercise `run.workerPools.update` or start an instance. Check the registration bootstrap, Task Queue binding, and update permission. → `references/gcp-cloud-run/diagnostics.md`. + ## Routing to reference files Most questions need 2–3 reference files. +### AWS Lambda + | User intent | Reference file(s) | |---|---| | What is a Serverless Worker? How do invocation and autoscaling work? What are the constraints? Serverless vs long-lived Workers? | `references/concepts.md` | @@ -236,6 +311,18 @@ Most questions need 2–3 reference files. | Worker not invoked, Workflows not progressing, inspect the WCI. | `references/wci.md` + `references//diagnostics.md` + the selected `references//sdk-.md` | | Long-running Activities and timeout relationships. Isolate Activities from resource exhaustion. | `references/concepts.md` (+ the selected `references//sdk-.md`) | +### GCP Cloud Run + +| User intent | Reference file(s) | +|---|---| +| Execution model, Activity duration, autoscaling, scale-in, graceful shutdown, or mixed fleets. | `references/gcp-cloud-run/constraints.md` + the selected SDK reference | +| Deploy, register, verify, or tear down a Worker Pool. | `references/gcp-cloud-run/setup.md` + the selected SDK reference | +| Permissions, runner vs invoker identity, or Terraform IAM setup. | `references/gcp-cloud-run/iam.md` | +| Update, publish a new build, or roll back. | `references/gcp-cloud-run/versioning.md` | +| Worker Pool not scaling, Worker not polling, or Workflows not progressing. | `references/gcp-cloud-run/diagnostics.md` | +| Metrics, tracing, logs, or scaling signals. | `references/gcp-cloud-run/observability.md` + the selected SDK reference | +| Self-hosted Temporal Service prerequisites. | `references/gcp-cloud-run/self-hosted.md` + `references/gcp-cloud-run/iam.md` | + ## Out of Scope - **General SDK development patterns** (Workflows, Activities, signals, queries, Worker Versioning concepts): see `skill-temporal-developer`. diff --git a/references/concepts.md b/references/concepts.md index a6483dc..c2665f6 100644 --- a/references/concepts.md +++ b/references/concepts.md @@ -2,15 +2,50 @@ -## Release status +## Shared serverless concepts + +### Release status **AWS Lambda — Public Preview since July 30, 2026.** Open to all Temporal Cloud customers. There is no access request, no support ticket, and no manual toggle to enable: a customer selects "AWS Lambda (Public Preview)" as the compute provider in the UI and sets up their Worker Deployment directly. Never route a user to support to "get access" for Lambda. -AWS Lambda is the only compute provider this skill supports. Do not adapt the Lambda material to any other provider. +**GCP Cloud Run — Public Preview.** Open to all Temporal Cloud customers with a GCP-hosted Namespace. There is no access request, support ticket, or manual toggle to enable: select GCP Cloud Run as the compute provider and proceed. Never route a user to support to "get access" for Cloud Run either. + +These are the two compute providers this skill supports. Do not adapt either provider's material to another provider. The [AWS Lambda invocation model](#aws-lambda-invocation-model) section below applies only to Lambda; Cloud Run's pool-based model, including what bounds an Activity and how scale-in works, is in `gcp-cloud-run/constraints.md`. Public Preview is not General Availability. APIs are still evolving and may be subject to backwards-incompatible changes between versions — pin SDK and CLI versions for anything long-lived, and read the installed package's real API surface rather than writing from memory. -## What is a Serverless Worker? +### Compute providers + +A compute provider is the configuration that tells Temporal how to start or scale compute for a Worker Deployment Version. It is set on the Worker Deployment Version and specifies the provider type, the compute target, and the identity Temporal uses to act on that target. + +For example, an AWS Lambda compute provider includes the Lambda function ARN and the IAM role that Temporal assumes to invoke the function. + +A GCP Cloud Run compute provider names the Worker Pool and the invoker service account that Temporal impersonates to resize it. + +Compute providers are only needed for Serverless Workers. Traditional long-lived Workers do not require a compute provider because the Worker process lifecycle is not managed by the Temporal server. + +#### Supported providers + + + +| Provider | Description | +|---|---| +| AWS Lambda | Temporal assumes an IAM role in your AWS account to invoke a Lambda function. | +| GCP Cloud Run | Temporal impersonates an invoker service account in your GCP project to resize a Cloud Run Worker Pool. | + +### Worker Versioning + +Serverless Workers require Worker Versioning: each Worker Deployment Version must point its compute provider at a stable, immutable build. → `references//versioning.md`. + +### Worker Controller Instance (WCI) + +Temporal coordinates serverless scaling through the WCI. See [Worker Controller Instance (WCI)](wci.md) for its lifecycle, Task Queue inputs, Workflow ID pattern, and inspection commands. + +## AWS Lambda invocation model + +The sections below describe how Serverless Workers behave on AWS Lambda, where Temporal invokes the Worker on demand. They do not apply to GCP Cloud Run. + +### What is a Serverless Worker? A Serverless Worker is a Temporal Worker that runs on serverless compute instead of a long-lived process. There is no always-on infrastructure to provision or scale. Temporal invokes the Worker when Tasks arrive on a Task Queue, and the Worker shuts down when the work is done. @@ -21,17 +56,13 @@ Serverless Workers require Worker Versioning. Each Serverless Worker must be ass Each Workflow must have an `AutoUpgrade` or `Pinned` versioning behavior, set per-Workflow or as a Worker-level default. -## How Serverless invocation works +### How Serverless invocation works With long-lived Workers, the Worker process starts, connects to Temporal, and polls a Task Queue for work. Temporal does not need to know anything about the Worker's infrastructure. With Serverless Workers, Temporal starts the Worker. -### Worker Controller Instance (WCI) - -Temporal coordinates serverless scaling through the WCI. See [Worker Controller Instance (WCI)](wci.md) for its lifecycle, Task Queue inputs, Workflow ID pattern, and inspection commands. - -### Invocation flow +#### Invocation flow The invocation flow works as follows: @@ -44,33 +75,33 @@ The invocation flow works as follows: -## Autoscaling +### Autoscaling The shared [WCI inputs](wci.md#inputs) cause Lambda function invocations when Tasks need Workers. Each Worker exits when its invocation finishes, allowing the Lambda fleet to scale to zero. -## Scaling with long-lived Workers +### Scaling with long-lived Workers Serverless Workers can share a Task Queue with long-lived Workers. Because Serverless Workers are only invoked on sync match failure, Serverless Workers only pick up Tasks that no long-lived Worker was available to handle. In practice, the Serverless Workers act as spillover capacity for the long-lived fleet. **Warning:** If you configure Serverless and long-lived Workers on the same Task Queue, do not enable dynamic scaling on the long-lived Workers. The two groups cannot coordinate their scaling behavior. If both scale dynamically, the long-lived Workers may scale up to handle the same Tasks that Temporal is simultaneously invoking Serverless Workers for, leading to unnecessary invocations and unpredictable scaling. -## Worker lifecycle +### Worker lifecycle A single Serverless Worker invocation has three phases: init, work, and shutdown. -### Init phase +#### Init phase The Worker initializes and establishes a client connection to Temporal. -### Work phase +#### Work phase The Worker polls the Task Queue and processes Tasks. -### Shutdown phase +#### Shutdown phase The Worker stops polling, waits for in-flight Tasks to finish, and runs any shutdown hooks (for example, OpenTelemetry telemetry flushes). Shutdown begins before the invocation deadline so the Worker can exit cleanly before the compute provider forcibly terminates the execution environment. -### Tuning for long-running Activities +#### Tuning for long-running Activities If your Worker handles long-running Activities, set these three values together: @@ -88,11 +119,11 @@ Raising only the shutdown deadline buffer makes the Worker stop polling earlier, Raising only the Worker stop timeout does not make the Worker stop polling earlier, which means the compute provider might terminate the Worker before the full stop timeout completes. -## Failure handling +### Failure handling Serverless Workers rely on Temporal's standard retry and timeout semantics to recover from failures. -### Worker crash +#### Worker crash If a Worker invocation crashes (out of memory, unhandled exception, etc.): @@ -100,7 +131,7 @@ If a Worker invocation crashes (out of memory, unhandled exception, etc.): - No manual intervention is required. -### Provider concurrency limit +#### Provider concurrency limit If the compute provider's concurrency limit is reached (for example, AWS Lambda account concurrency): @@ -108,7 +139,7 @@ If the compute provider's concurrency limit is reached (for example, AWS Lambda - Tasks remain in the Task Queue backlog. No data loss occurs. - Processing slows until concurrency frees up. -### Resource exhaustion across Activity slots +#### Resource exhaustion across Activity slots By default, a single Worker invocation may run multiple Activity slots. A crash or resource exhaustion in one Activity can affect other Activities running in the same invocation. @@ -119,7 +150,7 @@ To isolate Activities from each other: -## Constraints +### Constraints @@ -130,7 +161,7 @@ With single-slot configuration, each Activity gets a dedicated execution environ | Worker code | Same Temporal SDK Worker code, using the serverless Worker package for your SDK. | | Versioning | Worker Versioning is required. Each Workflow must have an `AutoUpgrade` or `Pinned` behavior, set per-Workflow or as a Worker-level default. | -## Worker Versioning with Serverless Workers +### Worker Versioning with Serverless Workers Serverless Workers require Worker Versioning, and the compute provider must invoke a stable, immutable build for each Worker Deployment Version. With AWS Lambda, this means aligning two versioning systems: @@ -153,23 +184,7 @@ The choice of Pinned or Auto-Upgrade controls how Workflows move between Worker See `aws-lambda/versioning.md` for the step-by-step `aws lambda publish-version` workflow and `aws-lambda/setup.md` (Step 4) for how to configure the compute provider with a versioned ARN. -## Compute providers - -A compute provider is the configuration that tells Temporal how to invoke a Serverless Worker. The compute provider is set on a Worker Deployment Version and specifies the provider type, the invocation target, and the credentials Temporal needs to trigger the invocation. - -For example, an AWS Lambda compute provider includes the Lambda function ARN and the IAM role that Temporal assumes to invoke the function. - -Compute providers are only needed for Serverless Workers. Traditional long-lived Workers do not require a compute provider because the Worker process lifecycle is not managed by the Temporal server. - -### Supported providers - - - -| Provider | Description | -|---|---| -| AWS Lambda | Temporal assumes an IAM role in your AWS account to invoke a Lambda function. | - -## Why use Serverless Workers? +### Why use Serverless Workers? @@ -178,7 +193,7 @@ Compute providers are only needed for Serverless Workers. Traditional long-lived - **Scale automatically.** The compute provider handles scaling natively. When traffic drops, instances scale down. When there is no work, there is no compute running. - **Pay only for what you use.** Workers run only when Tasks are available. For low or intermittent volume workloads, this pay-per-invocation model can significantly reduce compute costs. -## When to use Serverless Workers +### When to use Serverless Workers @@ -198,7 +213,7 @@ May not be ideal when: - Workloads require sustained high throughput. Long-lived Workers on dedicated compute may be more cost-effective and performant. - You need persistent connections. Some features require a persistent connection between the Worker and Temporal, which serverless invocations do not maintain. -## How Serverless Workers compare to long-lived Workers +### How Serverless Workers compare to long-lived Workers