# Cloudflare Workflows
Cloudflare Workflows is [[Cloudflare]]'s [[Durable Execution|durable execution]] engine. You write a class that extends `WorkflowEntrypoint`, put each unit of work in a `step.do()` call, and the platform persists every step's result. When a step fails, only that step is retried. If the machine dies, the instance picks up after the last completed step. And a "sleep 30 days" means nothing runs (and nothing is billed for compute) until day 30.
It runs on [[Cloudflare Workers]] and is built on SQLite-backed [[Cloudflare Durable Objects]]: one "Engine" Durable Object per instance, which executes your steps, stores their output and handles retries and sleeps. There's no cluster or worker fleet to operate. That's the main difference with Temporal-style systems, and also where most of its limits come from.
Timeline: announced during Developer Week in April 2024, open beta on October 24, 2024, generally available on April 7, 2025.
## Why it matters
Most backends end up needing "do A, then B, then wait for something, then C", and the classic answer is a queue, a cron job, a state column in a database and a lot of retry code. Workflows replaces that with plain code (TypeScript or [[Python]]) that reads top to bottom: no DSL, no JSON state machine, no YAML. Loops, `if` statements and steps generated at runtime all work.
In 2026 the focus moved to [[AI Agents]]. Cloudflare said so explicitly when it rebuilt the control plane in April 2026: instances are no longer created at human speed (one per signup) but at machine speed (an agent session starting dozens of them). An agent loop where every LLM call and tool call is a step gets retries, checkpoints and [[Human-in-the-Loop|human approval]] for free, and doesn't pay twice for tokens when something crashes halfway.
## The programming model
```ts
import { WorkflowEntrypoint, WorkflowEvent, WorkflowStep } from "cloudflare:workers";
import { NonRetryableError } from "cloudflare:workflows";
export class OrderWorkflow extends WorkflowEntrypoint<Env, { orderId: string }> {
async run(event: WorkflowEvent<{ orderId: string }>, step: WorkflowStep) {
const order = await step.do("load order", async () => loadOrder(event.payload.orderId));
await step.do(
"charge card",
{ retries: { limit: 5, delay: "10 seconds", backoff: "exponential" }, timeout: "5 minutes" },
async () => chargeCard(order),
);
const approval = await step.waitForEvent("wait for approval", { type: "approved", timeout: "24 hours" });
await step.sleep("cool-off", "1 day");
await step.do("ship", async () => ship(order, approval));
}
}
```
- **`step.do(name, config?, callback, rollback?)`**: runs the callback, persists its return value (anything structured-cloneable, up to 1 MiB; JavaScript steps can also return a `ReadableStream` for bigger binary output). On replay the cached value comes back instead of re-running the step
- **Retries**: default is 5 retries, 10 s delay, exponential backoff, 10-minute timeout per attempt. Backoff can be `constant`, `linear` or `exponential`, up to 10,000 retries per step. Since July 2026, `delay` can also be a function of the attempt and the error (e.g., wait longer on a 429, honor a `Retry-After`)
- **`NonRetryableError`**: throw it for permanent failures (bad input, auth failure) and the instance fails immediately
- **Step context**: the callback receives `ctx.attempt` (March 2026), `ctx.step.name`, `ctx.step.count` and the resolved `config` (April 2026)
- **Rollbacks** (June 5, 2026): attach a compensating handler to a `step.do()`. If the instance fails, handlers run in reverse step-start order. That's the saga pattern, with each compensation living next to the step it undoes
- **`step.sleep()` / `step.sleepUntil()`**: relative or absolute, up to 365 days. Sleeping instances hibernate
- **`step.waitForEvent()`**: pause until someone calls `instance.sendEvent({ type, payload })` from a Worker or the REST API. Timeout from 1 second to 365 days, 24 hours by default, and a timeout throws (wrap it in `try/catch` if you want to continue). Events sent before the instance reaches the wait are buffered
### Rules that will bite you
The `run()` method is replayed from the top whenever an instance wakes up, and step names are the cache key. Cloudflare's "Rules of Workflows" page boils down to this:
- Make each side effect [[Idempotency|idempotent]], because a step can run more than once (e.g., it succeeded but the result wasn't persisted)
- Keep state inside step return values, never in variables outside steps: the instance can hibernate and lose its memory at any point
- Name steps deterministically (no timestamps, no random values in names), and base `if` conditions on the payload or step outputs, never on `Date.now()` or `Math.random()`
- Always `await` your steps, and be careful with `Promise.race()` and `Promise.any()`
- Keep step timeouts at 30 minutes or less; use `waitForEvent()` for anything longer
## Managing instances
A Workflow binding (or, since September 27, 2026, `ctx.exports` for Workflows the Worker itself declares) exposes `create({ id?, params, retention? })`, `createBatch()` (up to 100 at once) and `get(id)`. An instance handle has `status()`, `pause()`, `resume()`, `restart()`, `terminate()`, `sendEvent()`, `delete()` and, since September 15, 2026, `subscribe()`, which streams the full event history and then live events (attempts, sleeps, waits, rollbacks) without polling. The statuses are `queued`, `running`, `paused`, `waiting`, `waitingForPause`, `errored`, `terminated`, `complete` and `unknown`.
The same operations are available through the REST API and `wrangler workflows`. Since April 2026, every Wrangler Workflows command accepts `--local` to target a `wrangler dev` session.
Ways to start an instance:
- From any Worker or Pages Function through a binding
- From the REST API (`POST .../workflows/{name}/instances`) or the CLI
- On a [[cron|cron]] schedule declared directly on the Workflow binding (June 2, 2026), no separate scheduled Worker needed
- From a [[Cloudflare Queues|Queues]] consumer, e.g., an R2 upload event lands in a queue and the consumer calls `createBatch()`
- Inside a step of another Workflow
- By an agent built with the [[Cloudflare Agents SDK]]
## Limits worth knowing (Workers Paid)
- **CPU**: 30 s per step by default, configurable up to 5 minutes (`limits.cpu_ms`). Wall-clock time per step is unlimited: waiting on I/O doesn't count
- **Steps**: 10,000 per instance by default, configurable up to 25,000 (March 2026; it was 1,024 before). Sleeps don't count toward that limit
- **State**: 1 MiB per step result and per event payload, 1 GB persisted per instance
- **Concurrency**: 50,000 running instances per account, 300 creations per second per account (100 per Workflow), 2 million queued instances (April 15, 2026). `waiting` instances (sleeping, waiting for an event or a retry) don't take a concurrency slot, so millions can sit idle at once
- **Subrequests**: 10,000 per instance by default, up to 10 million
- **Retention**: 30 days maximum. Workflows created since September 10, 2026 default to 7 days (older ones keep 30)
- **Free plan**: 10 ms CPU per invocation, 1,024 steps, 100 concurrent instances, 100 MB of state, 3-day retention
- Workflows can't be deployed into Workers for Platforms namespaces
## Pricing
Workflows is billed like Workers plus two extra dimensions:
| Dimension | Workers Free | Workers Paid |
| --- | --- | --- |
| Requests (instance creations) | 100,000/day, shared with Workers | 10M/month included, then $0.30 per million |
| CPU time | 10 ms per invocation | 30M CPU-ms/month included, then $0.02 per million CPU-ms |
| Storage | 1 GB-month | 1 GB-month included, then $0.20 per GB-month |
| Steps | 3,000/day | 500,000/month included, then $0.80 per 100,000 |
Step billing was announced on July 7, 2026, with step and storage billing starting no earlier than August 10, 2026 (storage had a price since GA, but billing for it was deferred until then). Retries and rollback handlers don't count as steps; according to the changelog, sleeps and event waits do. Idle time costs nothing in CPU, which is the whole point for workloads that mostly wait on LLMs and APIs.
A rough comparison: 1 million instances of a 10-step Workflow is 10M steps, i.e. about $76 in steps on top of the included quota. The same 10M state transitions on AWS Step Functions Standard ($0.025 per 1,000) would be $250. Steps and state transitions aren't the same unit, so take that as an order of magnitude only.
## Python Workflows
Python support landed in beta on August 22, 2025, on top of Python Workers (Pyodide). The docs no longer show a beta label, but I found no GA announcement. `step.do` becomes a decorator, `wait_for_event`, `sleep` and `sleep_until` are snake_case, and you need the `python_workers` and `python_workflows` compatibility flags. It adds something the TypeScript SDK doesn't have: DAG workflows, where a step declares its dependencies through its parameter names (older code used `depends=[...]`) and `concurrent=True` lets independent dependencies run in parallel.
## How it fits with the rest of the platform
- **[[Cloudflare Durable Objects]]**: the substrate. Rule of thumb from an HN thread I agree with: use a Durable Object to model an *entity* that lives indefinitely (a document, a user, a chat room), and a Workflow for a *process* that has a sequence of steps and ends
- **[[Cloudflare Agents SDK]]**: since v0.3.7 (February 3, 2026), an `AgentWorkflow` class gets typed access to its agent, and the agent gets `runWorkflow()`, `approveWorkflow()` / `rejectWorkflow()` for human-in-the-loop, plus pause, resume and terminate. The agent holds the WebSocket and the state; the Workflow does the long, retryable work and reports progress back. Project Think runs durable reasoning steps inside Workflows
- **[[Cloudflare Queues]]**: Workflows publishes lifecycle events (`instance.queued`, `instance.started`, `instance.paused`, completion and failure) through Queues event subscriptions, and a queue consumer is a natural way to fan instances out
- **Dynamic Workflows** (May 1, 2026): the `@cloudflare/dynamic-workflows` library runs Workflows whose code is loaded at runtime in a [[Cloudflare Dynamic Workers|Dynamic Worker]], so every tenant (or every agent) can ship its own durable workflow. Each instance carries routing metadata (e.g., a tenant ID), and the right code is reloaded every time the instance wakes up
- **[[Cloudflare Artifacts]]**: CI on push (August 4, 2026) is a Workflow written with `@cloudflare/ci`, triggered by `cf.artifacts.repo.pushed` events
- **[[Cloudflare R2]]** and **[[Cloudflare Workers AI]]**: the typical steps of a pipeline: read the file, call the model, write the result back
- **[[Wrangler]]**: since September 24, 2026, Workflows can be declared in the `exports` field of the config instead of a `workflows` binding (Wrangler 4.139.0+)
Tooling improved a lot in 2025–2026: Vitest test APIs (`introspectWorkflow()`, `introspectWorkflowInstance()`, September 2025), a dashboard visualizer that turns your code into a diagram by parsing its AST (beta, February 2026), and batch deletion of instances (September 17, 2026).
## Compared to the alternatives
- **Temporal**: the reference. Open source, self-hosted or Temporal Cloud (priced per million "actions"), many language SDKs, signals, queries, child workflows and an explicit versioning story. Its workflow code is replayed from an event history and must be deterministic; Cloudflare's model is lighter (only step results are memoized, by name) but has the same constraints in practice. Temporal is the safer bet for critical, high-throughput or multi-cloud work; Workflows wins when you're already on Workers and don't want to run anything
- **Inngest**: the closest API. `step.run`, `step.sleep` and `step.waitForEvent` came first there, and Inngest's founder claims Cloudflare's (and Vercel's) API is based on theirs. Inngest orchestrates functions that run on your own infrastructure (any cloud) and is event-driven by design. [[Vercel]]'s Workflow takes another path again, with `"use workflow"` / `"use step"` directives and a compile step
- **AWS Step Functions**: state machines written in a JSON DSL (Amazon States Language), deep integrations with AWS services, Standard executions up to a year, priced per state transition. Workflows is plain code, which is far easier to read, version and test, but has no catalog of integrations
- **Azure Durable Functions**: orchestrator functions replayed from an event-sourced history, with the same determinism rules as Temporal, tied to Azure Functions and its storage providers
- **DBOS and Restate**: lighter durable execution libraries that keep your existing infrastructure (DBOS checkpoints into Postgres). One HN user's split is a good summary: Restate for payments, DBOS when the workflow must commit in the same Postgres transaction, Cloudflare Workflows for cheap non-critical jobs like report generation
## Criticisms and caveats
- **Workers limits apply**: 128 MB of memory per isolate (so large files must be streamed) and the Workers cap on simultaneous open connections. One HN user who dropped Workflows in August 2025 described fetches failing once the open-connection limit was hit, low rate limits and incomplete Node.js compatibility. The limits went up a lot since then (concurrency went from 4,500 to 50,000, steps from 1,024 to 25,000) and Node.js compatibility is now on by default, but memory is still 128 MB
- **Start latency**: in March 2026, a user reported instances taking up to about 4 minutes to start, even when idle. The V2 control plane (April 2026) was designed to fix queueing and stuck instances, but I haven't found independent confirmation that it did
- **"GA" came early**: you couldn't delete a Workflow at GA (April 2025); deletion arrived on April 29, 2025, and deleting instances individually or in batches only on September 17, 2026. Several HN commenters see this as a pattern: Cloudflare ships new products faster than it rounds off existing ones
- **No documented versioning**: unlike Temporal's patching APIs, the docs don't explain what happens to sleeping instances when you deploy new code. Step names are the cache key, so renaming, removing or reordering steps that in-flight instances still have to replay is risky
- **Observability**: the dashboard was called buggy in 2025; `subscribe()`, the visualizer and Queues event subscriptions have since filled part of the gap
- **Lock-in**: Workflows only exists on Cloudflare's runtime. There is no self-hosted engine, only local emulation in `wrangler dev`
My take: if the rest of the app already lives on Workers, Workflows is the obvious default for anything that retries, waits or runs longer than a request, and the price is hard to beat. For money-critical flows, or when you need to leave Cloudflare one day, I'd look at Temporal, Restate or DBOS first.
## References
- Documentation: https://developers.cloudflare.com/workflows/
- Workers API: https://developers.cloudflare.com/workflows/build/workers-api/
- Rules of Workflows: https://developers.cloudflare.com/workflows/build/rules-of-workflows/
- Sleeping and retrying: https://developers.cloudflare.com/workflows/build/sleeping-and-retrying/
- Events and parameters: https://developers.cloudflare.com/workflows/build/events-and-parameters/
- Trigger Workflows: https://developers.cloudflare.com/workflows/build/trigger-workflows/
- Python Workflows: https://developers.cloudflare.com/workflows/python/
- Event subscriptions: https://developers.cloudflare.com/workflows/reference/event-subscriptions/
- Pricing: https://developers.cloudflare.com/workflows/reference/pricing/
- Limits: https://developers.cloudflare.com/workflows/reference/limits/
- Build a durable AI agent: https://developers.cloudflare.com/workflows/get-started/durable-agents/
- Changelog filtered to Workflows (RSS): https://developers.cloudflare.com/changelog/rss/workflows.xml
- Changelog, open beta (2024-10-24): https://developers.cloudflare.com/changelog/post/2024-10-24-workflows-beta/
- Changelog, GA (2025-04-07): https://developers.cloudflare.com/changelog/post/2025-04-07-workflows-ga/
- Changelog, Python Workflows beta (2025-08-22): https://developers.cloudflare.com/changelog/post/2025-08-22-workflows-python-beta/
- Changelog, Agents SDK v0.3.7 Workflows integration (2026-02-03): https://developers.cloudflare.com/changelog/post/2026-02-03-agents-workflows-integration/
- Changelog, 25,000 steps per instance (2026-03-03): https://developers.cloudflare.com/changelog/post/2026-03-03-step-limits-to-25k/
- Changelog, higher concurrency and creation limits (2026-04-15): https://developers.cloudflare.com/changelog/post/2026-04-15-workflows-limits-raised/
- Changelog, Dynamic Workflows (2026-05-01): https://developers.cloudflare.com/changelog/post/2026-05-01-dynamic-workflows/
- Changelog, cron schedules on the binding (2026-06-02): https://developers.cloudflare.com/changelog/post/2026-06-02-cron-workflows/
- Changelog, rollbacks (2026-06-05): https://developers.cloudflare.com/changelog/post/2026-06-05-saga-rollbacks/
- Changelog, step billing (2026-07-07): https://developers.cloudflare.com/changelog/post/2026-07-07-workflows-billing-updates/
- Changelog, retry delay functions (2026-07-09): https://developers.cloudflare.com/changelog/post/2026-07-09-dynamic-retry-delays/
- Changelog, 7-day default retention (2026-09-10): https://developers.cloudflare.com/changelog/post/2026-09-10-paid-retention-default/
- Changelog, `subscribe()` (2026-09-15): https://developers.cloudflare.com/changelog/post/2026-09-15-instance-event-subscriptions/
- Changelog, instance deletion (2026-09-17): https://developers.cloudflare.com/changelog/post/2026-09-17-instance-delete/
- Changelog, Workflows in `exports` (2026-09-24): https://developers.cloudflare.com/changelog/post/2026-09-24-workflow-exports/
- Changelog, `ctx.exports` (2026-09-27): https://developers.cloudflare.com/changelog/post/2026-09-27-workflow-ctx-exports/
- Build durable applications on Cloudflare Workers (Cloudflare blog, 2024-10-24): https://blog.cloudflare.com/building-workflows-durable-execution-on-workers/
- Cloudflare Workflows is now GA (Cloudflare blog, 2025-04-07): https://blog.cloudflare.com/workflows-ga-production-ready-durable-execution/
- Building a better testing experience for Workflows (Cloudflare blog, 2025-11-04): https://blog.cloudflare.com/better-testing-for-workflows/
- A closer look at Python Workflows (Cloudflare blog, 2025-11-10): https://blog.cloudflare.com/python-workflows/
- How we turn Workflows code into visual diagrams (Cloudflare blog, 2026-03-27): https://blog.cloudflare.com/workflow-diagrams/
- Rearchitecting the Workflows control plane (Cloudflare blog, 2026-04-15): https://blog.cloudflare.com/workflows-v2/
- Introducing Dynamic Workflows (Cloudflare blog, 2026-05-01): https://blog.cloudflare.com/dynamic-workflows/
- HN, "Anyone using Cloudflare Workflows in production?" (2026-03-11): https://news.ycombinator.com/item?id=47334792
- HN, limits criticism (2025-08-09): https://news.ycombinator.com/item?id=44845136
- HN, Restate vs DBOS vs Cloudflare Workflows (2026-05-28): https://news.ycombinator.com/item?id=48315400
- HN, Inngest founder on the step API (2025-10-23): https://news.ycombinator.com/item?id=45688794
- HN, GA without deletion (2026-06-25): https://news.ycombinator.com/item?id=48679522
- HN, Durable Objects vs durable execution (2026-08-07): https://news.ycombinator.com/item?id=49206181
- AWS Step Functions pricing: https://aws.amazon.com/step-functions/pricing/
- Temporal Cloud pricing: https://temporal.io/pricing
## Related
- [[Cloudflare]]
- [[Cloudflare Workers]]
- [[Cloudflare Durable Objects]]
- [[Cloudflare Queues]]
- [[Cloudflare Agents SDK]]
- [[Cloudflare Dynamic Workers]]
- [[Cloudflare Artifacts]]
- [[Cloudflare R2]]
- [[Cloudflare Workers AI]]
- [[Wrangler]]
- [[Durable Execution]]
- [[Idempotency]]
- [[Human-in-the-Loop]]
- [[AI Agents]]
- [[Vercel]]
- [[cron]]
- [[Python]]
- [[TypeScript]]