Durable execution in 2026: revisit Temporal, Restate and DBOS before you hand-roll retries, sagas and cron jobs
Picture an online pharmacy and diagnostics booking company in Ahmedabad (an invented example, not a real client). An order goes through five steps: hold the medicine stock, take the UPI payment, book a courier slot, send the customer an SMS, and, if the courier cannot pick up within four hours, release the stock and refund. Today this is a jobs table in PostgreSQL, a worker that polls it, retry counters in a JSON column and three cron jobs that look for orders stuck in odd states. Last month a rolling deploy restarted the workers in the middle of a busy evening. Some orders were refunded twice, some never got a courier, and the on-call engineer spent the night writing SQL. Now one architect wants Temporal, a second has read about Restate, and a third says DBOS can do it with the PostgreSQL they already run.
All three are durable execution engines. They record each step of a long-running function so that after a crash, a deploy or a timeout it continues from where it stopped. They differ in where that record lives, who drives the code, how they cope with code changes, and what licence you accept. This post compares Temporal, Restate and DBOS on those points, with short notes on Cadence, Argo Workflows and Airflow. Facts come from GitHub release pages, official docs, licence files and CVE records, checked on 12 October 2026. The drill near the end is my own run.
The short version
- Durable execution is not a queue with retries. The engine keeps a journal of completed steps and replays your function against it after a failure, so a careless code change can break running workflows. In my drill, all three engines caught such a change, each differently.
- Temporal is the mature, server-centred option: services backed by Cassandra, MySQL or PostgreSQL, and workers that poll task queues. Server 1.32.0 (11 September 2026) made Standalone Activities GA, 1.32.1 followed on 9 October, and the deprecated Worker Versioning APIs go in 1.33.
- Restate is a single binary with its own replicated log and RocksDB state that pushes invocations to your handlers over HTTP. Version 1.7.13 (1 October 2026) is current and 1.8.0 is at rc.2. The server is under the Business Source License 1.1, not an open source licence; the SDKs are MIT.
- DBOS is a library that checkpoints steps to PostgreSQL, with no orchestration server. DBOS Python 3.0 and TypeScript 5.0 (16 September 2026) changed the schema, and 2.x processes cannot process workflows created by 3.0. The libraries are MIT; the Conductor control plane is proprietary.
- Default retry behaviour differs a lot: Temporal retries activities without limit, Restate retries invocations 70 times and then pauses them, and DBOS does not retry a step unless you ask for it. Read the defaults before you trust them.
- Temporal published six CVEs between April and September 2026, one in the UI server, including CVE-2026-89139 (a namespace writer could run commands on the Worker Service host in 1.31.0 to 1.31.2). For self-hosted Temporal, 1.30.7, 1.31.3 or 1.32.x is the floor today.
Where things stand on 12 October 2026
Dates are from GitHub release pages and PyPI, converted to IST.
Project
Latest release
What it is
Licence
Temporal Server
1.32.1 on 9 October 2026 (1.32.0 on 11 September; 1.31.3 on 19 September; 1.30.7 on 18 September)
Workflow orchestration service with workers, task queues and event histories
MIT
Temporal CLI
1.9.1 on 15 September 2026, embedding Server 1.32.0 for start-dev
CLI and single-process development server
MIT
Temporal SDKs
Python SDK 1.34.0 on 1 October 2026
Go, Java, Python, TypeScript, .NET and others
Go, Python, TypeScript and .NET MIT; Java Apache 2.0
Restate
1.7.13 on 1 October 2026; 1.8.0-rc.2 on 9 October
Single-binary durable execution server with replicated log and embedded RocksDB
Server: Business Source License 1.1, changing to Apache 2.0 four years after each release
Restate SDKs
Python SDK 1.0.5 on 2 September 2026
TypeScript, Java and Kotlin, Python, Go, Rust
MIT
DBOS Transact
Python 3.2.0 and TypeScript 5.2 on 29 September 2026; Java 1.2.0 on 1 October; Go 1.6.0 on 7 October
Durable workflow library that checkpoints to PostgreSQL
MIT
DBOS Conductor
Self-hosted or DBOS-hosted
Control plane for recovery, dashboards and retention
Proprietary, licence key required
Cadence
1.4.2 on 3 October 2026
Uber’s workflow engine, the project Temporal was forked from
Apache 2.0; CNCF Sandbox since 22 May 2025
Date
Event
21 February 2017
Cadence repository created at Uber
17 October 2019
Temporal repository created; Temporal starts as a fork of Cadence
30 September 2020
Temporal Server 1.0.0
4 January 2023
Restate repository created
27 April 2023
Restate 0.1.0
12 July 2023
DBOS TypeScript library repository created
7 June 2024
Restate 1.0.0
19 July 2024
DBOS Python library repository created
15 February 2025
Restate 1.2.0: multi-node replicated clusters, restatectl, snapshots to object storage
22 May 2025
Cadence accepted into CNCF at Sandbox level
30 January 2026
Restate 1.6.0: pause and resume invocations, restart from a journal prefix
30 April 2026
Temporal Server 1.31.0: Task Queue Priority and Fairness GA, Cassandra 5 support
19 June 2026
Restate 1.7.0: flow control (concurrency rules), bounded invoker memory, UI 1.0, ingress request size limit
2 July 2026
DBOS Java 1.0.0
31 July 2026
DBOS Go 1.0.0
11 September 2026
Temporal Server 1.32.0: Standalone Activities GA, eager activity execution on by default, unified visibility query converter by default
16 September 2026
DBOS Python 3.0.0 and TypeScript 5.0: new storage schema, deprecated features removed
18 and 19 September 2026
Temporal 1.30.7 and 1.31.3 fix the September CVEs
1 October 2026
Restate 1.7.13 fixes Virtual Object state consistency across replicas with VQueues
9 October 2026
Temporal 1.32.1; Restate 1.8.0-rc.2
What durable execution actually means
You write a function that calls several steps. The engine records each step’s result in a durable journal before your code moves on. If the process dies, the function runs again from the top, but completed steps return their recorded results instead of running again. Timers are journal entries too, so a four-hour sleep survives a restart. Three consequences follow.
- The workflow function must be deterministic. On replay it must call the same steps in the same order. Clock reads, random numbers and API calls go inside steps.
- Steps run at least once, not exactly once. If the process dies after the payment API returned but before the result was recorded, the step runs again. “Exactly once” in vendor material means the recorded result is applied once; the external side effect still needs an idempotency key. My idempotency keys post covers that half, and nothing here replaces it.
- Code changes are a data migration. A running workflow has a journal written by the old code. If new code takes a different path, the journal no longer matches. Each engine has its own answer: replay checks and patching in Temporal, immutable deployments in Restate, application versions and patching in DBOS.
Three shapes of the same idea
Temporal
Restate
DBOS
What you deploy
A Temporal Service (frontend, history, matching and worker services) plus a database, plus your workers
One restate-server binary (single node or cluster), plus your service endpoints
Your application with the DBOS library, plus PostgreSQL
Where the journal lives
Event history in Cassandra, MySQL or PostgreSQL (SQLite for development)
Replicated log (Bifrost) with state in embedded RocksDB, snapshots to S3-compatible storage
Tables in a PostgreSQL “system database”
How work reaches your code
Workers poll task queues
Restate calls your handler over HTTP/2 (or Lambda and similar), streaming the journal
Your process runs the function itself; other processes pick up work from durable queues
Unit of work
Workflow and Activities
Service handler, Virtual Object (keyed, single writer) or Workflow
Workflow and steps
Replay check
SDK compares new commands with the history
SDK compares new journal entries with the recorded ones
Library compares step names by position
Not the same category
Cadence is the ancestor. Temporal’s README says it “originated as a fork of Uber’s Cadence” and is developed by the creators of Cadence. Cadence is still developed (1.4.2 on 3 October 2026) and joined the CNCF Sandbox in May 2025. If you run it, you have a choice, not an emergency.
Argo Workflows (4.1.5 on 9 October 2026) runs DAGs of containers on Kubernetes, and Apache Airflow (3.3.2 on 17 September 2026) schedules data pipelines. Both suit batch jobs; neither is built for thousands of short business transactions, each waiting on payments, humans or timers. For the nightly ETL, look there. For “this order must reach a final state whatever fails”, read on.
A message queue is also not durable execution. A queue delivers messages; it does not remember that step three of an order finished. My message queue post covers queues, and many teams will run both.
Temporal in 2026
The model
Temporal splits code into Workflows (deterministic orchestration) and Activities (anything that touches the outside world). Your workers poll task queues. On a workflow task, the SDK replays the event history through your code and sends back new commands, such as “schedule activity charge”. If they do not match the history, the SDK raises a nondeterminism error and the workflow task keeps failing until you fix the code, as my drill showed.
Hard limits shape design. An event history is limited to 51,200 events or 50 MB (warnings at 10,240 events or 10 MB), after which the execution is terminated. A payload warns at 256 KB and fails at 2 MB. A workflow may have at most 2,000 pending activities, child workflows, signals or cancellations of each type by default, and the docs suggest 500 or fewer. Entity workflows, such as one per customer, need Continue-As-New to start a fresh history.
Activities retry by default with exponential backoff from 1 second, coefficient 2.0, capped at 100 seconds, with unlimited attempts; workflows do not retry by default. A broken activity therefore retries quietly for days unless you set maximum_attempts or a schedule-to-close timeout. Recurring work belongs in Schedules, which have their own identity, rather than the older Temporal Cron Jobs.
What changed in 1.31 and 1.32
Server 1.31.0 (30 April 2026) made Task Queue Priority and Fairness generally available and added Cassandra 5 support. Server 1.32.0 (11 September 2026) changed several defaults:
- Standalone Activities are GA: an activity started from a client, without a workflow, still gets retries, timeouts and visibility. That suits a single durable call.
- Eager activity execution is on by default.
- The unified visibility query converter is the default, with breaking changes such as type-mismatched search attribute comparisons now being errors. The legacy converter goes in 1.33.
- Worker Versioning: 1.32 is the last release with the deprecated Build ID APIs. Before 1.33, move to Worker Deployment APIs and drain workflows that depend on legacy routing.
- Nexus callbacks now route by URL scheme, a change the notes say was made for security reasons, and
component.callbacks.allowedAddressesbecamecallback.allowedAddresses.
Versioning running workflows
Temporal gives you two tools. Patching puts the change behind workflow.patched("id") (named differently per SDK), which returns false while replaying old histories and true for new runs. Worker Versioning groups workers into Worker Deployment Versions. A Pinned workflow stays on the version it started on; an Auto-Upgrade workflow moves to the current version and must be kept replay-safe with patching. A patched version of my drill workflow:
@workflow.defn(name="Settle")
class Settle:
@workflow.run
async def run(self, wid: str) -> str:
await workflow.execute_activity(reserve, wid, start_to_close_timeout=T)
if workflow.patched("fraud-check"):
await workflow.execute_activity(fraud_check, wid, start_to_close_timeout=T)
txn = await workflow.execute_activity(charge, wid, start_to_close_timeout=T)
await workflow.sleep(8)
await workflow.execute_activity(notify, wid, start_to_close_timeout=T)
return txn
Running it yourself
The persistence page lists the versions Temporal tests: Cassandra 3.11, 4.0 and 5.0.4 or later, PostgreSQL 13.18, 14.15, 15.10 and 16.6, and MySQL 5.7 and 8.0.19 or later. SQLite is for development only. Advanced visibility runs on Elasticsearch, OpenSearch 2 (from Server 1.30.1) or the same SQL databases; the docs recommend Elasticsearch for anything beyond a few workflows. Notice that PostgreSQL 17 and 18 are not on the tested list, so test before you put Temporal on the cluster from my PostgreSQL HA post.
Upgrades must be sequential, one minor version at a time, after moving to the latest patch of your current minor. Skipping versions can leave old data formats the new code no longer understands.
Security is your job. If you do not configure an Authorizer, Temporal uses the noopAuthorizer, which allows every API request. Turn on mTLS and a claim mapper and authorizer before any second team shares the cluster; several of the CVEs below only matter once namespaces are a real boundary.
Restate in 2026
The model
Restate inverts the worker model. Your code runs as an ordinary HTTP service, Lambda function or similar. You register its URL, and Restate calls your handlers, streaming the journal to the SDK. During a long sleep Restate can suspend the invocation and call again when the timer fires, so on serverless platforms you do not pay for a sleeping workflow.
The server is one binary, which its architecture page calls a “partitioned, log-centric runtime”: each partition’s leader appends events to a replicated log (Bifrost) and a partition processor materialises state in embedded RocksDB. A step “happens” when a quorum acknowledges its log append. Snapshots go to S3-compatible storage to bound recovery time. Multi-node clusters arrived in 1.2.0 (February 2025).
There are three service types. A Service runs handlers in parallel. A Virtual Object is keyed, with key-value state and a single writer per key, which suits per-order state without a lock table. A Workflow runs its run handler once per ID. Sagas are code: keep a list of compensations and run them in reverse on a terminal error. There is no native cron; the docs suggest durable timers with a Virtual Object.
What changed in 1.5 to 1.7
- 1.5.0 (September 2025) introduced a new invocation retry policy.
- 1.6.0 (30 January 2026) added pause and resume, and restarting an invocation from a prefix of its journal. Deprecated SDK versions are now rejected.
- 1.7.0 (19 June 2026) added concurrency limits per scope (
restate rules, behindexperimental_enable_vqueues), a bounded invoker memory pool (1.5 GiB default), UI 1.0 and an ingress request size limit (32 MiB, returning 413). - 1.7.13 (1 October 2026) fixed Virtual Object state edited through the Admin API with VQueues being skipped on followers; affected replicated clusters must re-submit that state.
- 1.8.0 is at rc.2; deprecated root-level service-client options are removed in 1.8.
Retries, retention and defaults
The services configuration page documents the default invocation retry policy as exponential from 50 ms, factor 2, up to 60 seconds and 70 attempts, then pause until a human resumes it. The server configuration reference and --dump-config on my 1.7.13 binary show a 500 ms initial interval, so check your effective configuration rather than either page. Pause is the safer default for payments, because killing an invocation does not run your compensations.
Journals, idempotency records and finished workflow state are retained for 24 hours by default. Raise these for anything you will want to inspect next morning.
Versioning by immutable deployments
In Restate a deployment is immutable. You deploy new code to a new URL and register it; new invocations go to the latest deployment while in-flight ones finish where they started. No patch statements, but old deployments must run until they drain, which is why the docs warn against handlers that run for days. On FaaS platforms each version already has its own URL or ARN, and the Kubernetes operator automates registration, draining and scale to zero.
If you update code behind the same URL instead, which is what most Kubernetes rolling updates do, Restate detects the mismatch on replay and returns error 570 “Journal mismatch”. My drill showed this.
Operating it
Minor upgrades must not be skipped, and rollback is supported only to the previous minor and only if you have not used new features. In a cluster, wait for each node’s partitions to recover before restarting the next; a health check alone is not enough. The security page is direct: the server has no application-level authentication on its management interfaces by design. The admin port (9070) must be limited to operators, the ingress port (8080) belongs behind an authenticating proxy, and the fabric port (5122) must not leave the cluster. On my run, the fabric port listened on all interfaces by default while I had bound the other two to 127.0.0.1. Also note that disable-telemetry defaults to false; the reference says anonymous usage data goes to Scarf, and DO_NOT_TRACK=1 turns it off.
A Restate workflow from my drill, in the Python SDK:
wf = restate.Workflow("Settle")
@wf.main()
async def run(ctx: restate.WorkflowContext, wid: str) -> str:
await ctx.run_typed("reserve", mark, step="reserve", wid=wid)
txn = await ctx.run_typed("charge", mark, step="charge", wid=wid)
await ctx.sleep(timedelta(seconds=8))
await ctx.run_typed("notify", mark, step="notify", wid=wid)
return txn
app = restate.app([wf])
DBOS in 2026
The model
DBOS is a library. You decorate functions as workflows and steps, and it writes checkpoints to a PostgreSQL “system database”, with “no separate orchestration server”. The stated overhead is one write per step plus two per workflow. Large step outputs mean large writes, so return pointers to object storage, not files.
On a single node, DBOS recovers at startup: it re-runs workflows still PENDING for the current application version, skipping checkpointed steps. With several servers, something must decide who recovers a dead server’s workflows: DBOS Conductor, or your own process. Conductor stays off the execution path; servers open outbound WebSocket connections to it, and if it is unreachable, workflows keep running and recovery waits.
Work is spread with durable queues backed by PostgreSQL, with concurrency and rate limits. Scheduled workflows are cron schedules stored in the database, and the docs say each firing runs on exactly one process. A workflow ID set by the caller works as an idempotency key: start the same ID twice and it runs once.
DBOS steps do not retry by default. @DBOS.step() has retries_allowed=False; with retries_allowed=True the defaults are 3 attempts, 1-second interval and backoff rate 2.0. That is the opposite of Temporal’s default, and it is easy to miss when you move code.
What changed in 3.0
DBOS Python 3.0.0 and TypeScript 5.0 (16 September 2026) moved workflow inputs and outputs into their own tables and removed deprecated pieces, including @DBOS.transaction (replaced by datasources), in-memory queues and the admin server. The key rule from the upgrade guide: 3.0 can process workflows created by 2.x, but 2.x cannot process workflows created by 3.0. Do not run 2.x and 3.0 side by side on the same application version, and upgrade every DBOSClient user together. The Go and Java libraries reached 1.0.0 in July 2026.
Versioning
By default DBOS hashes your workflow source code into an application version, and a process recovers only workflows of its own version. Safe, with a surprise: change one line, redeploy, and the old version’s PENDING workflows are not recovered. In my drill they waited until I started the old code again. Either set application_version yourself and use patching, or keep old-version processes running until they drain, which Conductor’s version management helps with.
A DBOS workflow from my drill:
@DBOS.step()
def charge(wid: str) -> str:
return mark("charge", wid)
@DBOS.workflow()
def settle(wid: str) -> str:
reserve(wid)
txn = charge(wid)
DBOS.sleep(8)
notify(wid)
return txn
Conductor and the vendor benchmark
The library is MIT, but self-hosted Conductor is under a proprietary licence and needs a licence key; some features, such as metadata-only mode, need a paid DBOS plan. You can run DBOS without Conductor, but then multi-server recovery, dashboards and retention policies are yours to build.
DBOS’s own benchmark blog reports that one PostgreSQL server (AWS RDS db.m7i.24xlarge, 96 vCPUs, 384 GB RAM, 120K provisioned IOPS) sustained 144K small writes per second, about 43K workflows per second and about 12.1K queued workflows per second on one queue. That is the vendor’s run on a very large instance; on a smaller primary that also serves your application, measure on your own hardware.
Side by side
Temporal
Restate
DBOS
Extra infrastructure
Temporal Service and a database, usually Elasticsearch for visibility
Restate server or cluster, object storage for snapshots
None beyond PostgreSQL; Conductor optional
Languages
Go, Java, Python, TypeScript, .NET and others
TypeScript, Java and Kotlin, Python, Go, Rust
Python, TypeScript, Go, Java
Default step retries
Activities: unlimited, 1 s to 100 s backoff
Invocations: 70 attempts, then pause
None unless retries_allowed=True
Code changes
Patching, Worker Versioning (Pinned or Auto-Upgrade)
New immutable deployment per version
Application version hash, patching, version-aware recovery
Timers and cron
Durable timers, Schedules
Durable timers, delayed calls; no native cron
Durable sleep, cron schedules in the database
Per-key state
Workflow state, signals, updates
Virtual Objects with key-value state
Workflow events and streams; your own tables
Scaling limit to know
51,200 events or 50 MB per history
Partition count fixed at cluster creation today
Your PostgreSQL primary’s write capacity
Licence
MIT (Java SDK Apache 2.0)
Server BSL 1.1, SDKs MIT
Libraries MIT, Conductor proprietary
Managed option
Temporal Cloud
Restate Cloud
DBOS Cloud and hosted Conductor
Retries, sagas and cron, done properly
The Ahmedabad team’s real problem is that retry state, timeouts and compensation are scattered across a table, a worker and three cron jobs. An engine pulls them into one function, but some rules apply whichever you choose.
- Make every external call idempotent. Pass the order ID or a derived key to the payment and courier APIs. Engines re-run a step whose result was not recorded.
- Separate retryable from terminal errors. A card declined is terminal; a 503 is not. Temporal uses non-retryable error types, Restate uses
TerminalError, and DBOS usesshould_retryor simply catching the exception in the workflow. - Write compensations as code, in reverse order. Reserve stock, charge, book courier; if the courier step fails terminally, refund then release stock. Record each compensation as a step, so a crash during rollback does not refund twice.
- Bound everything. Set a maximum attempts or total time per step and per workflow. Unlimited retry on a broken endpoint is an outage that nobody sees.
- Replace sweeper cron jobs with timers. “Refund if no pickup in four hours” becomes a durable sleep or a race between a signal and a timer inside the workflow, not a query that runs every five minutes. Keep real calendar jobs, such as a daily settlement report, in Temporal Schedules, DBOS schedules or a Restate timer loop.
Security advisories
Every CVE ID below was checked against the CVE.org API (cveawg.mitre.org) on 12 October 2026 and returned a published record. Affected and fixed versions are from those records.
CVE
What it is
Affected
Fixed in
CVE-2026-89139
A namespace writer could make the Worker Service run a command of their choice through the subprocess compute provider
1.31.0 to 1.31.2
1.31.3
CVE-2026-87858
A completion callback with a source header could be re-targeted at the internal frontend
1.25.0 to 1.29.7, 1.30.0 to 1.30.6, 1.31.0 to 1.31.2
1.30.7, 1.31.3
CVE-2026-16652
Unbounded CPU use when computing a Schedule’s next action time
1.17.0 to 1.29.7, 1.30.0 to 1.30.6, 1.31.0 to 1.31.2
1.30.7, 1.31.3
CVE-2026-65655
Temporal UI Server could issue OAuth cookies without Secure behind a TLS-terminating proxy
UI Server 2.7.0 to before 2.53.2
UI Server 2.53.2
CVE-2026-5724
The streaming replication endpoint skipped authorisation
1.24.0 onwards on several lines
1.28.4, 1.29.6, 1.30.4, 1.31.2, 1.32.0
CVE-2026-5199
A writer in one namespace could signal, delete or reset workflows in another through the batch activity
1.29.0 to 1.29.4, 1.30.0 to 1.30.2
1.29.5, 1.30.3
CVE-2025-14987
Cross-namespace commands were not authorised for the target namespace
Up to 1.29.1
1.27.4, 1.28.2, 1.29.2
CVE-2025-14986
Multi-operation start validated against the wrong namespace
1.24.0 to 1.29.1
1.27.4, 1.28.2, 1.29.2
CVE-2025-8396
Memory exhaustion through the authorisation header
Before 1.26.3, 1.27.0 to 1.27.2, 1.28.0
1.26.3, 1.27.3, 1.28.1
The 1.29 line has no listed fix for CVE-2026-16652 or CVE-2026-87858, so 1.29 users should move to 1.30.7 or later. Most of these issues cross namespace boundaries, which exist only if you configured an authorizer; with the default noopAuthorizer, every caller already has full access.
Restate’s and DBOS’s GitHub security advisories pages, and Temporal’s own page, listed no published advisories when I checked on 12 October 2026; Temporal publishes through CVE records instead. For Restate, the bigger risk is the documented design choice of no authentication on admin and ingress. For DBOS, your security boundary is PostgreSQL: anyone who can write to the system database can change workflow state, so give the application role only what it needs.
Licences and terms
Component
Licence
What it means in practice
Temporal Server, CLI, UI, most SDKs
MIT
Use, modify and offer as a service freely
Temporal Java SDK
Apache 2.0
Permissive, with a patent grant
Restate server
Business Source License 1.1
Production use allowed, including internal platforms, except offering a “Public Restate Platform Service” where third parties register their own services against Restate’s APIs; each version becomes Apache 2.0 four years after release
Restate SDKs
MIT
Permissive
DBOS Transact (Python, TypeScript, Go, Java)
MIT
Permissive
DBOS Conductor (self-hosted)
Proprietary, licence key
A commercial dependency if you want its recovery and dashboards
Cadence
Apache 2.0
Permissive, CNCF Sandbox
The Restate licence’s own notice says it “is not an Open Source license”. For an internal order platform like the Ahmedabad team’s, the permitted-use list covers them explicitly. For a SaaS company that wants to let its customers register their own handlers, it does not. Have your legal team read the Additional Use Grant, not a summary of it.
Migration paths
From a jobs table and cron. Move one flow, such as courier booking, whole. Let in-flight orders drain from the jobs table and send new orders to the workflow. Use the order ID as the workflow ID so a duplicate request cannot start a second run.
From Cadence, or between engines. There is no portable journal format, so running workflows cannot move. Start new workflows on the new engine and let old ones finish. DBOS publishes a guide on migrating from Temporal; it is a vendor guide, but its concept mapping is useful.
Across versions. Temporal and Restate: one minor at a time. DBOS 2.x to 3.0: stop all 2.x processes before 3.0 processes start on the same version.
A decision guide
Your situation
A sensible starting point
Watch out for
Many teams, many languages, workflows that run for weeks, a platform team to run it
Temporal, self-hosted or Temporal Cloud
Database and Elasticsearch operations, history limits, sequential upgrades, authorizer setup
Services on serverless or Kubernetes, low-latency request-response flows, per-key state
Restate
BSL terms, no auth on admin and ingress, keeping old deployments until they drain
One or a few services, PostgreSQL already run well, small team
DBOS
Load on the primary, version-hash recovery after deploys, Conductor licence for multi-server recovery
Batch DAGs of containers or data pipelines
Argo Workflows or Airflow
Not designed for per-transaction business workflows
Already on Cadence
Stay unless you need what Temporal adds; plan a drain-style move if you do
Two diverged ecosystems
A single retried call with a timeout, no orchestration
A queue with idempotency, or Temporal Standalone Activities
Adding a platform for one call
For the Ahmedabad team, my suggestion has three steps. First, before any tool, write the order flow as one function on paper: steps, retries, timeouts, compensations and the idempotency key for each external call. Double refunds of the kind they saw usually come from compensation that is not recorded, and no engine fixes a wrong design. Second, with one main service and a PostgreSQL primary that has headroom, DBOS is the smallest change: no new server, and the journal sits next to data they already back up. They should set application_version explicitly, use patching for changes, and decide early whether to buy Conductor or build multi-server recovery. Third, if they expect more teams and languages, or flows that wait days for humans, Temporal is the safer long-term platform, provided someone owns its database and upgrades. Restate fits if they move to serverless; read the licence first.
A practical checklist
- List every hand-rolled retry loop, jobs table and sweeper cron job, with what it guards.
- For each workflow, write down the steps, which calls have side effects and the idempotency key for each.
- Check the engine’s default retry policy and set explicit attempts and timeouts per step.
- Mark terminal errors in code, so a declined payment does not retry for a day.
- Write compensations as recorded steps and test a failure in the middle of a rollback.
- Choose a versioning approach (patching, pinned versions or immutable deployments) before the first production deploy, not after the first breaking change.
- Where the SDK supports it, replay recorded histories against new code in CI before you deploy.
- Set retention for journals and histories long enough for next-day debugging.
- Lock down the control plane: Temporal authorizer and mTLS, Restate admin, ingress and fabric ports, DBOS system database roles.
- Run a crash drill: kill the worker, the server or the database in the middle of a workflow and watch it finish.
Common mistakes
- Calling an API from workflow code. It runs again on every replay. Put it in a step or activity.
- Assuming steps run exactly once. They run at least once; the external side still needs an idempotency key.
- Rolling update with a breaking change. Same URL or same version, new step order: in-flight workflows fail replay.
- Relying on DBOS’s default version hash without a plan. After a code change, the old version’s pending workflows wait for a process of that version.
- Copying Temporal habits to DBOS. A DBOS step without
retries_allowed=Truefails on the first error. - One generic step function for everything. My drill showed DBOS matching steps by function name and position, so a generic wrapper hid an inserted step and returned the wrong recorded result.
- Leaving defaults on the control plane. A Temporal Service with
noopAuthorizeror a Restate admin port on the network is an open door.
Drill: one order flow, three engines, crashes and a bad deploy
On 12 October 2026, from about 2:22 AM to 2:34 AM IST, on a shared Linux VM with 8 vCPUs, 15 GiB of RAM, Debian 13, kernel 6.12 and Python 3.13.5, without root and on loopback only, I checked each download. Temporal CLI 1.9.1 (which embeds Server 1.32.0, not 1.32.1) matched the release checksums.txt and GitHub’s SHA-256 digest; Restate server and CLI 1.7.13 matched their .sha256 files and GitHub digests. Python packages came from PyPI: temporalio 1.34.0, restate-sdk 1.0.5 (on hypercorn 0.18.0) and dbos 3.2.0. PostgreSQL 18.6 binaries came from my previous drill, checksums verified against the pgdg index. Afterwards I stopped every process and deleted the binaries, environments and data.
The same workflow ran on all three: reserve, charge, a durable sleep of 8 seconds, then notify. Each step appended a line to a ledger file, so I could count executions.
Part 1: kill the application process in the middle. In each case I killed the worker or service process with kill -9 right after charge was recorded, waited 15 seconds, and started it again.
Engine
Where it was during the kill
What happened
Restart to completion
Steps run
Temporal (dev server with SQLite file)
Worker killed during the timer
Timer fired on the server; new worker picked up the workflow task
3.02 s, mostly Python worker start-up
reserve 1, charge 1, notify 1
Restate (single node)
Service killed while the handler was still connected
Invocation showed backing-off, 5 retries, “Connection refused”
2.95 s, waiting for the next retry
reserve 1, charge 1, notify 1
DBOS (PostgreSQL 18.6)
Application killed during DBOS.sleep
New process logged “Recovering 1 workflows” at start-up
1.44 s
reserve 1, charge 1, notify 1
No step ran twice, and notify never waited a fresh 8 seconds. One harness lesson: my first Restate attempts killed a wrapper while hypercorn’s child process kept serving. I discarded those runs and served the app in one process.
Part 2: kill the engine. After charge, I killed the Temporal dev server with kill -9, waited 5 seconds and restarted it on the same database file: healthy in 0.18 s, workflow done 3.03 s after restart, notify 8.05 s after charge, so the timer kept its schedule. restate-server gave 0.13 s, 2.92 s and 8.01 s. For DBOS I stopped PostgreSQL in immediate mode; the app logged connection errors and the workflow finished 3.05 s after PostgreSQL returned, without restarting the app.
Part 3: a breaking change on redeploy. I started a workflow on the original code, killed the process after charge, and started new code that inserted a fraud_check step before charge.
Engine
What the new code saw
State of the old workflow
How I recovered
Temporal, unpatched change
“Nondeterminism error: Activity type of scheduled event ‘charge’ does not match activity type of activity command ‘fraud_check’”
Running, workflow task failing and retrying; notify never ran
Worker with workflow.patched("fraud-check"): old workflow finished without fraud_check in 5 s; a new workflow ran fraud_check
Restate, new code at the same URL
[570 Journal mismatch], “name: charge != fraud_check”
Backing off, retried 5 times in about 6 s
Old code back on the old URL: finished in 2 s. New code registered as a second deployment on another port: new invocations ran fraud_check, the old one stayed on its deployment
DBOS, new code, default version hash
“No workflows to recover from application version d1da…”
PENDING under the old version, untouched
Started the old code again: “Recovering 1 workflows”, completed
DBOS, new code, same version forced
“DBOS Error 11: During execution of workflow d4 step 2, function charge was recorded when fraud_check was expected”
ERROR, a final state
fork_workflow from step 3 under the old version: a new workflow ID replayed reserve and charge from checkpoints and ran sleep and notify, 9.0 s
Temporal and Restate kept the workflow alive and retrying until the code was fixed. DBOS’s version hash avoided the mismatch by not recovering the workflow, which is safe but silent; forcing the same version made it a terminal error that needed a fork. One more DBOS result surprised me. In a variant with one generic step function for all steps, DBOS matched by function name and position, returned charge’s recorded result for the inserted fraud_check without running it, and failed only one step later at the sleep. Name your steps.
Smaller observations: resident memory was about 127 MiB for the Temporal dev server after start (177 MiB at the end), 294 MiB for single-node Restate with defaults and 72 MiB for the DBOS application process. DBOS 3.2.0 created its system database and applied its schema migrations at first start; the database was under 9 MB after six workflows.
Honest limits: one VM, loopback only, the Temporal development server rather than a production Temporal Service, single-node Restate, one PostgreSQL instance, no Conductor, no clustering, no load and one run per scenario. Restart timings include Python start-up and retry backoff, so they show behaviour, not performance. I did not test Temporal Worker Versioning, Restate’s Kubernetes operator, DBOS queues or schedules, or any SDK other than Python. The only throughput numbers in this post are DBOS’s, labelled as the vendor’s.
What to unlearn and re-learn
- Unlearn “retries are a loop with a counter”. Re-learn that retries, timers and compensation belong in a recorded workflow, with the counter kept by the engine.
- Unlearn “durable execution means exactly once”. Re-learn that steps run at least once and the external side still needs idempotency keys.
- Unlearn “workflow engines need a big server”. Re-learn that the choice now runs from a full service (Temporal) through a single binary (Restate) to a library on PostgreSQL (DBOS).
- Unlearn “a deploy is just new code”. Re-learn that running workflows hold journals written by old code, so each deploy needs a versioning decision.
- Unlearn “the defaults are sensible for us”. Re-learn each engine’s retry, retention and security defaults, because they differ in ways that matter for payments.
Revisit your retries before you write another cron job
The Ahmedabad team does not need the most popular engine; it needs one place where an order’s progress is recorded, and a design that survives restarts and redeploys. Learn the durable execution model: journal, replay, deterministic workflow code and steps that run at least once. Unlearn the habit of adding a new cron job for every state an order can get stuck in, and the belief that a retry counter in a JSON column is a workflow. Re-learn the current Temporal, Restate and DBOS releases, their defaults and their licences from the release notes, because all three changed a lot in 2026. Practise by killing a worker, a server and a database in the middle of a workflow, and by deploying a deliberately breaking change, as I did on one VM. Then apply what you learn to one real flow first, with idempotent steps, explicit retry limits, recorded compensations and a versioning plan. Revisit it before your next big sale, and whenever one of these engines ships a new minor release.
Sources
- Temporal: repository, 1.32.1 release, 1.32.0 release, 1.31.3 release, 1.31.0 release, 1.30.7 release, 1.0.0 release, LICENSE, CLI 1.9.1 release, CLI server command, persistence, visibility, self-hosted defaults, workflow execution limits, retry policies, workflow definition, patching, Worker Versioning, Worker Versioning in production, Continue-As-New, Workflow ID, Schedules, upgrading the server, self-hosted security, security advisories, Python SDK on PyPI, Java SDK LICENSE
- Restate: repository, 1.7.13 release, 1.8.0-rc.2 release, 1.7.0 release, 1.6.0 release, 1.5.0 release, 1.2.0 release, 1.0.0 release, LICENSE, architecture, services, service configuration, server configuration reference, versioning, clusters, snapshots, server security, upgrading, timers and scheduling, microservice orchestration and sagas, security advisories, Python SDK on PyPI, TypeScript SDK LICENSE
- DBOS: Python library repository, Python 3.2.0 release, Python 3.0.0 release, TypeScript 5.0 release, Go 1.0.0 release, Java 1.0.0 release, Python LICENSE, architecture, workflows, steps, queues, scheduled workflows, upgrading workflow code, upgrading to 3.0, Conductor overview, self-hosting Conductor, migrating from Temporal, vendor benchmark on PostgreSQL, security advisories, Python package on PyPI
- Related projects: Cadence repository, Cadence 1.4.2 release, CNCF: Cadence Workflow, Argo Workflows 4.1.5 release, Apache Airflow 3.3.2 release
- CVE records (cveawg.mitre.org API): CVE-2026-89139, CVE-2026-87858, CVE-2026-16652, CVE-2026-65655, CVE-2026-5724, CVE-2026-5199, CVE-2025-14987, CVE-2025-14986, CVE-2025-8396
Releted Posts
PostgreSQL high availability: revisit Patroni and CloudNativePG before your next failover
Picture a lending company in Bengaluru (an invented example, not a real client). Its loan ledger runs on PostgreSQL 16 across three VMs, managed by Patroni with a three-node etcd cluster and HAProxy in front.
Read moreFluent Bit, Vector and the OpenTelemetry Collector: revisit your log pipeline before you ship more logs
A logistics company in Pune runs about 600 pods across three Kubernetes clusters. A Fluentd DaemonSet, set up in 2019 with a dozen Ruby plugins, ships about 25 crore log lines a day to Loki.
Read moreRabbitMQ, NATS and Kafka queues: revisit your message queue before you add another broker
A health insurance company in Hyderabad processes claims through RabbitMQ: three nodes on 3.13 with classic mirrored queues, about 40 lakh messages a day, from claim documents to OCR, OCR results to fraud checks and approvals to payments.
Read more