Durable execution in 2026: revisit Temporal, Restate and DBOS before you hand-roll retries, sagas and cron jobs

Picture an online pharmacy and diagnostics booking company in Ahmedabad (an invented example, not a real client). An order goes through five steps: hold the medicine stock, take the UPI payment, book a courier slot, send the customer an SMS, and, if the courier cannot pick up within four hours, release the stock and refund. Today this is a jobs table in PostgreSQL, a worker that polls it, retry counters in a JSON column and three cron jobs that look for orders stuck in odd states. Last month a rolling deploy restarted the workers in the middle of a busy evening. Some orders were refunded twice, some never got a courier, and the on-call engineer spent the night writing SQL. Now one architect wants Temporal, a second has read about Restate, and a third says DBOS can do it with the PostgreSQL they already run.

All three are durable execution engines. They record each step of a long-running function so that after a crash, a deploy or a timeout it continues from where it stopped. They differ in where that record lives, who drives the code, how they cope with code changes, and what licence you accept. This post compares Temporal, Restate and DBOS on those points, with short notes on Cadence, Argo Workflows and Airflow. Facts come from GitHub release pages, official docs, licence files and CVE records, checked on 12 October 2026. The drill near the end is my own run.

The short version

  • Durable execution is not a queue with retries. The engine keeps a journal of completed steps and replays your function against it after a failure, so a careless code change can break running workflows. In my drill, all three engines caught such a change, each differently.
  • Temporal is the mature, server-centred option: services backed by Cassandra, MySQL or PostgreSQL, and workers that poll task queues. Server 1.32.0 (11 September 2026) made Standalone Activities GA, 1.32.1 followed on 9 October, and the deprecated Worker Versioning APIs go in 1.33.
  • Restate is a single binary with its own replicated log and RocksDB state that pushes invocations to your handlers over HTTP. Version 1.7.13 (1 October 2026) is current and 1.8.0 is at rc.2. The server is under the Business Source License 1.1, not an open source licence; the SDKs are MIT.
  • DBOS is a library that checkpoints steps to PostgreSQL, with no orchestration server. DBOS Python 3.0 and TypeScript 5.0 (16 September 2026) changed the schema, and 2.x processes cannot process workflows created by 3.0. The libraries are MIT; the Conductor control plane is proprietary.
  • Default retry behaviour differs a lot: Temporal retries activities without limit, Restate retries invocations 70 times and then pauses them, and DBOS does not retry a step unless you ask for it. Read the defaults before you trust them.
  • Temporal published six CVEs between April and September 2026, one in the UI server, including CVE-2026-89139 (a namespace writer could run commands on the Worker Service host in 1.31.0 to 1.31.2). For self-hosted Temporal, 1.30.7, 1.31.3 or 1.32.x is the floor today.

Where things stand on 12 October 2026

Dates are from GitHub release pages and PyPI, converted to IST.

Project

Latest release

What it is

Licence

Temporal Server

1.32.1 on 9 October 2026 (1.32.0 on 11 September; 1.31.3 on 19 September; 1.30.7 on 18 September)

Workflow orchestration service with workers, task queues and event histories

MIT

Temporal CLI

1.9.1 on 15 September 2026, embedding Server 1.32.0 for start-dev

CLI and single-process development server

MIT

Temporal SDKs

Python SDK 1.34.0 on 1 October 2026

Go, Java, Python, TypeScript, .NET and others

Go, Python, TypeScript and .NET MIT; Java Apache 2.0

Restate

1.7.13 on 1 October 2026; 1.8.0-rc.2 on 9 October

Single-binary durable execution server with replicated log and embedded RocksDB

Server: Business Source License 1.1, changing to Apache 2.0 four years after each release

Restate SDKs

Python SDK 1.0.5 on 2 September 2026

TypeScript, Java and Kotlin, Python, Go, Rust

MIT

DBOS Transact

Python 3.2.0 and TypeScript 5.2 on 29 September 2026; Java 1.2.0 on 1 October; Go 1.6.0 on 7 October

Durable workflow library that checkpoints to PostgreSQL

MIT

DBOS Conductor

Self-hosted or DBOS-hosted

Control plane for recovery, dashboards and retention

Proprietary, licence key required

Cadence

1.4.2 on 3 October 2026

Uber’s workflow engine, the project Temporal was forked from

Apache 2.0; CNCF Sandbox since 22 May 2025

Date

Event

21 February 2017

Cadence repository created at Uber

17 October 2019

Temporal repository created; Temporal starts as a fork of Cadence

30 September 2020

Temporal Server 1.0.0

4 January 2023

Restate repository created

27 April 2023

Restate 0.1.0

12 July 2023

DBOS TypeScript library repository created

7 June 2024

Restate 1.0.0

19 July 2024

DBOS Python library repository created

15 February 2025

Restate 1.2.0: multi-node replicated clusters, restatectl, snapshots to object storage

22 May 2025

Cadence accepted into CNCF at Sandbox level

30 January 2026

Restate 1.6.0: pause and resume invocations, restart from a journal prefix

30 April 2026

Temporal Server 1.31.0: Task Queue Priority and Fairness GA, Cassandra 5 support

19 June 2026

Restate 1.7.0: flow control (concurrency rules), bounded invoker memory, UI 1.0, ingress request size limit

2 July 2026

DBOS Java 1.0.0

31 July 2026

DBOS Go 1.0.0

11 September 2026

Temporal Server 1.32.0: Standalone Activities GA, eager activity execution on by default, unified visibility query converter by default

16 September 2026

DBOS Python 3.0.0 and TypeScript 5.0: new storage schema, deprecated features removed

18 and 19 September 2026

Temporal 1.30.7 and 1.31.3 fix the September CVEs

1 October 2026

Restate 1.7.13 fixes Virtual Object state consistency across replicas with VQueues

9 October 2026

Temporal 1.32.1; Restate 1.8.0-rc.2

What durable execution actually means

You write a function that calls several steps. The engine records each step’s result in a durable journal before your code moves on. If the process dies, the function runs again from the top, but completed steps return their recorded results instead of running again. Timers are journal entries too, so a four-hour sleep survives a restart. Three consequences follow.

  1. The workflow function must be deterministic. On replay it must call the same steps in the same order. Clock reads, random numbers and API calls go inside steps.
  2. Steps run at least once, not exactly once. If the process dies after the payment API returned but before the result was recorded, the step runs again. “Exactly once” in vendor material means the recorded result is applied once; the external side effect still needs an idempotency key. My idempotency keys post covers that half, and nothing here replaces it.
  3. Code changes are a data migration. A running workflow has a journal written by the old code. If new code takes a different path, the journal no longer matches. Each engine has its own answer: replay checks and patching in Temporal, immutable deployments in Restate, application versions and patching in DBOS.

Three shapes of the same idea

Temporal

Restate

DBOS

What you deploy

A Temporal Service (frontend, history, matching and worker services) plus a database, plus your workers

One restate-server binary (single node or cluster), plus your service endpoints

Your application with the DBOS library, plus PostgreSQL

Where the journal lives

Event history in Cassandra, MySQL or PostgreSQL (SQLite for development)

Replicated log (Bifrost) with state in embedded RocksDB, snapshots to S3-compatible storage

Tables in a PostgreSQL “system database”

How work reaches your code

Workers poll task queues

Restate calls your handler over HTTP/2 (or Lambda and similar), streaming the journal

Your process runs the function itself; other processes pick up work from durable queues

Unit of work

Workflow and Activities

Service handler, Virtual Object (keyed, single writer) or Workflow

Workflow and steps

Replay check

SDK compares new commands with the history

SDK compares new journal entries with the recorded ones

Library compares step names by position

Not the same category

Cadence is the ancestor. Temporal’s README says it “originated as a fork of Uber’s Cadence” and is developed by the creators of Cadence. Cadence is still developed (1.4.2 on 3 October 2026) and joined the CNCF Sandbox in May 2025. If you run it, you have a choice, not an emergency.

Argo Workflows (4.1.5 on 9 October 2026) runs DAGs of containers on Kubernetes, and Apache Airflow (3.3.2 on 17 September 2026) schedules data pipelines. Both suit batch jobs; neither is built for thousands of short business transactions, each waiting on payments, humans or timers. For the nightly ETL, look there. For “this order must reach a final state whatever fails”, read on.

A message queue is also not durable execution. A queue delivers messages; it does not remember that step three of an order finished. My message queue post covers queues, and many teams will run both.

Temporal in 2026

The model

Temporal splits code into Workflows (deterministic orchestration) and Activities (anything that touches the outside world). Your workers poll task queues. On a workflow task, the SDK replays the event history through your code and sends back new commands, such as “schedule activity charge”. If they do not match the history, the SDK raises a nondeterminism error and the workflow task keeps failing until you fix the code, as my drill showed.

Hard limits shape design. An event history is limited to 51,200 events or 50 MB (warnings at 10,240 events or 10 MB), after which the execution is terminated. A payload warns at 256 KB and fails at 2 MB. A workflow may have at most 2,000 pending activities, child workflows, signals or cancellations of each type by default, and the docs suggest 500 or fewer. Entity workflows, such as one per customer, need Continue-As-New to start a fresh history.

Activities retry by default with exponential backoff from 1 second, coefficient 2.0, capped at 100 seconds, with unlimited attempts; workflows do not retry by default. A broken activity therefore retries quietly for days unless you set maximum_attempts or a schedule-to-close timeout. Recurring work belongs in Schedules, which have their own identity, rather than the older Temporal Cron Jobs.

What changed in 1.31 and 1.32

Server 1.31.0 (30 April 2026) made Task Queue Priority and Fairness generally available and added Cassandra 5 support. Server 1.32.0 (11 September 2026) changed several defaults:

  • Standalone Activities are GA: an activity started from a client, without a workflow, still gets retries, timeouts and visibility. That suits a single durable call.
  • Eager activity execution is on by default.
  • The unified visibility query converter is the default, with breaking changes such as type-mismatched search attribute comparisons now being errors. The legacy converter goes in 1.33.
  • Worker Versioning: 1.32 is the last release with the deprecated Build ID APIs. Before 1.33, move to Worker Deployment APIs and drain workflows that depend on legacy routing.
  • Nexus callbacks now route by URL scheme, a change the notes say was made for security reasons, and component.callbacks.allowedAddresses became callback.allowedAddresses.

Versioning running workflows

Temporal gives you two tools. Patching puts the change behind workflow.patched("id") (named differently per SDK), which returns false while replaying old histories and true for new runs. Worker Versioning groups workers into Worker Deployment Versions. A Pinned workflow stays on the version it started on; an Auto-Upgrade workflow moves to the current version and must be kept replay-safe with patching. A patched version of my drill workflow:

@workflow.defn(name="Settle")
class Settle:
    @workflow.run
    async def run(self, wid: str) -> str:
        await workflow.execute_activity(reserve, wid, start_to_close_timeout=T)
        if workflow.patched("fraud-check"):
            await workflow.execute_activity(fraud_check, wid, start_to_close_timeout=T)
        txn = await workflow.execute_activity(charge, wid, start_to_close_timeout=T)
        await workflow.sleep(8)
        await workflow.execute_activity(notify, wid, start_to_close_timeout=T)
        return txn

Running it yourself

The persistence page lists the versions Temporal tests: Cassandra 3.11, 4.0 and 5.0.4 or later, PostgreSQL 13.18, 14.15, 15.10 and 16.6, and MySQL 5.7 and 8.0.19 or later. SQLite is for development only. Advanced visibility runs on Elasticsearch, OpenSearch 2 (from Server 1.30.1) or the same SQL databases; the docs recommend Elasticsearch for anything beyond a few workflows. Notice that PostgreSQL 17 and 18 are not on the tested list, so test before you put Temporal on the cluster from my PostgreSQL HA post.

Upgrades must be sequential, one minor version at a time, after moving to the latest patch of your current minor. Skipping versions can leave old data formats the new code no longer understands.

Security is your job. If you do not configure an Authorizer, Temporal uses the noopAuthorizer, which allows every API request. Turn on mTLS and a claim mapper and authorizer before any second team shares the cluster; several of the CVEs below only matter once namespaces are a real boundary.

Restate in 2026

The model

Restate inverts the worker model. Your code runs as an ordinary HTTP service, Lambda function or similar. You register its URL, and Restate calls your handlers, streaming the journal to the SDK. During a long sleep Restate can suspend the invocation and call again when the timer fires, so on serverless platforms you do not pay for a sleeping workflow.

The server is one binary, which its architecture page calls a “partitioned, log-centric runtime”: each partition’s leader appends events to a replicated log (Bifrost) and a partition processor materialises state in embedded RocksDB. A step “happens” when a quorum acknowledges its log append. Snapshots go to S3-compatible storage to bound recovery time. Multi-node clusters arrived in 1.2.0 (February 2025).

There are three service types. A Service runs handlers in parallel. A Virtual Object is keyed, with key-value state and a single writer per key, which suits per-order state without a lock table. A Workflow runs its run handler once per ID. Sagas are code: keep a list of compensations and run them in reverse on a terminal error. There is no native cron; the docs suggest durable timers with a Virtual Object.

What changed in 1.5 to 1.7

  • 1.5.0 (September 2025) introduced a new invocation retry policy.
  • 1.6.0 (30 January 2026) added pause and resume, and restarting an invocation from a prefix of its journal. Deprecated SDK versions are now rejected.
  • 1.7.0 (19 June 2026) added concurrency limits per scope (restate rules, behind experimental_enable_vqueues), a bounded invoker memory pool (1.5 GiB default), UI 1.0 and an ingress request size limit (32 MiB, returning 413).
  • 1.7.13 (1 October 2026) fixed Virtual Object state edited through the Admin API with VQueues being skipped on followers; affected replicated clusters must re-submit that state.
  • 1.8.0 is at rc.2; deprecated root-level service-client options are removed in 1.8.

Retries, retention and defaults

The services configuration page documents the default invocation retry policy as exponential from 50 ms, factor 2, up to 60 seconds and 70 attempts, then pause until a human resumes it. The server configuration reference and --dump-config on my 1.7.13 binary show a 500 ms initial interval, so check your effective configuration rather than either page. Pause is the safer default for payments, because killing an invocation does not run your compensations.

Journals, idempotency records and finished workflow state are retained for 24 hours by default. Raise these for anything you will want to inspect next morning.

Versioning by immutable deployments

In Restate a deployment is immutable. You deploy new code to a new URL and register it; new invocations go to the latest deployment while in-flight ones finish where they started. No patch statements, but old deployments must run until they drain, which is why the docs warn against handlers that run for days. On FaaS platforms each version already has its own URL or ARN, and the Kubernetes operator automates registration, draining and scale to zero.

If you update code behind the same URL instead, which is what most Kubernetes rolling updates do, Restate detects the mismatch on replay and returns error 570 “Journal mismatch”. My drill showed this.

Operating it

Minor upgrades must not be skipped, and rollback is supported only to the previous minor and only if you have not used new features. In a cluster, wait for each node’s partitions to recover before restarting the next; a health check alone is not enough. The security page is direct: the server has no application-level authentication on its management interfaces by design. The admin port (9070) must be limited to operators, the ingress port (8080) belongs behind an authenticating proxy, and the fabric port (5122) must not leave the cluster. On my run, the fabric port listened on all interfaces by default while I had bound the other two to 127.0.0.1. Also note that disable-telemetry defaults to false; the reference says anonymous usage data goes to Scarf, and DO_NOT_TRACK=1 turns it off.

A Restate workflow from my drill, in the Python SDK:

wf = restate.Workflow("Settle")

@wf.main()
async def run(ctx: restate.WorkflowContext, wid: str) -> str:
    await ctx.run_typed("reserve", mark, step="reserve", wid=wid)
    txn = await ctx.run_typed("charge", mark, step="charge", wid=wid)
    await ctx.sleep(timedelta(seconds=8))
    await ctx.run_typed("notify", mark, step="notify", wid=wid)
    return txn

app = restate.app([wf])

DBOS in 2026

The model

DBOS is a library. You decorate functions as workflows and steps, and it writes checkpoints to a PostgreSQL “system database”, with “no separate orchestration server”. The stated overhead is one write per step plus two per workflow. Large step outputs mean large writes, so return pointers to object storage, not files.

On a single node, DBOS recovers at startup: it re-runs workflows still PENDING for the current application version, skipping checkpointed steps. With several servers, something must decide who recovers a dead server’s workflows: DBOS Conductor, or your own process. Conductor stays off the execution path; servers open outbound WebSocket connections to it, and if it is unreachable, workflows keep running and recovery waits.

Work is spread with durable queues backed by PostgreSQL, with concurrency and rate limits. Scheduled workflows are cron schedules stored in the database, and the docs say each firing runs on exactly one process. A workflow ID set by the caller works as an idempotency key: start the same ID twice and it runs once.

DBOS steps do not retry by default. @DBOS.step() has retries_allowed=False; with retries_allowed=True the defaults are 3 attempts, 1-second interval and backoff rate 2.0. That is the opposite of Temporal’s default, and it is easy to miss when you move code.

What changed in 3.0

DBOS Python 3.0.0 and TypeScript 5.0 (16 September 2026) moved workflow inputs and outputs into their own tables and removed deprecated pieces, including @DBOS.transaction (replaced by datasources), in-memory queues and the admin server. The key rule from the upgrade guide: 3.0 can process workflows created by 2.x, but 2.x cannot process workflows created by 3.0. Do not run 2.x and 3.0 side by side on the same application version, and upgrade every DBOSClient user together. The Go and Java libraries reached 1.0.0 in July 2026.

Versioning

By default DBOS hashes your workflow source code into an application version, and a process recovers only workflows of its own version. Safe, with a surprise: change one line, redeploy, and the old version’s PENDING workflows are not recovered. In my drill they waited until I started the old code again. Either set application_version yourself and use patching, or keep old-version processes running until they drain, which Conductor’s version management helps with.

A DBOS workflow from my drill:

@DBOS.step()
def charge(wid: str) -> str:
    return mark("charge", wid)

@DBOS.workflow()
def settle(wid: str) -> str:
    reserve(wid)
    txn = charge(wid)
    DBOS.sleep(8)
    notify(wid)
    return txn

Conductor and the vendor benchmark

The library is MIT, but self-hosted Conductor is under a proprietary licence and needs a licence key; some features, such as metadata-only mode, need a paid DBOS plan. You can run DBOS without Conductor, but then multi-server recovery, dashboards and retention policies are yours to build.

DBOS’s own benchmark blog reports that one PostgreSQL server (AWS RDS db.m7i.24xlarge, 96 vCPUs, 384 GB RAM, 120K provisioned IOPS) sustained 144K small writes per second, about 43K workflows per second and about 12.1K queued workflows per second on one queue. That is the vendor’s run on a very large instance; on a smaller primary that also serves your application, measure on your own hardware.

Side by side

Temporal

Restate

DBOS

Extra infrastructure

Temporal Service and a database, usually Elasticsearch for visibility

Restate server or cluster, object storage for snapshots

None beyond PostgreSQL; Conductor optional

Languages

Go, Java, Python, TypeScript, .NET and others

TypeScript, Java and Kotlin, Python, Go, Rust

Python, TypeScript, Go, Java

Default step retries

Activities: unlimited, 1 s to 100 s backoff

Invocations: 70 attempts, then pause

None unless retries_allowed=True

Code changes

Patching, Worker Versioning (Pinned or Auto-Upgrade)

New immutable deployment per version

Application version hash, patching, version-aware recovery

Timers and cron

Durable timers, Schedules

Durable timers, delayed calls; no native cron

Durable sleep, cron schedules in the database

Per-key state

Workflow state, signals, updates

Virtual Objects with key-value state

Workflow events and streams; your own tables

Scaling limit to know

51,200 events or 50 MB per history

Partition count fixed at cluster creation today

Your PostgreSQL primary’s write capacity

Licence

MIT (Java SDK Apache 2.0)

Server BSL 1.1, SDKs MIT

Libraries MIT, Conductor proprietary

Managed option

Temporal Cloud

Restate Cloud

DBOS Cloud and hosted Conductor

Retries, sagas and cron, done properly

The Ahmedabad team’s real problem is that retry state, timeouts and compensation are scattered across a table, a worker and three cron jobs. An engine pulls them into one function, but some rules apply whichever you choose.

  • Make every external call idempotent. Pass the order ID or a derived key to the payment and courier APIs. Engines re-run a step whose result was not recorded.
  • Separate retryable from terminal errors. A card declined is terminal; a 503 is not. Temporal uses non-retryable error types, Restate uses TerminalError, and DBOS uses should_retry or simply catching the exception in the workflow.
  • Write compensations as code, in reverse order. Reserve stock, charge, book courier; if the courier step fails terminally, refund then release stock. Record each compensation as a step, so a crash during rollback does not refund twice.
  • Bound everything. Set a maximum attempts or total time per step and per workflow. Unlimited retry on a broken endpoint is an outage that nobody sees.
  • Replace sweeper cron jobs with timers. “Refund if no pickup in four hours” becomes a durable sleep or a race between a signal and a timer inside the workflow, not a query that runs every five minutes. Keep real calendar jobs, such as a daily settlement report, in Temporal Schedules, DBOS schedules or a Restate timer loop.

Security advisories

Every CVE ID below was checked against the CVE.org API (cveawg.mitre.org) on 12 October 2026 and returned a published record. Affected and fixed versions are from those records.

CVE

What it is

Affected

Fixed in

CVE-2026-89139

A namespace writer could make the Worker Service run a command of their choice through the subprocess compute provider

1.31.0 to 1.31.2

1.31.3

CVE-2026-87858

A completion callback with a source header could be re-targeted at the internal frontend

1.25.0 to 1.29.7, 1.30.0 to 1.30.6, 1.31.0 to 1.31.2

1.30.7, 1.31.3

CVE-2026-16652

Unbounded CPU use when computing a Schedule’s next action time

1.17.0 to 1.29.7, 1.30.0 to 1.30.6, 1.31.0 to 1.31.2

1.30.7, 1.31.3

CVE-2026-65655

Temporal UI Server could issue OAuth cookies without Secure behind a TLS-terminating proxy

UI Server 2.7.0 to before 2.53.2

UI Server 2.53.2

CVE-2026-5724

The streaming replication endpoint skipped authorisation

1.24.0 onwards on several lines

1.28.4, 1.29.6, 1.30.4, 1.31.2, 1.32.0

CVE-2026-5199

A writer in one namespace could signal, delete or reset workflows in another through the batch activity

1.29.0 to 1.29.4, 1.30.0 to 1.30.2

1.29.5, 1.30.3

CVE-2025-14987

Cross-namespace commands were not authorised for the target namespace

Up to 1.29.1

1.27.4, 1.28.2, 1.29.2

CVE-2025-14986

Multi-operation start validated against the wrong namespace

1.24.0 to 1.29.1

1.27.4, 1.28.2, 1.29.2

CVE-2025-8396

Memory exhaustion through the authorisation header

Before 1.26.3, 1.27.0 to 1.27.2, 1.28.0

1.26.3, 1.27.3, 1.28.1

The 1.29 line has no listed fix for CVE-2026-16652 or CVE-2026-87858, so 1.29 users should move to 1.30.7 or later. Most of these issues cross namespace boundaries, which exist only if you configured an authorizer; with the default noopAuthorizer, every caller already has full access.

Restate’s and DBOS’s GitHub security advisories pages, and Temporal’s own page, listed no published advisories when I checked on 12 October 2026; Temporal publishes through CVE records instead. For Restate, the bigger risk is the documented design choice of no authentication on admin and ingress. For DBOS, your security boundary is PostgreSQL: anyone who can write to the system database can change workflow state, so give the application role only what it needs.

Licences and terms

Component

Licence

What it means in practice

Temporal Server, CLI, UI, most SDKs

MIT

Use, modify and offer as a service freely

Temporal Java SDK

Apache 2.0

Permissive, with a patent grant

Restate server

Business Source License 1.1

Production use allowed, including internal platforms, except offering a “Public Restate Platform Service” where third parties register their own services against Restate’s APIs; each version becomes Apache 2.0 four years after release

Restate SDKs

MIT

Permissive

DBOS Transact (Python, TypeScript, Go, Java)

MIT

Permissive

DBOS Conductor (self-hosted)

Proprietary, licence key

A commercial dependency if you want its recovery and dashboards

Cadence

Apache 2.0

Permissive, CNCF Sandbox

The Restate licence’s own notice says it “is not an Open Source license”. For an internal order platform like the Ahmedabad team’s, the permitted-use list covers them explicitly. For a SaaS company that wants to let its customers register their own handlers, it does not. Have your legal team read the Additional Use Grant, not a summary of it.

Migration paths

From a jobs table and cron. Move one flow, such as courier booking, whole. Let in-flight orders drain from the jobs table and send new orders to the workflow. Use the order ID as the workflow ID so a duplicate request cannot start a second run.

From Cadence, or between engines. There is no portable journal format, so running workflows cannot move. Start new workflows on the new engine and let old ones finish. DBOS publishes a guide on migrating from Temporal; it is a vendor guide, but its concept mapping is useful.

Across versions. Temporal and Restate: one minor at a time. DBOS 2.x to 3.0: stop all 2.x processes before 3.0 processes start on the same version.

A decision guide

Your situation

A sensible starting point

Watch out for

Many teams, many languages, workflows that run for weeks, a platform team to run it

Temporal, self-hosted or Temporal Cloud

Database and Elasticsearch operations, history limits, sequential upgrades, authorizer setup

Services on serverless or Kubernetes, low-latency request-response flows, per-key state

Restate

BSL terms, no auth on admin and ingress, keeping old deployments until they drain

One or a few services, PostgreSQL already run well, small team

DBOS

Load on the primary, version-hash recovery after deploys, Conductor licence for multi-server recovery

Batch DAGs of containers or data pipelines

Argo Workflows or Airflow

Not designed for per-transaction business workflows

Already on Cadence

Stay unless you need what Temporal adds; plan a drain-style move if you do

Two diverged ecosystems

A single retried call with a timeout, no orchestration

A queue with idempotency, or Temporal Standalone Activities

Adding a platform for one call

For the Ahmedabad team, my suggestion has three steps. First, before any tool, write the order flow as one function on paper: steps, retries, timeouts, compensations and the idempotency key for each external call. Double refunds of the kind they saw usually come from compensation that is not recorded, and no engine fixes a wrong design. Second, with one main service and a PostgreSQL primary that has headroom, DBOS is the smallest change: no new server, and the journal sits next to data they already back up. They should set application_version explicitly, use patching for changes, and decide early whether to buy Conductor or build multi-server recovery. Third, if they expect more teams and languages, or flows that wait days for humans, Temporal is the safer long-term platform, provided someone owns its database and upgrades. Restate fits if they move to serverless; read the licence first.

A practical checklist

  1. List every hand-rolled retry loop, jobs table and sweeper cron job, with what it guards.
  2. For each workflow, write down the steps, which calls have side effects and the idempotency key for each.
  3. Check the engine’s default retry policy and set explicit attempts and timeouts per step.
  4. Mark terminal errors in code, so a declined payment does not retry for a day.
  5. Write compensations as recorded steps and test a failure in the middle of a rollback.
  6. Choose a versioning approach (patching, pinned versions or immutable deployments) before the first production deploy, not after the first breaking change.
  7. Where the SDK supports it, replay recorded histories against new code in CI before you deploy.
  8. Set retention for journals and histories long enough for next-day debugging.
  9. Lock down the control plane: Temporal authorizer and mTLS, Restate admin, ingress and fabric ports, DBOS system database roles.
  10. Run a crash drill: kill the worker, the server or the database in the middle of a workflow and watch it finish.

Common mistakes

  • Calling an API from workflow code. It runs again on every replay. Put it in a step or activity.
  • Assuming steps run exactly once. They run at least once; the external side still needs an idempotency key.
  • Rolling update with a breaking change. Same URL or same version, new step order: in-flight workflows fail replay.
  • Relying on DBOS’s default version hash without a plan. After a code change, the old version’s pending workflows wait for a process of that version.
  • Copying Temporal habits to DBOS. A DBOS step without retries_allowed=True fails on the first error.
  • One generic step function for everything. My drill showed DBOS matching steps by function name and position, so a generic wrapper hid an inserted step and returned the wrong recorded result.
  • Leaving defaults on the control plane. A Temporal Service with noopAuthorizer or a Restate admin port on the network is an open door.

Drill: one order flow, three engines, crashes and a bad deploy

On 12 October 2026, from about 2:22 AM to 2:34 AM IST, on a shared Linux VM with 8 vCPUs, 15 GiB of RAM, Debian 13, kernel 6.12 and Python 3.13.5, without root and on loopback only, I checked each download. Temporal CLI 1.9.1 (which embeds Server 1.32.0, not 1.32.1) matched the release checksums.txt and GitHub’s SHA-256 digest; Restate server and CLI 1.7.13 matched their .sha256 files and GitHub digests. Python packages came from PyPI: temporalio 1.34.0, restate-sdk 1.0.5 (on hypercorn 0.18.0) and dbos 3.2.0. PostgreSQL 18.6 binaries came from my previous drill, checksums verified against the pgdg index. Afterwards I stopped every process and deleted the binaries, environments and data.

The same workflow ran on all three: reserve, charge, a durable sleep of 8 seconds, then notify. Each step appended a line to a ledger file, so I could count executions.

Part 1: kill the application process in the middle. In each case I killed the worker or service process with kill -9 right after charge was recorded, waited 15 seconds, and started it again.

Engine

Where it was during the kill

What happened

Restart to completion

Steps run

Temporal (dev server with SQLite file)

Worker killed during the timer

Timer fired on the server; new worker picked up the workflow task

3.02 s, mostly Python worker start-up

reserve 1, charge 1, notify 1

Restate (single node)

Service killed while the handler was still connected

Invocation showed backing-off, 5 retries, “Connection refused”

2.95 s, waiting for the next retry

reserve 1, charge 1, notify 1

DBOS (PostgreSQL 18.6)

Application killed during DBOS.sleep

New process logged “Recovering 1 workflows” at start-up

1.44 s

reserve 1, charge 1, notify 1

No step ran twice, and notify never waited a fresh 8 seconds. One harness lesson: my first Restate attempts killed a wrapper while hypercorn’s child process kept serving. I discarded those runs and served the app in one process.

Part 2: kill the engine. After charge, I killed the Temporal dev server with kill -9, waited 5 seconds and restarted it on the same database file: healthy in 0.18 s, workflow done 3.03 s after restart, notify 8.05 s after charge, so the timer kept its schedule. restate-server gave 0.13 s, 2.92 s and 8.01 s. For DBOS I stopped PostgreSQL in immediate mode; the app logged connection errors and the workflow finished 3.05 s after PostgreSQL returned, without restarting the app.

Part 3: a breaking change on redeploy. I started a workflow on the original code, killed the process after charge, and started new code that inserted a fraud_check step before charge.

Engine

What the new code saw

State of the old workflow

How I recovered

Temporal, unpatched change

“Nondeterminism error: Activity type of scheduled event ‘charge’ does not match activity type of activity command ‘fraud_check’”

Running, workflow task failing and retrying; notify never ran

Worker with workflow.patched("fraud-check"): old workflow finished without fraud_check in 5 s; a new workflow ran fraud_check

Restate, new code at the same URL

[570 Journal mismatch], “name: charge != fraud_check”

Backing off, retried 5 times in about 6 s

Old code back on the old URL: finished in 2 s. New code registered as a second deployment on another port: new invocations ran fraud_check, the old one stayed on its deployment

DBOS, new code, default version hash

“No workflows to recover from application version d1da…”

PENDING under the old version, untouched

Started the old code again: “Recovering 1 workflows”, completed

DBOS, new code, same version forced

“DBOS Error 11: During execution of workflow d4 step 2, function charge was recorded when fraud_check was expected”

ERROR, a final state

fork_workflow from step 3 under the old version: a new workflow ID replayed reserve and charge from checkpoints and ran sleep and notify, 9.0 s

Temporal and Restate kept the workflow alive and retrying until the code was fixed. DBOS’s version hash avoided the mismatch by not recovering the workflow, which is safe but silent; forcing the same version made it a terminal error that needed a fork. One more DBOS result surprised me. In a variant with one generic step function for all steps, DBOS matched by function name and position, returned charge’s recorded result for the inserted fraud_check without running it, and failed only one step later at the sleep. Name your steps.

Smaller observations: resident memory was about 127 MiB for the Temporal dev server after start (177 MiB at the end), 294 MiB for single-node Restate with defaults and 72 MiB for the DBOS application process. DBOS 3.2.0 created its system database and applied its schema migrations at first start; the database was under 9 MB after six workflows.

Honest limits: one VM, loopback only, the Temporal development server rather than a production Temporal Service, single-node Restate, one PostgreSQL instance, no Conductor, no clustering, no load and one run per scenario. Restart timings include Python start-up and retry backoff, so they show behaviour, not performance. I did not test Temporal Worker Versioning, Restate’s Kubernetes operator, DBOS queues or schedules, or any SDK other than Python. The only throughput numbers in this post are DBOS’s, labelled as the vendor’s.

What to unlearn and re-learn

  • Unlearn “retries are a loop with a counter”. Re-learn that retries, timers and compensation belong in a recorded workflow, with the counter kept by the engine.
  • Unlearn “durable execution means exactly once”. Re-learn that steps run at least once and the external side still needs idempotency keys.
  • Unlearn “workflow engines need a big server”. Re-learn that the choice now runs from a full service (Temporal) through a single binary (Restate) to a library on PostgreSQL (DBOS).
  • Unlearn “a deploy is just new code”. Re-learn that running workflows hold journals written by old code, so each deploy needs a versioning decision.
  • Unlearn “the defaults are sensible for us”. Re-learn each engine’s retry, retention and security defaults, because they differ in ways that matter for payments.

Revisit your retries before you write another cron job

The Ahmedabad team does not need the most popular engine; it needs one place where an order’s progress is recorded, and a design that survives restarts and redeploys. Learn the durable execution model: journal, replay, deterministic workflow code and steps that run at least once. Unlearn the habit of adding a new cron job for every state an order can get stuck in, and the belief that a retry counter in a JSON column is a workflow. Re-learn the current Temporal, Restate and DBOS releases, their defaults and their licences from the release notes, because all three changed a lot in 2026. Practise by killing a worker, a server and a database in the middle of a workflow, and by deploying a deliberately breaking change, as I did on one VM. Then apply what you learn to one real flow first, with idempotent steps, explicit retry limits, recorded compensations and a versioning plan. Revisit it before your next big sale, and whenever one of these engines ships a new minor release.

Sources

comments powered by Disqus

Releted Posts

PostgreSQL high availability: revisit Patroni and CloudNativePG before your next failover

Picture a lending company in Bengaluru (an invented example, not a real client). Its loan ledger runs on PostgreSQL 16 across three VMs, managed by Patroni with a three-node etcd cluster and HAProxy in front.

Read more

Fluent Bit, Vector and the OpenTelemetry Collector: revisit your log pipeline before you ship more logs

A logistics company in Pune runs about 600 pods across three Kubernetes clusters. A Fluentd DaemonSet, set up in 2019 with a dozen Ruby plugins, ships about 25 crore log lines a day to Loki.

Read more

RabbitMQ, NATS and Kafka queues: revisit your message queue before you add another broker

A health insurance company in Hyderabad processes claims through RabbitMQ: three nodes on 3.13 with classic mirrored queues, about 40 lakh messages a day, from claim documents to OCR, OCR results to fraud checks and approvals to payments.

Read more