Redis, Valkey, and Dragonfly operations: persistence and failover

A cache becomes a database the day nobody can rebuild it. After that day, persistence and failover are not settings copied from a blog. They are promises made to the people on call.

Part 1 of this series, Redis, Valkey, or Dragonfly: revisit the choice before you treat them as the same cache, covered licences, data file compatibility, and cluster basics. Part 2, Redis, Valkey, and Dragonfly benchmarks: read the method first, covered vendor speed claims. This part is about the boring hours: disk writes, replica catch-up, failover, and memory. Every fact is from the projects’ own docs, configs, release notes, or repositories, checked on 5 October 2026.

The short version

  • Redis and Valkey share one persistence model: RDB snapshots from a forked child, plus an optional append only file (AOF). Since Redis 7.0.0 the AOF is a set of files tracked by a manifest.
  • Dragonfly takes forkless, point in time snapshots in its own format by default, and can write a Redis-compatible RDB. Its docs say it does not support AOF.
  • All three replicate with REPLICAOF. Partial resync depends on a backlog you must size. Valkey’s dual-channel full sync is real, and off by default.
  • Redis and Valkey fail over through Sentinel or cluster mode. Dragonfly’s documented HA path is its Kubernetes operator, and its multi-shard cluster has no built-in failover.
  • Fork-based snapshots need memory headroom for copy-on-write. Measure it, do not guess it.

Persistence: what reaches the disk

RDB snapshots in Redis and Valkey

The server forks, the child writes a temporary RDB file, and the finished file replaces the old one. Both sample configs default to save 3600 1 300 100 60 10000: after an hour if at least 1 change, five minutes if 100, one minute if 10000. save "" turns snapshots off.

RDB suits backups and fast restarts, but a bad stop loses the minutes since the last snapshot. Forking a big dataset can pause clients for milliseconds, or up to a second on a very large dataset with a slow CPU.

Keep stop-writes-on-bgsave-error yes, the default, unless your monitoring would catch a failing disk faster. It refuses writes until a background save succeeds again.

AOF and its fsync policies

AOF logs every write and replays it on restart. appendfsync decides how often data is forced to disk:

  • always: fsync after each batch of writes, before replies. The docs call it “very very slow, very safe.”
  • everysec: the default. A background thread fsyncs once a second, so you can lose about one second of writes.
  • no: the kernel decides. The docs say Linux normally flushes every 30 seconds.

Since Redis 7.0.0 the AOF is multi-part: one base file and one or more incremental files in appenddirname, tracked by a manifest. During a rewrite the parent starts a new incremental file while the child writes a new base, and the manifest is swapped atomically. Before 7.0, writes during a rewrite were buffered in memory and written twice. The Valkey docs describe the same multi-part design.

Two restore facts. When both RDB and AOF are on, the server loads the AOF, because it is the most complete. And a truncated AOF tail loads anyway by default (aof-load-truncated yes), while corruption in the middle stops startup until redis-check-aof or valkey-check-aof repairs it, discarding everything after the bad offset.

Dragonfly: forkless snapshots, no AOF

Dragonfly’s snapshot design page describes a forkless, point in time snapshot. Each shard thread serialises its own data. If a write touches an entry not yet serialised, the old value is pushed into the snapshot first, then the change is applied.

By default --df_snapshot_format is true, so Dragonfly writes one .dfs file per shard plus a summary.dfs file. SAVE RDB writes a Redis-compatible RDB instead. It can load a Redis dump.rdb at startup from --dir, and since v2.0.0 (16 September 2026) also Valkey RDB files with the VALKEY080 header.

The flags to know:

  • --dir: where snapshots live. It accepts s3://bucket/prefix, which the docs mark as preview.
  • --dbfilename: name without extension, default dump-{timestamp}. Empty disables saving and loading.
  • --snapshot_cron: cron-style schedule, from Dragonfly 1.7.1. --save_schedule is deprecated.
  • A snapshot is also written on shutdown when dbfilename is set.

On AOF the docs are plain: “Currently, Dragonfly does not support AOF.” So after a crash you lose everything since the last snapshot, unless a replica has it. Also note that BGSAVE is documented as “equivalent to SAVE”, replies with a plain OK rather than the Redis text, and rejects concurrent saves.

Replication: how a replica catches up

PSYNC and the backlog

A replica tracks a replication ID and offset, and sends both with PSYNC on reconnect. If the primary’s backlog still holds the missing bytes, only the gap is sent. Otherwise the replica gets a full resync: a fresh RDB plus writes buffered during the transfer.

repl-backlog-size is 1mb in the Redis 8.10.2 sample config and 10mb in the Valkey 9.1.2 sample. A working rule, not from the docs: size it as replication bytes per second multiplied by the longest disconnect you want to survive. Then check sync_partial_ok and sync_full in INFO stats.

With repl-diskless-sync yes, set in both sample configs, the child streams the RDB straight to replica sockets. repl-diskless-sync-delay 5 waits five seconds so several replicas can share one transfer.

Full sync on two connections

Valkey 8.0 added dual-channel replication: the RDB and the stream of new writes travel on separate connections, and the buffer sits on the replica, not the primary. Valkey’s own announcement claims less memory pressure on the primary and sync time “cut by up to 50%” for heavy read workloads. That is the project’s own result. Enable it with dual-channel-replication-enabled yes on the primary and every replica, with repl-diskless-sync on at the primary. The replica then needs memory for that buffer.

Redis 8 has a separate design. Its config says the primary “may decide” to send the RDB and the stream in parallel, with replica-full-sync-buffer-limit capping the replica side. There is no switch for it.

Dragonfly replication

Dragonfly keeps REPLICAOF and ROLE, but its docs call the internals “vastly different.” Each shard has its own backlog. From v2.0.0 it is bounded by --shard_repl_backlog_time_ms (default 5000) and --shard_repl_backlog_max_bytes (default 0, meaning maxmemory divided by shard count divided by 200). If replicas full-sync after short drops, the docs say raise these. Watch the lag field in INFO REPLICATION.

The empty primary trap

The Redis replication docs describe this one. A primary with persistence off restarts automatically, comes back empty, and its replicas sync to the empty dataset. Keep persistence on, or do not auto-restart such a primary. Dragonfly’s --break_replication_on_master_restart (default false) breaks replication when the primary restarts, so the replica keeps its data.

Failover: who promotes the replica

Sentinel

Sentinel is the HA layer for non-cluster Redis and Valkey. The docs set the rules:

  • Run at least three Sentinels on machines that fail independently.
  • The quorum only marks a primary as down. A majority must still elect a leader to fail over.
  • Replication is asynchronous, so acknowledged writes can be lost. min-replicas-to-write and min-replicas-max-lag make an isolated primary stop taking writes, which narrows the window.
  • Clients must support Sentinel, and Docker port mapping breaks its auto discovery unless handled.

SENTINEL FAILOVER <name> forces a failover, which is what a drill needs.

Cluster mode

In Redis Cluster and Valkey Cluster, a primary unreachable for longer than cluster-node-timeout is treated as failing and a replica takes over. For planned work, run CLUSTER FAILOVER on a replica: clients switch only after it has processed the old primary’s full stream.

Valkey 8.0 added automatic failover for empty shards, synchronous replication of CLUSTER SETSLOT state to replicas, and slot migration state recovery after failover. Valkey 9.2.0-rc1 (16 September 2026), a release candidate, adds cluster-replica-priority and an if-empty value for cluster-replica-no-failover, so a replica that never received data refuses promotion.

Dragonfly

The Dragonfly high availability page covers the Dragonfly Operator. With replicas above 1 it runs one master, reconfigures replication when pods die, and points the service at the current master.

For Sentinel, there is a --replica_priority flag “for sentinel,” the migration docs use Redis Sentinel to promote a Dragonfly replica, and the repository has Sentinel failover tests. The docs give no guide for Sentinel over a Dragonfly-only fleet, so test it yourself.

In multi-shard cluster mode, the docs say Dragonfly provides only the data plane. Health checks, automatic failover, and rebalancing come from its cloud service, Dragonfly Swarm.

Memory headroom during snapshots

A forked child shares pages with the parent, and a page is copied only when one side changes it. So a BGSAVE or AOF rewrite needs extra memory in proportion to what you write during the save. Linux cannot predict that. The Redis FAQ example: with vm.overcommit_memory at 0, a 3 GB dataset with 2 GB free fails to fork. Both admin guides recommend vm.overcommit_memory = 1 and disabling Transparent Huge Pages. Neither gives a fixed headroom figure.

INFO reports current_cow_peak, rdb_last_cow_size, and aof_last_cow_size. Save under real write load, read the peak, and keep that much free with margin.

Dragonfly does not fork, so nothing is copied page by page. Its extra memory goes to serialisation buffers and old values pushed out during the snapshot. The docs give no sizing rule. Its README claims much lower peak memory than Redis during bgsave, a vendor test discussed in part 2. Measure RSS during SAVE on your data.

Valkey 9.2.0-rc1 adds opt-in forkless saves through forkless-infrastructure-enabled and bgsave-default-method forkless, at 4 additional bytes per key. It is a release candidate: staging only.

Backups and restore

Copying a Redis or Valkey RDB while the server runs is safe, because it is renamed into place atomically. Ship a copy off the machine daily. For AOF, the Redis docs say: set auto-aof-rewrite-percentage 0, wait for aof_rewrite_in_progress to be 0, copy the appenddirname directory, then restore the old value.

Redis 8.10 (GA 29 July 2026) added BACKUP START, LIST, SEAL, and CLEANUP. They produce a base RDB, an incremental AOF, and a manifest without pausing writes. Restore uses the startup-only preload-file, such as preload-file aof:/path/appendonly.aof.manifest. The Valkey docs I checked show no equivalent.

For Dragonfly, keep all .dfs files of one timestamp together, summary file included. The operator README lists automatic snapshots to PVCs and S3.

Side by side

Area

Redis

Valkey

Dragonfly

Snapshot

Forked RDB

Forked RDB; forkless opt-in in 9.2.0-rc1

Forkless .dfs by default; RDB on request

AOF

Multi-part since 7.0.0

Multi-part

Not supported

Partial resync

PSYNC and backlog

PSYNC and backlog

Per-shard backlog

Two-channel full sync

Redis 8, primary decides

Opt-in

Not described in docs

Failover

Sentinel, Redis Cluster

Sentinel, Valkey Cluster

Operator; none in multi-shard mode

Online backup

BACKUP since 8.10

Not documented

SAVE, --snapshot_cron

A starting configuration

These lines come from the sample configs. They are a start, not tuned values.

# redis.conf or valkey.conf
save 3600 1 300 100 60 10000
dbfilename dump.rdb
appendonly yes
appendfsync everysec
appenddirname "appendonlydir"
stop-writes-on-bgsave-error yes
repl-diskless-sync yes
min-replicas-to-write 1
min-replicas-max-lag 10
# Valkey only, on the primary and all replicas:
# dual-channel-replication-enabled yes

# Dragonfly equivalent, as command line flags:
# dragonfly --dir /var/lib/dragonfly --dbfilename "dump-{timestamp}" \
#   --snapshot_cron "0 */2 * * *" --break_replication_on_master_restart=true

Failure drills you can run

Use staging with production-sized data. Record time taken and data lost.

  1. Hard kill the primary with kill -9 under write load. Compare lost writes with your appendfsync policy or last Dragonfly snapshot.
  2. Restore yesterday’s backup to a fresh node. Time it, compare DBSIZE, and check whether AOF overrode your RDB.
  3. Save under load. Run BGSAVE at peak-like writes. Record current_cow_peak or RSS, and p99 latency.
  4. Cut a replica off for 10, 60, and 300 seconds. Note partial or full resync, and resize the backlog.
  5. Force a failover with SENTINEL FAILOVER, CLUSTER FAILOVER, or by deleting the master pod under the operator. Measure how long clients see errors.
  6. Restart an empty primary with persistence off, and watch the replicas. Repeat with the protections above.
  7. Fill the disk. Confirm writes stop as expected and alerts fire.

Alert on rdb_last_bgsave_status, aof_last_write_status, master_link_status, and sync_full for Redis and Valkey, and on replication lag for Dragonfly.

What to practise this week

  1. For each datastore, write two numbers: how much data you can lose, and how long you can be down. Check your settings match.
  2. Run drill 2. A backup never restored is not yet a backup.
  3. Run drill 3 and note the measured headroom next to the instance size.
  4. Confirm your client library reconnects after drill 5 without an application restart.

Next in this series

The next part looks at memory under load and upgrades: eviction, fragmentation, and moving between versions without a long outage.

Sources

comments powered by Disqus

Releted Posts

Redis, Valkey, and Dragonfly benchmarks: read the method first

A benchmark number without its method sounds complete, but it does not tell you what happened. The first note in this series, Redis, Valkey, or Dragonfly: revisit the choice before you treat them as the same cache, covered licences, history, data file compatibility, and cluster mode.

Read more

Redis, Valkey, or Dragonfly: revisit the choice before you treat them as the same cache

If your service already uses Redis, the name on the port has not changed. The product behind that name has. Treating Redis, Valkey, and Dragonfly as three labels for one cache is the habit worth unlearning.

Read more

How Redis Helps to Increase Service Performance in NodeJS

In modern backend applications, performance optimization is crucial for handling high traffic efficiently. Typically, a backend application consists of business logic and a database.

Read more