Redis, Valkey, and Dragonfly memory and upgrades: revisit before the next spike

A payments team runs its session cache with maxmemory 12gb on a 16 GB node and feels safe. On a festival sale evening a replica reconnects, a reporting job pulls a few huge hashes, and the kernel kills the process while used_memory still shows less than 12 GB. Two weeks later the team upgrades, turns on a new memory saving feature, hits a client bug and tries to roll back. The old binary refuses the new dump file. Both incidents come from one gap: knowing the settings, but not what they count and promise.

Earlier parts covered the choice and compatibility, benchmark claims, and persistence, failover and fork headroom. This part covers memory under load (maxmemory, eviction, fragmentation, defrag, big keys) and upgrades (rolling upgrades, RDB versions, rollback). Facts come from the projects’ docs, sample configs, release notes and source code at Redis 8.10.2, Valkey 9.1.2, Dragonfly v2.0.0 and Dragonfly Operator v1.7.0, checked on 6 October 2026. Drill numbers are from my own run.

The short version

  • maxmemory is compared with allocator memory minus replica and AOF buffers, not with RSS. Fragmentation, fork copy-on-write and the Lua VM can take you past the host limit.
  • Redis and Valkey default to noeviction. At the limit, writes fail with an OOM error while reads and deletes still work.
  • LRU and LFU are sampled approximations. Redis 8.6 added allkeys-lrm and volatile-lrm (least recently modified). Valkey does not have them.
  • mem_fragmentation_ratio mixes many overheads. Read allocator_frag_ratio before turning on active defrag.
  • Dragonfly does not evict unless started with --cache_mode=true. In my small drill, cache mode still returned Out of memory under fast pipelined writes.
  • Upgrade replicas first, then fail over. Rollback is the hard part: older servers refuse newer RDB versions (Redis 8.10 writes 15, Valkey 9 writes 80).

Versions in play on 6 October 2026

  • Redis: 8.10.2 (17 September 2026) is the newest stable release, with same-day patches 8.8.3, 8.6.7, 8.4.7 and 8.2.10. 8.10 went GA on 29 July 2026.
  • Valkey: 9.1.2 is current on the download page (dated 1 September 2026), with 8.1.10 and 7.2.14 for older lines. 9.2.0-rc1 (16 September 2026) is a release candidate.
  • Dragonfly: v2.0.0 (16 September 2026), and Dragonfly Operator v1.7.0 (28 September 2026).

How maxmemory really works

What is counted, and what is not

The Valkey docs state the rule: eviction starts when used_memory - mem_not_counted_for_evict > maxmemory. used_memory is what the server asked its allocator for. The excluded part is replica output buffers beyond the backlog size and the AOF buffer. Both projects exclude them on purpose: each evicted key creates a DEL for replicas and the AOF, and counting that buffer could start a loop that empties the dataset. Reading evict.c in both repositories, the backlog itself is counted up to its configured size.

  • Normal client buffers are counted. client-output-buffer-limit normal 0 0 0 (no limit) and maxmemory-clients 0 (off) are the defaults, so a slow reader of a huge reply can push keys out. maxmemory-clients 5% drops the most memory hungry clients first.
  • The Lua VM is outside. The INFO docs mark used_memory_vm_eval as “not part of used_memory”.
  • RSS is never compared. Fragmentation and fork copy-on-write (part 3 covered current_cow_peak) sit between maxmemory and what the kernel sees.
  • Replicas ignore maxmemory by default (replica-ignore-maxmemory yes) and delete keys only when the primary sends DELs. The sample config warns they can use more memory than the primary.
  • One command can overshoot, for example a big set intersection stored into a new key.

The docs suggest leaving free RAM when replicas are attached, without a fixed figure. Read mem_not_counted_for_evict under real load and size the rest from measurements.

noeviction and the OOM error

noeviction is the default. At the limit, commands that need memory fail: SET, INCR, HSET, LPUSH, SUNIONSTORE, SORT with STORE, and EXEC containing any of them. The exact reply is OOM command not allowed when used memory > 'maxmemory'. Reads and DEL keep working. The volatile-* policies behave like noeviction when no key has a TTL.

For alerts, INFO commandstats shows rejected_calls per command and INFO errorstats shows errorstat_OOM. Alert on these and on evicted_keys, not only on a memory percentage.

The policies

Redis and Valkey share eight: noeviction, allkeys-lru, allkeys-lfu, allkeys-random, volatile-lru, volatile-lfu, volatile-random and volatile-ttl. Redis 8.6.0 (10 February 2026) added allkeys-lrm and volatile-lrm, which update a key’s timestamp only on writes. The Redis docs suggest LRM for read heavy workloads where you want to evict data that is no longer being updated, however often it is read. Valkey 9.1.2 rejects allkeys-lrm at startup; I tried.

The docs call allkeys-lru a good default when unsure, and note that a TTL costs memory.

How LRU and LFU choose a victim

There is no global ordered list. The server samples maxmemory-samples keys (default 5, maximum 64) and keeps a pool of 16 candidates (EVPOOL_SIZE in the source). The docs say 10 samples come very close to true LRU, at more CPU cost.

Two details from the source matter:

  • LRU uses a 24-bit clock in seconds. Keys touched in the same second look equally recent. In my drill, writing at full speed while reading 1,000 hot keys after every batch, only 388 hot keys survived allkeys-lru. With a 0.2 second pause per batch, all 1,000 survived. My reading, not proven, is that the fast run fitted into a few seconds, so hot and new keys shared clock values.
  • LFU splits the 24 bits into 16 bits of minutes and an 8-bit logarithmic counter. New keys start at 5 (LFU_INIT_VAL). Defaults are lfu-log-factor 10 and lfu-decay-time 1 minute. In the same fast run, allkeys-lfu kept all 1,000 hot keys.

maxmemory-eviction-tenacity (default 10, range 0 to 100) trades eviction effort against latency. Freeing big values can block the main thread, so check lazy free: lazyfree-lazy-eviction is no in Redis 8.10.2 and yes in Valkey (sample config since 8.0, built-in default in 9.1.2).

Memory the allocator keeps: RSS, fragmentation and defrag

Read the right ratio

RSS rarely falls after deletes, because freed chunks share pages with live keys. The Redis memory optimisation page gives an example: fill 5 GB, delete 2 GB, and RSS probably stays near 5 GB. Provision for the peak. The INFO docs separate the numbers:

  • mem_fragmentation_ratio is RSS divided by used_memory, including code, libraries and allocator overheads. A high ratio over only a few megabytes is not an issue.
  • allocator_frag_ratio is what the docs call “the true (external) fragmentation metric”.
  • If used_memory is far above RSS, part of the process is swapped out. Expect latency.

Active defrag

Active defrag moves values into compact regions while the server runs. It needs the bundled jemalloc, the default on Linux (both binaries I used report jemalloc-5.3.0), and it is off by default. The sample configs say you never need it without a fragmentation problem. Defaults:

  • active-defrag-ignore-bytes 100mb and active-defrag-threshold-lower 10 (percent) decide when it starts. active-defrag-threshold-upper 100 is where maximum effort applies.
  • active-defrag-cycle-min 1 and active-defrag-cycle-max 25 are CPU percentages, not times.
  • Valkey 8.1 added active-defrag-cycle-us 500, a time slice. Its announcement says defrag was “improved to eliminate latencies greater than 1ms”, the project’s own statement.

Turn it on at runtime with CONFIG SET activedefrag yes and watch active_defrag_running and active_defrag_hits.

MEMORY commands and big keys

  • MEMORY STATS splits memory into overhead, dataset, backlog, client buffers and allocator numbers.
  • MEMORY DOCTOR gives plain advice. In my drill it flagged a past peak and allocator fragmentation, and suggested activedefrag.
  • MEMORY USAGE key samples 5 nested elements by default. SAMPLES 0 reads all of them, which is slow on a big key.
  • MEMORY PURGE asks jemalloc to release dirty pages.

One 2 GB hash hurts memory, replication and latency together. Find such keys with redis-cli --bigkeys, --memkeys or --keystats, or valkey-cli --bigkeys and --memkeys, preferably on a replica. Redis 8.6 also added a HOTKEYS command and key-memory-histograms, off by default and enabled only at startup. Delete big keys with UNLINK.

Newer memory features, briefly

  • Redis 8.10 compact hashes store shared field names once in a template. Off by default (hash-min-template-entries 0), and the config calls conversion “one-way” for the life of the key.
  • Valkey 8.1 rewrote its hash table and reports “roughly 20 byte reduction per key-value pair” without a TTL, the project’s own measurement. Valkey 9.1 raised the embedded string limit from 64 to 128 bytes. Part 2 discussed the 9.1 memory figures.
  • Valkey 9.2.0-rc1 adds maxmemory-scripts to cap cached EVAL scripts. It is a release candidate.

Dragonfly’s memory model

Dragonfly is a different engine, so do not copy Redis memory settings across.

  • maxmemory. --maxmemory defaults to 0. In the v2.0.0 source, 0 means 80% of available memory, and container cgroup limits are read. Each thread needs at least 256 MB, or the server exits at startup.
  • Allocator and default. Dragonfly uses mimalloc, so activedefrag does not apply. Without cache mode, INFO shows maxmemory_policy:noeviction and writes at the limit fail with Out of memory.
  • Cache mode. --cache_mode=true evicts near the limit, and INFO shows maxmemory_policy:eviction. In the v2.0.0 source maxmemory can change at runtime but cache_mode cannot. There is no policy choice. Dragonfly’s 2022 design post describes a 2Q style policy built into its Dashtable: new items sit in small probation buckets and move up when accessed again, and “No additional metadata is needed” per item. STICK protects chosen keys from eviction.
  • Background and RSS eviction are on by default (--enable_heartbeat_eviction, --enable_heartbeat_rss_eviction), with --max_eviction_per_heartbeat at 100.
  • RSS guard. --rss_oom_deny_ratio defaults to 1.25. In the source, when RSS exceeds 1.25 times maxmemory, commands marked DENYOOM fail and new connections on the main port are refused.
  • Defrag runs on thresholds, with no switch: --mem_defrag_threshold 0.7 of maxmemory, --mem_defrag_waste_threshold 0.2, checked every 60 seconds. v2.0.0 added burst and cooldown limits and MEMORY DEFRAGMENT-SEGMENTS. The MEMORY DEFRAGMENT help warns it “may temporarily monopolize” shard threads. There is no MEMORY DOCTOR.

Side by side

Area

Redis 8.10

Valkey 9.1

Dragonfly v2.0

Default at limit

noeviction

noeviction

No eviction unless --cache_mode

Eviction choices

10 policies, LRM since 8.6

8 policies

One built-in cache policy

Limit is checked against

used_memory minus buffers

used_memory minus buffers

Used memory, plus RSS checks

Allocator

jemalloc

jemalloc

mimalloc

Defrag

activedefrag, off

activedefrag, off

Threshold based, runs by itself

Lazy free on eviction

Off by default

On by default

Not documented

RDB version written

15

80

9 (.dfs is the default format)

A starting configuration

A start for a cache with replicas, not tuned values.

# redis.conf or valkey.conf
maxmemory 10gb                 # below the container limit, leave headroom
maxmemory-policy allkeys-lfu   # or allkeys-lru; allkeys-lrm on Redis 8.6+ only
maxmemory-samples 5
maxmemory-clients 5%           # drop the worst clients before keys go
lazyfree-lazy-eviction yes     # already the default in Valkey
# only after allocator_frag_ratio shows a real problem:
# activedefrag yes

# Dragonfly, as flags:
# dragonfly --maxmemory=10gb --cache_mode=true --proactor_threads=8

Upgrades without a long outage

Replica first, then switch

The Redis and Valkey admin docs give the same no-downtime path for one primary: start the new version as a replica (another host, or RAM for two instances), wait for the initial sync, compare key counts, move clients (optionally with CLIENT PAUSE on the old primary), promote with REPLICAOF NO ONE, then retire the old node. With Sentinel or cluster mode, upgrade replicas one by one and fail over manually. In a cluster, run CLUSTER FAILOVER on an upgraded replica, upgrade the old primary as a replica, and repeat per shard. Each step costs a full sync, so the headroom from part 3 applies.

Why rollback is hard: RDB versions

From rdb.h at each release tag:

Release

RDB version written

Redis 7.2

11

Redis 7.4 to 8.4

12

Redis 8.6

13

Redis 8.8

14

Redis 8.10

15

Valkey 7.2 and 8.x

11

Valkey 9.x

80, with a VALKEY header

Redis refuses any RDB newer than its own, and the same loader runs when a replica receives a full sync. Once the primary runs 8.10, an 8.8 replica cannot full sync from it and an 8.8 server cannot load its dump. Plan rollback as “restore the backup taken before the upgrade”, not “reinstall the old package”.

Valkey softens this. A Valkey 9 primary sends RDB 11 to a replica that reports an older version over diskless sync (replication.c), and since 9.0.0-rc2 rdb-version-check relaxed loads newer or foreign files “on a best-effort basis”. New features close that door: hash field expiration in Valkey 9.0 and compact hashes in Redis 8.10 write data that older versions cannot read. Keep them off until the rollback window closes.

On the Redis to Valkey boundary, covered in part 1: Valkey reads Redis files up to RDB 11 by default, so dumps from Redis 7.4 or later need the relaxed check or a logical migration. Dragonfly v2.0.0 loads RDB 5 to 12 and Valkey’s 80, so a Redis 8.6 or later dump will not load there either.

Dragonfly upgrades

Dragonfly’s switch command is REPLTAKEOVER <seconds>, run on a replica. In the v2.0.0 source it refuses until full sync is done and returns OK on a node that is already master. Dragonfly’s operator announcement says it locks the old master so ongoing operations complete first. The docs command reference does not list it, so test it before relying on it.

The operator v1.7.0 source shows the rollout. The StatefulSet uses OnDelete. When the pod spec changes, the operator deletes replicas one at a time and waits for each to be ready on the new version, runs REPLTAKEOVER 10000 on an updated replica, relabels it master and deletes the old master. With no replicas it just deletes the master: downtime. The docs I checked do not say whether an older Dragonfly can load a newer .dfs snapshot.

Client and config checks

  • Start the new binary with your old config in staging. Redis and Valkey stop on an unknown directive with Bad directive or wrong number of arguments. Dragonfly exits on an unknown flag.
  • Read changed defaults and deprecations, such as lazy free and Dragonfly’s --save_schedule (use --snapshot_cron).
  • Check version gates in code. Valkey 9.1.2 reports redis_version:7.2.4 in INFO server; the real number is in valkey_version.

Upgrade checklist

  1. Write down current and target versions, and the RDB version of each.
  2. Take a backup on the old version and prove it restores on the old version.
  3. Diff your config against the new sample config and start it in staging.
  4. Check memory headroom for one extra full sync per shard.
  5. Upgrade replicas one at a time. Watch master_link_status and sync_full.
  6. Fail over with CLUSTER FAILOVER, Sentinel or the operator. Measure client errors.
  7. Keep new data features (compact hashes, hash field TTLs) off until rollback is no longer needed.
  8. Afterwards, compare evicted_keys, allocator_frag_ratio, RSS and p99 latency with last week.

Common mistakes

  • Setting maxmemory equal to the container limit. Buffers, fork, fragmentation and Lua live above it.
  • Reading mem_fragmentation_ratio alone. After a mass delete it can be peak memory, not a leak.
  • Using volatile-lru without TTLs. It behaves like noeviction.
  • Copying allkeys-lrm into a Valkey config, or expecting Dragonfly to evict without --cache_mode=true.
  • Testing rollback by reinstalling the old package. It cannot read the new dump.
  • Running --bigkeys or MEMORY USAGE ... SAMPLES 0 on a busy primary.

Drill: eviction, fragmentation and an RDB boundary

Use a throwaway machine with persistence off (--save "" --appendonly no), and a small client script that writes in pipelines.

  1. Eviction. Set --maxmemory 64mb. Write 200,000 keys of 1,000 bytes, reading a fixed set of 1,000 “hot” keys after each batch. Count surviving hot and cold keys under allkeys-lru, allkeys-lfu and noeviction, at full speed and with a short pause per batch. On Redis 8.6 or later, add allkeys-lrm with keys that are only read, keys rewritten every batch and keys never touched.
  2. Fragmentation. Without maxmemory, load 400,000 keys of mixed sizes, UNLINK three of every four, read INFO memory and MEMORY DOCTOR, then enable activedefrag with lower thresholds and read again after 30 seconds.
  3. RDB boundary. Save a dump on each version and start the other version on it.
  4. Dragonfly. Repeat step 1 with and without --cache_mode=true.

My run, and its limits

I ran this on 6 October 2026 on a shared VM container: 8 vCPUs, 15 GB RAM (about 4.8 GB free), Debian 13, Linux 6.12, other workloads on the machine. Binaries: the official Valkey 9.1.2 Ubuntu Noble x86_64 build (checksum verified), Redis 8.10.2 built from its release tag, and the Dragonfly v2.0.0 release binary on one thread. The client was redis-py 6.1.0 on the same host. One run each: observations, not benchmarks.

Test (64 MiB limit, 200,000 writes of 1,000 bytes)

Result

Valkey allkeys-lru, full speed

0 errors, 138,860 evicted, 388 of 1,000 hot keys kept

Valkey allkeys-lru, 0.2 s pause per batch

1,000 of 1,000 hot keys kept, 0 of 1,000 cold

Valkey allkeys-lfu, full speed

1,000 of 1,000 hot keys kept, 46 of 1,000 cold

Valkey noeviction

138,453 writes rejected, 61,547 keys stored, GET still worked

Redis 8.10.2 allkeys-lru, paused

Read-only keys 1,000 kept, rewritten 1,000, untouched 0

Redis 8.10.2 allkeys-lrm, paused

Read-only keys 0 kept, rewritten 1,000, untouched 0

At the limit Valkey’s used_memory was 63.7 MiB while RSS was 79.1 MiB. The limit does not cap RSS.

Fragmentation on Valkey 9.1.2: after loading, used_memory 181.7 MiB and RSS 187.6 MiB. After deleting three quarters of the keys, used_memory fell to 49.1 MiB but RSS stood at 196.6 MiB, with mem_fragmentation_ratio 4.01 and allocator_frag_ratio 3.63. Thirty seconds after activedefrag yes, RSS was 65.0 MiB and allocator_frag_ratio 1.03, with 123,625 defrag hits.

RDB boundary: the Redis 8.10.2 dump began with REDIS0015, the Valkey dump with VALKEY080. Valkey refused the Redis file (Can't handle RDB format version 15) and Redis refused the Valkey file (Wrong signature trying to load DB from file). With rdb-version-check relaxed, Valkey loaded my one-key Redis file; a real dataset may not load.

Dragonfly v2.0.0, --maxmemory=256mb, 600,000 writes of 1,000 bytes: without cache mode, 318,000 writes failed with Out of memory and nothing was evicted. With cache mode at full speed, it evicted 51,500 keys and 267,554 writes still failed. With a 0.05 second pause per batch, it evicted 217,005 and 102,540 failed. A 1 GiB run with 1.5 million writes showed the same pattern. I did not tune the eviction flags, so I cannot say why. The narrow lesson: test your own write bursts at the limit before trusting cache mode.

What to unlearn and re-learn

  • Unlearn “maxmemory is the memory limit”. Re-learn it as a limit on counted allocator memory, and size the host for RSS.
  • Unlearn “LRU keeps my hot keys”. Re-learn that it is sampled with one second resolution, and that LFU or LRM may fit better.
  • Unlearn “a high fragmentation ratio means a leak”. Re-learn the allocator ratios and peak memory.
  • Unlearn “rollback is the upgrade in reverse”. Re-learn that RDB versions move one way.

Revisit your memory and upgrade plan

Memory trouble rarely starts at the limit you set. It starts in what that limit does not count, in a policy that has no keys to evict, and in a release that writes files the old one cannot read. Learn what maxmemory compares, unlearn the comfort of one ratio, re-learn eviction as sampling, practise the drill on your own data sizes, and apply the upgrade checklist before the next sale day, not during it.

Next in this series

Part 5 looks at clients: connection pools, timeouts and retries, RESP3, cluster redirects and client-side caching across Redis, Valkey and Dragonfly.

Sources

comments powered by Disqus

Releted Posts

Redis, Valkey, and Dragonfly benchmarks: read the method first

A benchmark number without its method sounds complete, but it does not tell you what happened. The first note in this series, Redis, Valkey, or Dragonfly: revisit the choice before you treat them as the same cache, covered licences, history, data file compatibility, and cluster mode.

Read more

Redis, Valkey, and Dragonfly operations: persistence and failover

A cache becomes a database the day nobody can rebuild it. After that day, persistence and failover are not settings copied from a blog.

Read more

Redis, Valkey, or Dragonfly: revisit the choice before you treat them as the same cache

If your service already uses Redis, the name on the port has not changed. The product behind that name has. Treating Redis, Valkey, and Dragonfly as three labels for one cache is the habit worth unlearning.

Read more