Redis, Valkey, and Dragonfly benchmarks: read the method first

A benchmark number without its method sounds complete, but it does not tell you what happened.

The first note in this series, Redis, Valkey, or Dragonfly: revisit the choice before you treat them as the same cache, covered licences, history, data file compatibility, and cluster mode. This note takes the published performance claims of all three projects and writes the machine, the version, and the client settings next to each figure. Where a page is silent, this note says so. Revisit.Tech ran no tests for this post. Every figure below is the vendor’s own run, from the vendor’s own site or repository.

The short version

  • Dragonfly’s “25X” line compares one Dragonfly process with one Redis process on a 64 vCPU machine. The README does not state the Redis version or the Dragonfly version used.
  • Redis answered in June 2022 with a 40-shard Redis 7.0 Cluster on the same machine type, and reported Redis ahead. That too was a vendor run, against a competitor.
  • Redis 8.x and Valkey 8.x both now publish I/O thread gains. Most of the Redis 8 throughput figures do not state the CPU model, the client tool, or the connection count.
  • Valkey 8.0’s headline of about 1.2 million requests per second appears with three different instance descriptions across three of its own posts.
  • None of these numbers can be placed side by side. A fair comparison has to be run by you, on one machine type, with your key sizes, and with persistence and replication switched on.

What to look for in any benchmark claim

Before reading the figures, keep this checklist open. A claim is only as good as the answers.

  1. Who ran it, and do they sell one of the products?
  2. Exact server version, and the date of the run.
  3. Server machine: instance type, vCPU count, CPU architecture, kernel.
  4. Server threads: io-threads for Redis and Valkey, proactor_threads for Dragonfly, or the number of shards.
  5. Client tool, client machine, and whether the client was on a separate host.
  6. Threads and connections from the client.
  7. Pipeline depth. A pipeline of 1 means one request at a time per connection.
  8. Value size and key count.
  9. Read and write ratio.
  10. Persistence and replication: on or off.
  11. Which latency percentiles were reported, and at what load.

Most published runs answer about half. That is usually launch writing, not bad faith. It still means the figure is not a fact about your system.

Dragonfly’s own numbers

The README and the 25X line

The Dragonfly README states that Dragonfly delivers “25X more throughput” compared to “legacy in-memory datastores.” The detail sits further down the same file.

  • Small instance: on an AWS m5.large, client on a separate c5n in the same zone, memtier_benchmark -c 20 --test-time 100 -t 4 -d 256. SETs: Redis 159K QPS, Dragonfly 173K. GETs: Redis 194K, Dragonfly 191K.
  • Slightly larger: on m5.xlarge with -t 6. SETs: Redis 190K, Dragonfly 279K. GETs: Redis 220K, Dragonfly 305K.
  • The 25X figure: on c6gn.16xlarge, “compared to a single Redis process,” crossing 3.8M QPS. Client was memtier on a separate c6gn.16xlarge, with -c 30 -n 200000 -d 256 and threads “tuned per server and instance type.”
  • With --pipeline=30, the README says 10M QPS for SET and 15M QPS for GET.
  • Memory: about 5GB loaded with debug populate 5000000 key 1024, then bgsave under update traffic. The README says Dragonfly was 30% more memory efficient at idle, and that Redis peaked at almost 3X Dragonfly’s memory during the snapshot.

Not stated: the Redis version, the Dragonfly version, the date of the run, whether Redis I/O threads were enabled, the exact memtier thread count for the 25X run, and the instance type for the memory test.

Read the m5.large result again. On the small instance the two servers are roughly level. The gap grows as cores are added, because one Redis process executes commands on one main thread. The 25X line measures that design choice on a 64 vCPU box. It is not a per-request speed difference.

The benchmark documentation page

Dragonfly’s benchmark page in its docs is more complete. As checked for this post (the page shows a last update of 10 August 2026), it records:

  • Server: Dragonfly v1.15.0, released on 4 March 2024, run as ./dragonfly --logtostderr --dbfilename=. An empty dbfilename means no snapshot file. Dragonfly used all vCPUs by default.
  • Server machine: c6gn.12xlarge (48 vCPUs), then c7gn.12xlarge (also 48 vCPUs).
  • Client: memtier_benchmark on a c7gn.16xlarge with 64 vCPUs, private IPs, same availability zone. Ubuntu 23.04, kernel 6.2.
  • Load: -t 60 -c 20 -n 200000, so 1,200 connections. Pipelined runs used -c 5 --pipeline=10.
  • Results on c6gn.12xlarge: writes about 4.2M ops/sec (p99 0.687 ms, p99.9 2.543 ms), reads about 4.1M (p99.9 0.903 ms), pipelined reads about 7.08M.
  • Results on c7gn.12xlarge: 5.2M writes, 6M reads, 8.9M pipelined reads.

Not stated: value size and key count. The commands pass no -d, so memtier’s default applies. In the current memtier source that default is 32 bytes. Few production caches store only 32-byte values. Also not stated: number of repeat runs, and any replication.

The same page compares Microsoft’s Garnet on c6in.12xlarge and reports Garnet ahead on pipelined reads, 25.4M against Dragonfly’s 6.9M. A vendor publishing a result it did not win is worth noting.

Dragonfly against Valkey on Google Cloud

On 4 March 2025 Dragonfly published a test against Valkey on GCP.

  • Versions: Dragonfly v1.26.1, Valkey 8.0.2 with --io-threads 10 --save ''.
  • Machines: GCP C4 with 16, 32, and 48 vCPUs, kernel 6.8. The client ran on “a larger machine,” type not named.
  • Client tool: dfly_bench, Dragonfly’s own load generator, not memtier.
  • Workload: SET and GET at 1:1, 64-byte values, 200 million key range, 20 minutes. Client threads and connections were set per server. Dragonfly got 24 to 48 client threads. Valkey got 10 to 12. Parameters were chosen to keep p99 under 0.5 ms.
  • Result: Dragonfly reports 2.4x Valkey’s throughput on 16 vCPUs and 4.5x on 48 vCPUs. On sorted sets with ZADD, it reports 29x on 48 vCPUs, and 12.6 KiB against 23.1 KiB per sorted set entry.

Not stated: the client machine type and the number of runs. Note the timing too. Valkey 8.1, with its new hash table, was announced on 2 April 2025, four weeks after this run.

Redis’s numbers

The 2022 reply

On 28 June 2022 Redis published “13 Years Later, Does Redis Need a New Architecture?” Their argument was that Redis is meant to scale by running many processes, so the fair comparison is a Redis Cluster on the same box.

  • Versions: Redis 7.0.0 built from source. Dragonfly built from source on 3 June 2022 at a named commit.
  • Machines: client and server both c6gn.16xlarge, 64 Arm cores, kernel 5.10, 126GB.
  • Redis setup: 40 primary shards on one VM, leaving 24 vCPUs idle. Two memtier processes in --cluster-mode, -t 24 -c 1, -d 256, 1 million keys, 180 seconds.
  • Dragonfly setup: -t 55 -c 30 -n 200000 -d 256.
  • Results reported: GET at pipeline 1, Redis 4.43M against Dragonfly 3.8M reproduced. GET at pipeline 30, 22.9M against 15.9M. SET at pipeline 1, 4.74M against 4M. SET at pipeline 30, 19.85M against 14M.
  • Headline: Redis 18% to 40% higher throughput.

Not stated: persistence settings, and how a normal team would run and fail over 40 shards on one machine. Dragonfly’s 30 April 2025 open letter says it had not replied to that post until then. The letter offered a managed service price comparison instead, with a 6.5M RPS figure for Dragonfly Swarm measured at --pipeline 10, 32-byte values, reads only. That is a cloud pricing claim, not an engine benchmark.

Redis 8.0 to 8.4

Redis’s own figures for each release, from its blog and release notes:

  • Redis 8.0 GA (1 May 2025): “up to 87%” lower command latency than Redis 7.2.5. The 8.0-M03 post explains the base: 149 tests, 90 improved, p50 reduction from 5.4% to 87.4%, median 16.7%.
  • Redis 8.0 I/O threads: with io-threads set to 8 “on a multi-core Intel CPU,” 37% to 112% more throughput, depending on the command. Not stated: CPU model, core count, client tool, connections, pipeline, value size.
  • Redis 8.0 replication: a 10 GB full sync with 26.84 million writes during the sync. Primary write rate 7.5% higher (471.9K against 438.8K ops/sec), sync 18% faster (101 against 123 seconds), peak primary buffer 35% lower (15.16 against 23.24 GB). Machine not stated.
  • Redis 8.2 GA (blog dated 8 August 2025): with 8 I/O threads, “up to 49%” more throughput than 8.0 at 20% writes and 80% reads, and over 1 million ops/sec on a single instance. Machine, client, and value size not stated. The same post claims 25% to 37% less memory for short string keys.
  • Redis 8.4 (released 18 November 2025): the release notes say over 30% more throughput than 8.2 for 10% SET and 90% GET on 4 cores. The blog adds 1 KB string values. CPU model and client settings not stated.

Each of these compares Redis with an older Redis. Useful for an upgrade decision, silent on Valkey or Dragonfly.

Valkey’s numbers

8.0 and the 1.19 million figure

The two “Unlock 1 Million RPS” posts were written by AWS engineers, and part 1 says the work grew out of their ElastiCache and MemoryDB experience.

  • 8.0-rc1 post (2 August 2024): up to 1.2 million QPS “on AWS’s r7g platform,” against a previous limit of 380K.
  • Part 1 deep dive (5 August 2024): 360K to 1.19M requests per second against Valkey 7.2, average latency 1.792 ms down to 0.542 ms. 8 I/O threads, 3 million keys, 512-byte values, 650 clients, sequential SET, on a C7g.16xlarge.
  • Part 2 (13 September 2024), the reproduction guide: a c7g.4xlarge with 16 cores, --io-threads 9 (main thread plus 8), saving disabled, network IRQs and the main thread pinned to chosen cores. Client from a separate instance: valkey-benchmark -t set -d 512 -r 3000000 -c 650 --threads 50.

So the same headline appears with r7g, C7g.16xlarge, and c7g.4xlarge. Treat the reproduction guide, the most specific of the three, as the method. Not stated: client instance type, and tail latency beyond the average. Note the manual IRQ and core pinning, which a default install will not do.

8.1, 9.0, and 9.1

  • 8.0 memory (4 September 2024): 16-byte values, 6,318,941 keys, cluster mode. used_memory fell from 693.64 MB on 7.2 to 550.56 MB on 8.0, about 20.6%. One of the two changes is the per-slot dictionary, which applies to cluster mode.
  • 8.1 (2 April 2025): the new hash table saves roughly 20 bytes per key without TTL and up to 30 bytes with TTL. About 10% more throughput than 8.0 “for pipeline workloads when I/O threading is not used.” Machine not stated.
  • 9.0 (deep dive dated 20 October 2025, release post 21 October 2025): over 1 billion requests per second across a 2,000-node cluster. Each shard had one primary and one replica, on r7g.2xlarge (8 cores, 64 GB), with io-threads 6 and saving disabled. Load came from 750 c7g.16xlarge client instances running valkey-benchmark SET with 512-byte values. That is roughly 1 million per primary across 1,000 primaries. It is a cluster scaling result, not a single-node one.
  • 9.1 memory (13 August 2026): string key overhead down 17% to 44%, about 26% on average, measured with 5 million items per test, 9.0 against 9.1. This is overhead, not total memory.

Why these numbers do not line up

Even when every figure is honest, they answer different questions.

  • Different machines. m5, c6gn, c7gn, c7g, r7g, and GCP C4 differ in cores, CPU generation, and network.
  • Different shapes. Single process, one multi-threaded process, 40 shards on one box, and 2,000 nodes across many boxes are four different systems.
  • Different thread settings. Each vendor tunes its own server well and the other server less well.
  • Pipelining. Pipeline depth 10 or 30 can multiply throughput. Your application may send one request at a time.
  • Value sizes. 32, 64, 256, 512, and 1,024 bytes all appear above. Memory, network, and copy costs change with size.
  • Client saturation. If the load generator runs out of CPU first, you are measuring the client.
  • Different clients. memtier, valkey-benchmark, and dfly_bench do not behave identically.
  • Persistence and replication off. Almost every run above disabled snapshots or had no replica. Production rarely does.
  • Latency at saturation. Peak ops/sec usually comes with latency you would not accept. A p99 at your real load is more useful.
  • “Up to.” This is the best case among many tests, not the typical one.

A fair test you can run yourself

  1. Write the question first. For example: “At 150K ops/sec with our key sizes, what is p99.9, with AOF on and one replica?” Peak throughput is a secondary number.
  2. Use the same instance type for each server, in one zone, freshly started. Put the client on a separate and larger machine. Watch client CPU during every run.
  3. Record versions. Save INFO server from each server, the exact binary or image digest, the kernel, the instance type, and the full config.
  4. Set threads deliberately. Try a few values of io-threads for Redis and Valkey and of proactor_threads for Dragonfly, all within the same core count. Record the setting next to the result.
  5. Use production sizes. Sample values with MEMORY USAGE, take the read and write mix from INFO commandstats, and copy the key count. Use a skewed key pattern if you have hot keys.
  6. Preload, then warm up. Fill the keyspace before measuring. Discard the first few minutes.
  7. Run long and repeat. Ten minutes or more, at least three runs. Report the median and the spread.
  8. Measure latency at your target rate, not only at saturation. Use --rate-limiting in memtier to hold a fixed rate per connection.
  9. Test your real pipeline depth, usually 1, plus the depth your client library really uses.
  10. Turn on what production uses. AOF or RDB, a replica attached, TLS if you use it. Trigger BGSAVE once during a run and watch latency and RSS.
  11. Measure memory per key. After loading, divide used_memory and RSS by the key count. Note fragmentation.
  12. Publish the method with the result, even if only inside your team.

An example preload and test with memtier_benchmark. Replace the address, sizes, and ratio with your own. The value sizes and weights below are placeholders, not a recommendation.

# Preload every key once
memtier_benchmark -s 10.0.1.20 -p 6379 --protocol=redis \
  --threads=4 --clients=10 --ratio=1:0 \
  --key-maximum=5000000 --key-pattern=P:P --requests=allkeys \
  --data-size-list=100:60,1024:30,8192:10 --hide-histogram

# Measured run: 10% writes, 90% reads, skewed keys, no pipelining
memtier_benchmark -s 10.0.1.20 -p 6379 --protocol=redis \
  --threads=8 --clients=25 --ratio=1:9 \
  --key-maximum=5000000 --key-pattern=G:G \
  --data-size-list=100:60,1024:30,8192:10 \
  --pipeline=1 --test-time=600 --run-count=3 \
  --distinct-client-seed --hide-histogram \
  --print-percentiles=50,99,99.9 \
  --json-out-file=run-01.json

Keep every flag the same across servers. Change one thing at a time.

What to practise this week

  1. Take one vendor figure quoted in your team and fill in the eleven-point checklist. Count the blanks.
  2. Pull INFO commandstats and sample 1,000 keys with MEMORY USAGE from a non-production copy. Write down your real read and write mix and value size spread.
  3. Run the preload and one measured run against the server you use today. Record p50, p99, and p99.9 at your current peak rate.
  4. Repeat once with persistence and a replica switched on, as in production.
  5. Only then add a second candidate on the same instance type.

Next in this series

Speed is one part. The next notes move to operations across Redis, Valkey, and Dragonfly: persistence and snapshots, replication and failover, memory behaviour under load, and upgrades.

Sources

comments powered by Disqus

Releted Posts

Redis, Valkey, or Dragonfly: revisit the choice before you treat them as the same cache

If your service already uses Redis, the name on the port has not changed. The product behind that name has. Treating Redis, Valkey, and Dragonfly as three labels for one cache is the habit worth unlearning.

Read more

MongoDB 8.0 Performance is 36% higher, but there is a catch…

TLDR: If your app is performance critical, think twice, thrice before upgrading to MongoDB 7.0 and 8.0. Here is why…

Read more

How Redis Helps to Increase Service Performance in NodeJS

In modern backend applications, performance optimization is crucial for handling high traffic efficiently. Typically, a backend application consists of business logic and a database.

Read more