Redis, Valkey, and Dragonfly clients: revisit your connection settings before the next failover
An order service runs on Valkey with Sentinel. At 2 AM the team runs a planned failover before patching the primary. The dashboard shows zero client errors, so everyone signs off. In the morning some order counters are lower than the payment records. Nothing crashed. The application kept writing to the old primary for several seconds after a replica was promoted, and those writes disappeared when the old primary was turned into a replica. The server did what its documentation says, and the client did what its defaults say. Nobody had read either.
Earlier parts covered the choice and compatibility, benchmark claims, persistence and failover and memory and upgrades. This part looks from the application side: client libraries and licences, RESP3, pools and multiplexing, timeouts, retries, failover, cluster redirects, pipelining, client-side caching and the server settings clients feel. Facts come from official docs, release notes and source code at the tags named below, checked on 6 October 2026. Drill numbers are from my own run.
The short version
- In 2026 most official clients moved to RESP3 by default: redis-py 8.0, node-redis 6.0, StackExchange.Redis 3.0, ioredis 6.0 and Jedis 8.0. Most keep old reply shapes, but test raw commands and scripts.
- Default timeouts range from 250 ms (Valkey GLIDE) to 60 seconds (Lettuce). Choose yours from your latency budget, not from the library.
- Retries hide a failover only when the retry budget is longer than the failover, and they can repeat writes. In redis-py 8.1.0, my Sentinel client had no retries unless I passed one.
- In my drill, a plain
SENTINEL FAILOVERleft the client writing to the old primary for about 11 seconds, and 1,069 acknowledged writes were lost. Valkey’sCOORDINATEDoption lost none. - Cluster clients must follow
MOVEDandASKand refresh the slot map. A plannedCLUSTER FAILOVERgave zero errors; killing a primary gave about 3 seconds of errors on that shard only. - Client-side caching needs RESP3 and real client support. redis-py 8.1.0 refused to enable it on Valkey 9.1.2 but used it on Redis 8.10.2 and Dragonfly v2.0.0.
The client landscape on 6 October 2026
Versions and dates are from each project’s GitHub releases or package registry. “Protocol” means what the client asks for when you do not set it.
Client
Latest release
Licence
Connection model
Default protocol
redis-py (Python)
8.1.0, 30 Jul 2026
MIT
Pool
RESP3 since 8.0
Jedis (Java)
8.0.1, 28 Aug 2026
MIT
Pool
HELLO 3, RESP2 fallback, since 8.0
Lettuce (Java)
7.8.0, 24 Sep 2026
MIT
Shared thread-safe connection
RESP3, RESP2 fallback
go-redis (Go)
9.23.0, 5 Oct 2026
BSD-2-Clause
Pool
RESP3 (Protocol: 3)
node-redis (Node.js)
6.3.0, 30 Sep 2026
MIT
One connection per client, pools optional
RESP3 since 6.0
ioredis (Node.js)
6.0.0, 31 Jul 2026
MIT
One connection per client
RESP3, RESP2 fallback, since 6.0
StackExchange.Redis (.NET)
3.3.1, 22 Sep 2026
MIT
Multiplexer
RESP3 in 3.x
Valkey GLIDE (many)
2.5.3, Sep 2026
Apache-2.0
Multiplexed, one per node
RESP3
valkey-go (Go)
1.0.78, 15 Sep 2026
Apache-2.0
Auto pipelining
Tries RESP3 first
valkey-py (Python)
6.1.1, 11 Aug 2025
MIT
Pool
RESP2
A few notes on this table:
- Licences. The server licence change in part 1 did not touch these clients. Redis maintained clients are MIT, except go-redis (BSD-2-Clause). GLIDE and valkey-go are Apache-2.0; the forks valkey-py, valkey-java and iovalkey are MIT.
- ioredis. Its README calls it “a stable project” maintained “on a best-effort basis” and recommends node-redis for new projects. It still shipped a major release with RESP3 in July 2026.
- The Valkey side. GLIDE has a Rust core with bindings for Python, Java, Node.js, Go and PHP; C#, C++ and Ruby are in development. valkey-java (a Jedis fork, last release October 2025) and iovalkey (an ioredis fork, 0.4.0 in July 2026) also exist. The valkey.io clients page shows older versions than the repositories, so check the repository.
- Dragonfly has no client of its own. Its SDK page says your Redis library “should work as expected” and lists nine officially supported ones, including redis-py, Jedis, go-redis, node-redis and ioredis. Its compatibility table matters for clients:
CLIENT TRACKINGis partial (noBCAST,PREFIXorREDIRECT), andCLIENT NO-EVICT,ASKINGandCLUSTER FAILOVERare unsupported.
RESP2, RESP3 and HELLO
Every connection starts in RESP2. A client sends HELLO 3 to switch, often with AUTH and SETNAME in the same command. RESP3 adds maps, doubles and push messages, so invalidations and pub/sub can share the data connection. Details of the 2026 switch:
- Fallback differs. Lettuce (
RedisHandshake), Jedis 8 and ioredis 6 fall back to RESP2 when the server answersNOPROTOor unknown command. valkey-go tries RESP3 first unlessAlwaysRESP2is set. In redis-py 8.1.0’sconnection.pyI found no fallback path, so a proxy that rejectsHELLOneedsprotocol=2. - Reply shapes. redis-py 8 keeps RESP2-style Python results unless you set
legacy_responses=False, and ioredis 6 keeps them throughreplyMapping: "legacy". StackExchange.Redis warns thatExecuteandScriptEvaluatecallers may see different structures. - Check what you got.
CLIENT LISTandCLIENT INFOshowresp=3orresp=2on Redis and Valkey. In my drill, redis-py 8.1.0 connected withresp=3and the older 6.1.0 on the same box withresp=2. - Server identity. Valkey 9.1.2 answers
HELLO 3withserver valkeyandversion 9.1.2, whileINFOstill saysredis_version:7.2.4(part 4). Dragonfly v2.0.0 answersserver redisandversion 7.4.0. Clients that gate features on these fields behave differently per server. - A Dragonfly quirk I measured. A bare
HELLO(no protocol argument) moved a RESP3 connection back to RESP2 on Dragonfly v2.0.0; Valkey kept RESP3. Clients send a version, but home-made health checks may not.
Pools or one shared connection
Pool clients (redis-py, Jedis, go-redis) lend each caller its own connection. Multiplexers (StackExchange.Redis, GLIDE, Lettuce by default) send everyone’s commands over one connection, often pipelined.
- Multiplexers cannot share blocking or stateful commands such as
BLPOPandMULTI/EXEC. GLIDE’s docs addWATCH(one thread’sWATCHaffects all threads on that client) and large values that delay small ones. Use a separate client for these. - Pools multiply. redis-py 8 defaults to
max_connections=100and go-redis to10 x GOMAXPROCS. Sixty pods at 100 each is 6,000 connections to one primary, before cron jobs and cluster clients that connect to every node. - The server has limits.
maxclientsis 10,000 on Redis and Valkey and 64,000 on Dragonfly. Redis lowers it at startup if the file descriptor limit is too small.
Timeouts: know which clock you are setting
Libraries mix three clocks under similar names: the connect timeout (TCP and TLS setup), the socket or read timeout (one read or write; a long BLPOP can trip it, as the redis-py 8.0 notes warn) and the command budget (one call including retries, like GLIDE’s request_timeout). Many clients have no command budget at all.
Client and version
Connect
Read or command
Retries and reconnect
redis-py 8.1.0
5 s
5 s socket
10 attempts, exponential with jitter, 10 ms to 1 s
Jedis 8.0.1
2 s
2 s socket
Docs show your own retry loop
Lettuce 7.8.0
10 s
60 s command
Auto reconnect, queues commands meanwhile
go-redis 9.23.0
5 s
5 s read and write
3 retries, 10 ms to 1 s backoff
node-redis 6.x
5 s
5 s command
Reconnect backoff up to 2 s plus 0 to 200 ms jitter
ioredis 6.0.0
10 s
None by default
Reconnect backoff up to 5 s plus jitter; queued commands fail after 20 attempts
StackExchange.Redis 3.x
5 s
5 s sync and async
3 connect retries, no command retries
Valkey GLIDE 2.5
2 s
250 ms total request
Built-in reconnect with jittered backoff
Docs can lag code: the go-redis production page still says 3 s reads and 8 to 512 ms backoff, while the v9.23.0 source sets 5 s and 10 ms to 1 s. Values also change between majors: redis-py had no socket timeout by default in 7.4 and has 5 s in 8.0.
Dead peers need TCP keepalive, not a read timeout. Server defaults are tcp-keepalive 300 and timeout 0 (never close idle clients) on Redis, Valkey and Dragonfly. redis-py 8 enables client keepalive (30 s idle, 5 s interval, 3 probes), and the Lettuce docs pair keepalive with TCP_USER_TIMEOUT.
Retries, backoff and in-flight commands
A retry bets that the error was temporary and the command did not run. The second half is not always true.
- Know what is retried. redis-py retries
ConnectionErrorandTimeoutErrorby default. go-redis retries EOF, dial errors and pool timeouts, and never a cancelled context. StackExchange.Redis retries connects only and suggests a library such as Polly for commands. - Know what is replayed. node-redis and ioredis queue commands while offline and send them after reconnecting; node-redis docs warn this can repeat a non-idempotent command and offer
disableOfflineQueue. Lettuce is at-least-once by default;autoReconnect(false)gives at-most-once and a replay filter (6.6 and later) skips chosen commands. - Add jitter. Thousands of clients reconnecting together is a second outage. redis-py, node-redis and ioredis 6 add jitter by default.
- Size the budget. Retries hide a failover only if attempts times backoff exceeds failover time, and the caller’s own timeout must allow that wait.
- Treat timeouts on writes as unknown. A timed-out
INCRmay or may not have run. Make writes idempotent where it matters, the same way as idempotency keys for HTTP retries.
One redis-py detail surprised me. The 10-attempt default lives in the Redis() constructor. A client from Sentinel.master_for() builds its own pool, and its connections had Retry(NoBackoff(), 0) when I inspected them. RedisCluster has its own cluster-level retry (10 attempts) while per-node connections also showed zero. Pass retry= explicitly and check the object. A starting point for redis-py 8.x, to adjust to your own budget:
from redis import Redis
from redis.retry import Retry
from redis.backoff import ExponentialWithJitterBackoff
from redis.sentinel import Sentinel
retry = Retry(ExponentialWithJitterBackoff(base=0.01, cap=1), 6)
common = dict(socket_connect_timeout=0.5, socket_timeout=0.5,
retry=retry, health_check_interval=10,
client_name="orders-api")
r = Redis(host="cache.internal", port=6379, **common)
sentinel = Sentinel([("s1", 26379), ("s2", 26379), ("s3", 26379)],
socket_timeout=0.5)
primary = sentinel.master_for("orders", **common) # pass retry here too
Failover: what the client must do
Sentinel-aware clients
The Sentinel client spec asks the library to try each Sentinel with a short timeout, ask SENTINEL get-master-addr-by-name, confirm with ROLE, and resolve again on every reconnect. Sentinel sends CLIENT KILL type normal to instances it reconfigures, forcing that.
The gap is the old primary. A plain SENTINEL FAILOVER treats the primary “as if” unreachable. It promotes the replica at once, but in Valkey’s sentinel.c an instance that still reports itself as master is converted only after a wait of four publish periods (2,000 ms each). Clients still connected to it keep writing, and those writes are discarded when it resyncs as a replica. That is the 1,069 lost writes in my drill.
Valkey 9.0 added SENTINEL FAILOVER <name> COORDINATED, which hands over with the FAILOVER command and kills client connections on both nodes. Valkey’s docs warn that it completes quickly, so clients must follow the full Sentinel protocol; clients that rely only on pub/sub messages may reconnect to the demoted node and get READONLY errors.
Cluster-aware clients
The cluster spec sets the rules:
MOVEDmeans the slot lives elsewhere now. Clients should refresh the whole slot map withCLUSTER SHARDS(or the deprecatedCLUSTER SLOTS), because a failover moves many slots at once.ASKmeans only this query goes elsewhere, preceded byASKING; later queries stay put. A client that cannot handleASKis “not a complete Redis Cluster client”.- Multi-key commands on keys split by a resharding return
TRYAGAIN. - Replica reads need
READONLYon that connection; otherwise the replica answersMOVED. Valkey 8.0 addedCLIENT CAPA redirectfor standalone replicas, which answer-REDIRECTwith the primary’s address.
The raw exchange from my drill, with valkey-cli in non-cluster mode:
$ valkey-cli -p 17301 SET user:1001 asha
MOVED 5712 127.0.0.1:17302
$ valkey-cli -p 17304 GET user:1001 # 17304 is a replica of 17302
MOVED 5712 127.0.0.1:17302
$ printf 'READONLY\nGET user:1001\n' | valkey-cli -p 17304
OK
asha
Lettuce’s guide recommends adaptive topology refresh and warming connections to every node before taking traffic, since they otherwise open lazily on first use.
Dragonfly
Dragonfly can emulate a one-node cluster for cluster clients. Its docs say emulated mode is the default when --cluster_mode is unset, but the v2.0.0 binary I ran replied “Cluster is disabled” until I passed --cluster_mode=emulated; then CLUSTER SLOTS showed one node owning all 16,384 slots. Multi-shard Dragonfly has no built-in failover (part 3), and CLUSTER FAILOVER and ASKING are unsupported.
Pipelining and transactions
Pipelining sends many commands before reading replies, saving round trips and system calls. It is not atomic: other clients’ commands can interleave. The server buffers replies in memory, so the Redis docs suggest batches (their example is 10k commands). Some clients pipeline for you: multiplexers by design, valkey-go by default and ioredis with enableAutoPipelining.
MULTI/EXEC runs queued commands without interleaving, but there is no rollback: a command that fails at runtime does not undo the others. WATCH adds check-and-set. In a cluster, all keys must share a slot (hash tags such as {user:1001}). On Dragonfly, scripts that touch keys not passed in KEYS fail with “script tried accessing undeclared key” unless allowed by a flag that makes Dragonfly stop other work while the script runs.
Client-side caching
With CLIENT TRACKING, the server remembers which keys a connection read and sends an invalidation when one changes. Default mode costs server memory (tracking-table-max-keys, 1,000,000 in the Redis sample config); broadcast mode tracks prefixes and costs none. Invalidations arrive on the same RESP3 connection, or with RESP2 through REDIRECT to a pub/sub connection. If the connection drops, the client must flush its local cache.
Support is uneven:
- Redis docs list redis-py 5.1.0, Jedis 5.2.0, node-redis 5.1.0 and go-redis 9.22.0 as supporting it. go-redis marks its API experimental, for standalone clients on database 0 only.
- Lettuce 7.8 deprecated its
ClientSideCachingclass as “legacy and incomplete”, with a redesign planned. - GLIDE 2.4 (May 2026) added caching in Python, Java, Node.js and Go, but the default is TTL only: entries can be stale until their TTL ends, even after your own write. Server-assisted invalidation exists only in Java and uses broadcast mode, which Dragonfly does not support.
- valkey-go builds it in through
DoCache()with a client-side TTL. - Name and version checks bite. redis-py 8.1.0 enables caching only when
HELLOreportsserver redisat version 7.4.0 or later. Valkey reportsserver valkey, so it refused; Dragonfly reportsredis 7.4.0, so it passed.
Server settings that clients feel
- Output buffers. By default normal clients have no output buffer limit; pub/sub clients are cut at 32 MB, or 8 MB for 60 s.
- Client eviction.
maxmemory-clientsis 0 (off) by default; the Redis docs suggest 5% as a start for large deployments. Protect your monitoring connection withCLIENT NO-EVICT on(Redis 7.0 and later; it worked on Valkey 9.1.2, Dragonfly rejected it). - Names. Set a client name per service; names appear in
SLOWLOGsince 4.0.CLIENT SETINFO(7.2 and later) records library name and version, and redis-py, Jedis, go-redis and node-redis send it on connect. - TLS. TLS is a build option, uses mutual TLS by default (
tls-auth-clients), and needstls-replicationandtls-clusterfor those links. The Redis docs say it reduces throughput per instance; Redis 8.0 supports I/O threads with TLS. Jedis 8.0 now enforces hostname verification by default.
# who is connected, with which library and protocol
CLIENT LIST
INFO clients
CONFIG GET maxclients
CONFIG GET timeout
CONFIG GET tcp-keepalive
CONFIG GET maxmemory-clients
# on the monitoring connection only
CLIENT SETNAME ops-monitor
CLIENT NO-EVICT on
# Valkey 9.0 or later: planned Sentinel failover without the old-primary window
SENTINEL FAILOVER mymaster COORDINATED
Practical checklist
- List every service with its client library, exact version and default protocol.
- Before a client major upgrade, run tests with the new default protocol and with RESP2 forced.
- Set connect, read and total timeouts explicitly, from your latency budget.
- Set retries explicitly, with jitter, for idempotent commands or with idempotency keys.
- Decide what happens to commands queued while disconnected.
- Give each service a client name; confirm
lib-name,lib-verandrespinCLIENT LIST. - Count connections (pods x pool size x nodes) against
maxclientsand file descriptor limits. - For clusters, enable topology refresh and warm up node connections before readiness.
- On Valkey 9 Sentinel, use
COORDINATEDfor planned failovers, and still test a hard kill. - Run the drill below after every client or server upgrade; record errors, longest call and lost writes.
Common mistakes
- Leaving Lettuce at 60 seconds, or GLIDE at 250 ms, without checking. Both are fine for someone; neither is your latency budget.
- Trusting zero errors during a failover. My cleanest looking run lost the most data.
- Retrying
INCRorLPUSHblindly. A retried write after a timeout can apply twice. - Using
BLPOPorWATCHon a shared multiplexed connection. - Assuming a Valkey or Dragonfly server unlocks the same client features. Feature checks look at server name and version.
- Assuming Dragonfly runs in emulated cluster mode without the flag. Test it on your version.
Drill: clients during failover, redirects and caching
Use a throwaway machine with persistence off, and a script that records each call’s start, duration and error.
- Defaults. Create a client with no options; print its timeouts, retry object and
CLIENT INFO. - Pipelining. Write 20,000 small keys one by one, in pipelines of 10, 100 and 1,000, and in
MULTI/EXECbatches of 100. - Sentinel. One primary, one replica, three Sentinels (
down-after-milliseconds 3000). RunINCRevery 10 ms with a 0.5 s socket timeout. Trigger a plainSENTINEL FAILOVER, aCOORDINATEDone and akill -9, with and without retries. Compare the final counter with acknowledged increments. - Cluster. Three primaries, three replicas (
cluster-node-timeout 2000). Write every 5 ms, kill one primary, bring it back, then runCLUSTER FAILOVERon it. - Caching. Read 100 keys 10,000 times with and without client-side caching, count
GETcalls on the server, then change a key from another client.
My run, and its limits
I ran this on 6 October 2026 on a shared VM container: 8 vCPUs, 15 GB RAM (about 4 GB free), Debian 13, Linux 6.12, with other workloads present. Servers: Valkey 9.1.2 (official build), Redis 8.10.2 (built from its tag) and Dragonfly v2.0.0 (release binary, 2 threads). Client: redis-py 8.1.0 on Python 3.13, everything on localhost. One run per failover case, so read these as observations, not benchmarks.
Defaults in redis-py 8.1.0 matched its source: 5 s connect and socket timeouts, keepalive on, pool of 100, 10 retries with jittered backoff from 10 ms to 1 s, and resp=3 on Valkey and Redis.
Pipelining on Valkey (single Python client, 100-byte values, median of three): 18,437 writes per second one by one, 74,992 with pipelines of 10, 117,810 with 100, 144,928 with 1,000, and 106,507 in MULTI/EXEC batches of 100. Dragonfly showed the same shape (17,272 one by one, 119,469 at 1,000). The Python client was the bottleneck, and loopback hides most of the network cost, so real gains depend on your round trip time.
Event (Valkey 9.1.2, Sentinel)
Client retries
Errors
Longest call
Acknowledged writes lost
Plain SENTINEL FAILOVER
None
0
4 ms
1,069
SENTINEL FAILOVER ... COORDINATED
None
38, over 0.9 s
0.50 s (timeout)
0
SENTINEL FAILOVER ... COORDINATED
10, jittered
1
5.7 s
0
kill -9 primary
None
399, over 4.3 s
5 ms
0
kill -9 primary
10, jittered
1
4.3 s
0
In the plain run the replica was primary 0.19 s after the command, but the old primary stayed primary until 11.3 s. After a kill, promotion took about 3.3 s. With retries, one call absorbed the whole wait and still failed. I did not investigate why retries in the coordinated run kept failing for 5.7 s.
Cluster, with RedisCluster defaults: killing the primary of slots 5461 to 10922 gave 74 errors over about 3.0 s, all on keys of that shard; other shards had no errors and a longest call of 0.26 s. A planned CLUSTER FAILOVER on the rejoined node gave zero errors and a longest call of 29 ms.
Client-side caching, redis-py 8.1.0
GETs sent
GETs reaching server
After an external write
Redis 8.10.2
10,000
100
New value on the first read
Dragonfly v2.0.0
10,000
100
New value on the first read
Valkey 9.1.2
Refused
Not run
“client-side caching is supported by Redis 7.4 or later”
Two smaller checks. With server timeout 2, a pooled client reconnected silently after 4 s idle, while a single long-lived connection without retries failed with “Connection closed by server”. Caching with protocol=2 was refused by the client on all three servers.
What to unlearn and re-learn
- Unlearn “the client handles failover”. Re-learn that the client only follows what the server tells it, and that a planned failover can leave writes on the old primary.
- Unlearn “retries make it safe”. Re-learn retries as a budget that costs latency and may repeat writes.
- Unlearn “RESP2 is what clients speak”. Re-learn that your next client upgrade probably moves to RESP3.
- Unlearn “same protocol, same features”. Re-learn that clients check server name and version before enabling features.
Revisit your client settings
Most cache incidents users notice pass through a client library, and most client settings in production are whatever the library shipped. Learn what your version does on connect, timeout and reconnect. Unlearn the comfort of zero errors. Re-learn retries, redirects and caching as contracts with the server. Practise the drill on staging with your real client and pool sizes, and apply the checklist before the next planned failover, not after the morning reconciliation.
Sources
- Redis docs: Client libraries and Connection pools and multiplexing
- Redis docs: Production usage for redis-py, Jedis, Lettuce, go-redis, node-redis, StackExchange.Redis
- Redis docs: Client-side caching introduction and reference
- Redis docs: HELLO, CLIENT TRACKING, CLIENT SETINFO, CLIENT NO-EVICT, CLIENT LIST, READONLY
- Redis docs: Cluster specification and Sentinel client spec
- Redis docs: Pipelining, Transactions, Client handling, TLS
- Redis 8.10.2 sample config
- Valkey clients page, Sentinel (coordinated failover), CLIENT CAPA, Client-side caching, Client handling
- Valkey 9.1.2 source: sentinel.c and release notes
- redis-py 8.0.0 release, 8.1.0 release, _defaults.py, connection.py, LICENSE
- Jedis 8.0.0 release, DefaultJedisClientConfig.java, LICENSE
- Lettuce 7.8.0 release, README, RedisHandshake.java, ClientSideCaching.java
- go-redis 9.23.0 release, options.go, error.go, README, LICENSE
- node-redis 6.0.0 release and client configuration
- ioredis README, 6.0.0 release, Upgrading from v5 to v6, RedisOptions.ts
- StackExchange.Redis: RESP3, Configuration, 3.0.0 release
- Valkey GLIDE README, 2.5.3 release, Connection management, Client-side caching, Python config.py
- valkey-go and valkey.go, valkey-py and connection.py, valkey-java, iovalkey
- Dragonfly docs: Command compatibility, SDKs, Cluster mode, Server flags, Scripting
- Dragonfly v2.0.0 release
Releted Posts
Redis, Valkey, and Dragonfly memory and upgrades: revisit before the next spike
A payments team runs its session cache with maxmemory 12gb on a 16 GB node and feels safe. On a festival sale evening a replica reconnects, a reporting job pulls a few huge hashes, and the kernel kills the process while used_memory still shows less than 12 GB.
Read moreRedis, Valkey, and Dragonfly benchmarks: read the method first
A benchmark number without its method sounds complete, but it does not tell you what happened. The first note in this series, Redis, Valkey, or Dragonfly: revisit the choice before you treat them as the same cache, covered licences, history, data file compatibility, and cluster mode.
Read moreRedis, Valkey, and Dragonfly operations: persistence and failover
A cache becomes a database the day nobody can rebuild it. After that day, persistence and failover are not settings copied from a blog.
Read more