PostgreSQL high availability: revisit Patroni and CloudNativePG before your next failover

Picture a lending company in Bengaluru (an invented example, not a real client). Its loan ledger runs on PostgreSQL 16 across three VMs, managed by Patroni with a three-node etcd cluster and HAProxy in front. It was set up in 2021 and has failed over only twice. Last month a hypervisor host froze for about a minute during patching. The cluster failed over as designed, but the next morning the reconciliation team found a few hundred EMI payments that the app had shown as successful and the database did not have. Now the platform team wants CloudNativePG on Kubernetes, the DBA wants to stay on Patroni and turn on synchronous replication, and the architect asks whether another operator would be safer.

Each option is reasonable, and none helps until the team knows what happened in that minute. High availability for PostgreSQL is not one feature. It is a set of decisions: who may become primary, how the old primary is stopped, how far behind a replica may be when promoted, how clients find the new primary, and how you recover when all of that fails. This post looks at those decisions in Patroni and CloudNativePG, with notes on Crunchy PGO, the Zalando operator, StackGres and Percona’s operator, plus backups, upgrades and security. Facts come from release notes, official docs, licence files and CVE records, checked on 11 October 2026. The drill near the end is my own run.

The short version

  • Patroni 4.1.5 (12 August 2026) and CloudNativePG 1.30.1 (23 September 2026) are current. Both default to asynchronous replication, so a failover can lose committed transactions. In my drill, a 4-second replication stall followed by a primary crash lost 341 acknowledged commits under Patroni’s defaults.
  • The bigger danger is an old primary that keeps running. When I killed only Patroni on the primary, with no watchdog, a replica was promoted about 40 seconds later while the old primary kept accepting writes: 3,298 acknowledged rows existed only there, and pg_rewind erased them when it rejoined.
  • CloudNativePG 1.30 added a Kubernetes Lease that an instance must hold before it promotes. The release note says plainly that the lease “is a promotion gate, not a fence”.
  • Synchronous replication protects acknowledged commits by stalling writes when no synchronous standby is available. Both projects offer quorum modes.
  • pgBackRest, used by Crunchy PGO and Percona’s operator, was declared unmaintained on 27 April 2026 and revived by a coalition of sponsors on 18 May. Version 2.59.3 (4 October) fixes possible weak encryption subkeys, and upgrading does not replace an existing weak one.
  • PostgreSQL 14 reaches end of life on 12 November 2026. The 13 August minor releases fixed 28 vulnerabilities, including CVE-2026-6471, where a role with REPLICATION can load an arbitrary library through logical decoding.

Where things stand on 11 October 2026

Dates are from GitHub release pages and project sites, converted to IST.

Project

Latest release

What it is

Licence

Patroni

4.1.5 and 4.0.11 on 12 August 2026

HA manager using etcd, Consul, ZooKeeper or the Kubernetes API

MIT

CloudNativePG

1.30.1 on 23 September 2026 (1.29 support ended 29 September)

Kubernetes operator with its own instance manager, no Patroni

Apache 2.0; CNCF Sandbox

Crunchy PGO

6.0.3 on 21 August 2026; 5.8.9 on 20 August

Operator built on Patroni and pgBackRest

Apache 2.0 code; Crunchy images under developer programme terms

Zalando postgres-operator

2.0.3 on 2 October 2026

Operator running Spilo images (Patroni plus PostgreSQL)

MIT

StackGres

1.19.3 on 2 October 2026; 1.18.10 on 8 October

OnGres operator with Patroni, Envoy and a web console

AGPL-3.0

Percona Operator for PostgreSQL

3.1.0 on 9 September 2026

Operator derived from Crunchy PGO

Apache 2.0

pgBackRest

2.59.3 on 4 October 2026

Backup, restore and WAL archiving

MIT

Barman

3.20.1 on 30 September 2026

EDB’s backup manager; its barman-cloud tools serve CloudNativePG

GPL-3.0

Barman Cloud plugin

0.15.1 on 30 September 2026

CNPG-I plugin replacing in-tree Barman Cloud support

Apache 2.0

etcd

3.7.2, 3.6.15 and 3.5.34 on 23 September 2026

Raft key-value store, a common Patroni DCS

Apache 2.0

PostgreSQL

18.6, 17.11, 16.15, 15.19, 14.24 on 13 August 2026; 19 Beta 4 in September

The database; 19 is planned for October 2026

PostgreSQL Licence

Date

Event

July 2015

Patroni repository created at Zalando, as a fork of Compose’s Governor

21 April 2022

CloudNativePG 1.15.0, first open source release of the operator built at 2ndQuadrant (later EDB)

29 August 2024

Patroni 4.0.0 adds quorum-based failover and drops “master” from its API

26 September 2024

PostgreSQL 17 adds failover-ready logical replication slots

21 January 2025

CloudNativePG accepted into CNCF at Sandbox level

24 April 2025

Kubernetes 1.33 deprecates the v1 Endpoints API in favour of EndpointSlices

2 June 2025

Snowflake announces it will acquire Crunchy Data

23 September 2025

Patroni 4.1.0 adds replication-aware readiness checks

25 September 2025

PostgreSQL 18 enables data checksums by default and keeps planner statistics through pg_upgrade

November 2025

CloudNativePG applies for CNCF Incubation (still open)

December 2025

Crunchy PGO 6.0.0 with a v1 PostgresCluster API

27 April 2026

pgBackRest’s maintainer announces it is no longer maintained

18 May 2026

pgBackRest “will continue”, funded by several sponsors

29 June 2026

CloudNativePG 1.30.0 adds the primary Lease and fixes two CVEs

29 July 2026

Zalando postgres-operator 2.0.1, the first usable 2.0 release

13 August 2026

PostgreSQL minor releases fix 28 vulnerabilities

4 October 2026

pgBackRest 2.59.3 fixes possible weak encryption subkeys

12 November 2026

PostgreSQL 14 end of life and the next scheduled minor releases

What actually happens during a failover

Every HA manager does the same five things in its own way.

  1. Detects that the primary is gone. A Patroni primary must renew a leader key in the distributed configuration store (DCS) with a ttl (default 30 seconds), and every node runs its loop every loop_wait (default 10). CloudNativePG’s operator watches the primary’s readiness and, since 1.30, a Kubernetes Lease (default leaseDurationSeconds 15).
  2. Elects one candidate. Patroni replicas race for the expired key; one lagging more than maximum_lag_on_failover (1 MB in the documented examples) cannot take part. CloudNativePG first stops the WAL receivers on all replicas, then promotes the most advanced.
  3. Stops the old primary writing. This is fencing, and most bad incidents start here. A primary that cannot renew its key should demote itself, but a frozen, killed or starved process cannot do anything.
  4. Moves clients, through load balancer health checks, a Kubernetes Service, a pooler or a multi-host connection string.
  5. Repairs the old primary, usually with pg_rewind, which throws away whatever it wrote after the fork.

Asynchronous replication and the data loss window

The primary acknowledges a commit after writing WAL locally; replicas catch up a little later. Patroni’s docs say that in asynchronous mode “the cluster is allowed to lose some committed transactions to ensure availability”, bounded in the worst case by maximum_lag_on_failover plus what was written in the last ttl seconds. CloudNativePG’s docs say the same: without synchronous replication, “some data loss is expected and accepted during failover”.

On a healthy network the window is small. The Bengaluru team’s missing payments fit another pattern: a frozen host, where replication stopped first and the primary died later. My drill reproduced it. Freezing the primary’s WAL senders for 4 seconds and then crashing the node lost 341 commits the client had seen succeed.

Synchronous replication and its cost

synchronous_commit decides how much a commit waits for: on (default) waits for the synchronous standby to flush WAL, remote_write for it to write, remote_apply for it to replay. synchronous_standby_names lists standbys by priority (FIRST n (...)) or quorum (ANY n (...)). Do not edit it yourself under either tool: Patroni manages it when synchronous_mode is on or quorum, and CloudNativePG derives it from .spec.postgresql.synchronous.

Three behaviours matter:

  • Writes stall without a synchronous standby, if you ask for it. Patroni’s synchronous_mode: on lets the primary run alone when no standby is suitable, unless synchronous_mode_strict is set. CloudNativePG’s dataDurability: required, the default once the stanza exists, blocks writes; preferred keeps writing and accepts possible loss.
  • Losing two nodes together can still lose data. If the primary and its one synchronous standby fail together, a third node may lack the latest commits. Patroni only lets the last leader or a listed synchronous standby promote; CloudNativePG’s failoverQuorum: true refuses to promote unless it can prove the candidate has every synchronously committed transaction.
  • A cancelled wait is not a rollback. PostgreSQL commits locally before waiting for the standby, so a cancelled wait leaves a visible but unreplicated transaction. In my drill, statement_timeout did not cancel a commit stuck in the SyncRep wait; the client simply hung. Clients need idempotent retries, as covered in the idempotency keys post.

Fencing and split-brain

Patroni lets a node run PostgreSQL as primary only while it can update the leader key, and demotes it before the key expires. That works while Patroni is alive. Its watchdog page lists when it is not: Patroni crashed or was killed, PostgreSQL stops too slowly, or the hypervisor paused the VM. For those cases Patroni can arm the Linux watchdog (/dev/watchdog), which resets the machine when keepalives stop. With mode: required, a node refuses to lead unless the watchdog is active.

The second valve is DCS failsafe mode. Without it, an etcd outage makes the primary demote itself, because one node cannot tell a partition from a store that is down. With failsafe_mode: true, the primary stays up if every known member answers its REST call. In my drill that was the difference between a 28-second write outage and none.

CloudNativePG uses Kubernetes itself. Since 1.27, a primary’s liveness probe fails when its instance manager can reach neither the API server nor any other instance, and the kubelet restarts it. That restart is a smart shutdown that lets open sessions keep committing for up to smartShutdownTimeout (180 seconds by default); the docs suggest 0 if that window matters. The Lease stops premature promotion. Manual fencing is the cnpg.io/fencedInstances annotation, which stops PostgreSQL but keeps the pod for investigation.

Replication slots and pg_rewind

Physical slots stop the primary from deleting WAL a replica still needs, and both tools manage them for members; CloudNativePG keeps “HA slots” advanced on standbys so they survive failover. The flip side is a dead replica pinning WAL until the disk fills, so set max_slot_wal_keep_size (unlimited by default). PostgreSQL 18 added idle_replication_slot_timeout, and Patroni 4.1.4 drops managed slots whose wal_status is lost. For CDC tools such as Debezium, PostgreSQL 17’s failover-ready logical slots are supported by both projects; test them through a switchover.

pg_rewind lets an old primary rejoin without a full copy. It needs wal_log_hints or data checksums, and PostgreSQL 18’s initdb now enables checksums by default. Patroni needs use_pg_rewind: true, which is off by default. Remember what a rewind does: anything the old primary wrote after the fork is gone, so copy the data directory first if you might need those rows.

Switchover is not failover

A switchover is planned: the old primary shuts down cleanly and archives its WAL before a chosen replica takes over. My Patroni switchover stalled writes for 1.56 seconds and lost nothing. Operator upgrades and many configuration changes end in a switchover, so you will see far more of them than failovers. CloudNativePG gives the old primary up to switchoverDelay (3,600 seconds by default) for that clean shutdown.

Patroni

From Governor to the common default

Patroni began at Zalando in 2015 as a fork of Compose’s Governor and now lives in its own GitHub organisation under MIT. Spilo in the Zalando operator, Crunchy PGO, Percona’s operator and StackGres all use it for failover. Older tools are quieter: repmgr (GPL-3.0, EDB) last released 5.5.0 in November 2024, and Stolon’s last release was 0.17.0 in September 2021.

Patroni’s DCS options are etcd v3, Consul, ZooKeeper, Exhibitor and the Kubernetes API, where keys live in Endpoints or ConfigMaps. With v1 Endpoints deprecated in Kubernetes 1.33, Zalando’s operator 2.0 made ConfigMaps the default “because endpoints will disappear”.

What 4.x changed

4.0.0 (August 2024) replaced “master” with “primary” in labels, API values and callbacks, and warns that upgrading to 4.x is reliable only from 3.1.0 or newer. It added synchronous_mode: quorum, which writes ANY n (...) and picks the candidate with the latest received transaction. 4.1.0 (September 2025) made replica readiness depend on replication, added lag columns to patronictl list, and added demote-cluster and promote-cluster. The patches since are worth reading: 4.1.1 copes with etcd’s security fixes in 3.6.9, 3.5.28 and 3.4.42, which require authentication for member discovery and lease keepalive; 4.1.4 drops lost slots and stops disabling the watchdog before client backends exit; 4.1.5 supports the new output_plugin_libraries setting from the August minor releases.

Timing and configuration

Patroni requires loop_wait + 2 * retry_timeout <= ttl. With defaults of 10, 10 and 30, a dead primary is replaced roughly 30 to 40 seconds after it stops renewing its key; my crash drills showed write gaps of 27 to 33 seconds. Lower values react faster but turn a slow DCS into false failovers. A small dynamic configuration, set with patronictl edit-config and stored in the DCS:

ttl: 30
loop_wait: 10
retry_timeout: 10
maximum_lag_on_failover: 1048576
failsafe_mode: true
synchronous_mode: quorum
synchronous_node_count: 1
postgresql:
  use_pg_rewind: true
  use_slots: true
  parameters:
    max_slot_wal_keep_size: 50GB

The watchdog goes in each node’s local patroni.yml; give the Patroni user access to the device first, or the node will never lead:

watchdog:
  mode: required
  device: /dev/watchdog
  safety_margin: 5

Routing clients to the primary

Patroni does not move an IP address. Clients find the primary through REST health checks: GET /primary returns 200 only on the running primary, GET /replica?lag=16MB only on a streaming replica within that lag, and /sync, /async and /quorum filter by role. HAProxy is the usual pattern:

listen pg_primary
    bind *:5000
    option httpchk GET /primary
    http-check expect status 200
    default-server inter 3s fall 3 rise 2 on-marked-down shutdown-sessions
    server pg1 10.0.1.11:5432 check port 8008
    server pg2 10.0.1.12:5432 check port 8008
    server pg3 10.0.1.13:5432 check port 8008

on-marked-down shutdown-sessions matters: without it, open connections to a demoted or orphaned primary stay alive. A libpq multi-host string with target_session_attrs=read-write is simpler but checks writability only when connecting; in my split-brain test it kept writing to the orphan. Poolers sit in front of either, as covered in the PgBouncer post.

CloudNativePG

A Kubernetes-native design

CloudNativePG was conceived at 2ndQuadrant, later acquired by EDB, and open sourced under Apache 2.0 in 2022. It uses no Patroni and no external store: an instance manager runs as PID 1 in each pod, the operator reconciles a Cluster resource, and the Kubernetes API is the source of truth. It joined CNCF at Sandbox level on 21 January 2025; its Incubation application from November 2025 was still open on 11 October 2026, with due diligence waiting on adopter interviews. EDB staff remain prominent, and EDB sells a supported build for OpenShift.

A minor release comes about every three months and is supported until three months after the next. So 1.30.x is the only supported line until 1.31, which the support page planned for about September 2026 and which was not out on 11 October. 1.30 supports Kubernetes 1.34 to 1.36 and PostgreSQL 14 to 18.

What 1.30 changed for failover

1.30.0 (29 June 2026) introduced the primary Lease. A clean shutdown releases it so a replica can take over at once; otherwise a replica waits for it to expire. The docs explain the risk it covers: a replica promoting while the old primary still holds unarchived WAL, which forks the timeline too early. 1.30.1 then fixed several failover bugs: it demotes an unreachable old primary immediately, no longer stalls when any instance is fenced, completes a pending failover instead of reverting it, and starts PostgreSQL only after the lease is held. If you run 1.29 or older, treat these as reasons to upgrade.

A manifest with synchronous replication and backups

Three instances, quorum replication with failover quorum, WAL archiving through the Barman Cloud plugin and a daily base backup. The schedule has six fields because CloudNativePG’s cron format starts with seconds.

apiVersion: barmancloud.cnpg.io/v1
kind: ObjectStore
metadata:
  name: ledger-store
spec:
  configuration:
    destinationPath: s3://pg-backups/ledger/
    s3Credentials:
      accessKeyId:
        name: ledger-s3
        key: ACCESS_KEY_ID
      secretAccessKey:
        name: ledger-s3
        key: ACCESS_SECRET_KEY
    wal:
      compression: gzip
  retentionPolicy: "30d"
---
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
  name: ledger
spec:
  instances: 3
  postgresql:
    parameters:
      max_slot_wal_keep_size: 50GB
    synchronous:
      method: any
      number: 1
      dataDurability: required
      failoverQuorum: true
  plugins:
  - name: barman-cloud.cloudnative-pg.io
    isWALArchiver: true
    parameters:
      barmanObjectName: ledger-store
  storage:
    size: 200Gi
---
apiVersion: postgresql.cnpg.io/v1
kind: ScheduledBackup
metadata:
  name: ledger-daily
spec:
  schedule: "0 30 1 * * *"
  backupOwnerReference: self
  cluster:
    name: ledger
  method: plugin
  pluginConfiguration:
    name: barman-cloud.cloudnative-pg.io

Clients use the ledger-rw Service for the primary, ledger-ro for replicas and ledger-r for any instance. The Pooler resource adds PgBouncer in front.

The other operators

Operator

HA engine

Backups

Worth knowing in 2026

Crunchy PGO 6.0

Patroni 4.1.4

pgBackRest 2.59.0

v1 PostgresCluster API; Crunchy Data is part of Snowflake; default images come from Crunchy’s portal under developer terms, or you build from PGDG packages

Zalando postgres-operator 2.0

Patroni in Spilo

WAL-G in Spilo, plus logical backups

Dropped PostgreSQL 13; SCRAM and ConfigMaps by default; 2.0.3 needs password_encryption: md5 if your pooler still uses MD5

StackGres 1.19

Patroni

Automated backups to object storage

AGPL-3.0, commercial licence on request; Envoy Postgres filter; web console

Percona Operator 3.x

Patroni

pgBackRest

3.0.0 moved the inherited Crunchy CRDs to its own API group so both can share a cluster, and documents migration from PGO

CloudNativePG 1.30

Own instance manager and Lease

Barman Cloud plugin, volume snapshots

In-tree Barman Cloud support to be removed in 1.31

The Patroni-based operators inherit Patroni’s mature logic and its timing rule, and they keep its keys in the Kubernetes API, which Patroni’s failsafe docs name as the DCS where lock update failures are more frequent. Crunchy’s and Percona’s backups depend on pgBackRest, whose funding scare this year is a reminder to check the bus factor of every component under your database. Zalando’s 2.0 is a true major release; read its migration guide first.

Backups and point-in-time recovery

HA protects against losing a server, not against a DELETE without a WHERE, a bad migration or ransomware, which replicas copy faithfully. You still need base backups, a continuous WAL archive and tested restores.

pgBackRest 2.59

Barman 3.20 and barman-cloud

Used by

Crunchy PGO, Percona, many VM setups

CloudNativePG (plugin), EDB users

Model

Repository with full, differential, incremental and block incremental backups

Central Barman server, or barman-cloud-* tools writing to object storage

Steward

David Steele, funded since May 2026 by sponsors including AWS, Supabase, pgEdge and Percona

EDB

Recent change

2.59.0 runs only restore as root by default; 2.59.3 weak subkey fix

3.20.0 needs Python 3.12 or later and adds S3 SSE-C

Two items need action now. First, pgBackRest’s 4 October advisory: before 2.59.3, if OpenSSL could not get random data (a seccomp policy blocking getrandom() is one cause), encryption subkeys and salts were silently built from memory. Only encrypted repositories are affected. The subkeys in archive.info and backup.info are created once by stanza-create, so the advisory gives a shell check, and a weak subkey means a new repository. I found no CVE ID for it. Second, CloudNativePG users on the in-tree barmanObjectStore should move to the plugin before 1.31.

CloudNativePG restores by bootstrapping a new Cluster from the object store with a target time, LSN, transaction ID or restore point. Whatever the tool, write down your recovery point and time objectives and time a restore of your largest database every quarter.

Upgrades

Minor versions keep the storage format but still need care: the August releases fixed 28 vulnerabilities, needed extra steps for parallel GIN builds, btree_gist and ltree, and skipped 18.5 because of a regression. The pattern under any HA manager is replicas first, then a switchover, then the old primary; CloudNativePG does this as a rolling update when you change the image.

Major versions mean dump and restore, logical replication into a new cluster (online), or pg_upgrade in place (offline). CloudNativePG runs pg_upgrade --link declaratively when you request a newer major image. Its docs warn that the whole cluster is down during the upgrade, replicas are recreated afterwards, and PITR does not cross the major version boundary, so take a new base backup at once. PGO 6.0 validates the environment before a major upgrade, and Percona 3.0 uses an official upgrade image. PostgreSQL 18 helps: pg_upgrade keeps planner statistics and offers a --swap mode. One trap: pg_upgrade needs matching checksum settings, and 18 enables checksums by default. If you are moving to 18, read the asynchronous I/O post before copying old settings.

Operator upgrades need the same respect. Upgrading CloudNativePG rolls every cluster to the new instance manager and ends with a switchover, automatically under the default primaryUpdateStrategy: unsupervised, unless you enable ENABLE_INSTANCE_MANAGER_INPLACE_UPDATES. For Patroni, upgrade replicas first, and reach 3.1 or later before moving to 4.x.

Security advisories

All rows were confirmed as PUBLISHED records through the cveawg.mitre.org API on 11 October 2026.

CVE

Component

What it allows

Fixed in

CVE-2026-44477

CloudNativePG

Metrics exporter session opened as superuser and demoted with SET ROLE allows escalation and OS command execution

1.28.3, 1.29.1

CVE-2026-55769

CloudNativePG

Operator connections without a fixed search_path let a database owner plant overloaded operators

1.28.4, 1.29.2, 1.30.0

CVE-2026-55765

CloudNativePG

Cleartext role passwords could be recorded by pg_stat_statements and recovered

1.28.4, 1.29.2, 1.30.0

CVE-2026-6471

PostgreSQL

A role with REPLICATION can make logical decoding load any file as a library

18.6, 17.11, 16.15, 15.19, 14.24

CVE-2026-33413

etcd

Authorization bypass in several APIs for untrusted gRPC clients

3.4.42, 3.5.28, 3.6.9

CVE-2026-33343

etcd

Nested transactions bypass key-range RBAC

3.4.42, 3.5.28, 3.6.9

CVE-2026-44283

etcd

PrevKv or lease attachment in transactions bypasses RBAC

3.4.44, 3.5.30, 3.6.11

CVE-2026-59818

etcd

CRL not enforced on the gRPC listener when listeners are split

3.5.32, 3.6.13

CVE-2026-73499

etcd

Watch API authorization bypass with open-ended ranges

3.5.33, 3.6.14, 3.7.1

CVE-2026-73500

etcd

Unbounded TLS handshake goroutines from idle connections

3.5.33, 3.6.14, 3.7.1

CloudNativePG 1.30.0 also fixed GHSA-7qwx-x8ff-3px9 (no CVE ID): control endpoints in the instance manager trusted network isolation instead of authenticating the operator. That fix was not backported, and the advisory tells older releases to restrict the status port with a NetworkPolicy. The common lessons: the database pod and the DCS are high-value targets, a tenant who owns a database can attack the operator’s superuser connections, and REPLICATION is more powerful than its name suggests.

A decision guide

Your situation

Reasonable first choice

PostgreSQL on VMs or bare metal

Patroni 4.1, three-node etcd, HAProxy with REST checks, watchdog required, pgBackRest or Barman

New Kubernetes platform, GitOps, no Patroni history

CloudNativePG 1.30 with the Barman Cloud plugin

Patroni skills, moving to Kubernetes

Crunchy PGO, Percona or Zalando, depending on images and support

Want a web console and all-in-one stack, AGPL acceptable

StackGres

No acknowledged commit may be lost

Synchronous quorum replication on three or more nodes, failover quorum or strict mode, idempotent clients

Small team, no on-call DBA

A managed PostgreSQL service may be the honest answer

For the Bengaluru team, I would not change platforms in a hurry; the same failure would happen anywhere without fencing. On the existing cluster: enable the watchdog in required mode, turn on failsafe_mode, add shutdown-sessions to HAProxy, and move the ledger to synchronous_mode: quorum with one synchronous standby. Make the payment API retry with idempotency keys, and drill a failover every quarter. A move to CloudNativePG can follow as its own project if the company is standardising on Kubernetes anyway.

A practical checklist

  1. Write down RPO and RTO per database, and whether writes may stall to protect them.
  2. Know your timers: ttl, loop_wait, retry_timeout in Patroni; primaryLease, failoverDelay, switchoverDelay, smartShutdownTimeout in CloudNativePG.
  3. Fence for real: Patroni watchdog in required mode; keep CloudNativePG’s isolation check on.
  4. Choose synchronous replication per database, in quorum mode with at least three nodes.
  5. Cap slot retention with max_slot_wal_keep_size and alert on inactive slots.
  6. Cut stale connections: HAProxy shutdown-sessions or the -rw Service, sane pooler timeouts, idempotent retries.
  7. Archive WAL continuously and restore to a point in time every quarter.
  8. Patch the whole stack: PostgreSQL, Patroni, etcd, the operator and the backup tool.
  9. Retire PostgreSQL 14 before 12 November 2026.
  10. Rehearse a switchover, a node crash, a DCS outage and a frozen primary while counting acknowledged writes.

Common mistakes

  • Reading “automatic failover” as “no data loss”. Both main tools default to asynchronous replication.
  • Running Patroni without a watchdog on VMs that can be paused. Its fencing needs Patroni alive.
  • Removing a server from the load balancer but leaving its sessions open.
  • Assuming a rewind is harmless. It discards the old primary’s writes after the fork.
  • Leaving slots unbounded. One dead replica can fill the primary’s disk.
  • Upgrading the operator on a Friday evening. It is a switchover of every cluster it manages.

Drill: failing over on purpose on one VM

On 11 October 2026, from about 2:34 PM to 2:53 PM IST, on a shared Linux VM with 8 vCPUs (Intel Xeon) and 15 GiB of RAM, without root or containers, I ran a three-node Patroni cluster on 127.0.0.1 only and stopped every process afterwards. PostgreSQL 18.6 came from the apt.postgresql.org trixie packages (SHA-256 matched the Packages index), unpacked into my own directory with Debian’s liburing2; etcd 3.6.15 ran as one member (SHA-256 matched SHA256SUMS); Patroni 4.1.5 came from PyPI with psycopg 3.3.6. No watchdog device was available.

Method: Patroni defaults (ttl 30, loop_wait 10, retry_timeout 10, maximum_lag_on_failover 1 MB), use_pg_rewind: true, checksums on. A Python writer inserted a numbered row about every 10 ms in autocommit mode through a multi-host string with target_session_attrs=read-write, and recorded only IDs whose commit succeeded. Afterwards I compared them with the table on the new primary. Lag and outages were simulated with SIGSTOP on the WAL senders or on etcd; a “node crash” was kill -9 on Patroni and all PostgreSQL processes of that node.

Scenario

Acknowledged writes

Lost after failover

Longest write gap

Planned switchover

3,306

0

1.56 s

Async, primary node crash, replicas caught up

3,952

0

29.24 s

Async, WAL senders frozen 4 s, then node crash

4,097

341

27.04 s

synchronous_mode: on, same freeze and crash

3,279

0

32.79 s, stalled from the freeze

etcd frozen 45 s, failsafe_mode off

3,719

0

28.15 s, primary demoted itself

etcd frozen 45 s, failsafe_mode on

5,728

0

0.13 s

Async, only Patroni killed on the primary

6,543, all on the old primary

3,298 existed only on the old primary

none; it kept writing

Crash failovers took about 30 seconds, as the timing rule predicts. With asynchronous replication, everything committed during the stall was lost; synchronous mode lost nothing, but the writer waited from the freeze until the new primary took over. Without failsafe mode, a frozen etcd became a half-minute write outage on a healthy database.

The last row is the one to remember. With Patroni killed and PostgreSQL still running, a replica was promoted about 40 seconds later, two servers were out of recovery, and my writer’s open connection kept committing to the old one. When it rejoined, pg_rewind reported diverged timelines and rewound it, and the 3,298 rows were gone. With synchronous mode on, the same test promoted a replica after about 29 seconds, and the writer’s commit on the old primary hung in the SyncRep wait, so nothing more was acknowledged there. A watchdog, or a load balancer that cuts sessions when the REST API stops answering, closes this gap.

Every old primary rejoined through pg_rewind without a fresh base backup. One harness error was mine: my first crash script killed the postmaster before its frozen WAL senders, which kept shared memory busy until I killed them by hand.

Honest limits: one VM, a single etcd member, loopback networking, tiny rows at about 100 a second, one run per scenario. SIGSTOP stands in for real network faults and paused VMs. I did not test HAProxy, poolers, Kubernetes, CloudNativePG or any operator, a watchdog, or real disks under load. The numbers describe this setup, not production.

What to unlearn and re-learn

  • Unlearn “automatic failover means no data loss”. Re-learn that the default is asynchronous, and decide per database whether to lose commits or stall writes.
  • Unlearn “the old primary stops when it loses the lock”. Re-learn that fencing needs a live process, a watchdog or the platform, and that open connections must be cut.
  • Unlearn “the operator handles backups”. Re-learn that HA and PITR are separate, and the backup tool has its own maintainers, licence and advisories.
  • Unlearn “operator upgrades are routine”. Re-learn that each one is a switchover of every cluster it manages.

Revisit your PostgreSQL HA before your next failover

The Bengaluru team does not need a new platform first; it needs to know which commits its current setup can lose, and why. Learn how your HA manager elects a primary, fences the old one and behaves when the store or network fails. Unlearn the comfort of a cluster that has failed over twice in five years without complaint. Re-learn synchronous replication, slots, pg_rewind and the backup chain from the current docs and release notes. Practise with a frozen replica, a crashed node and a killed Patroni, as I did on one VM, and count the acknowledged writes. Then apply what you find: a fencing method you have tested, a loss budget you can defend in writing, and a patch calendar covering PostgreSQL, the HA manager, the DCS and the backup tool. Revisit it whenever versions, platforms or the people on call change.

Sources

comments powered by Disqus

Releted Posts

Fluent Bit, Vector and the OpenTelemetry Collector: revisit your log pipeline before you ship more logs

A logistics company in Pune runs about 600 pods across three Kubernetes clusters. A Fluentd DaemonSet, set up in 2019 with a dozen Ruby plugins, ships about 25 crore log lines a day to Loki.

Read more

Vector search in 2026: revisit pgvector, Qdrant, Milvus, Weaviate before you add a vector database

An e-commerce company in Pune runs customer support on PostgreSQL. The database holds about 2 million help articles, product Q&A threads and resolved tickets, in English and a fair amount of Hinglish.

Read more

PgBouncer and its alternatives: revisit your PostgreSQL connection pooler

A team runs about forty services on Kubernetes against one PostgreSQL primary. During a sale, the autoscaler adds pods, each opens its own pool of ten connections, and PostgreSQL starts refusing logins with “remaining connection slots are reserved”.

Read more